Advancing Biostatistics Through Research and Collaboration
Secure Your RegistrationThe UP School of Statistics invites statisticians, biostatisticians, epidemiologists, students, practitioners of statistics in the industry and government, and other enthusiasts to the Biostatistics and Health Data Science Summit 2026. The event is a two-day conference that features talks from prominent researchers in biostatistics and epidemiology from prestigious international and local institutions, as well as early-career researchers in biostatistics.
Engage with international and local researchers advancing the fields of biostatistics
Professor
Associate Professor
Research Associate
Associate Professor
Assistant Professor
Assistant Professor
Assistant Professor
Assistant Professor
Assistant Professor
Assistant Professor
A timeline of interactive keynotes, technical sessions, and academic peer networking panels
Distribution of conference materials, badges, and credentials at the main lobby.
Introduction and welcoming for to mark the beginning of the conference.
Short pause to stretch your legs.
Presentation and talk from Dr. Thomas Mathew.
Short pause to stretch your legs.
Presentation and talk from Dr. Yehenew Kifle.
Taking of pictures and videos.
Time to eat lunch.
Presentation and talk from Dr. Thomas Mathew.
Short pause to stretch your legs.
Presentation and talk from Dr. Yehenew Kifle.
Short pause to stretch your legs.
Time to engage with others at the lobby!
For faculty and speakers
Distribution of conference materials, badges, and credentials at the main lobby.
Presentation and talk from Dr. Stephen Jun Villejo.
Short pause to stretch your legs.
Presentation and talk from Dr. Michael Daniel Lucagbo.
Short pause to stretch your legs.
Presentation and talk from Dr. Alvin Duke Sy.
Presentation and talk from Asst. Prof. Iana Michelle Garcia.
Time to eat lunch.
Presentation and talk from Asst. Prof. Mark Anthony Javelosa.
Presentation and talk from Asst. Prof. Lian Mae Tabien.
Short pause to stretch your legs.
Presentation and talk from Asst. Prof. John Gilbert Rabanal.
Presentation and talk from Asst. Prof. John Eustaquio.
Short pause to stretch your legs.
The official closing of the conference.
Showcasing academic studies driving meaningful innovation in biostatistics and its applications
In hypothesis testing problems, the null hypothesis is often a statement of equality. The two-sample t-test is a classic example where the null hypothesis states that two population means are equal. In some applications on the comparison of two means, it is more appropriate to test if the two means are equivalent; i.e., if they are close according to a specified threshold. Many applications call for testing equivalence, rather than equality, where the statement of equivalence is the alternative hypothesis so that equivalence is concluded by rejecting the null hypothesis. In the talk I will give a few applications where equivalence testing is required. These will include testing bioequivalence, testing the equivalence of several normal means relevant to the comparison of several treatments in a clinical trial, testing the equivalence of covariance matrices, comparing drug dissolution profiles, and equivalence testing for food safety assessment. All of the problems will be motivated and illustrated using practical applications.
READ ABSTRACTUnderstanding how large hydroelectric dams influence malaria risk is important for public health planning in malaria-endemic regions. This study investigates the association between household proximity to the Gilgel Gibe hydroelectric dam in Ethiopia and time to malaria infection. The analysis is complicated by clustering of individuals within villages and confounding between village membership and distance from the dam.
We compare marginal, fixed-effects, stratified, and frailty survival models for analyzing the clustered time-to-event data. The models produce notably different estimates of the effect of distance from the dam. In particular, the standard frailty model estimates a weighted combination of within- and between-village distance effects, which can be misleading when these effects differ. A between-within frailty formulation provides a useful approach for separating these sources of variation. The findings highlight the importance of appropriately accounting for clustering and covariate structure when analyzing spatially clustered infectious-disease survival data.
Identifying treatments or interventions that are cost-effective (more effective at a reasonable cost) is clearly important in health policy decision making, especially in the allocation of health care resources. Various measures of cost-effectiveness that are informative, intuitive and simple to explain have been suggested in the literature. Popular and widely used measures include the incremental cost-effectiveness ratio (ICER), defined as the ratio between the difference of average costs and the difference of average effectiveness in two populations receiving two treatments. The ICER is interpreted as the additional cost per unit of effectiveness gained. Yet another measure is the incremental net benefit (INB), which is the difference between the incremental cost and the incremental effectiveness after multiplying the latter with a “willingness-to-pay” amount. In the talk, I will provide a selected review of the statistical criteria and methodologies for cost-effectiveness analysis. In particular, some of the recently introduced probabilistic criteria will be discussed and examples will be given.
READ ABSTRACTCombining evidence across independent studies is a fundamental problem in biostatistics, particularly when studies share a common underlying effect but exhibit heterogeneous and unknown variances. Such settings arise naturally in meta-analysis, multi-center studies, laboratory investigations, and the integration of health data from multiple populations or data sources. This work investigates exact inference for a common mean across independent normal populations with unequal variances and compares the local power properties of six exact testing procedures.
We consider several approaches for combining study-specific evidence, including Tippett’s and Wilkinson’s procedures, Fisher’s method, the inverse normal method, and modified tests based on weighted combinations of component test statistics. Analytical expressions for the local power of the six procedures are derived under alternatives close to the null hypothesis. For equal sample sizes, the resulting expressions permit a uniform comparison of the tests irrespective of the unknown population variances. Numerical evaluations further examine their performance under both equal and unequal sample-size settings.
The results show that the modified exact tests consistently provide the strongest local power among the procedures considered, while the inverse normal method also performs competitively and offers computational simplicity. These findings provide practical guidance for combining evidence across heterogeneous studies and highlight the importance of selecting statistically efficient procedures when integrating information from multiple biomedical and health data sources.
This talk comprises two major parts, both centred around the overarching theme of real-time disease surveillance. The key statistical challenge I address is the spatial and temporal misalignment of the data. The first part considers the problem of estimating a latent spatially continuous process from aggregated outcome data. I adopt a block aggregation framework that separates the latent process model from the sampling model, and estimate it using a linearised version of Integrated Nested Laplace Approximation (INLA). I illustrate this approach by predicting SARS-CoV-2 viral load, measured as the number of gene copies, over arbitrary areal configurations in the United Kingdom using aggregated data from the catchment areas of a sparse network of sewage treatment works. To address the difficulty arising from the fact that log-Gaussian outcomes are not closed under aggregation, I use moment matching to estimate the sampling model parameters and propose an iterative procedure to adjust the non-linear predictor expression and the scaling of the likelihood precision. This modelling approach provides a coherent framework for inference, prediction over new areal units, and spatial disaggregation. The second part of my talk focuses on temporal misalignment in lagged covariate-response relationships, particularly in settings where the response variable is observed at a lower frequency than the predictors. In disease surveillance, for example, clinical outcomes may be recorded weekly and forecast at a one-week horizon, whereas environmental indicators such as pollution levels and wastewater viral loads are available daily. To address this, I extend the Mixed-Data Sampling framework to a spatial setting through a Spatial Distributed-Lag Mixed-Data Sampling (SDL-MIDAS) model. I discuss an efficient Bayesian computational tool for fitting these models, together with an accompanying R package that I am currently developing. An application from Wales shows that SDL-MIDAS models improve forecasts of weekly hospital admissions and positivity rates using daily wastewater signals as predictors. Overall, this talk discusses flexible and principled approaches for handling spatial and temporal misalignment, with important applications in integrating environmental surveillance data into public health monitoring systems.
READ ABSTRACTReference intervals are often considered the most widely used medical decision-making tool. They are used to assess the health status of individuals based on laboratory measurements. Individuals whose analyte values fall inside the reference interval are diagnosed as being healthy. On the other hand, those whose values fall outside the interval may be considered as having abnormal health status. Reference intervals are computed so that they contain the central 95% of the healthy population or reference population. Therefore, the endpoints of the reference interval are the 2.5th and 97.5th population percentiles. Since these percentiles are unknown in actual practice, the reference interval is estimated from a random sample. When using estimated percentiles, the resulting reference interval will cover less than 95% of the population. To address this problem of liberal coverage, a recommended solution is to use a 95% prediction interval.
When multiple analytes are needed to diagnose a complex condition accurately, simultaneous prediction intervals whose joint coverage is 95% are needed. In this presentation, we consider several approaches for constructing simultaneous prediction intervals under a broad range of settings, including the multivariate normal and nonparametric settings. The methods to be presented accommodate correlation among analytes, which is particularly relevant when multiple biomarkers are used to assess a complex medical condition. Furthermore, we extend the methodologies to the case where covariate information is available. Finally, we evaluate the methodologies using Monte Carlo simulations that report the estimated coverage probabilities.
Rare disease research presents methodological challenges that are seldom illustrated in methodological textbooks. Using a single-center Philippine gestational trophoblastic disease registry as a case study, this lecture reflects on a series of analytical decisions made while investigating time to remission under the constraints of limited data. Rather than comparing statistical methods exhaustively, the presentation highlights how different approaches addressed different questions encountered during the course of the research, and discusses the methodological considerations that informed those choices. The lecture is intended for researchers interested in the practical application of statistical methods in rare disease and small-sample clinical research.
READ ABSTRACTIn clinical laboratory medicine, analyte measurements are commonly evaluated using univariate reference intervals. However, when multiple analytes are required for diagnosis, applying separate univariate reference intervals increases the occurrence of false-positive results. Multivariate reference regions (MRRs) are more appropriate in this context because they account for the cross-correlations among analytes. Compared with conventional ellipsoidal MRRs, rectangular MRRs are easier to interpret and facilitate the identification of outlying analyte measurements. Prediction- and tolerance- based approaches have been proposed for constructing rectangular MRRs, with tolerance regions possessing a multiple-use property that makes them suitable for repeated application across multiple patients. In addition, analyte reference limits may be two-sided or one-sided depending on the biomarker and disease context. This paper focuses on regression-based rectangular MRRs constructed using multivariate tolerance criteria while accounting for covariate information. Specifically, we present regression-based multivariate tolerance regions, multivariate central tolerance regions, and mixed-sided multivariate tolerance regions. The performance of these proposed approaches is evaluated through simulation studies, and results show that all methods achieve coverage probabilities close to the nominal level. Finally, we apply the proposed approaches to a clinical laboratory dataset.
READ ABSTRACTEvery time a patient buys medicine at a local drugstore, a complex chain of data is set into motion. For pharmaceutical companies operating in the Philippines, tracking whether their products are succeeding requires navigating a massive, interconnected puzzle of information. In this talk, we shall be introduced to the fascinating world of commercial market intelligence and the diverse data ecosystems that power a multi-billion-peso industry.
We will unpack the primary data streams that companies use to monitor their Key Performance Indicators (KPIs) and outmaneuver competitors. The session will cover three essential data sources on which companies primarily rely for performance monitoring. These are the third-party syndicated data audits, primary distributor invoice logs and secondary point-of-sale data, and lastly, sales force and CRM data.
Decision limits play a crucial role in laboratory medicine by guiding diagnostic and clinical decision-making processes. Unlike reference intervals, decision limits are often one-sided thresholds, typically expressed as prediction or tolerance intervals. However, in complex diagnoses requiring multiple correlated analyte measurements, constructing separate univariate intervals may increase false positives. It is recommended to construct multivariate statistical regions to account for the cross- correlations among analytes. Hence, this research proposes a Bayesian multivariate regression framework for constructing combined statistical decision limits that incorporates covariates (e.g., age and sex). The methodology operates under a multivariate normal model and allows for prior specification using both non-informative and conjugate priors. Results from simulation studies indicate that limits derived under non-informative priors have highly satisfactory frequentist properties and outperform benchmark methods even under weak cross-correlations. In contrast, limits based on conjugate priors tend to underperform in small samples due to prior influence, although this effect diminishes with larger sample sizes. The proposed method is also computationally efficient, producing decision limits in under a minute, highlighting its practical feasibility for clinical application. Overall, this Bayesian multivariate approach provides a principled and adaptable framework for generating accurate decision limits in complex diagnostic contexts.
READ ABSTRACTQuantitative studies about young adolescent pregnancy in the Philippines had been done through the analysis of survey data. Official birth records provide a more advantageous data source in examining the dynamics of the young adolescent pregnancy due to regular recording of births. In this study, we propose a statistical methodology utilizing official birth records through a Bayesian hierarchical framework. The method accounts for the unique features of the aggregated young adolescent birth counts such as being a spatiotemporal data with known upper bounds and being characterized by the presence of excess zero counts. Our framework consists of the binomial hurdle model and Poisson hurdle model with offset, the Leroux model for spatial effects (Leroux et al., 2000), and an additive nonparametric trend model for temporal effects (Knorr-Held, 2000). We analyze and assess the performance of the proposed models using the quarterly data on young adolescent birth counts in Luzon, Philippines from 2006 to 2019 via Hamiltonian Monte Carlo. Areas that are more prone to observe a young adolescent birth and a temporal pattern about the occurrence of young adolescent births are identified.
READ ABSTRACTUnderstanding the causal mechanisms behind biomedical research oftenly requires detaching how exposures exert their effects on health outcomes over time through intermediate processes. However, in longitudinal studies, both exposures and confounders frequently vary with time which creates a complex feedback structures that would make standard mediation approaches invalid. This study developed a novel framework for causal mediation analysis in longitudinal biostatistical data with time-varying exposures and confounders by integrating marginal structural models (MSM) with structural nested models (SNM), and targeted maximum likelihood estimation (TMLE) in a single causal inference paradigm. A flexible estimation strategy was proposed that utilizes machine learning–based nuisance estimation to improve the robustness and efficiency but still maintaining valid causal interpretation under realistic longitudinal data structures. Simulation studies demonstrated that the proposed method provides unbiased estimates of direct and indirect effects even under complex time-dependent confounders. The methodology is applied to simulated data mimicking a large cohort study of HIV treatment dynamics, where the causal pathways linking antiretroviral therapy adherence, immune response, and virologic suppression are examined. The results shown substantial time- varying mediation effects which highlight the importance of accounting for temporal effect in biomedical causal inference.
READ ABSTRACTHosted within the state-of-the-art technical hubs of UP Diliman
UP School of Statistics, Quirino Ave. cor. T.M. Kalaw Street, UP Diliman, Quezon City, Philippines
The conference will be held at the UP School of Statistics Auditorium, a hub for statistical research and education in the Philippines, offering an engaging environment for scientific dialogue and knowledge sharing.
The host institutions bringing together global statistical innovation
The UP School of Statistics is a leading center for statistical research, education, and innovation in the Philippines. Through its academic programs and research initiatives, the School advances the development and application of statistical methods across diverse fields, including public health, economics, social sciences, natural sciences, data science, and more.
The Biostatistics Research Group of the UP School of Statistics focuses on the development and application of statistical methods to address challenges in health, medicine, and public health research. Through methodological innovation, interdisciplinary collaboration, and research training, the group advances the role of biostatistics in generating evidence and informing health-related decisions.
Answers to common questions regarding participation and administrative guidelines
The conference is open to researchers, faculty members, students, practitioners, and professionals interested in biostatistics, epidemiology, statistical computing, data science, and related fields. Participants from diverse backgrounds are welcome to engage in discussions, share insights, and build collaborations.
Participants can look forward to a two-day program featuring expert talks, research presentations, interactive sessions, and opportunities for networking and collaboration. The conference will highlight current developments and emerging challenges in biostatistics, epidemiology, and statistical computing.
For professionals, the registration fee is Php 1,000. Meanwhile, for students (undergraduate or graduate), the fee is Php 500. The registration fee covers snacks, lunch, and the conference kit. Available payment options are listed in the registration form.
Yes, participants who will attend at least one (1) day of the summit will be given certificates of participation.
As the inaugural BHDSS, this edition will feature invited presenters only. Future iterations of the conference will include opportunities for open submissions from the broader research community.
Reach out to the organizing committee for any inquiries regarding the conference
Seating capacities within the auditorium are limited. Secure your enrollment profile immediately.
Register for the Conference