Short Courses are offered as full- or half-day courses. The extended length allows attendees to gain an in-depth understanding of the topic. These courses often integrate seminar lectures covering foundational concepts with hands-on lab sessions that allow users to implement these concepts into practice.
SC1 | Empowering Your Study Through Federated Learning: the PDA framework
Required Level of Statistics or Programming: Entry-level statistical modeling, R program
Duration: Half-day
The rapid expansion of digital health data has transformed medical research, providing unprecedented opportunities to generate evidence from large-scale and multi-institutional electronic health records (EHR). Yet fully leveraging these data remains challenging: protecting patient privacy, managing high-dimensional and heterogeneous information, and enabling meaningful collaboration across institutions all pose substantial hurdles.
To address these challenges, our team has developed Privacy-preserving Distributed Algorithms (PDA) (https://pdamethods.org/) —a suite of federated learning tools that enable rigorous multi-site analyses without sharing individual patient data. PDA supports a broad range of statistical and machine learning tasks, spanning association studies, causal inference, clustering, and variable selection and prediction. PDA has been successfully deployed across data-centric networks such as OHDSI, PCORnet, the International Agency for Research on Cancer, and the NIH RECOVER Initiative, demonstrating its utility in pharmacoepidemiology, early disease prediction, subphenotyping, and hospital performance evaluation.
In this short course, we will also introduce our new framework MOSAiC (Multi-site One-Shot Aggregation of Compressed Risk Functions)—a unified paradigm for modern distributed research networks. Multi-site studies increasingly underpin biomedical discovery, yet inference across sites remains constrained by privacy regulations, data heterogeneity, sparse events in smaller centers, and the practical burden of multi-round communication. MOSAiC reframes federated learning as a mathematical problem of compressing and aggregating local risk functions, leveraging advances in tensor networks, a state-of-the-art technique for high-dimensional function approximation.
MOSAiC achieves four properties previously unattainable in federated algorithms beyond linear models: one-shot communication, lossless recovery of pooled-data estimates, inclusion of all sites regardless of size or event sparsity, and analytic submodel exploration without recontacting partners. We will illustrate MOSAiC’s validity and efficiency through applications in drug relabeling, drug repurposing, and post-market safety surveillance. Together, PDA and MOSAiC provide a powerful foundation for enabling privacy-preserving, scalable, and scientifically robust multi-institutional research.
This short course contains a 1.5-hour interactive session demonstrating the usage of the PDA over-the-air (OTA) platform to conduct an example federated learning project. Participants are encouraged to bring their laptops and learn to use the PDA software and OTA platform.
Instructors:
Yong Chen, University of Pennsylvania
Thomas Trikalinos, Brown University
Raymond Carroll, Texas A&M University
Chongliang Luo, Washington University in St. Louis
Instructor Biographies:

Yong Chen, PhD, is Professor of Biostatistics and Informatics in the Department of Biostatistics, Epidemiology, and Informatics at the University of Pennsylvania Perelman School of Medicine and founding director of the Center for Health AI and Synthesis of Evidence. His research develops methods for evidence synthesis, causal and federated learning, privacy-preserving distributed analysis, and real-world data integration. He is an elected Fellow of the ASA, ACMI, AMIA, ISI, and Society for Research Synthesis Methodology.

Thomas A. Trikalinos, MD, PhD, is Professor of Health Services, Policy and Practice, with a secondary appointment in Biostatistics, and director of the Center for Evidence Synthesis in Health at Brown University. His work advances systematic review, meta-analysis, comparative effectiveness research, decision analysis, and decision-making under deep uncertainty.

Raymond J. Carroll, PhD, is Distinguished Professor and the Jill and Stuart A. Harlin ’83 Chair in Statistics at Texas A&M University, where he directs the Institute for Applied Mathematics and Computational Science. He is internationally recognized for foundational contributions to measurement-error models, nonparametric and semiparametric regression, and statistical methods in nutritional epidemiology and genomics.

Chongliang (Jason) Luo, PhD, is a biostatistician and Assistant Professor of Surgery in the Division of Public Health Sciences at Washington University in St. Louis. His research focuses on data integration, multi-view learning, meta-analysis, and federated learning.
SC2 | Multimodal Integration and Multimodal Causal Inference using R/Bioconductor
Required Level of Statistics or Programming: Master's level knowledge of statistics or biostatistics (regression modeling and basic inference) and working knowledge of R. No advanced programming experience required.
Duration: Half-day
Course Description:
This half-day short course provides a practical, hands-on introduction to multimodal integration and multimodal causal inference.
The first half covers multimodal integration methods. The second half introduces causal inference in the multimodal setting. Throughout, methods are demonstrated in R and Bioconductor using publicly available multi-omics datasets, with reproducible code that attendees can adapt to their own studies.
Participants will be able to select an appropriate integration strategy for a given study design, implement it in R/Bioconductor, and understand when and how causal conclusions are warranted.
Instructor: Himel Mallick, Cornell University
Instructor Biography:

Himel Mallick is a tenure-track Principal Investigator at Cornell University. Prior to Cornell, he was an Associate Director at Merck Research Laboratories and a postdoctoral fellow at Harvard University. He is the creator of widely used open-source statistical software and multimodal AI tools, most notably MaAsLin, a community standard for microbiome multi-omics association analysis. Himel's research spans Bayesian statistics, machine learning, and multimodal AI for omics data science, including microbiome, single-cell, spatial omics, and digital pathology. He is a Fellow of the American Statistical Association, an Elected Member of the International Statistical Institute, and a recipient of multiple early career awards in data science and medicine.
SC3 | Modern AI for Biostatistics
Required Level of Statistics or Programming: Participants should have graduate-level training in statistics, biostatistics, data science, epidemiology, computer science, or a related quantitative field. Familiarity with regression, likelihood-based inference, prediction models, and basic machine learning concepts will be helpful. No prior deep learning, reinforcement learning, or large language model experience is required.
Basic programming experience in Python or R is recommended. The course will include optional hands-on examples in Python/PyTorch, but participants who do not wish to code can still follow the conceptual and methodological material. Participants interested in running the notebooks should bring a laptop and have access to a Google account or a local Python environment.
Duration: Half-day
Course Description:
Artificial intelligence is rapidly transforming biostatistics, biomedical data science, and precision health. Modern deep learning methods now provide powerful tools for modeling high-dimensional imaging, genomics, electronic health records, mobile health, and multimodal biomedical data, while reinforcement learning and large language models are creating new opportunities for sequential decision-making, adaptive interventions, clinical reasoning, and AI-assisted scientific discovery. This short course will introduce biostatisticians, statisticians, and quantitative biomedical researchers to the foundations, implementation, and responsible use of modern AI methods in biomedical applications.
The course will begin with a concise introduction to deep learning, covering neural network architectures, optimization, regularization, uncertainty, and practical implementation in Python/PyTorch. We will then discuss modern representation learning and generative modeling, including autoencoders, diffusion models, and transformer architectures, with examples drawn from biomedical imaging, multi-omics, and disease phenotyping. The second part of the course will focus on reinforcement learning for sequential treatment decisions, dynamic treatment regimes, mobile health interventions, and clinical policy evaluation, emphasizing the statistical challenges of confounding, off-policy evaluation, uncertainty quantification, and safety. We will then introduce large language models and biomedical agents, including prompting, retrieval-augmented generation, fine-tuning, alignment, tool use, and agentic workflows for literature review, statistical analysis, clinical documentation, and hypothesis generation.
Throughout the course, methodological concepts will be connected to case studies in biostatistics and biomedical research, such as medical image analysis, precision medicine, imaging genetics, adaptive treatment strategies, and biomedical knowledge extraction. Rather than presenting AI as a black-box prediction toolbox, the course will emphasize statistical thinking: study design, bias, validation, interpretability, reproducibility, calibration, causal considerations, privacy, and trustworthy deployment in health applications.
Participants will gain a working understanding of core deep learning, reinforcement learning, LLM, and agent methods, learn when these tools are appropriate for biomedical problems, and acquire practical guidance for implementing, evaluating, and communicating AI analyses. Hands-on examples and optional coding demonstrations will be provided through accessible Python notebooks. The course is intended to help ENAR participants bridge modern AI methodology with rigorous biostatistical practice.
Instructors:
Chengchun Shi, London School of Economics (LSE)
Bingxin Zhao, University of Pennsylvania
Hongtu Zhu, University of North Carolina at Chapel Hill
Instructor Biographies:

Chengchun Shi is an Associate Professor in the Department of Statistics at LSE. His work brings to light the relevance and significance of statistical learning in AI, and demonstrates the usefulness of RL as a framework for policy evaluation and A/B testing in two-sided marketplaces. His outstanding contributions have been recognized with the Peter Gavin Hall IMS Early Career Prize, IMS Tweedie Award and the RSS Research Prize.

Bingxin Zhao is an Associate Professor at the University of Pennsylvania whose work spans statistical genetics, imaging genomics, and AI for medicine. His honors include the 2025 IMS Tweedie New Researcher Award, recognition in the 2025 UK Biobank Scientific Impact Awards, the 2024 ICSA Outstanding Young Researcher Award, and UNC’s Dean’s Distinguished Dissertation Award. His heart-brain research was named among UK Biobank Imaging’s top ten findings.

Hongtu Zhu is the Kenan Distinguished Professor of Biostatistics, Statistics, Radiology, Computer Science and Genetics at the UNC. He was a DiDi Fellow and Chief Scientist of Statistics at DiDi Chuxing and held the Endowed Bao-Shan Jing Professorship in Diagnostic Imaging at MD Anderson Cancer Center between 2016 and 2018. He received an established investigator award from the Cancer Prevention Research Institute of Texas in 2016, the INFORMS Daniel H. Wagner Prize for Excellence in Operations Research Practice in 2019, the IMS 2027 Medallion award and Lecture, and the COPSS 2025 Snedecor Award.

Zhangzhi (Fred) Peng is a PhD student at Duke University working on generative models and protein design.
SC4 | scorcher: A Tidy, Composable Framework for Building AI/ML Models Without Extensive Coding
Required Level of Statistics or Programming: Participants should have basic familiarity with:
- R and the tidyverse.
- Regression modeling or supervised prediction.
- Train/test splits or cross-validation.
No prior deep learning experience is required. Familiarity with torch, neural networks, or GPU computing is helpful but not necessary.
Duration: Half-day
Course Description:
Modern health and policy research increasingly relies on artificial intelligence and machine learning (AI/ML), yet many scientists, applied biostatisticians, and community-engaged researchers face substantial barriers to developing these models themselves. AI/ML software is often written in Python, requires familiarity with object-oriented programming, and can feel disconnected from the tidy, modular workflows commonly used in R-based data science. This creates a gap between experts who understand the scientific problem and the technical infrastructure required to build, evaluate, and interpret such models.
This short course introduces scorcher, an R framework for building deep learning models using tidy, composable, pipe-friendly syntax. The goal of scorcher is to make deep learning more accessible by presenting neural networks as modular statistical models composed of inputs, transformations, layers, losses, and outputs. Participants will learn how to move from familiar modeling concepts, such as predictors, outcomes, model formulas, fitted values, loss functions, validation, and prediction, to flexible neural network architectures for biomedical data.
The course combines conceptual explanation, live coding, and hands-on exercises. Participants will build simple feed-forward neural networks, extend them to multiple-input and multiple-output architectures, evaluate model performance, and interpret predictions in scientifically meaningful terms. Modules will work through prediction, image-derived classification, and diffusion tasks. We will emphasize reproducible workflows, modular code, and transparent model evaluation.
Instructors:
Stephen Salerno, Washington University in St. Louis
Awan Afiaz, University of Washington
Instructor Biographies:

Stephen Salerno, PhD, is an Assistant Professor at the Bursky School of Public Health at Washington University in St. Louis. His research develops statistical methods, machine learning tools, and open-source software for biomedical and public health applications, with an emphasis on making modern statistical and AI methods more accessible, interpretable, and reproducible. He is a co-developer of scorcher, an R package that provides a tidy, composable interface for building deep learning models with torch.

Awan Afiaz, MS, is a PhD student in Biostatistics at the University of Washington, advised by Jeffrey T. Leek at the Fred Hutch Cancer Center. His research lies at the intersection of statistical inference and AI/ML, including methods for inference with predicted data and multimodal deep learning for clinical applications. He develops open-source statistical software and is a co-developer of scorcher, with a particular interest in making deep learning accessible to researchers without specialized programming backgrounds.
SC5 | Generative AI for Biomedical Data Science: Synthetic Data, Protein Design, and EHR Applications
Required Level of Statistics or Programming: No prior experience with generative models is required, although familiarity with basic statistical modeling and machine learning will be helpful.
Duration: Full-day
Course Description:
Generative AI is rapidly changing biomedical research, creating new tools for synthetic data generation, scientific simulation, and biological design. At the same time, biomedical applications raise statistical challenges that are not fully addressed by standard machine learning workflows, including limited sample sizes, rare outcomes, irregular longitudinal measurements, distribution shift, privacy constraints, fairness concerns, and the need for scientifically meaningful validation. This short course will introduce modern generative modeling methods for biomedical data science, with emphasis on both statistical principles and emerging biological applications. We will begin with core ideas behind generative models, including variational autoencoders, generative adversarial networks, diffusion models, score-based models, and flow-matching methods. The course will then focus on applications in modern biomedical research. Topics will include synthetic data generation for electronic health records and irregular time series, data augmentation for rare outcomes and imbalanced samples, fairness and bias in synthetic biomedical data, generative modeling for single-cell trajectories, and flow-based approaches for protein structure and biological sequence generation. Across these examples, we will discuss when generative models improve downstream learning, when they may introduce bias or instability, and how their outputs should be evaluated. A central theme of the course is responsible and statistically grounded use of generative AI. The course is designed for biostatisticians, biomedical data scientists, and quantitative researchers interested in using, evaluating, or developing generative models for biomedical studies. Participants will leave with a conceptual understanding of major generative modeling frameworks, practical guidance for biomedical applications, and a critical perspective on opportunities in this rapidly evolving area.
Instructors:
Anru Zhang, Duke University
Alex Tong, Aithyra
Fred Zhangzhi Peng, Duke University
Instructor Biographies:

Anru Zhang, PhD, is the Eugene Anson Stead, Jr. M.D. Associate Professor of Biostatistics and Bioinformatics, Computer Science, and Statistical Science at Duke University. His research focuses on high-dimensional statistical inference, tensor learning, non-convex optimization, and generative AI for health data science. He is a recipient of the COPSS Emerging Leader Award, NSF CAREER Award, and Bernoulli Society New Researcher Award.

Alex Tong, PhD, is a Director of Machine Learning at Aithyra and a former postdoctoral fellow at McGill University and Mila - Quebec AI Institute. His research specializes in geometric deep learning, generative diffusion models, and machine learning applications in structural biology and single-cell genomics.

Fred Zhangzhi Peng is a PhD candidate in Biostatistics at Duke University. His research focuses on statistical machine learning, foundation models for electronic health records, and privacy-preserving synthetic data generation.
SC6 | Estimating Correlations and Causal Relationships with Linked Data Sources
Required Level of Statistics or Programming: Participants should have prior experience using R and be comfortable conducting statistical analyses, as the workshop will involve coding exercises and interpretation of methodological output. Prior experience with record linkage is not required.
Duration: Half-day
Course Description:
Informed decision-making depends on complete and accurate information, yet data relevant to health and policy decisions are often fragmented across health providers, insurance systems, private organizations, and government agencies. Record linkage can integrate these sources by identifying records that correspond to the same individual, enabling estimation of associations and causal relationships that may not be possible from any single dataset. Privacy regulations, however, often restrict access to unique identifiers such as Social Security numbers, names, and addresses, requiring linkage based on incomplete, inconsistent, or error-prone identifying information. As a result, linked datasets are inherently subject to linkage error. These challenges are particularly important when overlap between data sources is limited or linkage quality varies across subpopulations, potentially leading to differential error rates. Even modest levels of linkage error can introduce bias and distort marginal and conditional associations.
This workshop will provide an overview of modern statistical methods for conducting inference with linked data while accounting for linkage error and uncertainty. We will consider both primary analyses, where the analyst has access to the original source datasets, and secondary analyses, where only the linked dataset is available. We will discuss methods for estimating both causal effects and associations, emphasizing their assumptions, practical advantages and limitations, and the implications of linkage uncertainty for inference and decision-making. Methods will be illustrated using health-related examples and implemented in R, providing hands-on experience with realistic data settings. The workshop will equip participants with both a conceptual understanding of statistical inference with linked data and practical skills for applying current methods in research.
Instructor: Roee Gutman, Brown University
Instructor Biography:

Roee Gutman, PhD, is an Associate Professor in the Department of Biostatistics at Brown University. His research focuses on causal inference, file matching and record linkage, missing data analysis, and Bayesian computation, with extensive applications in health services research, Medicare claims analysis, and comparative effectiveness studies. He has published widely on methods for combining disparate data sources and correcting for linkage errors in observational studies.
SC7 | Causal Machine Learning for Discovering Heterogeneous Treatment Effects
Required Level of Statistics or Programming: Introductory to medium knowledge in statistics, causal inference and R programming.
Duration: Half-day
Course Description:
This short course provides a comprehensive overview of current state-of-the-art approaches for causal machine learning for the discovery of heterogeneous treatment effects. As precision medicine and personalized policy interventions become increasingly central to research and practice, understanding who benefits most from treatments is crucial for optimizing resource allocation and improving outcomes.
We begin with a focused review of heterogeneous treatment effects estimation methods, distinguishing between approaches that estimate conditional effects versus those that discover (interpretable) subgroups and effect modifiers. The course then introduces methods for discovering heterogeneous treatment effects, emphasizing interpretable machine learning approaches including tree-based methods (causal trees, causal forests, policy trees), rule-based approaches, and recent advances in causal distillation frameworks.
The course will comprise both theoretical discussions of the statistics behind the different methods discussed, as well as a hands-on component, where students will apply the different methods on relevant real-world datasets.
Instructors:
Falco J. Bargagli-Stoffi, UCLA
Ana Kenney, UCI
Melody Huang, Yale
Instructor Biographies:

Falco J. Bargagli-Stoffi, Assistant Professor, UCLA Biostatistics, has extensive teaching experience in causal inference, machine learning, and biostatistics across multiple countries and languages. At UCLA, he teaches causal inference at the PhD, Master's, and undergraduate levels, and has designed short courses for UNU-MERIT United Nations University, the University of Maastricht, the Flemish Ministry of Education, and the University of Florence and Bruno Kessler Foundation.

Ana Kenney, Assistant Professor, UCI Statistics, has over a decade of teaching experience in mathematics, statistics, and interdisciplinary data-science training. At the University of Colorado Denver, she taught college algebra, calculus, and introductory statistics to a diverse population of working professionals and first-generation students. At UC Irvine, she now teaches the second statistical methods sequence at both undergraduate and graduate levels.

Melody Huang, Assistant Professor, Yale Political Science and Statistics & Data Science, has extensive teaching experience in causal inference, statistics, and data science across disciplines. She recently designed and now teaches a PhD-level Causal Inference course at Yale for an interdisciplinary audience, pairing theoretical rigor with hands-on replication exercises using published political science papers.
SC8 | Introduction to Bayesian Methods for Clinical Trial Design and Sample Size Determination for Time-to-Event Data
Required Level of Statistics or Programming: MS in Statistics or Biostatistics
Duration: Full-day
Course Description:
This full-day short course is designed to give biostatisticians and data scientists a comprehensive overview of the use of Bayesian methods for clinical trial design for time to event data as well as training on how these methods can be implemented using standard software. Applications will be demonstrated using Stan, R, and SAS.
The first part of the course gives a broad overview of Bayesian sample size determination with a focus on phase II/III trials. Focus is paid to four concepts that govern sample size determination. Applications will be provided based on actual Phase II/III trials using software implementations written as Stan, SAS (macros) or using R.
The second part of the course will focus broadly on advanced Bayesian trial designs for time to event data. The types of designs considered fall into two broad categories: (1) designs that borrow information through the use of an informative prior specified a priori, and (2) designs that seek to borrow information across subgroups within a single trial. Examples designs of type (1) include trials where the goal may be to show that a next-generation medical device (e.g., a coronary stent) is non-inferior or superior to a previous generation of the same device and designs that seek to extrapolate information on treatment efficacy from adult to pediatric disease settings. In both cases, there are often multiple historical datasets that could inform the prior and attention will be given to techniques that can accommodate multiple information sources and to situations where historical dataset(s) are composed of controls. These are just a few of the real data settings that will be illustrated. Bayesian designs will be developed and illustrated for a wide variety of survival models including parametric exponential, Weibull, and log-normal models, semiparametric piecewise exponential models, parametric and semiparametric cure rate models, models for recurrent events, joint models of longitudinal and survival data, multivariate survival models, survival models for delayed treatment effects, and models for meta-experimental design. For all of these models, case studies will be presented to illustrate Bayesian designs either in the presence of or absence of historical data. In the case of the presence of historical data, the power prior, and normalized power prior, or Robust MAP will be utilized in the Bayesian design.
Instructor:
Joseph Ibrahim, University of North Carolina at Chapel Hill
Instructor Biography:

Dr. Ibrahim is an Alumni Distinguished Professor of biostatistics at UNC Chapel Hill, Director of the Biostatistics and Data Management Core at UNC’s Lineberger Cancer Center, and has several grants funded from NCI.
His research areas are Bayesian inference, missing data problems, and cancer research. With over 35 years’ experience working in clinical trials, he directs the UNC Laboratory for Innovative Clinical Trials. He is also the Director of Graduate Studies in UNC’s Department of
Biostatistics, as well as the Program Director of the cancer genomics training grant in the department.
He was the Editor of JASA – Applications and Case Studies (JASA-ACS) from 2013-2015 and is currently an Associate Editor for JASA-ACS.
He has published over 400 research papers, most in top statistical journals. He has also published two advanced graduate-level books on Bayesian survival analysis and Monte Carlo methods in Bayesian computation. He is an elected fellow of the American Statistical Association, Institute of Mathematical Statistics, International Society for Bayesian Analysis, International Statistical Institute, and was recently awarded the Samuel S. Wilks memorial award in 2024.
For the past 35 years, he has been deeply engaged in Bayesian clinical trials and genomics research, including methods for meta-analysis and network meta-analysis, adaptive methods for clinical trials, and informative prior elicitation based on historical data for Bayesian clinical trials design and analysis.