Tutorials are roughly 2 hours in length and focus on a particular topic or software package. The sessions are more interactive than a standard lecture, often encouraging participants active engagement and hands-on participation.
T1 | Data-Driven Storytelling with Quarto: From Research Question to Accessible Reporting
Required Level of Statistics or Programming: Basic exposure to R computing environment
Course Description:
Effective scientific communication relies on our ability to take data and tell a compelling story to drive impact and make actionable decisions. This short course introduces “data-driven storytelling” using Quarto, an open-source platform that seamlessly combines data analysis, visualization, and communication in a single, reproducible document.
Working with a provided dataset, participants will walk through a complete data analysis workflow. We will begin by refining a topic idea into a “SMART” research aim and outline a statistical analysis plan. Next, participants will learn how to implement this analysis in R using Quarto, including basic data wrangling, exploratory data analysis, and the proposed research analysis.
The course will then demonstrate how to communicate our statistical results clearly using an audience-appropriate output. Participants will create informative tables and visualizations and learn practical strategies for making figures intuitive and accessible. Finally, we will draft a brief report or presentation in Quarto that weaves together text, code, and results into a coherent narrative aimed at non-technical or broadly interdisciplinary audiences.
Throughout, the focus is on conceptual understanding and communication rather than advanced programming. Participants should have basic familiarity with R (e.g. working in RStudio, running simple scripts, reading in data), but no prior experience with Quarto is expected. All Quarto features will be introduced step-by-step, with code examples provided and explained. Before the course begins, instructors will administer a brief survey to gauge participants’ computing and statistical backgrounds. This information will be used to calibrate the level of technical detail, examples, and pacing. By the end of the course, participants will have:
Instructors:
Andrea Lane, Duke University
Kim Webb, University of Pittsburgh
Instructor Biographies:

Kimberly A. H. Webb, PhD, is an Assistant Professor of Medicine in the Division of General Internal Medicine. Prior to joining the University of Pittsburgh, Dr. Webb developed novel bias-correction methods for misclassified outcome and mediator variables in observational studies. She used these methods to study patterns of misdiagnosis in diseases like myocardial infarction and gestational hypertension, as well as algorithmic fairness concerns in the pretrial detention system. In addition to her methodological contributions, Dr. Webb is interested in applied research in social determinants of health and healthcare decision-making. Dr. Webb was a varsity swimmer in college and now likes to stay active by hiking, running, and swimming with the Pittsburgh Elite Aquatics “Masters” group. She also enjoys baking and trying new restaurants.

Andrea Lane is an Assistant Professor of the Practice in the Social Science Research Institute at Duke University. Andrea teaches courses in the Master in Interdisciplinary Data Science (MIDS) program and the Department of Statistical Sciences. She completed her PhD in Biostatistics at Emory University, and her research focuses on analyzing complex health data. She has collaborated with teams in a variety of biomedical application areas and clinical domains, including anesthesiology, “omics” data, critical care, and community interventions.
T2 | What Works for Whom: Estimating and Interpreting Heterogeneous Treatment Effects in Single- and Multi-Study Data
Required Level of Statistics or Programming: Baseline understanding of causal inference principles and working knowledge of R/R Markdown
Course Description:
Modern clinical, public health, and policy research increasingly seeks to move beyond average treatment effects and identify for whom interventions are most effective. Understanding heterogeneous treatment effects (HTE) has become central to advancing patient-centered care and evidence generation in real-world settings but requires substantially larger sample sizes and can be challenging to identify given multiple possible moderators. The increased use and availability of large-scale observational data and pooled data resources across multiple studies has created new opportunities to investigate HTE. Methodological development has followed suit, with a rapidly expanding set of flexible, non-parametric machine learning methods available for HTE estimation. Despite their flexibility, these techniques can be challenging to interpret and validate in practice, leading researchers to often fall back on traditional subgroup analyses and leave many potential moderators unexplored due to multiple testing concerns. This interactive tutorial will provide a practical introduction to modern machine learning methods for HTE estimation, with an emphasis on interpretation of findings. The session will begin with a brief conceptual overview of HTE, key assumptions, and differences between traditional subgroup analyses and modern machine learning approaches. Participants will then learn how to implement some of the most popular and readily available modern approaches: causal forests and Bayesian Additive Regression Trees (BART). We will incorporate hands-on coding exercises in R/RStudio using R Markdown files and a workable dataset. Participants will work through examples of HTE estimation in both single-study and multi-study settings, highlighting how information can be leveraged across studies and methods adapted accordingly to improve estimation and support broader evidence generation. Tools for interpretation of results will include visualization of conditional average treatment effect estimates and their uncertainty, post hoc descriptive approaches such as regression trees and linear projections, variable importance metrics, and methods for distinguishing meaningful heterogeneity from noise. Participants will leave the tutorial with practical skills for implementing modern HTE methods, interpreting treatment effect estimates, conducting checks for robustness of identified moderators, and applying these approaches to clinical and policy-relevant questions involving real-world data from individual or multiple studies.
Instructors:
Carly Brantner, Duke University
Elizabeth Stuart, Johns Hopkins Bloomberg School of Public Health
Instructor Biographies:

Carly L. Brantner, Ph.D. is an Assistant Professor in Biostatistics and Bioinformatics at Duke University and the Duke Clinical Research Institute. Her research centers on drawing conclusions from real-world data, ranging from electronic health records, mobile health apps, to wearable devices. She is particularly focused on leveraging modern statistical methods to estimate how treatment effects vary across individuals, because medical care is rarely one-size-fits-all. Dr. Brantner collaborates primarily in women’s health, pediatric health, and team science, drawing on data from EHR systems, the PCORnet® network, and platforms like Natural Cycles and Oura.

Elizabeth A. Stuart, Ph.D. is the Frank Hurley and Catharine Dorrier Chair and Bloomberg Professor of American Health in the Department of Biostatistics at the Johns Hopkins Bloomberg School of Public Health, with joint appointments in the Departments of Mental Health and Health Policy and Management. She was previously Executive Vice Dean for Academic Affairs at the School. Her research interests are in design and analysis approaches for estimating causal effects in experimental and non-experimental studies, including questions around the external validity of randomized trials and the internal validity of non-experimental studies, as well as methods for combining data sources to assess treatment effect heterogeneity and for evidence synthesis. She is an elected fellow of the National Academy of Medicine, the American Statistical Association, and the American Association for the Advancement of Science.
T3 | Crafting Presentations That Connect: A Tutorial for Statisticians
Required Level of Statistics or Programming: Participants should have a basic understanding of statistical inference and regression modeling at the graduate level. Prior exposure to Bayesian statistics will be helpful but is not required. Familiarity with R will be useful for following the software demonstrations, but advanced programming skills are not required.
Course Description:
This tutorial will provide an accessible and practical introduction to Bayesian information borrowing methods for clinical trials. We will begin with an overview of widely used approaches, including the power prior, commensurate prior, meta-analytic prior, and self-adapting mixture (SAM) prior, highlighting their key ideas, assumptions, and operating characteristics. We will then discuss extensions of these methods to more complex clinical trial settings, including longitudinal outcomes, missing data, ordered external data sources, and the joint modeling of multiple outcomes. These extensions will be motivated by real clinical trial examples to illustrate when and how information borrowing can be effectively implemented in practice. Software tools and practical implementation strategies will also be demonstrated.
Instructor:
Yong Zang, Indiana University School of Medicine
Instructor Biography:

Yong Zang, PhD, is a Showalter Scholar Associate Professor in the Department of Biostatistics and Health Data Science at the Indiana University School of Medicine, where he also serves as Co-Director of Clinical Research for the Biostatistics and Data Management Core at the IU Simon Comprehensive Cancer Center. He received his Ph.D. in Statistics from the University of Hong Kong and completed postdoctoral training at The University of Texas MD Anderson Cancer Center. His research focuses on clinical trial design and health data informatics, with over 90 peer-reviewed publications supported by the National Institutes of Health, the Showalter Trust, and Eli Lilly.
T4 | NIH R01 Applications: What I Wish Someone Had Told Me
Required Level of Statistics or Programming: None
Course Description:
NIH R01 applications are among the most consequential and least transparently discussed challenges in an academic career. Existing resources tend to focus on formatting requirements and review criteria, which, while useful, leave many applicants unprepared for the subtler decisions that often determine success or failure.
This tutorial draws from direct, accumulated experience on the applicant side. I have submitted NIH grants as PI or MPI more than 20 times over the past decade, with outcomes spanning the full range: strong scores that went unfunded, cycles with three simultaneous submissions, and lessons from applications I would approach very differently today. Having also served as a reviewer, I draw from both perspectives to offer practical, candid guidance grounded in what the process actually looks like from where most attendees will be sitting.
The tutorial covers five interconnected areas. Idea development: how to assess whether a project is genuinely R01-ready and what preliminary data actually needs to look like. Strategy and logistics: team assembly, institute and study section selection, and time management across the submission cycle. Grantsmanship: the structure, depth, and content choices that distinguish competitive applications. Interpreting reviews: how to read summary statements, understand reviewer language, and have a productive conversation with your program officer. Resubmission: when to revise, when to let go, and how to carry ideas forward across applications.
The goal is not to prescribe a formula but to share the specific lessons, including the mistakes, that accumulated over years of applying. Some perspectives will be unconventional. All are grounded in direct experience.
This tutorial is intended for postdoctoral fellows and junior faculty preparing their first NIH method grant who want honest, experience-based guidance that goes beyond what the funding opportunity announcement already says.
Instructor:
Gen Li, University of Michigan
Instructor Biography:

Dr. Gen Li is an Associate Professor in the Department of Biostatistics and a co-director of the Cancer Data Science Shared Resources at the University of Michigan. Before joining Michigan, he was an Assistant Professor and Sanford Bolton Faculty Scholar in the Department of Biostatistics at Columbia University. Dr. Li’s research interests lie in the development of new statistical methods for complex biomedical data, including multi-way tensor array data, multi-view data, and compositional data, with applications to omics studies. His research is supported by several NIH grants.