Health data science, learned and built in the open.
In-depth guides to biostatistics, epidemiology and R; end-to-end projects on synthetic Alberta health data; and the tools that make the work repeatable.
What the lab covers
Every guide is a self-contained page with worked examples and runnable code. Most also have a Chinese edition.
Biostatistics
From descriptive statistics to causal inference: a full applied-statistics sequence for health research, each guide with runnable R code.
14 guides →Epidemiology
A population view of health: disease frequency, study design, bias, and modern causal methods.
3 guides →Data Visualization
Static, publication-ready graphics and interactive browser charts in R.
2 guides →Reproducible Reporting
Narrative, code and results in one document that rebuilds itself.
1 guide →Data Platforms & GIS
Working with data where it lives: cloud warehouses and spatial analysis.
2 guides →R & Python Tooling
VLabR, a small R package for cleaning, summarizing and de-identifying health data — plus Python practice notes.
Explore tools →Featured project
Synthetic Alberta emergency & inpatient data
A consulting-style demonstration built on fully synthetic NACRS and DAD records — no real patients — studying two linked questions in order.
- Does triage acuity (CTAS) and time to disposition predict hospitalization?
- Among those admitted, what drives inpatient length of stay?
- Seven consulting memos (A00–A06) and runnable R examples
Recently updated
- VLabRR package for data cleaning, summary tables and de-identification
- Project 001Synthetic NACRS & DAD consulting study, memos A00–A06
- Python LabBeginner Python exercises
- Getting Started with ArcGISNew bilingual GIS course
- Biostatistics guides14 in-depth statistics guides