article Open AccessTop 1% cited
Calculating the sample size required for developing a clinical prediction model
BMJ · 2020 · Vol. 368 · pp. m441–m441
Richard D Riley✉(Keele University)Joie Ensor(Keele University)Kym I E Snell(Keele University)Frank E. Harrell(Vanderbilt University)Glen P. Martin(Manchester Academic Health Science Centre)Johannes B. Reitsma(University Medical Center Utrecht)Karel G.M. Moons(University Medical Center Utrecht)Gary S. Collins(Nuffield Orthopaedic Centre)Maarten van Smeden(Leiden University Medical Center)
Abstract
Clinical prediction models aim to predict outcomes in individuals, to inform diagnosis or prognosis in healthcare. Hundreds of prediction models are published in the medical literature each year, yet many are developed using a dataset that is too small for the total number of participants or outcome events. This leads to inaccurate predictions and consequently incorrect healthcare decisions for some individuals. In this article, the authors provide guidance on how to calculate the sample size required to develop a clinical prediction model.
Meta-analysis and systematic reviewsStatistical Methods in EpidemiologySepsis Diagnosis and TreatmentSample size determinationComputer scienceSample (material)Outcome (game theory)Predictive modellingHealth careData scienceData miningStatisticsMachine learning
MeSH terms
Clinical Decision-MakingForecastingHumansModels, TheoreticalSample Size
Funding
- Georgia Clinical and Translational Science Alliance
- National Institute for Health and Care Research
- Department of Health and Social Care
- Nederlandse Organisatie voor Wetenschappelijk Onderzoek
- National Institutes of Health
- NIHR School for Primary Care Research
- National Center for Advancing Translational Sciences
Citations
2,352
FWCI
170.79
field-weighted impact
References
80
Percentile
100%
vs. same field & year
Citations per year
Cited by
Risk stratification of patients admitted to hospital with covid-19 using the ISARIC WHO Clinical Characterisation Protocol: development and validation of the 4C Mortality Score
BMJ · 2020 · 1,087 citations
TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods
BMJ · 2024 · 1,560 citations
Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal
BMJ · 2020 · 3,177 citations
References
Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD): Explanation and Elaboration
Annals of Internal Medicine · 2015 · 5,044 citations
A prognostic index in primary breast cancer
British Journal of Cancer · 1982 · 716 citations
Importance of events per independent variable in proportional hazards regression analysis II. Accuracy and precision of regression estimates
Journal of Clinical Epidemiology · 1995 · 2,095 citations
Cardiovascular disease risk profiles
American Heart Journal · 1991 · 2,171 citations
A simulation study of the number of events per variable in logistic regression analysis
Journal of Clinical Epidemiology · 1996 · 8,674 citations
The Nottingham prognostic index in primary breast cancer
Breast Cancer Research and Treatment · 1992 · 1,067 citations
Internal validation of predictive models
Journal of Clinical Epidemiology · 2001 · 2,572 citations
Regularization and Variable Selection Via the Elastic Net
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2005 · 20,431 citations
Citation Network
How this paper connects to the literature. Drag to explore, click any node to open that paper.
