article Open AccessTop 10% cited
Bias in error estimation when using cross-validation for model selection
BMC Bioinformatics · 2006 · Vol. 7(1) · pp. 91–91
Sudhir Varma✉(National Cancer Institute)Richard Simon(National Cancer Institute)
Abstract
We show that using CV to compute an error estimate for a classifier that has itself been tuned using CV gives a significantly biased estimate of the true error. Proper use of CV for estimating true error of a classifier developed using a well defined algorithm requires that all steps of the algorithm, including classifier parameter tuning, be repeated in each CV loop. A nested CV procedure provides an almost unbiased estimate of the true error.
Gene expression and cancer classificationFault Detection and Control SystemsMachine Learning and Data ClassificationClassifier (UML)Pattern recognition (psychology)Support vector machineCentroidStatisticsComputer scienceCross-validationMathematicsBayes error rateWord error rate
MeSH terms
AlgorithmsArtificial IntelligenceComputer SimulationData Interpretation, StatisticalModels, GeneticPattern Recognition, AutomatedSensitivity and SpecificityReproducibility of ResultsModels, StatisticalBiasOligonucleotide Array Sequence AnalysisGene Expression Profiling
Citations
1,888
FWCI
6.92
field-weighted impact
References
13
Percentile
98%
vs. same field & year
Citations per year
Cited by
Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD): Explanation and Elaboration
Annals of Internal Medicine · 2015 · 5,044 citations
Machine learning algorithm validation with a limited sample size
PLoS ONE · 2019 · 1,570 citations
References
Statistical Learning Theory
Technometrics · 1999 · 26,915 citations
Citation Network
How this paper connects to the literature. Drag to explore, click any node to open that paper.
