Scinovex
article Open AccessTop 1% cited

Machine learning algorithm validation with a limited sample size

PLoS ONE · 2019 · Vol. 14(11) · pp. e0224365–e0224365
Andrius VabalasEmma GowenEllen PoliakoffAlexander J. Casson

Abstract

Advances in neuroimaging, genomic, motion tracking, eye-tracking and many other technology-based data collection methods have led to a torrent of high dimensional datasets, which commonly have a small number of samples because of the intrinsic high cost of data collection involving human participants. High dimensional data with a small number of samples is of critical importance for identifying biomarkers and conducting feasibility and pilot work, however it can lead to biased machine learning (ML) performance estimates. Our review of studies which have applied ML to predict autistic from non-autistic individuals showed that small sample size is associated with higher reported classification accuracy. Thus, we have investigated whether this bias could be caused by the use of validation methods which do not sufficiently control overfitting. Our simulations show that K-fold Cross-Validation (CV) produces strongly biased performance estimates with small sample sizes, and the bias is still evident with sample size of 1000. Nested CV and train/test split approaches produce robust and unbiased performance estimates regardless of sample size. We also show that feature selection if performed on pooled training and testing data is contributing to bias considerably more than parameter tuning. In addition, the contribution to bias by data dimensionality, hyper-parameter space and number of CV folds was explored, and validation methods were compared with discriminable data. The results suggest how to design robust testing methodologies when working with small datasets and how to interpret the results of other studies based on what validation method was used.

Gene expression and cancer classificationMachine Learning and Data ClassificationCell Image Analysis TechniquesOverfittingSample size determinationComputer scienceArtificial intelligenceCross-validationData collectionMachine learningSelection biasSample (material)Statistics

MeSH terms

Machine LearningAlgorithmsData Interpretation, StatisticalHumansSample SizeBiomedical Research

Funding

  • University of Manchester
  • Engineering and Physical Sciences Research Council
Citations
1,570
FWCI
44.37
field-weighted impact
References
33
Percentile
100%
vs. same field & year
Citations per year
References
Machine learning applications in genetics and genomics
Nature Reviews Genetics · 2015 · 1,956 citations
Cross-Validatory Choice and Assessment of Statistical Predictions
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1974 · 10,283 citations
A review of feature selection techniques in bioinformatics
Bioinformatics · 2007 · 5,229 citations
Libsvm : A library for support vector machines
Medical Entomology and Zoology · 2008 · 10,114 citations
Related articles
An Overview of Overfitting and its Solutions
Journal of Physics Conference Series · 2019 · 2,169 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.