Scinovex
article Open AccessTop 1% cited

Missing value estimation methods for DNA microarrays

Bioinformatics · 2001 · Vol. 17(6) · pp. 520–525
Olga G. TroyanskayaMichael CantorGavin SherlockPat BrownTrevor HastieRobert TibshiraniDavid BotsteinRuss B. Altman

Abstract

We present a comparative study of several methods for the estimation of missing values in gene microarray data. We implemented and evaluated three methods: a Singular Value Decomposition (SVD) based method (SVDimpute), weighted K-nearest neighbors (KNNimpute), and row average. We evaluated the methods using a variety of parameter settings and over different real data sets, and assessed the robustness of the imputation methods to the amount of missing data over the range of 1--20% missing values. We show that KNNimpute appears to provide a more robust and sensitive method for missing value estimation than SVDimpute, and both SVDimpute and KNNimpute surpass the commonly used row average method (as well as filling missing values with zeros). We report results of the comparative experiments and provide recommendations and tools for accurate estimation of missing microarray data under a variety of conditions.

Gene expression and cancer classificationBioinformatics and Genomic NetworksData Mining Algorithms and ApplicationsMissing dataImputation (statistics)Data miningCluster analysisComputer scienceSingular value decompositionRobustness (evolution)Pattern recognition (psychology)AlgorithmArtificial intelligence

MeSH terms

AlgorithmsCell CycleData DisplayData Interpretation, StatisticalMultigene FamilyMathematical ComputingSaccharomyces cerevisiaeSensitivity and SpecificitySoftwareGene ExpressionCluster AnalysisOligonucleotide Array Sequence Analysis
Citations
4,180
FWCI
17.71
field-weighted impact
References
21
Percentile
100%
vs. same field & year
Citations per year
References
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.