Scinovex
article Open AccessTop 1% cited

SMOTE for high-dimensional class-imbalanced data

BMC Bioinformatics · 2013 · Vol. 14(1) · pp. 106–106
Rok BlagusLara Lusa

Abstract

In practice, in the high-dimensional setting only k-NN classifiers based on the Euclidean distance seem to benefit substantially from the use of SMOTE, provided that variable selection is performed before using SMOTE; the benefit is larger if more neighbors are used. SMOTE for k-NN without variable selection should not be used, because it strongly biases the classification towards the minority class.

Imbalanced Data Classification TechniquesFinancial Distress and Bankruptcy PredictionArtificial Intelligence in HealthcareUndersamplingOversamplingRandom forestComputer scienceClass (philosophy)Artificial intelligenceClustering high-dimensional dataMachine learningPattern recognition (psychology)Data mining

MeSH terms

AlgorithmsClassificationComputer SimulationGene Expression ProfilingSupport Vector Machine
Citations
1,050
FWCI
16.94
field-weighted impact
References
47
Percentile
99%
vs. same field & year
Citations per year
Cited by
References
Classification and Regression Trees.
Biometrics · 1984 · 23,850 citations
A molecular signature of metastasis in primary solid tumors
Nature Genetics · 2002 · 2,483 citations
Support-Vector Networks
Machine Learning · 1995 · 32,108 citations
Random Forests
Machine Learning · 2001 · 121,242 citations
Classification and Regression Trees.
Journal of the American Statistical Association · 1986 · 21,013 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.