Scinovex
article Open AccessTop 1% cited

Boosting the margin: a new explanation for the effectiveness of voting methods

The Annals of Statistics · 1998 · Vol. 26(5)
Peter L. BartlettYoav FreundWee Sun LeeRobert E. Schapire

Abstract

One of the surprising recurring phenomena observed in experiments with boosting is that the test error of the generated classifier usually does not increase as its size becomes very large, and often is observed to decrease even after the training error reaches zero. In this paper, we show that this phenomenon is related to the distribution of margins of the training examples with respect to the generated voting classification rule, where the margin of an example is simply the difference between the number of correct votes and the maximum number of votes received by any incorrect label. We show that techniques used in the analysis of Vapnik’s support vector classifiers and of neural networks with small weights can be applied to voting methods to relate the margin distribution to the test error. We also show theoretically and experimentally that boosting is especially effective at increasing the margins of the training examples. Finally, we compare our explanation to those based on the bias-variance decomposition.

Imbalanced Data Classification TechniquesMachine Learning and AlgorithmsAdvanced Bandit Algorithms ResearchBoosting (machine learning)MathematicsVotingMargin (machine learning)Pattern recognition (psychology)Artificial intelligenceSupport vector machineMachine learningClassifier (UML)Statistics
Citations
2,317
FWCI
87.50
field-weighted impact
References
47
Percentile
100%
vs. same field & year
Citations per year
Cited by
On combining classifiers
IEEE Transactions on Pattern Analysis and Machine Intelligence · 1998 · 5,240 citations
An introduction to kernel-based learning algorithms
IEEE Transactions on Neural Networks · 2001 · 3,478 citations
Cost-sensitive boosting for classification of imbalanced data
Pattern Recognition · 2007 · 1,416 citations
Improved Boosting Algorithms Using Confidence-rated Predictions
Machine Learning · 1999 · 1,951 citations
Soft Margins for AdaBoost
Machine Learning · 2001 · 1,298 citations
References
Classification and Regression Trees.
Biometrics · 1984 · 23,850 citations
Arcing classifier (with discussion and a rejoinder by the author)
The Annals of Statistics · 1998 · 1,094 citations
Heuristics of instability and stabilization in model selection
The Annals of Statistics · 1996 · 1,152 citations
UCI Repository of machine learning databases
Medical Entomology and Zoology · 1998 · 10,524 citations
The Strength of Weak Learnability
Machine Learning · 1990 · 3,302 citations
Support-Vector Networks
Machine Learning · 1995 · 32,108 citations
What Size Net Gives Valid Generalization?
Neural Computation · 1989 · 1,550 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.