Scinovex
article Open AccessTop 1% cited

An experimental comparison of classification algorithms for imbalanced credit scoring data sets

Expert Systems with Applications · 2011 · Vol. 39(3) · pp. 3446–3453
Iain BrownChristophe Mues

Abstract

In this paper, we set out to compare several techniques that can be used in the analysis of imbalanced credit scoring data sets. In a credit scoring context, imbalanced data sets frequently occur as the number of defaulting loans in a portfolio is usually much lower than the number of observations that do not default. As well as using traditional classification techniques such as logistic regression, neural networks and decision trees, this paper will also explore the suitability of gradient boosting, least square support vector machines and random forests for loan default prediction. Five real-world credit scoring data sets are used to build classifiers and test their performance. In our experiments, we progressively increase class imbalance in each of these data sets by randomly under-sampling the minority class of defaulters, so as to identify to what extent the predictive power of the respective techniques is adversely affected. The performance criterion chosen to measure this effect is the area under the receiver operating characteristic curve (AUC); Friedman’s statistic and Nemenyi post hoc tests are used to test for significance of AUC differences between techniques. The results from this empirical study indicate that the random forest and gradient boosting classifiers perform very well in a credit scoring context and are able to cope comparatively well with pronounced class imbalances in these data sets. We also found that, when faced with a large class imbalance, the C4.5 decision tree algorithm, quadratic discriminant analysis and k-nearest neighbours perform significantly worse than the best performing classifiers.

Financial Distress and Bankruptcy PredictionImbalanced Data Classification TechniquesCredit Risk and Financial RegulationsMachine learningArtificial intelligenceDecision treeComputer scienceSupport vector machineBoosting (machine learning)Random forestGradient boostingLinear discriminant analysisContext (archaeology)

Funding

  • Engineering and Physical Sciences Research Council
Citations
702
FWCI
21.84
field-weighted impact
References
38
Percentile
99%
vs. same field & year
Citations per year
Cited by
Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research
European Journal of Operational Research · 2015 · 1,087 citations
Learning from class-imbalanced data: Review of methods and applications
Expert Systems with Applications · 2016 · 2,276 citations
References
Neural networks for pattern recognition
Choice Reviews Online · 1994 · 18,690 citations
Greedy function approximation: A gradient boosting machine.
The Annals of Statistics · 2001 · 27,794 citations
Applied logistic regression
Choice Reviews Online · 1990 · 35,647 citations
Benchmarking state-of-the-art classification algorithms for credit scoring
Journal of the Operational Research Society · 2003 · 866 citations
Random Forests
Machine Learning · 2001 · 121,242 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.