Scinovex
articleTop 1% cited

Strictly Proper Scoring Rules, Prediction, and Estimation

Journal of the American Statistical Association · 2007 · Vol. 102(477) · pp. 359–378
Tilmann GneitingAdrian E. Raftery

Abstract

Scoring rules assess the quality of probabilistic forecasts, by assigning a numerical score based on the predictive distribution and on the event or value that materializes. A scoring rule is proper if the forecaster maximizes the expected score for an observation drawn from the distributionF if he or she issues the probabilistic forecast F, rather than G ≠ F. It is strictly proper if the maximum is unique. In prediction problems, proper scoring rules encourage the forecaster to make careful assessments and to be honest. In estimation problems, strictly proper scoring rules provide attractive loss and utility functions that can be tailored to the problem at hand. This article reviews and develops the theory of proper scoring rules on general probability spaces, and proposes and discusses examples thereof. Proper scoring rules derive from convex functions and relate to information measures, entropy functions, and Bregman divergences. In the case of categorical variables, we prove a rigorous version of the Savage representation. Examples of scoring rules for probabilistic forecasts in the form of predictive densities include the logarithmic, spherical, pseudospherical, and quadratic scores. The continuous ranked probability score applies to probabilistic forecasts that take the form of predictive cumulative distribution functions. It generalizes the absolute error and forms a special case of a new and very general type of score, the energy score. Like many other scoring rules, the energy score admits a kernel representation in terms of negative definite functions, with links to inequalities of Hoeffding type, in both univariate and multivariate settings. Proper scoring rules for quantile and interval forecasts are also discussed. We relate proper scoring rules to Bayes factors and to cross-validation, and propose a novel form of cross-validation known as random-fold cross-validation. A case study on probabilistic weather forecasts in the North American Pacific Northwest illustrates the importance of propriety. We note optimum score approaches to point and quantile estimation, and propose the intuitively appealing interval score as a utility function in interval estimation that addresses width as well as coverage.

Meteorological Phenomena and SimulationsHydrology and Drought AnalysisForecasting Techniques and ApplicationsScoring ruleProbabilistic logicCategorical variableMathematicsQuantileUnivariateInterpretabilityImprecise probabilityProbability distributionArtificial intelligence
Citations
5,266
FWCI
44.79
field-weighted impact
References
116
Percentile
100%
vs. same field & year
Citations per year
Cited by
Approximate Bayesian Inference for Latent Gaussian models by using Integrated Nested Laplace Approximations
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2009 · 5,256 citations
Probabilistic Forecasts, Calibration and Sharpness
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2007 · 1,691 citations
References
On a Notion of Data Depth Based on Random Simplices
The Annals of Statistics · 1990 · 796 citations
General notions of statistical depth function
The Annals of Statistics · 2000 · 831 citations
VERIFICATION OF FORECASTS EXPRESSED IN TERMS OF PROBABILITY
Monthly Weather Review · 1950 · 5,108 citations
Theory of Point Estimation
Technometrics · 1999 · 4,285 citations
Probabilistic Forecasts, Calibration and Sharpness
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2007 · 1,691 citations
Statistical Methods in the Atmospheric Sciences
Technometrics · 1996 · 7,691 citations
Using Bayesian Model Averaging to Calibrate Forecast Ensembles
Monthly Weather Review · 2005 · 1,928 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.