Scinovex
article Open AccessTop 1% cited

Least angle regression

The Annals of Statistics · 2004 · Vol. 32(2)
Bradley EfronTrevor HastieIain M. JohnstoneRobert Tibshirani

Abstract

The purpose of model selection algorithms such as All Subsets, Forward Selection and Backward Elimination is to choose a linear model on the basis of the same set of data to which the model will be applied. Typically we have available a large collection of possible covariates from which we hope to select a parsimonious set for the efficient prediction of a response variable. Least Angle Regression (LARS), a new model selection algorithm, is a useful and less greedy version of traditional forward selection methods. Three main properties are derived: (1) A simple modification of the LARS algorithm implements the Lasso, an attractive version of ordinary least squares that constrains the sum of the absolute regression coefficients; the LARS modification calculates all possible Lasso estimates for a given problem, using an order of magnitude less computer time than previous methods. (2) A different LARS modification efficiently implements Forward Stagewise linear regression, another promising new model selection method; this connection explains the similar numerical results previously observed for the Lasso and Stagewise, and helps us understand the properties of both methods, which are seen as constrained versions of the simpler LARS algorithm. (3) A simple approximation for the degrees of freedom of a LARS estimate is available, from which we derive a Cp estimate of prediction error; this allows a principled choice among the range of possible LARS estimates. LARS and its variants are computationally efficient: the paper describes a publicly available algorithm that requires only the same order of magnitude of computational effort as ordinary least squares applied to the full set of covariates.

Statistical Methods and InferenceAdvanced Statistical Methods and ModelsControl Systems and IdentificationLasso (programming language)Ordinary least squaresMathematicsAlgorithmSelection (genetic algorithm)Set (abstract data type)Model selectionElastic net regularizationLinear regressionRegression

Funding

  • National Science Foundation
  • National Institutes of Health
Citations
9,400
FWCI
118.54
field-weighted impact
References
64
Percentile
100%
vs. same field & year
Citations per year
Cited by
On the use of cross-validation for time series predictor evaluation
Information Sciences · 2012 · 1,028 citations
Adaptive sparse polynomial chaos expansion based on least angle regression
Journal of Computational Physics · 2010 · 1,409 citations
The Adaptive Lasso and Its Oracle Properties
Journal of the American Statistical Association · 2006 · 7,497 citations
Stability Selection
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2010 · 2,054 citations
Bayesian Compressive Sensing
IEEE Transactions on Signal Processing · 2008 · 2,375 citations
Simultaneous analysis of Lasso and Dantzig selector
The Annals of Statistics · 2009 · 2,504 citations
References
Classification and Regression Trees.
Biometrics · 1984 · 23,850 citations
Greedy function approximation: A gradient boosting machine.
The Annals of Statistics · 2001 · 27,794 citations
Asymptotics for lasso-type estimators
The Annals of Statistics · 2000 · 1,317 citations
Applied Linear Regression
Technometrics · 1987 · 2,855 citations
Estimation of the Mean of a Multivariate Normal Distribution
The Annals of Statistics · 1981 · 2,732 citations
Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing
Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1995 · 106,483 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.