Elastic net is a method in statistics and machine learning that combines the penalties of lasso regression and ridge regression. It performs coefficient shrinkage and feature selection, while encouraging correlated predictors to receive similar coefficients. Introduced by Hui Zou and Trevor Hastie in 2005, it was designed particularly for problems with many predictors, including cases where their number exceeds the number of observations. Its name refers to a flexible “net” capable of retaining groups of related variables. (doi.org)
Mathematical formulation
For linear regression, suppose there are observations, predictors, responses , predictor vectors , an intercept , and coefficients . A common elastic-net objective function is
Here, and , using standard norm notation. The first term is a squared-error loss function; the second supplies regularization. The intercept is ordinarily excluded from the penalty. (stat.ethz.ch)
Two hyperparameters control the fit. The nonnegative determines overall penalty strength, while determines its composition:
- gives the lasso penalty.
- gives the ridge penalty.
- combines both.
- removes the penalty, recovering the ordinary least-squares objective.
Parameter names differ between implementations: in scikit-learn, alpha denotes overall strength, corresponding to above, while l1_ratio corresponds to the mixing parameter . Numerical values therefore cannot be compared solely by matching parameter names. (scikit-learn.org)
Sparsity and correlated predictors
The absolute-value penalty can set coefficients exactly to zero, producing a sparse model. The quadratic penalty shrinks coefficients without itself imposing a threshold at zero. Their combination balances sparse selection with the stabilization provided by ridge regression. With strongly correlated predictors, lasso may retain one variable while excluding others that convey similar information; elastic net encourages these predictors to enter or leave the model together. (doi.org)
This grouping effect is a tendency rather than a guarantee of selecting an entire predefined group. For identical predictor columns and a positive quadratic penalty, symmetry and uniqueness imply equal fitted coefficients. For merely similar columns, coefficients need not be identical. The effect depends on predictor scaling, penalty strength, and the balance between the two penalties. (web.stanford.edu)
Elastic net also avoids a restriction associated with standard lasso solutions: it can retain more predictors than there are observations. This makes it relevant to high-dimensional data containing many related measurements, such as gene-expression datasets. The original paper demonstrated the method through simulations and a microarray classification example. These examples do not establish that elastic net universally outperforms either lasso or ridge. (doi.org)
Optimization and computation
The objective belongs to convex optimization. When and , its quadratic penalty makes it strictly convex in the penalized coefficients, yielding a unique coefficient solution even when predictor columns are linearly dependent. This addresses degeneracies that can arise in a pure lasso fit with highly redundant variables. (web.stanford.edu)
A widely used algorithm is coordinate descent, which updates one coefficient at a time while holding the others fixed. For centered predictors scaled so that , the update is
where is the residual excluding predictor . The soft-thresholding operator is . The numerator implements sparse selection; the denominator adds ridge-like shrinkage. (rdrr.io)
Software often computes a regularization path across decreasing values of , reusing each solution as the starting point for the next. These “warm starts” make fitting a sequence of models efficient. The original elastic-net paper instead presented the LARS-EN path algorithm. (web.stanford.edu)
Scaling and model selection
Because penalties operate on coefficient magnitudes, predictor units matter. Feature scaling, commonly centering and division by a standard deviation, makes the penalty more comparable across features measured in different units. Standardization changes the coordinates in which regularization is applied; fitted coefficients can subsequently be expressed on the original scale. (github.com)
The penalty parameters are commonly selected by cross-validation, comparing predictive error over candidate settings. For regression, mean squared error is one available criterion. Implementations such as ElasticNetCV can search both overall strength and the mixing ratio, then refit the selected configuration using the complete training dataset. (scikit-learn.org)
Evaluation must distinguish parameter selection from measurement of generalization. Repeatedly choosing parameters using a test set can produce overfitting to that set. Likewise, preprocessing statistics calculated from held-out observations create data leakage. In cross-validation, transformations are learned from each training fold and applied to its held-out fold; final performance is assessed separately from tuning. (github.com)
Extensions and terminology
Elastic-net penalties are not restricted to squared-error regression. They can accompany the negative log-likelihood objectives of generalized linear models, including logistic regression and multinomial regression. The loss changes with the response model, while the combined coefficient penalty remains. (web.stanford.edu)
A historical distinction concerns coefficient rescaling. Zou and Hastie called the direct combined-penalty estimator the naïve elastic net and proposed a subsequent rescaling to counter additional shrinkage. Many later publications and software packages use “elastic net” for the direct penalized objective shown above, without that correction. Consequently, reproducing a reported fit requires checking its objective, normalization, parameter definitions, and any post-fitting rescaling. (doi.org)
References
- Regularization Paths for Generalized Linear Modelsweb.stanford.edu
- An Introduction to glmnetstat.ethz.ch
- glmnet: vignettes/glmnet.Rmdrdrr.io
- ElasticNetCV — scikit-learn documentationscikit-learn.org
- scikit-learn/doc/modules/preprocessing.rstgithub.com
- scikit-learn/doc/modules/cross_validation.rstgithub.com