Skip to main content

Lund University Publications

LUND UNIVERSITY LIBRARIES

The Choice of Normalization Influences Shrinkage in Regularized Regression

Larsson, Johan LU orcid and Wallin, Jonas LU (2025) In Transactions on Machine Learning Research 2025-November.
Abstract

Regularized models are often sensitive to the scales of the features in the data and it has therefore become standard practice to normalize (center and scale) the features before fitting the model. But there are many different ways to normalize the features and the choice may have dramatic effects on the resulting model. In spite of this, there has so far been no research on this topic. In this paper, we begin to bridge this knowledge gap by studying normalization in the context of lasso, ridge, and elastic net regression. We focus on binary features and show that their class balances (proportions of ones) directly influences the regression coefficients and that this effect depends on the combination of normalization and regularization... (More)

Regularized models are often sensitive to the scales of the features in the data and it has therefore become standard practice to normalize (center and scale) the features before fitting the model. But there are many different ways to normalize the features and the choice may have dramatic effects on the resulting model. In spite of this, there has so far been no research on this topic. In this paper, we begin to bridge this knowledge gap by studying normalization in the context of lasso, ridge, and elastic net regression. We focus on binary features and show that their class balances (proportions of ones) directly influences the regression coefficients and that this effect depends on the combination of normalization and regularization methods used. We demonstrate that this effect can be mitigated by scaling binary features with their variance in the case of the lasso and standard deviation in the case of ridge regression, but that this comes at the cost of increased variance of the coefficient estimates. For the elastic net, we show that scaling the penalty weights, rather than the features, can achieve the same effect. Finally, we also tackle mixes of binary and normal features as well as interactions and provide some initial results on how to normalize features in these cases.

(Less)
Please use this url to cite or link to this publication:
author
and
organization
publishing date
type
Contribution to journal
publication status
published
subject
in
Transactions on Machine Learning Research
volume
2025-November
external identifiers
  • scopus:105024860428
ISSN
2835-8856
language
English
LU publication?
yes
additional info
Publisher Copyright: © 2025, Transactions on Machine Learning Research. All rights reserved.
id
98ef555f-ca18-4409-9f45-184448f7a756
alternative location
https://openreview.net/pdf?id=6xKyDBIwQ5
date added to LUP
2026-03-03 15:09:26
date last changed
2026-03-03 15:10:00
@article{98ef555f-ca18-4409-9f45-184448f7a756,
  abstract     = {{<p>Regularized models are often sensitive to the scales of the features in the data and it has therefore become standard practice to normalize (center and scale) the features before fitting the model. But there are many different ways to normalize the features and the choice may have dramatic effects on the resulting model. In spite of this, there has so far been no research on this topic. In this paper, we begin to bridge this knowledge gap by studying normalization in the context of lasso, ridge, and elastic net regression. We focus on binary features and show that their class balances (proportions of ones) directly influences the regression coefficients and that this effect depends on the combination of normalization and regularization methods used. We demonstrate that this effect can be mitigated by scaling binary features with their variance in the case of the lasso and standard deviation in the case of ridge regression, but that this comes at the cost of increased variance of the coefficient estimates. For the elastic net, we show that scaling the penalty weights, rather than the features, can achieve the same effect. Finally, we also tackle mixes of binary and normal features as well as interactions and provide some initial results on how to normalize features in these cases.</p>}},
  author       = {{Larsson, Johan and Wallin, Jonas}},
  issn         = {{2835-8856}},
  language     = {{eng}},
  series       = {{Transactions on Machine Learning Research}},
  title        = {{The Choice of Normalization Influences Shrinkage in Regularized Regression}},
  url          = {{https://openreview.net/pdf?id=6xKyDBIwQ5}},
  volume       = {{2025-November}},
  year         = {{2025}},
}