Regularization

Regularization

What is regularization

Alternative intuition for deep neural networks:
Regularization reduces overfitting by letting the weight of units decay and get closer to 0 (given that λ are usually large). If the weights almost zero, than the networks becomes almost linear and will avoid overfitting.

Cost function with regularization

When you choose regularization, a regularization term will be added to the cost function.
See Cost Functions#Cost function with regularization.

Types of Techniques

Shrinking

L1/Lasso regularization

Sparse regularized models are often used for feature selection stability and Stability Selection

L2/Ridge regularization

The Hundred-Page Machine Learning Book

  • If your only goal is to maximize the performance of the model on the holdout data, then L2 usually gives better results. L2 also has the advantage of being differentiable, so gradient descent can be used for optimizing the objective function.

Elastic Net

Dropout regularization

Batch-normalization

Data augmentation

Early stopping

Pasted image 20230316144212.png300

Interesting

But long-term training may lead to flip in large models, see here

Regularization in Bayesian framework