[Quick recap]Deep learning 4 Regularization

What is regularization?
Any aspect of a learning algorithm that is intended to lower the generalization erro but not the training error.

Norm-based Regularization

Standard regularization method (convex models) is:
[math]J_\Omega (\theta;S)=J(\theta;S)+\Omega(\theta)[/math],
where [math]\Omega[/math] is independent on training data. A common choice is the L2/Frobenius-norm penalty for deep networks,
[math]\Omega(\theta) = \frac{1}{2}\sum_{l=1}^{L}\mu^l||\mathbb{W}^l||_F^2, \mu^l\geq0[/math]
Here we usually
* only penalize weights, not biases.
* one [math]\mu^l[/math] per layer.

Weight Decay

Regularization based on L2-norm is also called weight-decay as [math]\frac{\partial\Omega}{\partial{w_ij^l}}=\mu^lw_{ij}^l[/math]
Gradient descent gets modified as [math]\theta(t+1)=(1-\mu)\theta(t)-n * \triangledown_{\theta}J[/math]
New_Theta = weight_decayed_theta – step_size * original_gradient
*Quadratic (Taylor) approxiamation of J around J-optimal:
[math]J(\theta) \approx{} J(\theta^\ast)+\frac{1}{2}(\theta – \theta^\ast)^T\mathbb{H}(\theta – \theta^\ast)[/math], where H is the Hessian Matrix.

Leave a Reply

Your email address will not be published. Required fields are marked *