LASSO and Ridge Regression
馃Regularization
Regularization modifies the learning objective to:
馃憠 Purpose:
- Control model complexity
- Prevent overfitting
- Improve generalization
馃敺 Ridge Regression (L2 Regularization)
馃搻 Objective Function
馃攳 Key Idea
- Penalizes square of coefficients
- Larger weights → heavily penalized
馃幆 Effects
- Shrinks coefficients toward zero
- No coefficient becomes exactly zero
- Keeps all features
馃З Properties
- Produces smooth models
- Handles multicollinearity well
- Reduces model variance
馃搳 Intuition
Think of it as:
“Discouraging large weights, but not eliminating features”
馃敹 LASSO Regression (L1 Regularization)
Least Absolute Shrinkage and Selection Operator
馃搻 Objective Function
馃攳 Key Idea
- Penalizes absolute value of coefficients
馃幆 Effects
- Some coefficients become exactly zero
- Performs automatic feature selection
✂️ Sparsity
馃憠 Key property:
- Produces sparse models
- Removes irrelevant features
馃搳 Intuition
Think of it as:
“Forcing the model to keep only the most important features”
⚖️ Core Differences (Geometric Insight)
| Aspect | Ridge (L2) | LASSO (L1) |
|---|---|---|
| Constraint shape | Circle | Diamond |
| Boundary | Smooth | Sharp corners |
| Result | Shrink weights | Zero out weights |
馃憠 Sharp corners in LASSO → sparsity
馃搳 Behavior Comparison
馃敼 Coefficient Shrinkage
| Feature | Ridge | LASSO |
|---|---|---|
| x₁ | 2.2 | 2.1 |
| x₂ | 1.6 | 1.5 |
| x₃ | 0.8 | 0 |
| x₄ | 0.6 | 0 |
馃敼 Key Observations
- Ridge → all features retained
- LASSO → some features removed
⚖️Detailed Comparison Table
| Criterion | Ridge (L2) | LASSO (L1) |
|---|---|---|
| Penalty | | | |w| |
| Sparsity | No | Yes |
| Feature selection | No | Yes |
| Stability | High | Lower |
| Correlated features | Shares weights | Picks one feature |
| Computation | Closed form | Iterative |
馃Bias–Variance Perspective
| Model | Bias | Variance |
|---|---|---|
| Ridge | Moderate | Low |
| LASSO | Higher | Low |
馃攳 Effect of Regularization Parameter (位)
馃敼 Small 位
- Both behave like ordinary regression
馃敼 Medium 位
- Ridge → shrink coefficients
- LASSO → shrink + remove some
馃敼 Large 位
- Ridge → all weights very small
- LASSO → most weights = 0
馃И Practical Considerations
馃敼 Feature Scaling (Important!)
- Required for both methods
馃敼 Model Selection
-
Choose 位 using:
- Validation set
- Cross-validation
馃Intuitive Analogies
Ridge:
“Keep all features, but reduce their influence”
LASSO:
“Select only the most important features”
馃煟When to Use Which?
馃敺 Use Ridge when:
- Many correlated features
- All features are important
- Need stable model
馃敹 Use LASSO when:
- Many irrelevant features
- Need feature selection
- Want interpretable model
馃敆 Connection: Elastic Net
Combines both:
馃憠 Useful when:
- Features are correlated
- Need both sparsity and stability
馃幆Key Takeaways
- Both methods reduce overfitting
- Ridge → smooth shrinkage
- LASSO → sparse solutions
| Use Case | Preferred Method |
|---|---|
| Many correlated features | Ridge |
| Feature selection needed | LASSO |
| Balanced approach | Elastic Net |

Comments
Post a Comment