Multivariate Linear Regression using Gradient Descent
Multivariate Linear Regression using Gradient Descent
🔹 1. Model (Hypothesis Function)
For multiple features:
🔹2. Matrix Form (Very Important)
Add a column of 1s for intercept:
📌 Parameter Vector
📌 Model Equation
🔹 3. Cost Function (MSE)
👉 Goal: Minimize this cost
🔹 4. Gradient of Cost Function
For all parameters together:
👉 This gives partial derivatives for all coefficients
🔹 5. Gradient Descent Update Rule
🔸 Substitute Gradient
🔹 6. Expanded Form (Component-wise)
For each parameter :
🔹 7. Algorithm Steps
📌 Gradient Descent Procedure
- Initialize
-
Compute predictions:
-
Compute error:
-
Compute gradient:
-
Update parameters:
- Repeat until convergence
🔹 8. Intuition
- Each coefficient is adjusted based on its contribution to error
- All parameters are updated simultaneously
🔹 9. Role of Learning Rate (α)
| Learning Rate | Effect |
|---|---|
| Small | Slow convergence |
| Large | May diverge |
| Optimal | Fast convergence |
🔹 10. Why Use Gradient Descent?
👉 Preferred when:
- Large number of features
- Large dataset
- Matrix inversion is costly
🔹 11. Important Notes
- Feature scaling is very important
- Helps faster convergence
- Works for high-dimensional data
🔹 12. Comparison with Normal Equation
| Method | Advantage | Limitation |
|---|---|---|
| Normal Equation | Exact solution | Expensive for large data |
| Gradient Descent | Scalable | Needs tuning |
🔹 13. Summary
👉 Multivariate GD is just an extension of simple GD:
- Works on vectors instead of scalars
- Updates all parameters together
👉 Gradient Descent:
- Iteratively minimizes error
- Finds optimal regression coefficients
- Works efficiently for large-scale problems
Appendix
Goal
We want to derive:
🔹 1. Start with Cost Function
👉 Let:
So:
🔹 2. Expand the Expression
🔹 3. Take Gradient w.r.t β
We differentiate term by term.
🔸 Term 1:
🔸 Term 2:
🔸 Term 3:
🔹 4. Combine Results
🔹 5. Simplify
Cancel 2:
Comments
Post a Comment