Hinge Loss in SVM and Finding the Decision Boundary Using Gradient Descent
Hinge Loss in SVM and Finding the Decision Boundary Using Gradient Descent
The basic idea is:
SVM finds the decision boundary by minimizing a loss function consisting of two parts:
- Hinge loss → penalizes classification errors and margin violations.
- Regularization term → encourages a large margin.
Let's understand this step by step.
1. Recall the SVM Decision Function
For a Linear SVM:
For two features:
The decision boundary is where:
The prediction is:
For SVM mathematics, we normally use:
2. What is Hinge Loss?
Hinge loss measures whether a point is:
- correctly classified and safely outside the margin,
- correctly classified but inside the margin,
- or incorrectly classified.
The hinge loss is:
Since:
we get:
3. Understanding
The important quantity is:
Let:
This is sometimes called the functional margin.
There are three important cases.
Case 1:
The point is:
- correctly classified ✓
- outside the margin ✓
Therefore:
No hinge loss.
Case 2:
The point is:
- correctly classified ✓
- but inside the margin ✗
Therefore:
SVM applies a penalty.
Case 3:
The point is:
- incorrectly classified ✗
Therefore:
The penalty becomes even larger.
4. Simple Hinge Loss Table
| Situation | Hinge Loss | |
|---|---|---|
| Correct and outside margin | ||
| On margin | ||
| Correct but inside margin | ||
| On decision boundary | ||
| Misclassified |
5. Example of Calculating Hinge Loss
Suppose:
and the SVM gives:
Then:
The hinge loss is:
The point is correctly classified, but it lies inside the margin.
Another Example
Suppose:
and:
Then:
Therefore:
This point is incorrectly classified.
6. Why is Hinge Loss Used?
Remember the hard-margin constraint:
Hard-margin SVM requires every point to satisfy this condition.
But real data may contain noise.
Instead of enforcing this constraint strictly, hinge loss gives a penalty whenever:
Thus, hinge loss provides a way to formulate SVM as an optimization problem.
7. The Complete SVM Loss Function
To train an SVM, we use two components.
Part 1: Regularization
This encourages a smaller .
Since the margin is:
minimizing means maximizing the margin.
Part 2: Hinge Loss
For training examples:
8. Complete SVM Objective Function
A common form is:
where:
- → maximizes the margin
- hinge loss → penalizes violations
- → controls the importance of classification errors
9. How Does Gradient Descent Find the Decision Boundary?
The main idea is simple.
Initially, we do not know:
So we start with some initial values.
For example:
Then repeatedly:
- Calculate predictions.
- Calculate hinge loss.
- Calculate the gradient.
- Update and .
- Repeat.
Eventually, the values converge toward a good decision boundary.
10. Gradient Descent Rule
The general gradient descent rule is:
where:
is the learning rate.
For SVM:
and:
11. Gradient of the SVM Objective
Recall:
We consider two cases.
Case 1: Point Correctly Classified Outside the Margin
If:
then:
Only the regularization term contributes.
The gradient with respect to is:
The gradient with respect to is:
Therefore, the update is:
and:
Case 2: Point is Inside the Margin or Misclassified
If:
then the hinge loss is active.
The gradient is:
and:
Therefore:
and:
Simplifying:
This update moves the decision boundary in a direction that improves classification.
12.Complete Algorithm for Linear SVM Using Gradient Descent
Step 1: Convert labels
Step 2: Initialize
Step 3: Repeat for several epochs
For each training example :
Calculate:
If:
update:
Otherwise:
update:
and:
Step 4: Final Classifier
After training:
The decision boundary is:
Final Summary
Hinge Loss
It penalizes:
- incorrectly classified points
- correctly classified points inside the margin
SVM Objective
Gradient/Subgradient Updates
If:
If:
Final Decision Boundary
and :
Comments
Post a Comment