Linear SVM with Soft Margin
Linear SVM with Soft Margin
1. Why Do We Need Soft-Margin SVM?
In Hard-Margin SVM, we assume that the two classes can be perfectly separated.
For example:
Class 0 | Margin | Class 1 ● ● | | ○ ○ ● | | ○
Every point must be on the correct side of the decision boundary.
But in real-world datasets, we may have:
- noisy data,
- overlapping classes,
- outliers,
- wrongly labeled data.
For example:
○ Class 1 ● ● ● ● ↑ Outlier ------------------------------- Decision Boundary ○ ○ ○ Class 1
A Hard-Margin SVM may not find a solution when the data cannot be perfectly separated.
👉 Therefore, we use Soft-Margin SVM.
2. Main Idea of Soft-Margin SVM
Soft-Margin SVM says:
It is acceptable for some points to be inside the margin or even on the wrong side of the decision boundary.
But such violations receive a penalty.
So, Soft-Margin SVM tries to balance two objectives:
Objective 1: Find a large margin
Objective 2: Reduce classification errors
3. A Simple Example
Consider the following data:
| Class | ||
|---|---|---|
| 1 | 2 | 0 |
| 2 | 3 | 0 |
| 3 | 3 | 0 |
| 4 | 4 | 1 |
| 5 | 5 | 1 |
| 3 | 5 | 1 |
| 4 | 2 | 0 |
Suppose the data looks approximately like this:
X₂ ↑ 6 | 5 | ○ ○ 4 | ○ 3 | ● ● 2 | ● ● +------------------------------→ X₁ ● = Class 0 ○ = Class 1
Sometimes, one or two points may be close to the other class.
Instead of requiring:
Soft-Margin SVM allows some violations.
4. Understanding the Three Important Regions
Suppose we have:
Class 0 Margin Margin Class 1 ● ● | | ○ ○ | | | Decision | | Boundary |
Mathematically:
Negative margin boundary
Decision boundary
Positive margin boundary
Case 1: Correctly Classified and Outside Margin
For a Class point:
Example:
Then:
This point is:
✅ Correctly classified
✅ Outside the margin
✅ No penalty
Case 2: Correctly Classified but Inside the Margin
Suppose:
and:
Then:
The point is:
✅ Correctly classified
But:
❌ Inside the margin
Therefore, Soft-Margin SVM gives a penalty.
Case 3: Incorrectly Classified
Suppose:
but:
Then:
This means the point is:
❌ Incorrectly classified.
A larger penalty is given.
5. Introducing the Slack Variable
Soft-Margin SVM introduces a new variable called the:
The Hard-Margin constraint was:
The Soft-Margin constraint becomes:
where:
The value of tells us how much the point violates the margin.
6. Meaning of Different Values of
| Slack Variable | Meaning |
|---|---|
| Correctly classified and outside/on margin | |
| Correctly classified but inside margin | |
| On the decision boundary | |
| Incorrectly classified |
This is a very useful table for students.
7. Mathematical Formulation of Soft-Margin SVM
The optimization problem is:
Subject to:
and:
8. Understanding the Objective Function
The objective is:
It has two parts.
Part 1
This tries to:
Part 2
This tries to:
Therefore, Soft-Margin SVM balances:
9. What is the Role of ?
The parameter:
controls how strongly we penalize errors.
Large
If:
errors are heavily penalized.
The SVM tries very hard to classify every training point correctly.
Large C "Don't make mistakes!"
This may result in a smaller margin.
Small
If:
the SVM allows more violations.
Small C "A few mistakes are acceptable."
This can result in a wider margin.
Simple Comparison
| Value of | Effect |
|---|---|
| Large | Fewer training errors, smaller margin |
| Small | More violations allowed, wider margin |
10. Simple Numerical Example Using Slack Variable
Suppose we have a point:
with class:
Suppose the SVM decision function gives:
Then:
The Soft-Margin condition is:
Substituting:
Therefore:
So:
This means:
The point is correctly classified, but it violates the margin by .
11. Example of a Misclassified Point
Suppose:
but:
Then:
The constraint is:
Therefore:
Thus:
Since:
the point is misclassified.
12. Soft Margin and Hinge Loss
Instead of explicitly calculating slack variables, we can use hinge loss.
The hinge loss is:
Notice that this is equivalent to the amount of margin violation.
Therefore:
So the Soft-Margin SVM objective can also be written as:
This is the formulation commonly used when implementing SVM using gradient descent.
13. Hard Margin vs Soft Margin
| Hard-Margin SVM | Soft-Margin SVM |
|---|---|
| Requires perfectly separable data | Allows some violations |
| No classification errors allowed | Some errors are allowed |
| Sensitive to outliers | More robust to outliers |
| Uses strict constraints | Uses slack variables |
| Best for ideal datasets | Better for real-world datasets |
Hard-Margin SVM does not allow violations. Soft-Margin SVM allows violations but penalizes them.
Final Summary
Hard Margin
No violations allowed.
Soft Margin
Violations are allowed.
Comments
Post a Comment