Linear SVM with Soft Margin

 

Linear SVM with Soft Margin 

1. Why Do We Need Soft-Margin SVM?

In Hard-Margin SVM, we assume that the two classes can be perfectly separated.

For example:

Class 0       | Margin |       Class 1

  ●  ●        |        |        ○  ○
  ●           |        |           ○

Every point must be on the correct side of the decision boundary.

But in real-world datasets, we may have:

  • noisy data,
  • overlapping classes,
  • outliers,
  • wrongly labeled data.

For example:

                    ○ Class 1
       ● ● ●      ●
                 ↑
              Outlier

-------------------------------
      Decision Boundary

                 ○ ○ ○
              Class 1

A Hard-Margin SVM may not find a solution when the data cannot be perfectly separated.

👉 Therefore, we use Soft-Margin SVM.


2. Main Idea of Soft-Margin SVM

Soft-Margin SVM says:

It is acceptable for some points to be inside the margin or even on the wrong side of the decision boundary.

But such violations receive a penalty.

So, Soft-Margin SVM tries to balance two objectives:

Objective 1: Find a large margin

Objective 2: Reduce classification errors


3. A Simple Example

Consider the following data:

X1X_1X2X_2Class
120
230
330
441
551
351
420

Suppose the data looks approximately like this:

X₂
↑

6 |

5 |             ○       ○

4 |                  ○

3 |      ●     ●

2 |   ●                 ●

  +------------------------------→ X₁

    ● = Class 0
    ○ = Class 1

Sometimes, one or two points may be close to the other class.

Instead of requiring:

Every point must be perfectly separated\text{Every point must be perfectly separated}

Soft-Margin SVM allows some violations.


4. Understanding the Three Important Regions

Suppose we have:

Class 0        Margin       Margin       Class 1

   ● ●             |             |          ○ ○
                   |             |
                   | Decision    |
                   | Boundary    |

Mathematically:

Negative margin boundary

wTx+b=−1w^Tx+b=-1

Decision boundary

wTx+b=0\boxed{w^Tx+b=0}

Positive margin boundary

wTx+b=+1w^Tx+b=+1

Case 1: Correctly Classified and Outside Margin

For a Class +1+1 point:

y(wTx+b)>1y(w^Tx+b)>1

Example:

y=+1,f(x)=2y=+1,\qquad f(x)=2

Then:

yf(x)=2y f(x)=2

This point is:

✅ Correctly classified
✅ Outside the margin
✅ No penalty


Case 2: Correctly Classified but Inside the Margin

Suppose:

y=+1y=+1

and:

f(x)=0.5f(x)=0.5

Then:

yf(x)=0.5yf(x)=0.5

The point is:

✅ Correctly classified

But:

❌ Inside the margin

Therefore, Soft-Margin SVM gives a penalty.


Case 3: Incorrectly Classified

Suppose:

y=−1y=-1

but:

f(x)=2f(x)=2

Then:

yf(x)=(−1)(2)=−2yf(x)=(-1)(2)=-2

This means the point is:

❌ Incorrectly classified.

A larger penalty is given.


5. Introducing the Slack Variable ξ\xi

Soft-Margin SVM introduces a new variable called the:

Slack Variable ξi\boxed{\text{Slack Variable } \xi_i}

The Hard-Margin constraint was:

yi(wTxi+b)≥1\boxed{ y_i(w^Tx_i+b)\geq1 }

The Soft-Margin constraint becomes:

yi(wTxi+b)≥1−ξi\boxed{ y_i(w^Tx_i+b)\geq1-\xi_i }

where:

ξi≥0\boxed{\xi_i\geq0}

The value of ξi\xi_i tells us how much the point violates the margin.


6. Meaning of Different Values of ξ\xi

Slack VariableMeaning
ξi=0\xi_i=0    Correctly classified and outside/on margin
0<ξi<10<\xi_i<1    Correctly classified but inside margin
ξi=1\xi_i=1    On the decision boundary
ξi>1\xi_i>1    Incorrectly classified

This is a very useful table for students.


7. Mathematical Formulation of Soft-Margin SVM

The optimization problem is:

Minimize 12∥w∥2+C∑i=1nξi\boxed{ \text{Minimize } \frac{1}{2}\|w\|^2+ C\sum_{i=1}^{n}\xi_i }

Subject to:

yi(wTxi+b)≥1−ξi\boxed{ y_i(w^Tx_i+b)\geq1-\xi_i }

and:

ξi≥0\boxed{ \xi_i\geq0 }

8. Understanding the Objective Function

The objective is:

12∥w∥2+C∑ξi\boxed{ \frac{1}{2}\|w\|^2+ C\sum\xi_i }

It has two parts.

Part 1

12∥w∥2\frac12\|w\|^2

This tries to:

maximize the margin\boxed{\text{maximize the margin}}

Part 2

C∑ξiC\sum\xi_i

This tries to:

reduce margin violations\boxed{\text{reduce margin violations}}

Therefore, Soft-Margin SVM balances:

Large MarginvsFew Errors\boxed{ \text{Large Margin} \quad\text{vs}\quad \text{Few Errors} }

9. What is the Role of CC?

The parameter:

C\boxed{C}

controls how strongly we penalize errors.


Large CC

If:

C=1000C=1000

errors are heavily penalized.

The SVM tries very hard to classify every training point correctly.

Large C

"Don't make mistakes!"

This may result in a smaller margin.


Small CC

If:

C=0.1C=0.1

the SVM allows more violations.

Small C

"A few mistakes are acceptable."

This can result in a wider margin.


Simple Comparison

Value of CCEffect
Large CC    Fewer training errors, smaller margin
Small CC    More violations allowed, wider margin

10. Simple Numerical Example Using Slack Variable

Suppose we have a point:

x=(2,3)x=(2,3)

with class:

y=+1y=+1

Suppose the SVM decision function gives:

wTx+b=0.6w^Tx+b=0.6

Then:

y(wTx+b)y(w^Tx+b) =(+1)(0.6)=(+1)(0.6) =0.6=0.6

The Soft-Margin condition is:

y(wTx+b)≥1−ξy(w^Tx+b)\geq1-\xi

Substituting:

0.6≥1−ξ0.6\geq1-\xi

Therefore:

ξ≥0.4\xi\geq0.4

So:

ξ=0.4\boxed{\xi=0.4}

This means:

The point is correctly classified, but it violates the margin by 0.40.4.


11. Example of a Misclassified Point

Suppose:

y=−1y=-1

but:

wTx+b=0.5w^Tx+b=0.5

Then:

y(wTx+b)y(w^Tx+b) =(−1)(0.5)=(-1)(0.5) =−0.5=-0.5

The constraint is:

−0.5≥1−ξ-0.5\geq1-\xi

Therefore:

ξ≥1.5\xi\geq1.5

Thus:

ξ=1.5\boxed{\xi=1.5}

Since:

ξ>1\xi>1

the point is misclassified.


12. Soft Margin and Hinge Loss

Instead of explicitly calculating slack variables, we can use hinge loss.

The hinge loss is:

Li=max⁡(0,1−yi(wTxi+b))\boxed{ L_i=\max(0,1-y_i(w^Tx_i+b)) }

Notice that this is equivalent to the amount of margin violation.

Therefore:

ξi=max⁡(0,1−yi(wTxi+b))\boxed{ \xi_i=\max(0,1-y_i(w^Tx_i+b)) }

So the Soft-Margin SVM objective can also be written as:

J(w,b)=12∥w∥2+C∑i=1nmax⁡(0,1−yi(wTxi+b))\boxed{ J(w,b) = \frac12\|w\|^2 + C\sum_{i=1}^{n} \max(0,1-y_i(w^Tx_i+b)) }

This is the formulation commonly used when implementing SVM using gradient descent.


13. Hard Margin vs Soft Margin

Hard-Margin SVMSoft-Margin SVM
Requires perfectly separable data    Allows some violations
No classification errors allowed    Some errors are allowed
Sensitive to outliers    More robust to outliers
Uses strict constraints    Uses slack variables
Best for ideal datasets    Better for real-world datasets


Hard-Margin SVM does not allow violations. Soft-Margin SVM allows violations but penalizes them.


Final Summary

Hard Margin

yi(wTxi+b)≥1\boxed{ y_i(w^Tx_i+b)\geq1 }

No violations allowed.


Soft Margin

yi(wTxi+b)≥1−ξi\boxed{ y_i(w^Tx_i+b)\geq1-\xi_i }

Violations are allowed.


Objective Function

min⁡12∥w∥2+C∑iξi\boxed{ \min \frac12\|w\|^2+ C\sum_i\xi_i }

Equivalent Hinge Loss Form

min⁡12∥w∥2+C∑imax⁡(0,1−yi(wTxi+b))\boxed{ \min \frac12\|w\|^2+ C\sum_i \max(0,1-y_i(w^Tx_i+b)) }
Soft-Margin SVM tries to find the widest possible separating boundary while allowing some points to violate the margin or even be misclassified. The parameter C controls the trade-off between a large margin and classification errors.

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming