Kernel Trick in SVM
Kernel Trick in SVM
The Kernel Trick is one of the most important ideas in Non-Linear SVM.
If the data cannot be separated by a straight line in the original space, SVM imagines the data in a higher-dimensional space where it can be separated by a straight line. A kernel helps us do this without explicitly calculating the higher-dimensional coordinates.
1. Start with a Simple Problem
Consider a dataset with one feature .
| Class | |
|---|---|
| -3 | |
| -2 | |
| -1 | |
| 0 | |
| 1 | |
| 2 | |
| 3 |
On a number line:
Class +1 Class -1 Class +1 ● ● × × × ● ● -3 -2 -1 0 1 2 3
Here:
- Class is on both sides
- Class is in the middle
A single straight boundary cannot separate the two classes.
So, this is a non-linear classification problem in the original space.
2. The Main Idea: Transform the Feature
Let us create a new feature:
This is called a feature transformation.
Now calculate .
| Class | ||
|---|---|---|
| -3 | 9 | |
| -2 | 4 | |
| -1 | 1 | |
| 0 | 0 | |
| 1 | 1 | |
| 2 | 4 | |
| 3 | 9 |
Now look only at the transformed feature:
Class -1 Class +1 × × × | ● ● ● ● 0 1 1 4 4 9 9 ↑ Decision boundary
Now the classes can be separated by a straight boundary.
This is the important idea:
3. What is Feature Mapping?
We can represent the transformation as:
Here:
- = original feature
- = transformed feature
For example:
becomes:
Similarly:
becomes:
4. But Why Do We Need the Kernel Trick?
For a simple example like:
we can easily calculate the transformation.
But imagine data with:
- 100 features,
- thousands of training examples,
- very complex transformations.
The transformed feature space could have:
- thousands,
- millions,
- or even infinitely many dimensions.
Explicitly calculating all these transformed features may be difficult.
This is where the Kernel Trick becomes useful.
5. The Important Mathematical Idea
SVM does not always need the actual transformed coordinates.
During training, many calculations involve the dot product:
Instead of explicitly calculating:
and:
we use a kernel function:
This is the Kernel Trick.
6. A Simple Mathematical Example
Suppose we choose the feature mapping:
This transforms a one-dimensional input into a two-dimensional feature space.
Let's take:
Then:
Now take:
Then:
The dot product in the transformed space is:
7. The Kernel Does This Directly
Instead of calculating:
and:
separately, we can define:
Now substitute:
We get exactly the same result!
Therefore:
8. Why is This Called a "Trick"?
Because we get the result of calculations in a higher-dimensional space:
without explicitly creating all the transformed features.
So the SVM behaves as if it is working in a higher-dimensional space.
Original Data x │ │ ▼ Kernel Function │ │ ▼ Effect of Higher-Dimensional Feature Space │ ▼ Linear SVM
That is why it is called the:
9. A More Visual 2-Dimensional Example
Consider this dataset:
○ ○ ○ ○ ○ ○ ○ ● ● ● ● ○ ○ ○ ○ ○ ○ ○
Suppose:
- ● = Class 0
- ○ = Class 1
The Class 0 points are in the center.
The Class 1 points surround them.
Can a straight line separate them?
Add a New Dimension
Suppose we calculate:
This measures the squared distance of a point from the origin.
Now the data is represented using:
where:
Points near the center have small values of .
Points far from the center have large values of .
Therefore, in the higher-dimensional space, it may become possible to separate the classes using a plane.
Higher-dimensional idea: Class 1 ○ ○ ○ ------------------------ ← Linear separating plane ● ● Class 0
When we look back at the original two-dimensional space, the boundary may appear as a:
10. Common Kernel Functions
1. Linear Kernel
Used when the data is approximately linearly separable.
No complicated transformation is needed.
2. Polynomial Kernel
This allows SVM to learn polynomial-shaped boundaries.
For example:
- quadratic boundaries,
- cubic boundaries.
3. RBF Kernel
The Radial Basis Function (RBF) kernel is very popular.
It can handle complex non-linear patterns.
The basic idea is:
Points that are close together are considered more similar.
11. Simple Comparison
Without Kernel
Input Data ↓ Try to draw a straight line ↓ Cannot separate the classes
With Kernel
Input Data ↓ Kernel Function ↓ Conceptually map to a higher dimension ↓ Find a linear separating hyperplane ↓ Appears as a curved boundary in the original space
12. The Most Important Formula
The key formula students should remember is:
where:
- are original data points.
- represents transformation into a higher-dimensional feature space.
- calculates the dot product in that feature space.
Final Summary
Problem
Data cannot be separated by a straight line.
⬇️
Solution
Conceptually transform the data:
⬇️
In the New Space
Find a linear separating hyperplane.
⬇️
Kernel Trick
Instead of explicitly calculating , use:
Comments
Post a Comment