RBF Kernel in SVM

 

RBF Kernel in SVM

The RBF (Radial Basis Function) kernel is commonly used in SVM to handle non-linearly separable data.

Instead of using a straight-line decision boundary, the RBF kernel helps SVM create curved and complex decision boundaries.


1. RBF Kernel Formula

The RBF kernel is computed as:

K(xi,xj)=exp⁡(−γ∥xi−xj∥2)\boxed{ K(x_i,x_j)=\exp\left(-\gamma\|x_i-x_j\|^2\right) }

It calculates the similarity between two data points.


2. Parameters and Terms in the Formula

SymbolMeaning
xix_i    First data point
xjx_j    Second data point
∥xi−xj∥2\|x_i-x_j\|^2    Squared Euclidean distance between the points
γ\gamma    Controls the influence range of a data point
exp⁡()\exp()    Exponential function e(⋅)e^{(\cdot)}
K(xi,xj)K(x_i,x_j)    Kernel similarity value

3. How Does the RBF Kernel Compute Similarity?

Let us take two points:

xi=(1,2)x_i=(1,2)

and

xj=(3,5)x_j=(3,5)

Suppose:

γ=0.1\gamma=0.1

Step 1: Calculate the Difference

xi−xj=(1−3,  2−5)x_i-x_j=(1-3,\;2-5) =(−2,−3)=(-2,-3)

Step 2: Calculate the Squared Euclidean Distance

∥xi−xj∥2\|x_i-x_j\|^2 =(−2)2+(−3)2=(-2)^2+(-3)^2 =4+9=4+9 13\boxed{13}

Step 3: Multiply by −γ-\gamma

−γ∥xi−xj∥2-\gamma\|x_i-x_j\|^2 =−0.1(13)=-0.1(13) =−1.3=-1.3

Step 4: Apply the Exponential Function

K(xi,xj)=e−1.3K(x_i,x_j)=e^{-1.3} K(xi,xj)≈0.273\boxed{K(x_i,x_j)\approx0.273}

Therefore, the similarity between these two points is approximately:

0.273\boxed{0.273}

4. Interpretation of the Kernel Value

The RBF kernel value lies between:

0<K(xi,xj)≤1\boxed{0<K(x_i,x_j)\leq1}

If two points are identical

xi=xjx_i=x_j

Then:

∥xi−xj∥2=0\|x_i-x_j\|^2=0

Therefore:

K(xi,xj)=e0K(x_i,x_j)=e^0 K(xi,xj)=1\boxed{K(x_i,x_j)=1}

So, identical points have maximum similarity.


If two points are far apart

The distance becomes large:

∥xi−xj∥2→large\|x_i-x_j\|^2\rightarrow\text{large}

Therefore:

e−γ∥xi−xj∥2→0e^{-\gamma\|x_i-x_j\|^2}\rightarrow0

So:

Far points have similarity close to 0\boxed{\text{Far points have similarity close to }0}

5. The Role of γ\gamma

The parameter:

γ\boxed{\gamma}

controls how quickly similarity decreases as distance increases.

Small γ\gamma

Suppose:

γ=0.01\gamma=0.01

The similarity decreases slowly.

Therefore, even relatively distant points can influence each other.

Result:

Smoother and simpler decision boundary\boxed{\text{Smoother and simpler decision boundary}}

Large γ\gamma

Suppose:

γ=1\gamma=1

The similarity decreases quickly.

Only nearby points strongly influence each other.

Result:

More complex decision boundary\boxed{\text{More complex decision boundary}}

6. Example Showing the Effect of γ\gamma

Suppose the squared distance between two points is:

∥xi−xj∥2=4\|x_i-x_j\|^2=4

Case 1: Small γ=0.1\gamma=0.1

K=e−0.1(4)K=e^{-0.1(4)} =e−0.4=e^{-0.4}
K≈0.670
\boxed{K\approx0.670}

Case 2: Large γ=1\gamma=1

K=e−1(4)K=e^{-1(4)} =e−4=e^{-4}
K≈0.018
\boxed{K\approx0.018}

Comparison

γ\gamma    Kernel ValueMeaning
0.1        0.670        Points still have considerable similarity
1        0.018        Points have very little similarity

Thus:

Large γ⇒similarity decreases faster\boxed{\text{Large }\gamma\Rightarrow\text{similarity decreases faster}}

7. Another Important Parameter: CC

In an SVM using an RBF kernel, another important parameter is:

C\boxed{C}

The parameter CC controls the penalty for classification errors.

Small CC

The model allows more classification errors.

It generally produces a smoother decision boundary.

Large CC

The model strongly tries to classify all training points correctly.

It may produce a more complex boundary and can lead to overfitting.


8. Difference Between CC and γ\gamma

This is an important distinction for students.

ParameterControls
CC    Penalty for classification errors
γ\gamma    Influence range of each data point

Easy way to remember:

C→Cost of errors\boxed{ C\rightarrow\text{Cost of errors} }

γ→Geographical range of influence
\boxed{ \gamma\rightarrow\text{Geographical range of influence} }

9. How RBF Kernel is Used in SVM

During SVM training, the kernel function is computed between training points.

Conceptually:

K(x1,x2)K(x_1,x_2) K(x1,x3)K(x_1,x_3) K(x2,x3)K(x_2,x_3)

and so on.

These values form a kernel matrix.

For example:

K=[K(x1,x1)K(x1,x2)K(x1,x3)K(x2,x1)K(x2,x2)K(x2,x3)K(x3,x1)K(x3,x2)K(x3,x3)]K= \begin{bmatrix} K(x_1,x_1) & K(x_1,x_2) & K(x_1,x_3)\\ K(x_2,x_1) & K(x_2,x_2) & K(x_2,x_3)\\ K(x_3,x_1) & K(x_3,x_2) & K(x_3,x_3) \end{bmatrix}

The SVM uses these similarity values to find the support vectors and construct a non-linear decision boundary.


10. Simple Computational Flow

You can explain the process to students as:

Two data points
      ↓
Calculate distance between them
      ↓
Square the distance
      ↓
Multiply by -γ
      ↓
Apply exponential function
      ↓
Get similarity value K(xi, xj)
      ↓
SVM uses similarities to find
a non-linear decision boundary

⭐ Summary

The RBF kernel calculates the similarity between two data points using their distance. Nearby points produce a kernel value close to 1, while distant points produce a value close to 0. The parameter γ controls how quickly this similarity decreases, and C controls how strongly the SVM penalizes classification errors.

The RBF kernel is computed as:

K(xi,xj)=exp⁡(−γ∥xi−xj∥2)
​​​

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming