Introduction to Non-Linear SVM

 

Introduction to Non-Linear SVM

A Support Vector Machine (SVM) is used for classification by finding a decision boundary that separates different classes.

1. Linear SVM

In a Linear SVM, the classes can be separated using a straight line.

For example:

Class 0              Class 1

 ●  ●  ●       |       ○  ○  ○
 ●  ●          |       ○  ○

              ↑
      Straight decision boundary

In two dimensions, the decision boundary is:

w1x1+w2x2+b=0\boxed{w_1x_1+w_2x_2+b=0}

2. The Problem: What if a Straight Line Cannot Separate the Data?

Consider the following data:

             ○  ○  ○

         ○             ○

              ● ●
              ● ●

         ○             ○

             ○  ○  ○

Here:

  • ● = Class 0
  • ○ = Class 1

Class 0 is surrounded by Class 1.

A straight line cannot separate these two classes.

This type of data is called:

Non-linearly separable data\boxed{\text{Non-linearly separable data}}

3. Non-Linear SVM

A Non-Linear SVM is used when the data cannot be separated by a straight line in the original feature space.

Instead, it finds a non-linear decision boundary, such as:

  • a circle,
  • a curve,
  • an ellipse,
  • or another complex shape.

For the above example, the decision boundary may look like:

             ○  ○  ○

         ○             ○

             _______
            /       \
           |  ● ● ●  |
           |  ● ● ●  |
            \_______/

         ○             ○

             ○  ○  ○

4. How Does Non-Linear SVM Work?

The main idea is:

Transform the data into a higher-dimensional space where it can be separated using a linear boundary.

This transformation is represented by:

x→ϕ(x)\boxed{x\rightarrow\phi(x)}

Then SVM finds a linear separating hyperplane in the new feature space.

So:

Non-linear in original space→Linear in higher-dimensional space\boxed{ \text{Non-linear in original space} \rightarrow \text{Linear in higher-dimensional space} }

When we look at the boundary again in the original space, it appears non-linear.


5. The Kernel Trick

Actually creating many new features can be computationally expensive.

Therefore, SVM uses a kernel function.

The kernel allows SVM to perform calculations as if the data had been transformed into a higher-dimensional space.

The basic idea is:

K(xi,xj)=ϕ(xi)Tϕ(xj)\boxed{ K(x_i,x_j)=\phi(x_i)^T\phi(x_j) }

This is called the Kernel Trick.


6. Common Kernels Used in Non-Linear SVM

Polynomial Kernel

Useful for polynomial-shaped decision boundaries.

RBF Kernel

Useful for complex non-linear patterns and curved decision boundaries.

The RBF kernel is one of the most commonly used kernels in practice.


Linear SVM vs Non-Linear SVM

Linear SVMNon-Linear SVM
Uses a straight boundary    Uses a curved or complex boundary
Works for linearly separable data    Works for non-linear patterns
Simple and fast    More flexible
No kernel or linear kernel    Uses kernels such as RBF or Polynomial

⭐

A Non-Linear SVM is used when the classes cannot be separated by a straight line. It uses kernel functions to handle complex patterns and create non-linear decision boundaries.

The key idea to remember:

If we cannot separate with a straight line⇒Use a non-linear SVM with a suitable kernel

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming