Perceptron Training Rule


The objective of perceptron learning is to find a set of weights that makes the perceptron correctly classify all training examples, provided that the training data are linearly separable.

For a training example x\mathbf{x}, the perceptron computes

z=wTx+bz=\mathbf{w}^{T}\mathbf{x}+b

and produces the output

y={+1,z≥0−1,z<0y= \begin{cases} +1, & z\geq0\\ -1, & z<0 \end{cases}

where:

  • x\mathbf{x} = input vector
  • w\mathbf{w} = weight vector
  • bb = bias
  • yy = output produced by the perceptron
  • tt = target (desired) output
  • η\eta = learning rate

1. Perceptron Weight Update Rule

The perceptron training rule updates each weight according to

wi←wi+η(t−y)xi\boxed{ w_i \leftarrow w_i+\eta(t-y)x_i }

The bias can similarly be updated as

b←b+η(t−y)\boxed{ b\leftarrow b+\eta(t-y) }

Therefore, in vector form:

w←w+η(t−y)x\boxed{ \mathbf{w}\leftarrow\mathbf{w}+\eta(t-y)\mathbf{x} }

and

b←b+η(t−y)\boxed{ b\leftarrow b+\eta(t-y) }

2. Meaning of the Error Term

The quantity

t−y\boxed{t-y}

determines how the weights should be changed.

Case 1: Correct classification

If

t=yt=y

then

t−y=0t-y=0

Therefore,

Δwi=0\Delta w_i=0

and

Δb=0\Delta b=0

So no weight update is made.


Case 2: Target +1+1, predicted −1-1

Suppose

t=+1,y=−1t=+1,\qquad y=-1

Then

t−y=1−(−1)=2t-y=1-(-1)=2

Therefore,

wi←wi+2ηxiw_i\leftarrow w_i+2\eta x_i

The weights associated with positive inputs increase, helping the perceptron produce a larger value of zz for this example.


Case 3: Target −1-1, predicted +1+1

Suppose

t=−1,y=+1t=-1,\qquad y=+1

Then

t−y=−1−(+1)=−2t-y=-1-(+1)=-2

Therefore,

wi←wi−2ηxiw_i\leftarrow w_i-2\eta x_i

The weights associated with positive inputs decrease, helping reduce zz for this example.


3. Example of Weight Update

Suppose

x=[1,2]\mathbf{x}=[1,2] w=[0.3,0.4]\mathbf{w}=[0.3,0.4] b=0.1b=0.1

and

η=0.1\eta=0.1

Suppose the target is

t=+1t=+1

but the perceptron produces

y=−1y=-1

Then:

t−y=1−(−1)=2t-y=1-(-1)=2

For the first weight:

w1new=0.3+(0.1)(2)(1)w_1^{new}=0.3+(0.1)(2)(1) w1new=0.5\boxed{w_1^{new}=0.5}

For the second weight:

w2new=0.4+(0.1)(2)(2)w_2^{new}=0.4+(0.1)(2)(2) w2new=0.8\boxed{w_2^{new}=0.8}

Bias:

bnew=0.1+(0.1)(2)b^{new}=0.1+(0.1)(2) bnew=0.3\boxed{b^{new}=0.3}

Thus,

wnew=[0.5,0.8],bnew=0.3\boxed{ \mathbf{w}_{new}=[0.5,0.8],\qquad b_{new}=0.3 }

4. Perceptron Training Algorithm

The complete training process can be summarized as:

Initialize weights and bias
        ↓
Select a training example
        ↓
Calculate z = wᵀx + b
        ↓
Calculate perceptron output y
        ↓
Compare y with target t
        ↓
Is y = t ?
   ↙          ↘
 YES           NO
  ↓             ↓
No update    Update weights
              and bias
                  ↓
          Next training example
                  ↓
          Repeat for epochs
                  ↓
      All examples classified
             correctly?
                  ↓
                STOP




5. Convergence Condition

The important theoretical result emphasized in the perceptron learning approach is:

If the training data are linearly separable, the perceptron learning algorithm converges.\boxed{ \text{If the training data are linearly separable, the perceptron learning algorithm converges.} }

That means after a finite number of updates, the perceptron can find a weight vector that correctly classifies all training examples.

However, if the data are not linearly separable, convergence is not guaranteed.

For example:

XOR cannot be learned by a single perceptron\boxed{\text{XOR cannot be learned by a single perceptron}}

because XOR is not linearly separable.


🧠 Student takeaway

Perceptron learning=Predict→Compare→Update if wrong→Repeat\boxed{ \text{Perceptron learning} = \text{Predict} \rightarrow \text{Compare} \rightarrow \text{Update if wrong} \rightarrow \text{Repeat} }

The central rule to remember is:

winew=wiold+η(t−y)xi\boxed{ w_i^{new}=w_i^{old}+\eta(t-y)x_i }

and

bnew=bold+η(t−y)\boxed{ b^{new}=b^{old}+\eta(t-y) }

This is the perceptron training rule.

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming