Perceptron Learning Using Gradient Descent

 

🧑‍🏫Perceptron Learning Using Gradient Descent

The important point is that there are two closely related ideas:

  • Perceptron learning rule → uses the classification error to update weights.
  • Delta rule / gradient descent → uses a continuous output and minimizes a loss function.

So, when we say "single-layer perceptron learning using gradient descent," we usually mean training a single-layer linear unit using gradient descent. This is the foundation for understanding how gradient descent later trains multilayer networks.


1. Single-Layer Network

Consider a neuron with two inputs:

   



The neuron first calculates the weighted sum:

z=w1x1+w2x2+b\boxed{ z=w_1x_1+w_2x_2+b }

For a linear unit:

y=z\boxed{y=z}

Therefore:

y=w1x1+w2x2+b\boxed{ y=w_1x_1+w_2x_2+b }

2. Why Use Gradient Descent?

Suppose the target output is tt, but the neuron produces yy.

There is an error:

t−yt-y

We want to find weights w1,w2w_1,w_2 that minimize this error.

We define the squared error:

E=12(t−y)2\boxed{ E=\frac{1}{2}(t-y)^2 }

The objective is:

Minimize E\boxed{\text{Minimize }E}

Gradient descent provides a systematic way to find the weights that minimize EE.


3. The Gradient Descent Idea

Imagine the error as a surface.

            

The gradient tells us the direction in which the error increases.

Therefore, we move in the opposite direction:

wnew=wold−η∂E∂w\boxed{ w_{\text{new}} = w_{\text{old}} - \eta \frac{\partial E}{\partial w} }

where η\eta is the learning rate.


4. Deriving the Weight Update Rule

This is the most important part.

We have:

E=12(t−y)2E=\frac{1}{2}(t-y)^2

and:

y=w1x1+w2x2+by=w_1x_1+w_2x_2+b

We want:

∂E∂wi\frac{\partial E}{\partial w_i}

Using the chain rule:

∂E∂wi=∂E∂y∂y∂wi\frac{\partial E}{\partial w_i} = \frac{\partial E}{\partial y} \frac{\partial y}{\partial w_i}

First:

∂E∂y=−(t−y)\frac{\partial E}{\partial y}=-(t-y)

and:

∂y∂wi=xi\frac{\partial y}{\partial w_i}=x_i

Therefore:

∂E∂wi=−(t−y)xi\frac{\partial E}{\partial w_i} =-(t-y)x_i

Now apply gradient descent:

winew=wiold−η∂E∂wiw_i^{new} = w_i^{old} - \eta \frac{\partial E}{\partial w_i}

Substituting:

winew=wiold−η[−(t−y)xi]w_i^{new} = w_i^{old} - \eta[-(t-y)x_i]

Therefore:

winew=wiold+η(t−y)xi\boxed{ w_i^{new} = w_i^{old} + \eta(t-y)x_i }

This is the delta rule for a single linear unit.

Similarly, for the bias:

bnew=bold+η(t−y)\boxed{ b^{new}=b^{old}+\eta(t-y) }

5. Complete Learning Rule

For each training example:

Step 1: Calculate output

y=w1x1+w2x2+b\boxed{ y=w_1x_1+w_2x_2+b }

Step 2: Calculate error

e=t−y\boxed{ e=t-y }

Step 3: Update weights

wi←wi+ηexi\boxed{ w_i\leftarrow w_i+\eta e x_i }

Step 4: Update bias

b←b+ηe\boxed{ b\leftarrow b+\eta e }

Repeat this process for all training examples.


6. Worked Example

Let's take a very simple example.

Suppose:

x1=1,x2=2x_1=1,\qquad x_2=2

Initial weights:

w1=0.2,w2=0.3w_1=0.2,\qquad w_2=0.3

Bias:

b=0.1b=0.1

Target:

t=2t=2

Learning rate:

η=0.1\eta=0.1

Step 1: Calculate the output

y=w1x1+w2x2+by=w_1x_1+w_2x_2+b y=(0.2)(1)+(0.3)(2)+0.1y=(0.2)(1)+(0.3)(2)+0.1 y=0.2+0.6+0.1y=0.2+0.6+0.1 y=0.9\boxed{y=0.9}

Step 2: Calculate the error

e=t−ye=t-y e=2−0.9e=2-0.9 e=1.1\boxed{e=1.1}

The prediction is too small, so the weights should increase appropriately.


Step 3: Update w1w_1

The update rule is:

w1new=w1+ηex1w_1^{new}=w_1+\eta e x_1

Substitute:

w1new=0.2+(0.1)(1.1)(1)w_1^{new}=0.2+(0.1)(1.1)(1) w1new=0.2+0.11w_1^{new}=0.2+0.11 w1new=0.31\boxed{w_1^{new}=0.31}

Step 4: Update w2w_2

w2new=w2+ηex2w_2^{new}=w_2+\eta e x_2 w2new=0.3+(0.1)(1.1)(2)w_2^{new}=0.3+(0.1)(1.1)(2) w2new=0.3+0.22w_2^{new}=0.3+0.22 w2new=0.52\boxed{w_2^{new}=0.52}

Step 5: Update the bias

bnew=b+ηeb^{new}=b+\eta e bnew=0.1+(0.1)(1.1)b^{new}=0.1+(0.1)(1.1) bnew=0.21\boxed{b^{new}=0.21}

Therefore, after one update:

w1=0.31,w2=0.52,b=0.21\boxed{ w_1=0.31,\qquad w_2=0.52,\qquad b=0.21 }

7. What Happens in the Next Iteration?

We use the updated parameters:

w1=0.31,w2=0.52,b=0.21w_1=0.31,\quad w_2=0.52,\quad b=0.21

Calculate the output again:

y=(0.31)(1)+(0.52)(2)+0.21y=(0.31)(1)+(0.52)(2)+0.21 y=0.31+1.04+0.21y=0.31+1.04+0.21 y=1.56\boxed{y=1.56}

Previously:

y=0.9y=0.9

Now:

y=1.56y=1.56

The target is:

t=2t=2

So the error has reduced:

∣2−0.9∣=1.1|2-0.9|=1.1

whereas now:

∣2−1.56∣=0.44|2-1.56|=0.44

Excellent! The weight update moved the prediction closer to the target.


8. Another Update

Error:

e=2−1.56=0.44e=2-1.56=0.44

Update w1w_1:

w1=0.31+(0.1)(0.44)(1)w_1=0.31+(0.1)(0.44)(1) w1=0.354\boxed{w_1=0.354}

Update w2w_2:

w2=0.52+(0.1)(0.44)(2)w_2=0.52+(0.1)(0.44)(2) w2=0.608\boxed{w_2=0.608}

Update bias:

b=0.21+(0.1)(0.44)b=0.21+(0.1)(0.44) b=0.254\boxed{b=0.254}

Again, the prediction moves toward the target.

What is the Gradient in This Case?

For our single-layer linear neuron, the gradient tells us:

How much will the error change if we slightly change each weight?

Suppose the neuron is

y=w1x1+w2x2+by=w_1x_1+w_2x_2+b

and the squared error is

E=12(t−y)2E=\frac{1}{2}(t-y)^2

The gradient of the error with respect to the weights is:

∇wE=[∂E∂w1∂E∂w2]\boxed{ \nabla_{\mathbf w}E= \begin{bmatrix} \frac{\partial E}{\partial w_1}\\[4pt] \frac{\partial E}{\partial w_2} \end{bmatrix} }

We calculate each component separately.

For w1w_1

Using the chain rule:

∂E∂w1=∂E∂y∂y∂w1\frac{\partial E}{\partial w_1} = \frac{\partial E}{\partial y} \frac{\partial y}{\partial w_1}

Since

∂E∂y=−(t−y)\frac{\partial E}{\partial y}=-(t-y)

and

∂y∂w1=x1\frac{\partial y}{\partial w_1}=x_1

we get:

∂E∂w1=−(t−y)x1\boxed{ \frac{\partial E}{\partial w_1}=-(t-y)x_1 }

Similarly,

∂E∂w2=−(t−y)x2\boxed{ \frac{\partial E}{\partial w_2}=-(t-y)x_2 }

Therefore, the gradient is

∇wE=−(t−y)[x1x2]\boxed{ \nabla_{\mathbf w}E = -(t-y) \begin{bmatrix} x_1\\ x_2 \end{bmatrix} }


🔢 Using our previous example

We had:

x1=1,x2=2x_1=1,\qquad x_2=2 t=2,y=0.9t=2,\qquad y=0.9

Therefore:

t−y=1.1t-y=1.1

The gradient is:

∇wE=−(1.1)[12]\nabla_{\mathbf w}E = -(1.1) \begin{bmatrix} 1\\ 2 \end{bmatrix}

So:

∇wE=[−1.1−2.2]\boxed{ \nabla_{\mathbf w}E= \begin{bmatrix} -1.1\\ -2.2 \end{bmatrix} }

This means:

∂E∂w1=−1.1\frac{\partial E}{\partial w_1}=-1.1

and

∂E∂w2=−2.2\frac{\partial E}{\partial w_2}=-2.2

What does this tell us?

The negative signs tell us that increasing these weights will decrease the error in this particular situation.

Gradient descent moves in the opposite direction of the gradient:

wnew=wold−η∇wE\boxed{ \mathbf w_{\text{new}} = \mathbf w_{\text{old}} -\eta\nabla_{\mathbf w}E }

Thus:

wnew=[0.20.3]−0.1[−1.1−2.2]\mathbf w_{\text{new}} = \begin{bmatrix} 0.2\\ 0.3 \end{bmatrix} - 0.1 \begin{bmatrix} -1.1\\ -2.2 \end{bmatrix}

giving:

wnew=[0.310.52]\mathbf w_{\text{new}} = \begin{bmatrix} 0.31\\ 0.52 \end{bmatrix}

which is exactly the weight update we calculated earlier.

🧠 In simple words

Think of the gradient as an arrow pointing uphill on the error surface:

Gradient=direction of greatest increase in error\boxed{\text{Gradient}=\text{direction of greatest increase in error}}

Therefore:

−Gradient=direction of greatest decrease in error\boxed{-\text{Gradient}=\text{direction of greatest decrease in error}}

And gradient descent simply says:

"Look at the gradient and take a small step in the opposite direction."


9. What Is Happening Geometrically?

Every weight vector represents a particular model.

Initially:

(w1,w2)=(0.2,0.3)(w_1,w_2)=(0.2,0.3)

The error is relatively high.

Gradient descent changes the weights:

(0.2,0.3)→(0.31,0.52)→(0.354,0.608)→⋯(0.2,0.3) \rightarrow (0.31,0.52) \rightarrow (0.354,0.608) \rightarrow\cdots

The weights gradually move toward a region where the error is smaller.

Conceptually:







10. For Multiple Training Examples

In practice, we have many training examples:

(x1(1),x2(1),t(1))(x_1^{(1)},x_2^{(1)},t^{(1)}) (x1(2),x2(2),t(2))(x_1^{(2)},x_2^{(2)},t^{(2)}) ⋮\vdots (x1(N),x2(N),t(N))(x_1^{(N)},x_2^{(N)},t^{(N)})

The total squared error can be written as:

E=12∑d=1N(t(d)−y(d))2\boxed{ E= \frac{1}{2} \sum_{d=1}^{N} (t^{(d)}-y^{(d)})^2 }

Gradient descent attempts to minimize this total error.


11. Learning Process

The complete algorithm is:

Initialize weights and bias
          ↓
Choose a training example
          ↓
Calculate:
y = w₁x₁ + w₂x₂ + b
          ↓
Calculate error:
e = t - y
          ↓
Update:
wᵢ = wᵢ + ηexᵢ
b  = b + ηe
          ↓
Next training example
          ↓
Repeat for all examples
          ↓
Repeat for several epochs
          ↓
Error becomes small

12. Important Point: Delta Rule vs Perceptron Rule

This distinction is important when teaching from Alpaydin.

Perceptron

The output is thresholded:

y={+1,z≥0−1,z<0y= \begin{cases} +1,&z\geq0\\ -1,&z<0 \end{cases}

and the perceptron rule updates weights based on classification errors.

Delta rule

The neuron is unthresholded:

y=wTx+by=\mathbf{w}^T\mathbf{x}+b

and gradient descent minimizes the squared error:

E=12∑d(t(d)−y(d))2\boxed{ E=\frac12\sum_d(t^{(d)}-y^{(d)})^2 }

with update:

wi←wi+η(t−y)xi\boxed{ w_i\leftarrow w_i+\eta(t-y)x_i }

This distinction is important because gradient descent requires a differentiable objective, whereas the hard threshold of the classic perceptron is not differentiable at its threshold.


13. Why Is This Important for MLP?

This single-neuron example gives us the foundation for multilayer networks.

For a single linear unit:

Gradient descent→Update weights\boxed{ \text{Gradient descent} \rightarrow \text{Update weights} }

For an MLP:

Backpropagation→Calculate gradients\boxed{ \text{Backpropagation} \rightarrow \text{Calculate gradients} }

followed by:

Gradient descent→Update all weights\boxed{ \text{Gradient descent} \rightarrow \text{Update all weights} }

So the progression is:

Single Linear Unit→Delta Rule→Gradient Descent→Backpropagation→MLP\boxed{ \text{Single Linear Unit} \rightarrow \text{Delta Rule} \rightarrow \text{Gradient Descent} \rightarrow \text{Backpropagation} \rightarrow \text{MLP} }

🎓 Summary


A single-layer linear neuron learns by comparing its predicted output with the target, calculating the error, and changing its weights in the direction that reduces the error. Gradient descent determines this direction using the gradient of the error function. For a linear unit with squared error, this gives the delta rule wi←wi+η(t−y)xiw_i\leftarrow w_i+\eta(t-y)x_i. Repeating this process over training examples gradually moves the weights toward values that minimize the overall training error. This simple idea forms the foundation for training multilayer neural networks using backpropagation and gradient descent.

⭐ Three formulas students should remember

​Output:

y=∑iwixi+b\boxed{y=\sum_iw_ix_i+b}

Error:

E=12(t−y)2\boxed{E=\frac12(t-y)^2}

Weight update:

wi←wi+η(t−y)xi\boxed{w_i\leftarrow w_i+\eta(t-y)x_i}

​​​

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming