Perceptron Learning Using Gradient Descent
🧑🏫Perceptron Learning Using Gradient Descent
The important point is that there are two closely related ideas:
- Perceptron learning rule → uses the classification error to update weights.
- Delta rule / gradient descent → uses a continuous output and minimizes a loss function.
So, when we say "single-layer perceptron learning using gradient descent," we usually mean training a single-layer linear unit using gradient descent. This is the foundation for understanding how gradient descent later trains multilayer networks.
1. Single-Layer Network
Consider a neuron with two inputs:
The neuron first calculates the weighted sum:
For a linear unit:
Therefore:
2. Why Use Gradient Descent?
Suppose the target output is , but the neuron produces .
There is an error:
We want to find weights that minimize this error.
We define the squared error:
The objective is:
Gradient descent provides a systematic way to find the weights that minimize .
3. The Gradient Descent Idea
Imagine the error as a surface.
The gradient tells us the direction in which the error increases.
Therefore, we move in the opposite direction:
where is the learning rate.
4. Deriving the Weight Update Rule
This is the most important part.
We have:
and:
We want:
Using the chain rule:
First:
and:
Therefore:
Now apply gradient descent:
Substituting:
Therefore:
This is the delta rule for a single linear unit.
Similarly, for the bias:
5. Complete Learning Rule
For each training example:
Step 1: Calculate output
Step 2: Calculate error
Step 3: Update weights
Step 4: Update bias
Repeat this process for all training examples.
6. Worked Example
Let's take a very simple example.
Suppose:
Initial weights:
Bias:
Target:
Learning rate:
Step 1: Calculate the output
Step 2: Calculate the error
The prediction is too small, so the weights should increase appropriately.
Step 3: Update
The update rule is:
Substitute:
Step 4: Update
Step 5: Update the bias
Therefore, after one update:
7. What Happens in the Next Iteration?
We use the updated parameters:
Calculate the output again:
Previously:
Now:
The target is:
So the error has reduced:
whereas now:
Excellent! The weight update moved the prediction closer to the target.
8. Another Update
Error:
Update :
Update :
Update bias:
Again, the prediction moves toward the target.
What is the Gradient in This Case?
For our single-layer linear neuron, the gradient tells us:
How much will the error change if we slightly change each weight?
Suppose the neuron is
and the squared error is
The gradient of the error with respect to the weights is:
We calculate each component separately.
For
Using the chain rule:
Since
and
we get:
Similarly,
Therefore, the gradient is
🔢 Using our previous example
We had:
Therefore:
The gradient is:
So:
This means:
and
What does this tell us?
The negative signs tell us that increasing these weights will decrease the error in this particular situation.
Gradient descent moves in the opposite direction of the gradient:
Thus:
giving:
which is exactly the weight update we calculated earlier.
🧠 In simple words
Think of the gradient as an arrow pointing uphill on the error surface:
Therefore:
And gradient descent simply says:
"Look at the gradient and take a small step in the opposite direction."
9. What Is Happening Geometrically?
Every weight vector represents a particular model.
Initially:
The error is relatively high.
Gradient descent changes the weights:
The weights gradually move toward a region where the error is smaller.
Conceptually:
10. For Multiple Training Examples
In practice, we have many training examples:
The total squared error can be written as:
Gradient descent attempts to minimize this total error.
11. Learning Process
The complete algorithm is:
Initialize weights and bias ↓ Choose a training example ↓ Calculate: y = w₁x₁ + w₂x₂ + b ↓ Calculate error: e = t - y ↓ Update: wᵢ = wᵢ + ηexᵢ b = b + ηe ↓ Next training example ↓ Repeat for all examples ↓ Repeat for several epochs ↓ Error becomes small
12. Important Point: Delta Rule vs Perceptron Rule
This distinction is important when teaching from Alpaydin.
Perceptron
The output is thresholded:
and the perceptron rule updates weights based on classification errors.
Delta rule
The neuron is unthresholded:
and gradient descent minimizes the squared error:
with update:
This distinction is important because gradient descent requires a differentiable objective, whereas the hard threshold of the classic perceptron is not differentiable at its threshold.
13. Why Is This Important for MLP?
This single-neuron example gives us the foundation for multilayer networks.
For a single linear unit:
For an MLP:
followed by:
So the progression is:
🎓 Summary
A single-layer linear neuron learns by comparing its predicted output with the target, calculating the error, and changing its weights in the direction that reduces the error. Gradient descent determines this direction using the gradient of the error function. For a linear unit with squared error, this gives the delta rule . Repeating this process over training examples gradually moves the weights toward values that minimize the overall training error. This simple idea forms the foundation for training multilayer neural networks using backpropagation and gradient descent.
⭐ Three formulas students should remember
Output:

Comments
Post a Comment