Batch Gradient Descent
Batch Gradient Descent
1. What is Batch Gradient Descent?
In Batch Gradient Descent, we do not update the weights after every individual training example.
Instead:
We process the entire training dataset, calculate the total gradient, and then update the weights once.
The basic idea is:
This complete process is called one iteration, or commonly one epoch when the entire training dataset has been processed.
2. Why do we need Batch Gradient Descent?
Suppose we have training examples:
For each example, the neuron produces:
and has error:
If we change the weights immediately after every example, the direction of the update is based on only one example.
Instead, Batch Gradient Descent asks:
What direction will reduce the total error over the entire training dataset?
Therefore, we calculate the total error:
and find the gradient of this total error.
3. Mathematical derivation
For a linear neuron:
The total squared error is:
The derivative with respect to is:
Therefore, gradient descent gives:
Substituting the gradient:
Similarly, for the bias:
This is the mathematical form of Batch Gradient Descent for the linear unit.
4. Algorithm
The algorithm can be written as:
BATCH-GRADIENT-DESCENT(training_examples, η) Initialize w1, w2, ..., wn and b Repeat until termination condition is met: Initialize: Δwi ← 0 Δb ← 0 For each training example (x, t): Calculate: y ← Σ wi xi + b Calculate error: e ← t - y For each weight wi: Δwi ← Δwi + η e xi Δb ← Δb + η e Update weights: wi ← wi + Δwi Update bias: b ← b + Δb
Notice the important point:
The weights are not updated inside the training-example loop.
They are updated only after all training examples have been processed.
5. Simple example
Suppose we have three training examples:
| Example | ||
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 6 |
Suppose initially:
and
Example 1
Weight change:
Example 2
Using the same original weight :
Accumulated:
Example 3
Total:
Only now do we update:
So:
The important observation is that all three examples contributed to one weight update.
6. Why is this called "Batch"?
The word batch means the complete set of training examples.
Suppose we have:
training examples.
In Batch Gradient Descent:
Then the process repeats for the next epoch:
7. Batch vs Single-example Update
This is an important distinction for students.
Single-example / Online update
For every example:
calculate , calculate the error, and immediately update:
So:
Batch Gradient Descent
Process all examples first:
So:
8. Advantages of Batch Gradient Descent
1. Stable and consistent direction
Because the gradient is calculated using all training examples, the update direction is based on the complete dataset.
This generally makes the optimization path smoother.
2. Less noisy updates
With one training example, the update can be strongly influenced by that particular example.
Batch Gradient Descent considers:
so individual examples have less influence on the direction of a single update.
3. Exact gradient of the training loss
For the chosen loss function, Batch Gradient Descent calculates the gradient of the entire training dataset loss:
Thus, the update follows the true gradient of the empirical training loss.
4. Deterministic updates
If the dataset and learning rate remain unchanged, processing the same dataset produces the same gradient and therefore the same update.
This makes it relatively easy to analyze and understand mathematically.
5. Good for understanding gradient descent
For teaching the concept of gradient descent, Batch Gradient Descent is particularly useful because students can clearly see:
9. Disadvantages of Batch Gradient Descent
1. Computationally expensive for large datasets
Suppose there are:
training examples.
A single update requires processing all one million examples.
Therefore:
2. Requires more memory
The complete batch may need to be available for computing the gradient, especially in straightforward implementations.
For very large datasets, this can be inconvenient.
3. Slow parameter updates
The weights are updated only after the complete dataset has been processed.
With a very large dataset, the model may have to process a lot of data before making even one update.
4. Not ideal for continuously arriving data
Suppose new training examples are continuously arriving.
Batch Gradient Descent is less convenient because it is designed around repeatedly processing a defined training set.
10. Comparison with other approaches
This leads naturally to three important forms of gradient descent:
| Method | Examples used for one update | Weight updates |
|---|---|---|
| Batch Gradient Descent | Entire dataset | 1 per epoch |
| Stochastic Gradient Descent (SGD) | 1 example | per epoch |
| Mini-batch Gradient Descent | Small batch | per epoch |
For example, if:
and mini-batch size:
then:
updates are made per epoch.
11. Visual intuition
Think of the error surface as a mountain landscape.
Batch Gradient Descent:
SGD:
Mini-batch:
12. Summary
Batch Gradient Descent calculates the gradient using all training examples and updates the weights only once after the entire training dataset has been processed.
Mathematically:
and
The central idea is:
This is also a very natural bridge from the delta rule to mini-batch gradient descent and backpropagation in MLPs.
Comments
Post a Comment