ReLU Activation Function
- Get link
- X
- Other Apps
๐ต ReLU Activation Function
ReLU stands for Rectified Linear Unit. It is one of the most widely used activation functions in modern neural networks, particularly in hidden layers.
1. Definition of ReLU
ReLU is defined as:
Equivalently, we can write it as a piecewise function:
In simple words:
ReLU keeps positive values and converts negative values to zero.
Examples
So:
| ReLU() | |
|---|---|
| -5 | 0 |
| -2 | 0 |
| -1 | 0 |
| 0 | 0 |
| 1 | 1 |
| 2 | 2 |
| 5 | 5 |
2. How ReLU works inside a neuron
As with other activation functions, first calculate the weighted sum:
Then apply ReLU:
Therefore,
Example
Suppose
and
First calculate :
Now apply ReLU:
3. What happens for positive and negative inputs?
The easiest way to understand ReLU is:
๐ด Negative input
If
then
The neuron produces zero output.
๐ข Positive input
If
then
The neuron passes the value directly to the next layer.
Therefore:
4. Graph of ReLU
The graph has two very simple regions:
and
So it looks like this conceptually:
The important point is that ReLU is nonlinear, even though its positive portion is a straight line.
5. Derivative of ReLU
For backpropagation, we need the derivative.
For
the derivative is
At , ReLU is not differentiable in the ordinary sense because the slope changes abruptly from to .
In practical neural-network implementations, a subgradient convention is chosen, commonly either or .
6. Why is ReLU useful for neural networks?
The most important reason is that ReLU introduces nonlinearity.
Consider:
Without an activation function, a neuron simply performs a linear transformation.
ReLU gives:
This allows a neural network with multiple layers to model nonlinear relationships.
For example, a multilayer network can learn complex patterns such as:
- image patterns
- handwritten digits
- speech patterns
- object features
- nonlinear classification boundaries
7. Why is ReLU often preferred over sigmoid and tanh in hidden layers?
This is an important comparison for students.
Sigmoid
Its derivative is
The maximum derivative is only
For large positive or negative , the derivative becomes very small.
Tanh
Its derivative is
Its maximum derivative is
but it also suffers from saturation for large .
ReLU
and
For positive , the derivative is exactly 1.
Therefore, positive activations do not suffer from the same saturation problem as sigmoid and tanh.
8. ReLU and the vanishing-gradient problem
This is one of the major advantages of ReLU.
For sigmoid:
when becomes very large or very negative.
For tanh:
when
But for ReLU:
for all positive .
Therefore, when the neuron is active:
This was an important reason ReLU became popular in deep neural networks.
9. The "dying ReLU" problem
ReLU has its own important limitation.
For negative inputs:
and
Suppose a neuron consistently receives negative values.
Then:
and
Consequently, the neuron may stop receiving useful gradient updates and effectively become inactive.
This is called the:
10. Leaky ReLU: a solution to dying ReLU
One solution is Leaky ReLU.
It is defined as:
where is a small positive number, such as
For example:
instead of
Thus, Leaky ReLU allows a small gradient for negative inputs.
11. Why ReLU is so popular
The main advantages are:
✅ 1. Very simple
There is no exponential calculation.
✅ 2. Computationally efficient
It is much simpler to calculate than sigmoid or tanh.
✅ 3. Introduces nonlinearity
It allows neural networks to learn nonlinear functions.
✅ 4. Better gradient behavior for positive inputs
for .
✅ 5. Sparse activation
Negative inputs produce zero:
Therefore, many neurons may be inactive for a particular input. This creates sparse activations.
12. Where is ReLU used?
ReLU is particularly common in the hidden layers of neural networks.
For example:
It is widely used in architectures for:
- ๐ผ️ Image classification
- ๐️ Computer vision
- ๐ Speech processing
- ๐ Natural language processing
- ๐ค Deep neural networks
- ๐ง Convolutional neural networks
Modern networks also frequently use ReLU variants such as Leaky ReLU, PReLU, ELU, and GELU depending on the architecture.
๐ Summary
"ReLU is a gate. If the input is negative, it closes the gate and produces 0. If the input is positive, it opens the gate and passes the input unchanged."
Mathematically:
๐ง Memory trick
or simply:
This one equation captures the entire basic idea of ReLU.
ReLU (Rectified Linear Unit) is a widely used nonlinear activation function in neural networks, especially in hidden layers. It is defined as , meaning negative inputs are converted to zero while positive inputs are passed unchanged. Its derivative is for negative inputs and for positive inputs, which helps reduce the vanishing-gradient problem seen with sigmoid and tanh. ReLU is simple, computationally efficient, and introduces the nonlinearity needed to learn complex patterns. However, neurons can become permanently inactive for negative inputs, known as the dying ReLU problem.
- Get link
- X
- Other Apps
Comments
Post a Comment