ReLU Activation Function

 

๐Ÿ”ต ReLU Activation Function

ReLU stands for Rectified Linear Unit. It is one of the most widely used activation functions in modern neural networks, particularly in hidden layers.


1. Definition of ReLU

ReLU is defined as:

ReLU⁡(z)=max⁡(0,z)\boxed{\operatorname{ReLU}(z)=\max(0,z)}

Equivalently, we can write it as a piecewise function:

ReLU⁡(z)={0,z<0z,z≥0\boxed{ \operatorname{ReLU}(z)= \begin{cases} 0, & z<0\\ z, & z\geq0 \end{cases} }

In simple words:

ReLU keeps positive values and converts negative values to zero.

Examples

ReLU⁡(−5)=0\operatorname{ReLU}(-5)=0 ReLU⁡(−2)=0\operatorname{ReLU}(-2)=0 ReLU⁡(0)=0\operatorname{ReLU}(0)=0 ReLU⁡(3)=3\operatorname{ReLU}(3)=3 ReLU⁡(10)=10\operatorname{ReLU}(10)=10

So:

zz   ReLU(zz)
-50
-20
-10
00
11
22
55

2. How ReLU works inside a neuron

As with other activation functions, first calculate the weighted sum:

z=∑i=1nwixi+bz=\sum_{i=1}^{n}w_ix_i+b

Then apply ReLU:

a=ReLU⁡(z)a=\operatorname{ReLU}(z)

Therefore,

a=max⁡(0,∑i=1nwixi+b)\boxed{ a=\max\left(0,\sum_{i=1}^{n}w_ix_i+b\right) }

Example

Suppose

x1=2,x2=3x_1=2,\qquad x_2=3 w1=0.5,w2=0.2w_1=0.5,\qquad w_2=0.2

and

b=−2b=-2

First calculate zz:

z=(0.5)(2)+(0.2)(3)−2z=(0.5)(2)+(0.2)(3)-2 z=1+0.6−2z=1+0.6-2 z=−0.4z=-0.4

Now apply ReLU:

a=ReLU⁡(−0.4)a=\operatorname{ReLU}(-0.4) a=0\boxed{a=0}

3. What happens for positive and negative inputs?

The easiest way to understand ReLU is:

๐Ÿ”ด Negative input

If

z<0z<0

then

ReLU⁡(z)=0\operatorname{ReLU}(z)=0

The neuron produces zero output.

๐ŸŸข Positive input

If

z>0z>0

then

ReLU⁡(z)=z\operatorname{ReLU}(z)=z

The neuron passes the value directly to the next layer.

Therefore:

Negative→0Positive→unchanged\boxed{ \text{Negative}\rightarrow0 \qquad \text{Positive}\rightarrow\text{unchanged} }

4. Graph of ReLU

The graph has two very simple regions:

y=0for x<0y=0 \quad \text{for }x<0

and

y=xfor x≥0y=x \quad \text{for }x\geq0

So it looks like this conceptually:



The important point is that ReLU is nonlinear, even though its positive portion is a straight line.


5. Derivative of ReLU

For backpropagation, we need the derivative.

For

ReLU⁡(z)={0,z<0z,z>0\operatorname{ReLU}(z)= \begin{cases} 0,&z<0\\ z,&z>0 \end{cases}

the derivative is

ReLU⁡′(z)={0,z<01,z>0\boxed{ \operatorname{ReLU}'(z)= \begin{cases} 0,&z<0\\ 1,&z>0 \end{cases} }

At z=0z=0, ReLU is not differentiable in the ordinary sense because the slope changes abruptly from 00 to 11.

In practical neural-network implementations, a subgradient convention is chosen, commonly either 00 or 11.


6. Why is ReLU useful for neural networks?

The most important reason is that ReLU introduces nonlinearity.

Consider:

z=w1x1+w2x2+bz=w_1x_1+w_2x_2+b

Without an activation function, a neuron simply performs a linear transformation.

ReLU gives:

a=max⁡(0,w1x1+w2x2+b)a=\max(0,w_1x_1+w_2x_2+b)

This allows a neural network with multiple layers to model nonlinear relationships.

For example, a multilayer network can learn complex patterns such as:

  • image patterns
  • handwritten digits
  • speech patterns
  • object features
  • nonlinear classification boundaries

7. Why is ReLU often preferred over sigmoid and tanh in hidden layers?

This is an important comparison for students.

Sigmoid

ฯƒ(z)=11+e−z\sigma(z)=\frac{1}{1+e^{-z}}

Its derivative is

ฯƒ′(z)=ฯƒ(z)(1−ฯƒ(z))\sigma'(z)=\sigma(z)(1-\sigma(z))

The maximum derivative is only

ฯƒ′(0)=0.25\sigma'(0)=0.25

For large positive or negative zz, the derivative becomes very small.


Tanh

tanh⁡(z)=ez−e−zez+e−z\tanh(z)=\frac{e^z-e^{-z}}{e^z+e^{-z}}

Its derivative is

tanh⁡′(z)=1−tanh⁡2(z)\tanh'(z)=1-\tanh^2(z)

Its maximum derivative is

tanh⁡′(0)=1\tanh'(0)=1

but it also suffers from saturation for large ∣z∣|z|.


ReLU

ReLU⁡(z)=max⁡(0,z)\operatorname{ReLU}(z)=\max(0,z)

and

ReLU⁡′(z)={0,z<01,z>0\operatorname{ReLU}'(z)= \begin{cases} 0,&z<0\\ 1,&z>0 \end{cases}

For positive zz, the derivative is exactly 1.

Therefore, positive activations do not suffer from the same saturation problem as sigmoid and tanh.


8. ReLU and the vanishing-gradient problem

This is one of the major advantages of ReLU.

For sigmoid:

ฯƒ′(z)→0\sigma'(z)\rightarrow0

when zz becomes very large or very negative.

For tanh:

tanh⁡′(z)→0\tanh'(z)\rightarrow0

when

∣z∣→∞|z|\rightarrow\infty

But for ReLU:

ReLU⁡′(z)=1\operatorname{ReLU}'(z)=1

for all positive zz.

Therefore, when the neuron is active:

gradient can pass through without being multiplied by a small activation derivative\boxed{\text{gradient can pass through without being multiplied by a small activation derivative}}

This was an important reason ReLU became popular in deep neural networks.


9. The "dying ReLU" problem

ReLU has its own important limitation.

For negative inputs:

ReLU⁡(z)=0\operatorname{ReLU}(z)=0

and

ReLU⁡′(z)=0\operatorname{ReLU}'(z)=0

Suppose a neuron consistently receives negative zz values.

Then:

a=0a=0

and

∂a∂z=0\frac{\partial a}{\partial z}=0

Consequently, the neuron may stop receiving useful gradient updates and effectively become inactive.

This is called the:

Dying ReLU problem\boxed{\text{Dying ReLU problem}}

10. Leaky ReLU: a solution to dying ReLU

One solution is Leaky ReLU.

It is defined as:

LeakyReLU⁡(z)={ฮฑz,z<0z,z≥0\boxed{ \operatorname{LeakyReLU}(z)= \begin{cases} \alpha z,&z<0\\ z,&z\geq0 \end{cases} }

where ฮฑ\alpha is a small positive number, such as

ฮฑ=0.01\alpha=0.01

For example:

LeakyReLU⁡(−5)=0.01(−5)=−0.05\operatorname{LeakyReLU}(-5)=0.01(-5)=-0.05

instead of

ReLU⁡(−5)=0\operatorname{ReLU}(-5)=0

Thus, Leaky ReLU allows a small gradient for negative inputs.




11. Why ReLU is so popular

The main advantages are:

✅ 1. Very simple

ReLU⁡(z)=max⁡(0,z)\operatorname{ReLU}(z)=\max(0,z)

There is no exponential calculation.

✅ 2. Computationally efficient

It is much simpler to calculate than sigmoid or tanh.

✅ 3. Introduces nonlinearity

It allows neural networks to learn nonlinear functions.

✅ 4. Better gradient behavior for positive inputs

ReLU⁡′(z)=1\operatorname{ReLU}'(z)=1

for z>0z>0.

✅ 5. Sparse activation

Negative inputs produce zero:

z<0⇒a=0z<0\Rightarrow a=0

Therefore, many neurons may be inactive for a particular input. This creates sparse activations.


12. Where is ReLU used?

ReLU is particularly common in the hidden layers of neural networks.

For example:

Input→Linear→ReLU→Linear→ReLU→Output\boxed{ \text{Input} \rightarrow \text{Linear} \rightarrow \text{ReLU} \rightarrow \text{Linear} \rightarrow \text{ReLU} \rightarrow \text{Output} }

It is widely used in architectures for:

  • ๐Ÿ–ผ️ Image classification
  • ๐Ÿ‘️ Computer vision
  • ๐Ÿ”Š Speech processing
  • ๐Ÿ“ Natural language processing
  • ๐Ÿค– Deep neural networks
  • ๐Ÿง  Convolutional neural networks

Modern networks also frequently use ReLU variants such as Leaky ReLU, PReLU, ELU, and GELU depending on the architecture.


๐ŸŽ“ Summary


"ReLU is a gate. If the input is negative, it closes the gate and produces 0. If the input is positive, it opens the gate and passes the input unchanged."

Mathematically:

z<0⇒0z≥0⇒z\boxed{ z<0\Rightarrow0 \qquad z\geq0\Rightarrow z }

๐Ÿง  Memory trick

ReLU = Keep Positive, Remove Negative\boxed{\text{ReLU = Keep Positive, Remove Negative}}

or simply:

ReLU⁡(z)=max⁡(0,z)\boxed{\operatorname{ReLU}(z)=\max(0,z)}

This one equation captures the entire basic idea of ReLU.


ReLU (Rectified Linear Unit) is a widely used nonlinear activation function in neural networks, especially in hidden layers. It is defined as ReLU⁡(z)=max⁡(0,z)\operatorname{ReLU}(z)=\max(0,z), meaning negative inputs are converted to zero while positive inputs are passed unchanged. Its derivative is 00 for negative inputs and 11 for positive inputs, which helps reduce the vanishing-gradient problem seen with sigmoid and tanh. ReLU is simple, computationally efficient, and introduces the nonlinearity needed to learn complex patterns. However, neurons can become permanently inactive for negative inputs, known as the dying ReLU problem.

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming