Introdcution to Activation Functions in Neural Networks
Activation Functions in Neural Networks
An activation function is a mathematical function applied to the weighted sum of a neuron. It determines the output (activation) of that neuron and, most importantly, allows a neural network to learn non-linear relationships.
A neuron first calculates a weighted sum:
The activation function then transforms this value:
where:
- = input
- = weight
- = bias
- = weighted input
- = activation function
- = output/activation of the neuron
1. Why Do We Need Activation Functions?
The most important purpose of an activation function is to introduce non-linearity into a neural network.
Suppose we construct a network with several layers but use only linear functions.
For example:
and
Substituting :
which can be rewritten as:
So, even though we have multiple layers, the entire network is still equivalent to one linear transformation.
Therefore:
Stacking linear layers without nonlinear activation functions does not give the network the ability to learn nonlinear functions.
This is why activation functions are essential in hidden layers.
2. An Intuitive Example
Consider the XOR problem we discussed earlier.
The four XOR points are:
| XOR | ||
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
A single linear decision boundary cannot separate these points.
A neural network needs to create nonlinear decision regions.
Activation functions allow hidden neurons to transform the original input into a new representation in which the problem can become separable.
Weights determine what information is learned, while activation functions determine how the learned information is transformed.
3. Activation Function and Biological Neurons
The biological analogy is useful, but it should not be taken literally.
A biological neuron receives signals through dendrites, integrates them in the cell body, and may produce an output signal when sufficient stimulation is received.
Similarly, an artificial neuron:
Inputs ↓ Weighted Sum ↓ Add Bias ↓ Activation Function ↓ Output
The activation function is therefore a mathematical abstraction of the neuron's response.
4. Activation Functions and Learning
During training, a neural network adjusts its weights and biases to reduce prediction error.
For gradient-based learning, we need to know:
How much does the output change when a weight changes?
This is obtained using derivatives/gradients.
For example:
and these gradients are propagated backward through the network during backpropagation.
Therefore, activation functions need to be differentiable, or at least have useful derivatives/subgradients for the optimization method being used.
5. Two Major Categories
Activation functions can broadly be divided into:
Activation Functions | ┌──────────┴──────────┐ ↓ ↓ Linear Non-linear | | Mainly output Mainly hidden layers layers | ┌─────────────┼─────────────┐ ↓ ↓ ↓ Sigmoid Tanh ReLU
6. Linear Activation Function
The simplest linear activation function is:
A commonly used simplified form is:
Graph
The output can take any value from:
Characteristics
| Property | Linear Activation |
|---|---|
| Function | |
| Output range | |
| Non-linearity | No |
| Derivative | Constant |
| Mainly used | Output layer |
| Hidden layers | Generally not useful |
7. Where Is Linear Activation Useful?
Linear activation is particularly useful in the output layer of regression networks.
Example
Suppose we want to predict:
- house price
- temperature
- salary
- sales revenue
- exam marks
These values are continuous and may lie over a wide range.
We therefore often use:
at the output.
For example:
Input → Hidden Layer → Hidden Layer → Output ↓ Linear activation ↓ House Price = ₹52.5 lakh
9. Non-Linear Activation Functions
Nonlinear activation functions are the key to the expressive power of neural networks.
Common examples include:
- Sigmoid
- Tanh
- ReLU
- Leaky ReLU
- ELU
- Softmax — commonly used for multiclass output
9.1 Sigmoid Function
The sigmoid function is:
Its output lies between:
Characteristics
- Smooth and differentiable
- Output between 0 and 1
- Useful when an output can be interpreted as a probability
- Historically important in neural networks
Limitation
For very large positive or negative values, the gradient becomes very small. This is associated with the vanishing-gradient problem.
10. Tanh Activation
The hyperbolic tangent function is:
Its output lies between:
Compared with sigmoid, tanh is zero-centered, which can be useful during optimization.
However, it can also suffer from vanishing gradients when its input becomes very large or very negative.
11. ReLU Activation
The Rectified Linear Unit (ReLU) is:
It gives:
- for negative inputs
- for positive inputs
Why is ReLU important?
ReLU is widely used in hidden layers because it is:
- computationally simple
- easy to differentiate
- effective for deep networks
- less prone to vanishing gradients in its positive region than sigmoid/tanh
Problem: Dying ReLU
If a neuron consistently receives negative inputs, its output can remain zero and its gradient can also be zero.
This led to alternatives such as Leaky ReLU.
12. Leaky ReLU
Leaky ReLU allows a small negative output instead of completely becoming zero:
where is a small positive value such as 0.01.
The idea is:
Don't completely shut down the neuron for negative inputs.
13. Softmax Activation
Softmax is particularly useful in the output layer of multiclass classification networks.
Suppose we have three classes:
Output Layer ┌───────────┐ │ Cat │ │ Dog │ │ Horse │ └───────────┘
Softmax converts the output scores into values that sum to 1.
For class :
For example:
| Class | Probability |
|---|---|
| Cat | 0.10 |
| Dog | 0.75 |
| Horse | 0.15 |
The predicted class would be Dog.
14. Choosing an Activation Function
A useful undergraduate summary is:
| Problem | Typical Output Activation |
|---|---|
| Regression | Linear |
| Binary classification | Sigmoid |
| Multiclass classification | Softmax |
| Hidden layers in many modern networks | ReLU / variants |
| Older neural networks / some specialized models | Tanh / Sigmoid |
The exact choice can depend on the architecture, loss function, and optimization method.
15. Summary
Linear Activation
Linear activation does not introduce non-linearity. Therefore, it is generally used in the output layer for regression problems.
Nonlinear Activation
Nonlinear activation functions allow neural networks to learn complex and nonlinear relationships. They are therefore essential in hidden layers.
⭐ The Most Important Concept
Students should remember this sequence:
And the central idea is:
Without nonlinear activation functions, a multilayer neural network is still just a linear model, regardless of how many linear layers it contains.
This is the fundamental reason activation functions are so important in neural networks and deep learning.
Comments
Post a Comment