Introdcution to Activation Functions in Neural Networks

 

Activation Functions in Neural Networks

An activation function is a mathematical function applied to the weighted sum of a neuron. It determines the output (activation) of that neuron and, most importantly, allows a neural network to learn non-linear relationships.

A neuron first calculates a weighted sum:

The activation function then transforms this value:

where:

  • π‘₯𝑖 = input
  • 𝑀𝑖 = weight
  • 𝑏 = bias
  • 𝑧 = weighted input
  • πœ™ = activation function
  • π‘Ž = output/activation of the neuron

1. Why Do We Need Activation Functions?

The most important purpose of an activation function is to introduce non-linearity into a neural network.

Suppose we construct a network with several layers but use only linear functions.

For example:

β„Ž=π‘Š1π‘₯+𝑏1

and

𝑦=π‘Š2β„Ž+𝑏2

Substituting β„Ž:

𝑦=π‘Š2(π‘Š1π‘₯+𝑏1)+𝑏2

which can be rewritten as:

𝑦=π‘Šπ‘₯+𝑏

So, even though we have multiple layers, the entire network is still equivalent to one linear transformation.

Therefore:

Stacking linear layers without nonlinear activation functions does not give the network the ability to learn nonlinear functions.

This is why activation functions are essential in hidden layers.


2. An Intuitive Example

Consider the XOR problem we discussed earlier.

The four XOR points are:

π‘₯1π‘₯2XOR
000
011
101
110

A single linear decision boundary cannot separate these points.

A neural network needs to create nonlinear decision regions.

Activation functions allow hidden neurons to transform the original input into a new representation in which the problem can become separable.


Weights determine what information is learned, while activation functions determine how the learned information is transformed.


3. Activation Function and Biological Neurons

The biological analogy is useful, but it should not be taken literally.

A biological neuron receives signals through dendrites, integrates them in the cell body, and may produce an output signal when sufficient stimulation is received.

Similarly, an artificial neuron:

Inputs
   ↓
Weighted Sum
   ↓
Add Bias
   ↓
Activation Function
   ↓
Output

The activation function is therefore a mathematical abstraction of the neuron's response.


4. Activation Functions and Learning

During training, a neural network adjusts its weights and biases to reduce prediction error.

For gradient-based learning, we need to know:

How much does the output change when a weight changes?

This is obtained using derivatives/gradients.

For example:

∂π‘Ž∂𝑧

and these gradients are propagated backward through the network during backpropagation.

Therefore, activation functions need to be differentiable, or at least have useful derivatives/subgradients for the optimization method being used.



5. Two Major Categories

Activation functions can broadly be divided into:

                 Activation Functions
                         |
              ┌──────────┴──────────┐
              ↓                     ↓
           Linear               Non-linear
              |                     |
        Mainly output          Mainly hidden
           layers                 layers
                                  |
                    ┌─────────────┼─────────────┐
                    ↓             ↓             ↓
                  Sigmoid        Tanh          ReLU

6. Linear Activation Function

The simplest linear activation function is:

𝑓(π‘₯)=π‘Žπ‘₯+𝑏

A commonly used simplified form is:

𝑓(π‘₯)=π‘₯

Graph

The output can take any value from:

−∞ to +∞

Characteristics

PropertyLinear Activation
Function   π‘“(π‘₯)=π‘Žπ‘₯+𝑏
Output range  (−∞,+∞)
Non-linearity  No
Derivative  Constant
Mainly used  Output layer
Hidden layers  Generally not useful

7. Where Is Linear Activation Useful?

Linear activation is particularly useful in the output layer of regression networks.

Example

Suppose we want to predict:

  • house price
  • temperature
  • salary
  • sales revenue
  • exam marks

These values are continuous and may lie over a wide range.

We therefore often use:

𝑦=𝑓(𝑧)=𝑧

at the output.

For example:

Input → Hidden Layer → Hidden Layer → Output
                              ↓
                       Linear activation
                              ↓
                       House Price = ₹52.5 lakh

9. Non-Linear Activation Functions

Nonlinear activation functions are the key to the expressive power of neural networks.

Common examples include:

  1. Sigmoid
  2. Tanh
  3. ReLU
  4. Leaky ReLU
  5. ELU
  6. Softmax — commonly used for multiclass output

9.1 Sigmoid Function

The sigmoid function is:

𝑓(π‘₯)=11+𝑒−π‘₯

Its output lies between:

0<𝑓(π‘₯)<1

Characteristics

  • Smooth and differentiable
  • Output between 0 and 1
  • Useful when an output can be interpreted as a probability
  • Historically important in neural networks

Limitation

For very large positive or negative values, the gradient becomes very small. This is associated with the vanishing-gradient problem.


10. Tanh Activation

The hyperbolic tangent function is:

𝑓(π‘₯)=tanh⁡(π‘₯)

Its output lies between:

−1 and +1

Compared with sigmoid, tanh is zero-centered, which can be useful during optimization.

However, it can also suffer from vanishing gradients when its input becomes very large or very negative.


11. ReLU Activation

The Rectified Linear Unit (ReLU) is:

𝑓(π‘₯)=max⁡(0,π‘₯)

It gives:

  • 0 for negative inputs
  • π‘₯ for positive inputs

Why is ReLU important?

ReLU is widely used in hidden layers because it is:

  • computationally simple
  • easy to differentiate
  • effective for deep networks
  • less prone to vanishing gradients in its positive region than sigmoid/tanh

Problem: Dying ReLU

If a neuron consistently receives negative inputs, its output can remain zero and its gradient can also be zero.

This led to alternatives such as Leaky ReLU.


12. Leaky ReLU

Leaky ReLU allows a small negative output instead of completely becoming zero:

𝑓(π‘₯)={π‘₯π‘₯>0𝛼π‘₯π‘₯≤0

where 𝛼 is a small positive value such as 0.01.

The idea is:

Don't completely shut down the neuron for negative inputs.


13. Softmax Activation

Softmax is particularly useful in the output layer of multiclass classification networks.

Suppose we have three classes:

              Output Layer
              ┌───────────┐
              │   Cat     │
              │   Dog     │
              │   Horse   │
              └───────────┘

Softmax converts the output scores into values that sum to 1.

For class 𝑖:

𝑃𝑖=𝑒𝑧𝑖∑𝑗𝑒𝑧𝑗

For example:

Class    Probability
Cat    0.10
Dog    0.75
Horse    0.15

The predicted class would be Dog.


14. Choosing an Activation Function

A useful undergraduate summary is:

ProblemTypical Output Activation
Regression    Linear
Binary classification    Sigmoid
Multiclass classification    Softmax
Hidden layers in many modern networks    ReLU / variants
Older neural networks / some specialized models    Tanh / Sigmoid

The exact choice can depend on the architecture, loss function, and optimization method.


15.  Summary 

Linear Activation

Linear activation does not introduce non-linearity. Therefore, it is generally used in the output layer for regression problems.

Nonlinear Activation

Nonlinear activation functions allow neural networks to learn complex and nonlinear relationships. They are therefore essential in hidden layers.


⭐ The Most Important Concept

Students should remember this sequence:

Input→Weighted Sum→Bias→Activation→Output

And the central idea is:

Without nonlinear activation functions, a multilayer neural network is still just a linear model, regardless of how many linear layers it contains.

This is the fundamental reason activation functions are so important in neural networks and deep learning.

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming