Sigmoid / Logistic Activation Function
🟢 Sigmoid / Logistic Activation Function
1. 🧠 Introduction
The sigmoid function, also called the logistic function, is one of the classical nonlinear activation functions used in artificial neural networks.
It is particularly useful for binary classification, because it converts any real-valued input into a value between 0 and 1.
The sigmoid function is defined as:
where:
- = input to the activation function
- = Euler's number, approximately
- = output of the sigmoid function
⚙️ 2. Sigmoid in an Artificial Neuron
Before applying the activation function, a neuron calculates the weighted sum of its inputs and adds a bias.
For inputs:
or, in expanded form:
The sigmoid function is then applied to :
Therefore:
This is the complete computation performed by a sigmoid neuron.
📈 3. Shape of the Sigmoid Function
The sigmoid function has an S-shaped curve.
Conceptually:
There are three important regions:
🔹 When is very negative
then:
🔹 When
Since :
🔹 When is very positive
then:
Therefore:
🔢 4. Numerical Examples
Let's calculate the sigmoid for different values of .
Example 1:
Example 2:
Example 3:
Example 4:
Example 5:
Summary
| -5 | 0.007 |
| -2 | 0.119 |
| -1 | 0.269 |
| 0 | 0.500 |
| 1 | 0.731 |
| 2 | 0.881 |
| 5 | 0.993 |
This table clearly shows that sigmoid compresses any real-valued input into the interval (0,1).
🎯 5. Why Is Sigmoid Useful for Classification?
Suppose a neural network is predicting whether an email is:
- 0 → Not Spam
- 1 → Spam
The neuron calculates:
Suppose:
Then:
The output is approximately 0.90.
This can be interpreted as an estimated probability-like output for the positive class:
So the model gives a high estimated probability to the positive class.
🚦 6. Converting Sigmoid Output into a Class
The sigmoid output itself is continuous.
For example:
To convert this into a binary class, we normally choose a threshold.
A common threshold is:
The classification rule is:
Because:
the same decision rule can also be written as:
🧮 7. Complete Example of a Sigmoid Neuron
Suppose we have two inputs:
Weights:
Bias:
Step 1: Calculate weighted sum
Step 2: Apply sigmoid
Since:
the predicted class is:
🔄 8. Sigmoid and Logistic Regression
An important connection for students is that logistic regression uses exactly this sigmoid transformation.
First calculate:
Then:
Substituting :
Therefore, a single neuron with a sigmoid activation function has essentially the same mathematical form as logistic regression for binary classification.
This is a very useful connection between classical machine learning and neural networks.
📐 9. Derivative of the Sigmoid Function
The derivative of sigmoid is particularly convenient.
Starting with:
the derivative is:
This can also be written as:
This derivative is used during backpropagation to calculate gradients.
🔄 10. Derivation of the Sigmoid Derivative
Starting with:
Differentiate:
Therefore:
Now observe:
and:
Therefore:
Hence:
📊 11. Maximum Value of the Gradient
The derivative is:
Let:
Then:
This is maximum when:
Therefore:
Thus, the derivative of sigmoid can never be greater than 0.25.
⚠️ 12. Vanishing Gradient Problem
This is the major disadvantage of sigmoid.
Consider a very large positive input:
Then:
The derivative becomes:
This is extremely small.
Similarly, for:
we obtain a very small derivative.
Therefore:
🧠 13. Why Does This Cause a Problem?
Backpropagation uses gradients to update weights.
The basic gradient-descent update is:
where:
- = loss
- = learning rate
- = gradient
If the gradient becomes very small:
then:
Therefore, the weights change very slowly.
This is called the:
🌡️ 14. Saturation
Sigmoid saturates at both ends.
For large negative values:
For large positive values:
The curve becomes almost flat in these regions.
Output 1 | ───────── | ____/ | __/ 0.5 |---------------● | __/ | __/ 0 |──────── +──────────────────────────→ z
A flat curve means a small gradient.
Therefore:
🎯 15. Applications of Sigmoid
📨 15.1 Binary Classification
The most common application is binary classification.
Examples:
- Spam / Not Spam
- Fraud / Not Fraud
- Pass / Fail
- Disease / No Disease
- Defective / Non-defective
- Churn / No Churn
✅ 16. Advantages of Sigmoid
🟢 1. Output is Between 0 and 1
This makes it naturally suitable for probability-like outputs.
🟢 2. Smooth Function
Sigmoid is continuous and differentiable everywhere.
🟢 3. Simple Derivative
🟢 4. Excellent for Binary Classification
It provides a natural single-output representation for a binary target.
🟢 5. Historically Important
Sigmoid was extensively used in traditional neural networks and remains appropriate for many binary-output applications.
❌ 17. Disadvantages of Sigmoid
🔴 1. Vanishing Gradient
This can make deep networks learn slowly.
🔴 2. Saturation
The output becomes almost constant near 0 and 1.
🔴 3. Not Zero-Centered
The output satisfies:
so it never produces negative activation values.
🔴 4. Maximum Gradient Is Only 0.25
This can contribute to shrinking gradients across multiple layers.
🔴 5. More Computationally Expensive Than ReLU
Sigmoid requires an exponential operation:
whereas ReLU is simply:
⚖️ 18. Sigmoid vs Linear Activation
| Property | Linear | Sigmoid |
|---|---|---|
| Formula | ||
| Shape | Straight line | S-shaped |
| Output range | ||
| Nonlinear | ❌ No | ✅ Yes |
| Binary classification | ❌ Generally unsuitable | ✅ Suitable |
| Regression output | ✅ Common | ❌ Generally unsuitable |
| Derivative | ||
| Saturation | ❌ No | ✅ Yes |
| Vanishing gradient | Not from sigmoid saturation | ⚠️ Yes |
| Zero-centered | Yes for | ❌ No |
🏆 19. Summary
Students can remember sigmoid using the following 5 points:
🟢 1. Formula
🟢 2. Range
🟢 3. Shape
🟢 4. Main Application
🔴 5. Main Limitation
🎓
The sigmoid or logistic activation function is a nonlinear function that maps any real-valued input to a value between 0 and 1. It is widely used in the output layer of binary classification neural networks because its output can be interpreted as a probability-like value. Its derivative is , which makes it suitable for gradient-based learning. However, sigmoid suffers from saturation, is not zero-centered, and can cause vanishing gradients, making it generally unsuitable for the hidden layers of modern deep neural networks.
Comments
Post a Comment