Sigmoid / Logistic Activation Function

 

🟢 Sigmoid / Logistic Activation Function

1. 🧠 Introduction

The sigmoid function, also called the logistic function, is one of the classical nonlinear activation functions used in artificial neural networks.

It is particularly useful for binary classification, because it converts any real-valued input into a value between 0 and 1.

The sigmoid function is defined as:

σ(z)=11+e−z\boxed{\sigma(z)=\frac{1}{1+e^{-z}}}

where:

  • zz = input to the activation function
  • ee = Euler's number, approximately 2.718282.71828
  • σ(z)\sigma(z) = output of the sigmoid function




⚙️ 2. Sigmoid in an Artificial Neuron

Before applying the activation function, a neuron calculates the weighted sum of its inputs and adds a bias.

For nn inputs:

z=∑i=1nwixi+b\boxed{z=\sum_{i=1}^{n}w_i x_i+b}

or, in expanded form:

z=w1x1+w2x2+⋯+wnxn+bz=w_1x_1+w_2x_2+\cdots+w_nx_n+b

The sigmoid function is then applied to zz:

y=σ(z)\boxed{y=\sigma(z)}

Therefore:

y=11+e−(∑i=1nwixi+b)\boxed{ y=\frac{1}{1+e^{-\left(\sum_{i=1}^{n}w_ix_i+b\right)}} }

This is the complete computation performed by a sigmoid neuron.


📈 3. Shape of the Sigmoid Function

The sigmoid function has an S-shaped curve.

σ(z)=11+e−z\boxed{\sigma(z)=\frac{1}{1+e^{-z}}}

Conceptually:

    
   

There are three important regions:

🔹 When zz is very negative

z→−∞z\rightarrow-\infty

then:

σ(z)→0\boxed{\sigma(z)\rightarrow0}

🔹 When z=0z=0

σ(0)=11+e0\sigma(0)=\frac{1}{1+e^0}

Since e0=1e^0=1:

σ(0)=0.5\boxed{\sigma(0)=0.5}

🔹 When zz is very positive

z→+∞z\rightarrow+\infty

then:

σ(z)→1\boxed{\sigma(z)\rightarrow1}

Therefore:

0<σ(z)<1\boxed{0<\sigma(z)<1}

🔢 4. Numerical Examples

Let's calculate the sigmoid for different values of zz.

Example 1: z=0z=0

σ(0)=11+e0\sigma(0)=\frac{1}{1+e^0} =11+1=\frac{1}{1+1} σ(0)=0.5\boxed{\sigma(0)=0.5}

Example 2: z=1z=1

σ(1)=11+e−1\sigma(1)=\frac{1}{1+e^{-1}} σ(1)≈0.731\boxed{\sigma(1)\approx0.731}

Example 3: z=−1z=-1

σ(−1)=11+e1\sigma(-1)=\frac{1}{1+e^1} σ(−1)≈0.269\boxed{\sigma(-1)\approx0.269}

Example 4: z=5z=5

σ(5)=11+e−5\sigma(5)=\frac{1}{1+e^{-5}} σ(5)≈0.993\boxed{\sigma(5)\approx0.993}

Example 5: z=−5z=-5

σ(−5)=11+e5\sigma(-5)=\frac{1}{1+e^5} σ(−5)≈0.007\boxed{\sigma(-5)\approx0.007}

Summary

zzσ(z)\sigma(z)
-5        0.007
-20.119
-10.269
00.500
10.731
20.881
50.993

This table clearly shows that sigmoid compresses any real-valued input into the interval (0,1).


🎯 5. Why Is Sigmoid Useful for Classification?

Suppose a neural network is predicting whether an email is:

  • 0 → Not Spam
  • 1 → Spam

The neuron calculates:

z=wTx+bz=w^Tx+b

Suppose:

z=2.2z=2.2

Then:

σ(2.2)=11+e−2.2\sigma(2.2)=\frac{1}{1+e^{-2.2}} σ(2.2)≈0.900\boxed{\sigma(2.2)\approx0.900}

The output is approximately 0.90.

This can be interpreted as an estimated probability-like output for the positive class:

P(y=1∣x)≈0.90\boxed{P(y=1|x)\approx0.90}

So the model gives a high estimated probability to the positive class.


🚦 6. Converting Sigmoid Output into a Class

The sigmoid output itself is continuous.

For example:

0.12,0.37,0.51,0.82,0.970.12,\quad0.37,\quad0.51,\quad0.82,\quad0.97

To convert this into a binary class, we normally choose a threshold.

A common threshold is:

T=0.5\boxed{T=0.5}

The classification rule is:

y^={1if σ(z)≥0.50if σ(z)<0.5\boxed{ \hat y= \begin{cases} 1 & \text{if }\sigma(z)\geq0.5\\ 0 & \text{if }\sigma(z)<0.5 \end{cases} }

Because:

σ(0)=0.5\sigma(0)=0.5

the same decision rule can also be written as:

y^={1z≥00z<0\boxed{ \hat y= \begin{cases} 1 & z\geq0\\ 0 & z<0 \end{cases} }

🧮 7. Complete Example of a Sigmoid Neuron

Suppose we have two inputs:

x1=2,x2=3x_1=2,\qquad x_2=3

Weights:

w1=0.8,w2=−0.4w_1=0.8,\qquad w_2=-0.4

Bias:

b=−0.5b=-0.5

Step 1: Calculate weighted sum

z=w1x1+w2x2+bz=w_1x_1+w_2x_2+b

z=(0.8)(2)+(−0.4)(3)−0.5
z=(0.8)(2)+(-0.4)(3)-0.5
z=1.6−1.2−0.5z=1.6-1.2-0.5 z=−0.1\boxed{z=-0.1}

Step 2: Apply sigmoid

y=11+e−(−0.1)y=\frac{1}{1+e^{-(-0.1)}} y=11+e0.1y=\frac{1}{1+e^{0.1}} y≈0.475\boxed{y\approx0.475}

Since:

0.475<0.50.475<0.5

the predicted class is:

y^=0\boxed{\hat y=0}

🔄 8. Sigmoid and Logistic Regression

An important connection for students is that logistic regression uses exactly this sigmoid transformation.

First calculate:

z=wTx+b\boxed{z=w^Tx+b}

Then:

P(y=1∣x)=11+e−z\boxed{ P(y=1|x)=\frac{1}{1+e^{-z}} }

Substituting zz:

P(y=1∣x)=11+e−(wTx+b)\boxed{ P(y=1|x)= \frac{1}{1+e^{-(w^Tx+b)}} }

Therefore, a single neuron with a sigmoid activation function has essentially the same mathematical form as logistic regression for binary classification.

This is a very useful connection between classical machine learning and neural networks.


📐 9. Derivative of the Sigmoid Function

The derivative of sigmoid is particularly convenient.

Starting with:

σ(z)=11+e−z\sigma(z)=\frac{1}{1+e^{-z}}

the derivative is:

dσ(z)dz=σ(z)(1−σ(z))\boxed{ \frac{d\sigma(z)}{dz} = \sigma(z)(1-\sigma(z)) }

This can also be written as:

σ′(z)=σ(z)[1−σ(z)]\boxed{ \sigma'(z)=\sigma(z)[1-\sigma(z)] }

This derivative is used during backpropagation to calculate gradients.


🔄 10. Derivation of the Sigmoid Derivative

Starting with:

σ(z)=(1+e−z)−1\sigma(z)=(1+e^{-z})^{-1}

Differentiate:

σ′(z)=−(1+e−z)−2(−e−z)\sigma'(z) = -(1+e^{-z})^{-2}(-e^{-z})

Therefore:

σ′(z)=e−z(1+e−z)2\sigma'(z)=\frac{e^{-z}}{(1+e^{-z})^2}

Now observe:

σ(z)=11+e−z\sigma(z)=\frac{1}{1+e^{-z}}

and:

1−σ(z)=1−11+e−z1-\sigma(z) = 1-\frac{1}{1+e^{-z}} =e−z1+e−z=\frac{e^{-z}}{1+e^{-z}}

Therefore:

σ(z)(1−σ(z))=11+e−ze−z1+e−z\sigma(z)(1-\sigma(z)) = \frac{1}{1+e^{-z}} \frac{e^{-z}}{1+e^{-z}} =e−z(1+e−z)2=\frac{e^{-z}}{(1+e^{-z})^2}

Hence:

σ′(z)=σ(z)(1−σ(z))\boxed{\sigma'(z)=\sigma(z)(1-\sigma(z))}

📊 11. Maximum Value of the Gradient

The derivative is:

σ′(z)=σ(z)(1−σ(z))\sigma'(z)=\sigma(z)(1-\sigma(z))

Let:

σ(z)=p\sigma(z)=p

Then:

σ′(z)=p(1−p)\sigma'(z)=p(1-p)

This is maximum when:

p=0.5p=0.5

Therefore:

σ′(z)=0.5(1−0.5)\sigma'(z)=0.5(1-0.5) =0.5×0.5=0.5\times0.5 σ′(z)max⁡=0.25\boxed{\sigma'(z)_{\max}=0.25}

Thus, the derivative of sigmoid can never be greater than 0.25.





⚠️ 12. Vanishing Gradient Problem

This is the major disadvantage of sigmoid.

Consider a very large positive input:

z=10z=10

Then:

σ(10)≈0.99995\sigma(10)\approx0.99995

The derivative becomes:

σ′(10)=0.99995(1−0.99995)\sigma'(10) = 0.99995(1-0.99995) σ′(10)≈0.00005\boxed{\sigma'(10)\approx0.00005}

This is extremely small.

Similarly, for:

z=−10z=-10

we obtain a very small derivative.

Therefore:

σ′(z)→0whenz→±∞\boxed{\sigma'(z)\rightarrow0 \quad\text{when}\quad z\rightarrow\pm\infty}

🧠 13. Why Does This Cause a Problem?

Backpropagation uses gradients to update weights.

The basic gradient-descent update is:

wnew=wold−η∂L∂w\boxed{ w_{\text{new}} = w_{\text{old}} - \eta \frac{\partial L}{\partial w} }

where:

  • LL = loss
  • η\eta = learning rate
  • ∂L∂w\frac{\partial L}{\partial w} = gradient

If the gradient becomes very small:

∂L∂w≈0\frac{\partial L}{\partial w}\approx0

then:

wnew≈woldw_{\text{new}}\approx w_{\text{old}}

Therefore, the weights change very slowly.

This is called the:

Vanishing Gradient Problem\boxed{\text{Vanishing Gradient Problem}}

🌡️ 14. Saturation

Sigmoid saturates at both ends.

For large negative values:

z→−∞z\rightarrow-\infty σ(z)→0\sigma(z)\rightarrow0

For large positive values:

z→+∞z\rightarrow+\infty σ(z)→1\sigma(z)\rightarrow1

The curve becomes almost flat in these regions.

Output
  1 |                         ─────────
    |                    ____/
    |                 __/
0.5 |---------------●
    |            __/
    |         __/
  0 |────────
    +──────────────────────────→ z

A flat curve means a small gradient.

Therefore:

Saturation→small gradient→slow learning\boxed{\text{Saturation}\rightarrow\text{small gradient}\rightarrow \text{slow learning}}

🎯 15. Applications of Sigmoid

📨 15.1 Binary Classification

The most common application is binary classification.

Examples:

  • Spam / Not Spam
  • Fraud / Not Fraud
  • Pass / Fail
  • Disease / No Disease
  • Defective / Non-defective
  • Churn / No Churn

✅ 16. Advantages of Sigmoid

🟢 1. Output is Between 0 and 1

0<σ(z)<1\boxed{0<\sigma(z)<1}

This makes it naturally suitable for probability-like outputs.

🟢 2. Smooth Function

Sigmoid is continuous and differentiable everywhere.

🟢 3. Simple Derivative

σ′(z)=σ(z)(1−σ(z))\boxed{\sigma'(z)=\sigma(z)(1-\sigma(z))}

🟢 4. Excellent for Binary Classification

It provides a natural single-output representation for a binary target.

🟢 5. Historically Important

Sigmoid was extensively used in traditional neural networks and remains appropriate for many binary-output applications.


❌ 17. Disadvantages of Sigmoid

🔴 1. Vanishing Gradient

σ′(z)→0for large ∣z∣\boxed{\sigma'(z)\rightarrow0 \quad\text{for large }|z|}

This can make deep networks learn slowly.

🔴 2. Saturation

The output becomes almost constant near 0 and 1.

🔴 3. Not Zero-Centered

The output satisfies:

0<σ(z)<1\boxed{0<\sigma(z)<1}

so it never produces negative activation values.

🔴 4. Maximum Gradient Is Only 0.25

σ′(z)≤0.25\boxed{\sigma'(z)\leq0.25}

This can contribute to shrinking gradients across multiple layers.

🔴 5. More Computationally Expensive Than ReLU

Sigmoid requires an exponential operation:

e−ze^{-z}

whereas ReLU is simply:

ReLU⁡(z)=max⁡(0,z)\boxed{\operatorname{ReLU}(z)=\max(0,z)}

⚖️ 18. Sigmoid vs Linear Activation

Property LinearSigmoid
Formula                           f(z)=z                   f(z)=zf(z)=11+e−z\displaystyle f(z)=\frac{1}{1+e^{-z}}
ShapeStraight lineS-shaped
Output range(−∞,+∞)(-\infty,+\infty)(0,1)(0,1)
Nonlinear❌ No✅ Yes
Binary classification❌ Generally unsuitable✅ Suitable
Regression output✅ Common❌ Generally unsuitable
Derivative11σ(z)(1−σ(z))\sigma(z)(1-\sigma(z))
Saturation❌ No✅ Yes
Vanishing gradientNot from sigmoid saturation⚠️ Yes
Zero-centeredYes for f(z)=zf(z)=z❌ No

🏆 19. Summary

Students can remember sigmoid using the following 5 points:

🟢 1. Formula

σ(z)=11+e−z\boxed{\sigma(z)=\frac{1}{1+e^{-z}}}

🟢 2. Range

0<σ(z)<1\boxed{0<\sigma(z)<1}

🟢 3. Shape

S-shaped\boxed{\text{S-shaped}}

🟢 4. Main Application

Binary Classification\boxed{\text{Binary Classification}}

🔴 5. Main Limitation

Vanishing Gradient\boxed{\text{Vanishing Gradient}}

🎓

The sigmoid or logistic activation function is a nonlinear function that maps any real-valued input to a value between 0 and 1. It is widely used in the output layer of binary classification neural networks because its output can be interpreted as a probability-like value. Its derivative is σ′(z)=σ(z)(1−σ(z))\sigma'(z)=\sigma(z)(1-\sigma(z)), which makes it suitable for gradient-based learning. However, sigmoid suffers from saturation, is not zero-centered, and can cause vanishing gradients, making it generally unsuitable for the hidden layers of modern deep neural networks.

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming