Tanh / Hyperbolic Tangent Activation Function
- Get link
- X
- Other Apps
🔵 Tanh / Hyperbolic Tangent Activation Function
The tanh (hyperbolic tangent) function is another important nonlinear activation function used in artificial neural networks.
It is closely related to the sigmoid function, but it has one important advantage: its output is centered around zero.
The tanh function is particularly useful for understanding the evolution from sigmoid-based neural networks to modern activation functions such as ReLU.
🧠 1. What Is the Tanh Function?
The hyperbolic tangent function is defined as:
It can also be expressed in terms of the sigmoid function:
where:
The output of tanh lies between:
−1<tanh(z)<1So, conceptually:
Tanh=scaled sigmoid−shift📈 2. Shape of the Tanh Function
The tanh function has an S-shaped curve, similar to sigmoid.
There are three important regions:
🔹 Large negative input
Then:
🔹 Zero input
Then:
🔹 Large positive input
Then:
Therefore:
🔢 3. Numerical Examples
Let's calculate some tanh values.
| -5 | -0.9999 |
| -2 | -0.9640 |
| -1 | -0.7616 |
| 0 | 0 |
| 1 | 0.7616 |
| 2 | 0.9640 |
| 5 | 0.9999 |
This shows that tanh compresses an arbitrary real-valued input into:
⚙️ 4. Tanh in an Artificial Neuron
Just like any other activation function, tanh is applied after calculating the weighted sum.
First:
Then:
Therefore, the complete neuron is:
🧮 5. Numerical Example
Suppose:
Weights:
Bias:
Step 1: Calculate weighted sum
Step 2: Apply tanh
Therefore, the neuron produces:
y≈0.834
🧠 6. Why Is Tanh Better Than Sigmoid in Some Situations?
This is an important comparison.
Sigmoid:
Tanh:
The major difference is that tanh is zero-centered.
For example:
Sigmoid: 0 ──────────────── 1 ↑ all outputs are positive Tanh: -1 ─────── 0 ─────── +1 ↑ zero-centered
🎯 7. Zero-Centered Property
Suppose the input is positive:
Then:
If the input is negative:
then:
And:
Therefore, tanh produces both positive and negative activations.
This can make optimization more favorable than sigmoid in some settings.
📊8. Sigmoid vs Tanh
| Property | Sigmoid | Tanh |
|---|---|---|
| Formula | ||
| Output range | ||
| Shape | S-shaped | S-shaped |
| Zero-centered | ❌ No | ✅ Yes |
| Nonlinear | ✅ Yes | ✅ Yes |
| Saturates | ✅ Yes | ✅ Yes |
| Vanishing gradients | ⚠️ Yes | ⚠️ Yes |
| Binary output | Very useful | Less direct |
| Historical hidden-layer use | Common | Common |
| Modern deep hidden layers | Less common | Less common |
🧮 9. Derivative of Tanh
The derivative of tanh has a particularly simple form.
Therefore:
This derivative is important during backpropagation.
🔍 10. Deriving the Tanh Derivative
Start with:
Using the quotient rule:
After simplifying:
This is equivalent to:
The second form is much more convenient when discussing neural-network learning.
📈 11. Maximum Gradient of Tanh
We have:
At:
we know:
Therefore:
Thus the maximum derivative of tanh is:
Compare this with sigmoid:
This is one reason tanh can have stronger gradients around zero than sigmoid.
⚠️ 12. Tanh and the Vanishing Gradient Problem
Although tanh has some advantages over sigmoid, it still suffers from vanishing gradients.
Consider a large positive value:
We have:
Therefore:
which is approximately:
The derivative is very small.
Similarly, for a large negative value:
and again:
Thus:
🌡️ 13. Saturation of Tanh
The tanh function saturates near and .
Output +1 |──────────────────── | ___ | __/ 0 |────────────●──────────── | __/ | __/ -1 |────── +────────────────────→ z
In the saturated regions:
Therefore:
So:
🔄 14. Tanh and Backpropagation
During backpropagation, gradients are propagated through the network.
The derivative of the activation function contributes to this gradient.
For tanh:
The weight update is generally:
where:
- = loss
- = learning rate
- = gradient
If tanh is saturated:
then the gradient can become very small.
🧩 15. Applications of Tanh
🧠 Hidden Layers of Classical Neural Networks
Tanh was historically a popular activation function for hidden layers.
It is still useful in certain architectures where zero-centered activations are desirable.
It is also used in RNN and LSTM🎛️ 16. Advantages of Tanh
🟢 1. Nonlinear
Tanh allows a neural network to learn complex nonlinear relationships.
🟢 2. Zero-Centered
Its output range is:
This is an important advantage over sigmoid.
🟢 3. Stronger Gradient Around Zero
At:
we have:
compared with sigmoid's maximum derivative:
🟢 4. Produces Positive and Negative Activations
This can be useful when the representation naturally requires both directions.
🟢 5. Useful in RNN/LSTM Architectures
Tanh has historically been and remains an important component of recurrent architectures.
❌ 19. Disadvantages of Tanh
🔴 1. Vanishing Gradients
For large :
This can slow training in deep networks.
🔴 2. Saturation
The function approaches:
and:
at the extremes.
🔴 3. Computationally More Expensive Than ReLU
Tanh involves exponential calculations:
whereas ReLU is:
🔴 4. Not Usually the First Choice for Modern Deep Hidden Layers
Modern deep networks often use ReLU-family or other newer activation functions because they can optimize more effectively in many settings.
📌 20. Where Should We Use Tanh?
A useful guideline is:
| Location / Problem | Tanh |
|---|---|
| Classical hidden layers | ✅ Used historically |
| RNN hidden state | ✅ Common |
| LSTM candidate/state | ✅ Common |
| Binary classification output | ❌ Usually sigmoid |
| Multiclass classification output | ❌ Usually softmax |
| Regression output | ❌ Usually linear |
| Modern deep hidden layers | ⚠️ Often replaced by ReLU-family functions |
⭐ 21. Key Points for Students
Remember these six points:
① Formula
② Range
③ Shape
④ Zero-centered
⑤ Derivative
⑥ Main limitation
🎓 22. Summary
The tanh (hyperbolic tangent) activation function is a nonlinear activation function that maps a real-valued input to the range . Unlike sigmoid, tanh is zero-centered, which can make optimization more favorable in some situations. Its derivative is , with a maximum value of 1 at . However, tanh saturates for large positive and negative inputs, causing its gradient to approach zero and potentially producing the vanishing-gradient problem. Tanh has been widely used in recurrent neural networks and LSTM architectures and was historically common in hidden layers of neural networks.
- Get link
- X
- Other Apps
Comments
Post a Comment