Posts

Showing posts from February, 2026

Tanh vs Sigmoid vs ReLU

  ⚖️  Tanh vs Sigmoid vs ReLU Property Sigmoid Tanh ReLU Formula 1 1 + e − z \frac{1}{1+e^{-z}} e z − e − z e z + e − z \frac{e^z-e^{-z}}{e^z+e^{-z}} max ⁡ ( 0 , z ) \max(0,z) Range ( 0 , 1 ) (0,1) ( − 1 , 1 ) (-1,1) [ 0 , ∞ ) [0,\infty) Nonlinear Yes Yes Yes Zero-centered ❌ ✅ ❌ Saturates Both sides Both sides Positive side: No Vanishing gradient Yes Yes Reduced for positive inputs Main issue Vanishing gradients Vanishing gradients Dying ReLU Typical use Binary output Some hidden/RNN layers Hidden layers

Batch Gradient Descent

  Batch Gradient Descent 1. What is Batch Gradient Descent? In Batch Gradient Descent , we do not update the weights after every individual training example . Instead: We process the entire training dataset, calculate the total gradient, and then update the weights once. The basic idea is: All training examples → Calculate total gradient → Update weights \boxed{ \text{All training examples} \rightarrow \text{Calculate total gradient} \rightarrow \text{Update weights} } This complete process is called one iteration , or commonly one epoch when the entire training dataset has been processed. 2. Why do we need Batch Gradient Descent? Suppose we have N N training examples: ( x 1 , t 1 ) , ( x 2 , t 2 ) , … , ( x N , t N ) (x_1,t_1),(x_2,t_2),\ldots,(x_N,t_N) For each example, the neuron produces: y d y_d and has error: e d = t d − y d e_d=t_d-y_d If we change the weights immediately after every example, the direction of the update is based on only one example . Ins...

ReLU Activation Function

Image
  🔵 ReLU Activation Function ReLU stands for Rectified Linear Unit . It is one of the most widely used activation functions in modern neural networks, particularly in hidden layers . 1. Definition of ReLU ReLU is defined as: ReLU ⁡ ( z ) = max ⁡ ( 0 , z ) \boxed{\operatorname{ReLU}(z)=\max(0,z)} Equivalently, we can write it as a piecewise function: ReLU ⁡ ( z ) = { 0 , z < 0 z , z ≥ 0 \boxed{ \operatorname{ReLU}(z)= \begin{cases} 0, & z<0\\ z, & z\geq0 \end{cases} } In simple words: ReLU keeps positive values and converts negative values to zero. Examples ReLU ⁡ ( − 5 ) = 0 \operatorname{ReLU}(-5)=0 ReLU ⁡ ( − 2 ) = 0 \operatorname{ReLU}(-2)=0 ReLU ⁡ ( 0 ) = 0 \operatorname{ReLU}(0)=0 ReLU ⁡ ( 3 ) = 3 \operatorname{ReLU}(3)=3 ReLU ⁡ ( 10 ) = 10 \operatorname{ReLU}(10)=10 So: z z    ReLU( z z ) -5 0 -2 0 -1 0 0 0 1 1 2 2 5 5 2. How ReLU works inside a neuron As with other activation functions, first calculate the weighted sum: z = ∑ ...