Multilayer Perceptron (MLP)-Multilayer Feed-Forward Neural Network

 

🧠 Multilayer Perceptron (MLP)

Multilayer Feed-Forward Neural Network

🎓 Definition: An MLP is a multilayer feed-forward neural network consisting of an input layer, one or more hidden layers, and an output layer, in which information flows from input to output through weighted connections and nonlinear activation functions.


1. 🔗 From a Perceptron to an MLP


A single neuron computes:

z=w1x1+w2x2+⋯+wnxn+bz=w_1x_1+w_2x_2+\cdots+w_nx_n+b

and then applies an activation function:

y=ϕ(z)y=\phi(z)

Graphically:

                 🧠 SINGLE PERCEPTRON

 

A Multilayer Perceptron connects many such neurons together in layers.

                 🧠 MULTILAYER PERCEPTRON

                Features Learned Prediction                   representation

So the fundamental idea is:

Many neurons+multiple layers⇒MLP\boxed{ \text{Many neurons}+\text{multiple layers} \Rightarrow \text{MLP} }

2. 🧩 Why is it called a "Multilayer Perceptron"?

Let's break the name into three parts.

🔹 Multi

There are multiple layers of neurons.

🔹 Layer

Neural network is  organized into layers:

Input Layer→Hidden Layer(s)→Output Layer\boxed{ \text{Input Layer} \rightarrow \text{Hidden Layer(s)} \rightarrow \text{Output Layer} }

🔹 Perceptron

Each processing neuron performs a weighted combination of its inputs followed by an activation function.

Therefore:





Multilayer Perceptron=MLP






\boxed{\text{Multilayer Perceptron} = \text{MLP}}

What does "Multilayer Feed-Forward" mean?

Let's break the name into three parts.

🔹 Multi-layer

There are multiple layers of neurons:

Input Layer → Hidden Layer(s) → Output Layer

🔹 Feed-forward

Information moves only in one direction:

Input→Hidden→Output\boxed{ \text{Input} \rightarrow \text{Hidden} \rightarrow \text{Output} }

There are no connections going backward or forming cycles during the forward computation.

🔹 Neural network

Each processing node is an artificial neuron, and neurons are connected through weighted connections.


3. 🏗️ Basic Architecture of an MLP

Consider an MLP with:

  • 3 input neurons
  • 4 hidden neurons
  • 1 output neuron
                         🧠 MULTILAYER PERCEPTRON

        INPUT LAYER          HIDDEN LAYER             OUTPUT LAYER

        Features             Representation            Prediction

       



The information moves:

Input→Hidden→Output\boxed{ \text{Input} \rightarrow \text{Hidden} \rightarrow \text{Output} }

This is why it is called a feed-forward neural network.


4. 🟢 Input Layer

The input layer receives the features.

Suppose we want to predict whether a student will pass an examination.

We could use:

InputFeature
x1x_1Attendance
x2x_2Internal marks
x3x_3Assignment score

Then:

x=[x1x2x3]\mathbf{x} = \begin{bmatrix} x_1\\ x_2\\ x_3 \end{bmatrix}

The input layer has three input units.

🎓 Important point

The input layer generally does not perform the main neural computation. It supplies the feature values to the first hidden layer.


5. 🟡 Hidden Layer — Where the Learning Representation Begins

The hidden layer is where neurons combine the input information.

Consider hidden neuron h1h_1.

It receives:

x1,x2,x3x_1,x_2,x_3

with corresponding weights:

w11,w12,w13w_{11},w_{12},w_{13}

It calculates:

z1=w11x1+w12x2+w13x3+b1z_1=w_{11}x_1+w_{12}x_2+w_{13}x_3+b_1

Then applies an activation function:

h1=ϕ(z1)\boxed{ h_1=\phi(z_1) }

Similarly:

h2=ϕ(w21x1+w22x2+w23x3+b2)h_2=\phi(w_{21}x_1+w_{22}x_2+w_{23}x_3+b_2)

and so on.

💡 Important idea

The hidden layer transforms the original input into a new representation.

Raw features→learned representation\boxed{ \text{Raw features} \rightarrow \text{learned representation} }

6. 🔥 Why Does an MLP Need Hidden Layers?


A single perceptron can learn a linear decision boundary.

For example:

       Class 1
    ●  ●  ●
    ●  ●

    ───────────────  ← Linear decision boundary

    ●  ●
    ●  ●  ●
       Class 0

But many real-world problems are nonlinear.

The classic example is XOR.

                 XOR

          x₂
           ↑
        1  │  1       0  Class 1 (1s)
           │
        0  │  0       1  Class 0 (0s)
           │
           └────────────→ x₁
              0      1

      The classes cannot be separated
      by one straight line.

A single perceptron cannot solve XOR.

An MLP can.

Hidden layers + nonlinear activation⇒Nonlinear decision boundaries\boxed{ \text{Hidden layers + nonlinear activation} \Rightarrow \text{Nonlinear decision boundaries} }

7. ⚙️ What Happens Inside One Hidden Neuron?

Every hidden neuron essentially performs two operations.

Step 1 — Weighted sum

z=∑iwixi+b\boxed{ z=\sum_iw_ix_i+b }

Step 2 — Activation

h=ϕ(z)\boxed{ h=\phi(z) }

This same basic operation is repeated by every neuron.


8. 🔵 Output Layer

The output layer produces the final prediction.

Suppose the hidden layer produces:

h1,h2,h3,h4h_1,h_2,h_3,h_4

The output neuron calculates:

zo=v1h1+v2h2+v3h3+v4h4+boz_o=v_1h_1+v_2h_2+v_3h_3+v_4h_4+b_o

and then:

y=ψ(zo)\boxed{ y=\psi(z_o) }

The activation function depends on the problem.

🎯 Task    Output activation
Regression    Linear
Binary classification    Sigmoid
Multiclass classification    Softmax

For binary classification:

y=11+e−zo\boxed{ y=\frac{1}{1+e^{-z_o}} }

9. 🔄 How Does Information Flow Through an MLP?

This process is called Forward Propagation.

          🚀 FORWARD PROPAGATION

Input
  │
  ▼
┌──────────────┐
│ Input Layer  │
└──────┬───────┘
       │
       ▼
┌──────────────┐
│ Hidden Layer │
│              │
└──────┬───────┘
       │
       ▼
┌──────────────┐
│ Output Layer │
└──────┬───────┘
       │
       ▼
   Prediction y

Mathematically:

First hidden layer

z(1)=W(1)x+b(1)\boxed{ \mathbf{z}^{(1)} = W^{(1)}\mathbf{x} + \mathbf{b}^{(1)} }

Activation

h(1)=ϕ(z(1))\boxed{ \mathbf{h}^{(1)} = \phi(\mathbf{z}^{(1)}) }

Output layer

z(2)=W(2)h(1)+b(2)\boxed{ \mathbf{z}^{(2)} = W^{(2)}\mathbf{h}^{(1)} + \mathbf{b}^{(2)} }

Output

y=ψ(z(2))\boxed{ \mathbf{y} = \psi(\mathbf{z}^{(2)}) }

Therefore:

x→z(1)→h(1)→z(2)→y\boxed{ \mathbf{x} \rightarrow \mathbf{z}^{(1)} \rightarrow \mathbf{h}^{(1)} \rightarrow \mathbf{z}^{(2)} \rightarrow \mathbf{y} }

10. 🧮 MLP with More Than One Hidden Layer

An MLP can contain several hidden layers.

For example:

       

For layer ll:

z(l)=W(l)a(l−1)+b(l)\boxed{ \mathbf{z}^{(l)} = W^{(l)}\mathbf{a}^{(l-1)} + \mathbf{b}^{(l)} }

and:

a(l)=ϕ(z(l))\boxed{ \mathbf{a}^{(l)} = \phi(\mathbf{z}^{(l)}) }

This equation is extremely important for understanding modern neural networks.


11. 🚨 Why Must We Use Nonlinear Activation Functions?

Suppose every layer were linear.

First layer:

h=W1x+b1\mathbf{h} = W_1\mathbf{x}+\mathbf{b}_1

Second layer:

y=W2h+b2\mathbf{y} = W_2\mathbf{h}+\mathbf{b}_2

Substituting:

y=W2(W1x+b1)+b2\mathbf{y} = W_2(W_1\mathbf{x}+\mathbf{b}_1)+\mathbf{b}_2

Therefore:

y=W2W1x+W2b1+b2\mathbf{y} = W_2W_1\mathbf{x} + W_2\mathbf{b}_1 + \mathbf{b}_2

which is still:

y=Wx+b\boxed{ \mathbf{y}=W\mathbf{x}+\mathbf{b} }

So multiple linear layers can still behave like one linear layer.

💡 Therefore:

Nonlinear activation functions give MLP its nonlinear modeling power.\boxed{ \text{Nonlinear activation functions give MLP its nonlinear modeling power.} }

Common choices include:

  • ReLU
  • Sigmoid
  • Tanh

12. 🧠 What Does the MLP Actually Learn?

This is an important conceptual question.

The MLP does not explicitly receive rules such as:

"If attendance is greater than 75%, classify as pass."

Instead, it learns:

Weights and biases\boxed{ \text{Weights and biases} }

during training.

Initially:

W,b→small/random valuesW,b \rightarrow \text{small/random values}

During training:

W←W−η∂E∂W\boxed{ W\leftarrow W-\eta\frac{\partial E}{\partial W} }

and:

b←b−η∂E∂b\boxed{ b\leftarrow b-\eta\frac{\partial E}{\partial b} }

Repeated updates allow the network to find parameter values that reduce the loss.


13. 🔁 How Does an MLP Learn?

his is the complete learning cycle that students should understand.

                    🧠 MLP TRAINING

                         INPUT
                           │
                           ▼
                  🚀 FORWARD PASS
                           │
                           ▼
                     Prediction y
                           │
                           ▼
                    🎯 Target t
                           │
                           ▼
                    📉 Calculate Loss
                           │
                           ▼
                 ⬅️ BACKPROPAGATION
                           │
                           ▼
                    ∇ Gradients
                           │
                           ▼
                 ⚙️ GRADIENT DESCENT
                           │
                           ▼
                  Update W and b
                           │
                           └───────────┐
                                       │
                                       ▼
                                     Repeat

The fundamental learning cycle is:

Forward Pass→Loss→Backpropagation→Gradient Descent→Update Parameters→Repeat\boxed{ \text{Forward Pass} \rightarrow \text{Loss} \rightarrow \text{Backpropagation} \rightarrow \text{Gradient Descent} \rightarrow \text{Update Parameters} \rightarrow \text{Repeat} }

14. 🔍 Backpropagation vs Gradient Descent

Students frequently confuse these two.

🔄 Backpropagation

Backpropagation calculates:

∂E∂w\boxed{ \frac{\partial E}{\partial w} }

for all the weights and biases.

It uses the chain rule.

⚙️ Gradient Descent

Gradient descent uses those gradients to update the parameters:

wnew=wold−η∂E∂w\boxed{ w_{\text{new}} = w_{\text{old}} - \eta \frac{\partial E}{\partial w} }

Therefore:

Backpropagation calculates the gradients; gradient descent uses the gradients to learn the parameters.


15. 🧩 MLP Solving XOR

This is a beautiful example to connect your previous lectures.

Recall:

XOR=(OR) AND (NAND)XOR=(OR)\ AND\ (NAND)

An MLP can implement this conceptually:

                     HIDDEN LAYER

 x₁ ───────────────► [ OR ] ─────────┐
                                      │
                                      ▼
                                    [ AND ] ───► XOR
                                      ▲
                                      │
 x₁ ───────────────► [ NAND ] ───────┘

The hidden neurons construct intermediate information:

h1=OR(x1,x2)h_1=OR(x_1,x_2) h2=NAND(x1,x2)h_2=NAND(x_1,x_2)

Then:

y=AND(h1,h2)y=AND(h_1,h_2)

This demonstrates the power of hidden layers.


16. 🌟 MLP as a Function Approximator

One of the most important theoretical properties of MLPs is their ability to approximate complex functions.

A simple neuron can represent:

y=w1x1+w2x2+by=w_1x_1+w_2x_2+b

An MLP can represent much more complex mappings:

y=f(x)\boxed{ \mathbf{y}=f(\mathbf{x}) }

For example:

Simple model:

Input ───────────────► Output
       Linear


MLP:

Input ──► Hidden ──► Hidden ──► Output
          🧠          🧠
       Nonlinear    Nonlinear

       Complex function approximation

With suitable nonlinear activation functions and enough hidden units, an MLP can approximate a broad class of continuous functions.

⚠️ Important qualification

Universal approximation does not mean:

  • training is always easy,
  • the network will automatically find the solution,
  • fewer neurons are always sufficient,
  • the network will automatically generalize well.

It is primarily a statement about representational capability.


17. ⚖️ Perceptron vs MLP

Feature🟢 Perceptron🧠 MLP
Input layer    ✓                ✓
Hidden layers    ✗                ✓
Multiple layers    ✗                ✓
Nonlinear modeling    Limited                ✓
XOR    ❌                ✓
Complex functions    Limited                ✓
Training    Perceptron rule            Backpropagation + gradient descent
Learned parameters    Weights & bias            Weights & biases
Representation learning    Limited                ✓

18. 🌟 Advantages of MLP

✅ 1. Nonlinear modeling

MLPs can learn nonlinear relationships.

✅ 2. Complex decision boundaries

They can create much more complex decision boundaries than a single perceptron.

✅ 3. Function approximation

They can approximate many nonlinear functions.

✅ 4. Automatic representation learning

Hidden layers transform the input into useful intermediate representations.

✅ 5. Versatile

MLPs can perform:

  • 🎯 Classification
  • 📈 Regression
  • 🔢 Function approximation
  • 🔍 Pattern recognition
  • 📊 Prediction

19. ⚠️ Limitations of MLP

❌ 1. Can overfit

A sufficiently large network may memorize training data.

❌ 2. Training can be computationally expensive

Large MLPs may contain a very large number of parameters.

❌ 3. Hyperparameter selection is important

For example:

Learning rate,  hidden layers,  neurons,  batch size,  epochs\boxed{ \text{Learning rate} ,\; \text{hidden layers} ,\; \text{neurons} ,\; \text{batch size} ,\; \text{epochs} }

❌ 4. Optimization difficulties

Training can encounter problems such as:

  • Vanishing gradients
  • Exploding gradients
  • Poor local optimization behavior
  • Slow convergence

❌ 5. Interpretability

It can be difficult to explain exactly how a large MLP arrives at a particular prediction.


20. 🎓  Summary 


                 🧠 MULTILAYER PERCEPTRON (MLP)

       INPUT              HIDDEN LAYERS              OUTPUT

     Features          Learned representations       Prediction

       x₁ ● ───────► ● ───────► ● ─────────────► ●
       x₂ ● ───────► ● ───────► ● ─────────────► y
       x₃ ● ───────► ● ───────► ● ─────────────► ●

             │
             │   Weighted sum + Activation
             ▼
          Nonlinear
       transformation
             │
             ▼
        Forward Pass
             │
             ▼
        Prediction y
             │
             ▼
          Loss E
             │
             ▼
     ⬅️ Backpropagation
             │
             ▼
        Gradients ∇E
             │
             ▼
     ⚙️ Gradient Descent
             │
             ▼
      Update W and b
             │
             └────────── 🔄 Repeat

⭐ Five points students should remember

  1. MLP = Multilayer Perceptron, commonly described as a multilayer feed-forward neural network.
  2. It consists of input, one or more hidden layers, and output layers.
  3. Each neuron performs a weighted sum + bias + activation function.
  4. Nonlinear activation functions allow the MLP to learn nonlinear relationships.
  5. MLP training follows:
Forward Propagation→Loss→Backpropagation→Gradient Descent→Weight Update\boxed{ \text{Forward Propagation} \rightarrow \text{Loss} \rightarrow \text{Backpropagation} \rightarrow \text{Gradient Descent} \rightarrow \text{Weight Update} }

🎯 

A Multilayer Perceptron (MLP) is a feed-forward neural network consisting of an input layer, one or more hidden layers, and an output layer, where neurons are connected by weighted connections and nonlinear activation functions enable the network to learn complex nonlinear mappings

Comments

Popular posts from this blog

Machine Learning PCCST503 Semester5 KTU CS 2024 Scheme - Dr Binu V P

Introduction to Machine Learning (ML)

Distinguishing Machine Learning from Traditional Programming