Skip to main content

Multilayer Perceptron

A hidden layer builds intermediate features before producing an answer. In this tiny network, two neurons measure how the inputs differ. Switch their nonlinearity off to see why adding layers alone is not enough.

Try it yourself

Why does XOR need a hidden layer?

XOR means “one input is 1, but not both”. Toggle the inputs, then remove ReLU: the same weights stop solving the problem.

Two inputs, two hidden units and one output with fixed weights[0,1] → [0,1] → 1InputsReLUOutput+1−1−1+1+1+10x₁1x₂0h₁1h₂1ŷEdge labels are fixed weights · dashed edges have negative weights

[0, 1] → h = [0, 1] → Output = 1

ReLU keeps positive differences and clips negative ones to zero. Adding the two hidden values detects unequal inputs.

These weights are chosen by hand, not trained. This demonstrates a working representation, not a guarantee that training will find it.
Compare all four inputs
x₁x₂TargetOutput
0000
0111
1011
1100
Open the code · NumPy / JAX

The interaction above runs in JavaScript. These Python examples use the same data and equations, starting from the defaults—not your current controls. JAX can compute the loss gradient with jax.grad. JAX autodiff guide

In a new local project, save the example as demo.py. CPU setup: uv init ai-lab → cd ai-lab → uv add numpy jax → uv run demo.py

import jax.numpy as jnp
# NumPy: replace the import above with: import numpy as jnp

x = jnp.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
w1 = jnp.array([[1., -1.], [-1., 1.]])
w2 = jnp.array([1., 1.])
h = jnp.maximum(x @ w1, 0.)
print(h @ w2) # [0., 1., 1., 0.]
print((x @ w1) @ w2) # [0., 0., 0., 0.]: no ReLU

An MLP alternates affine maps with elementwise nonlinearities. For input x∈Rd\mathbf{x}\in\mathbb{R}^{d}, hidden width hh, and output width kk:

z1=W1x+b1,h=ϕ(z1),o=W2h+b2,\mathbf{z}_1=W_1\mathbf{x}+\mathbf{b}_1, \quad \mathbf{h}=\phi(\mathbf{z}_1), \quad \mathbf{o}=W_2\mathbf{h}+\mathbf{b}_2,

with W1∈Rh×dW_1\in\mathbb{R}^{h\times d} and W2∈Rk×hW_2\in\mathbb{R}^{k\times h}. Without ϕ\phi, consecutive affine maps collapse into one; depth alone would not enlarge the function class.

Heads, Losses, and Assumptions​

ProblemOutputCommon lossCheck
real-valued regressionlinear valueMSE, MAE, or likelihood-derived losstails, scale, heteroscedasticity
binary classificationone logitbinary cross-entropythreshold, costs, calibration
exclusive multiclasskk logitssoftmax cross-entropyclasses really are exclusive
multilabelkk independent logitsindependent BCElabel dependence and thresholds

The hidden activation and output semantics serve different purposes. Do not select a library loss first and retrofit the question afterward.

What Backpropagation Computes​

For one loss LL, let δ2=∂L/∂o\boldsymbol{\delta}_2=\partial L/\partial\mathbf{o}. The chain rule gives

∂L∂W2=δ2h⊤,δ1=(W2⊤δ2)⊙ϕ′(z1),\frac{\partial L}{\partial W_2}=\boldsymbol{\delta}_2\mathbf{h}^{\top}, \qquad \boldsymbol{\delta}_1=(W_2^{\top}\boldsymbol{\delta}_2)\odot\phi'(\mathbf{z}_1), ∂L∂W1=δ1x⊤.\frac{\partial L}{\partial W_1}=\boldsymbol{\delta}_1\mathbf{x}^{\top}.

Autodiff removes hand-coded derivatives, not shape, numerical, or objective assumptions. A zero gradient may indicate an optimum, saturation, a dead ReLU, an incorrect mask, or a disconnected graph.

One Scalar Backward Pass​

Take x=2x=2, w1=1w_1=1, b1=−1b_1=-1, w2=3w_2=3, b2=0b_2=0, target y=1y=1, and L=(o−y)2/2L=(o-y)^2/2. Forward evaluation gives z1=1z_1=1, h=ReLU⁡(1)=1h=\operatorname{ReLU}(1)=1, o=3o=3, L=2L=2. Backward evaluation gives δ2=2\delta_2=2, ∂L/∂w2=2\partial L/\partial w_2=2, δ1=6\delta_1=6, and ∂L/∂w1=12\partial L/\partial w_1=12. The bias gradients are ∂L/∂b2=2\partial L/\partial b_2=2 and ∂L/∂b1=6\partial L/\partial b_1=6. All derivatives use the same forward-pass weights; updating w2w_2 before computing δ1\delta_1 would differentiate a different computation.

Backpropagation computes gradients; an optimizer applies the update. For a mean minibatch loss, average per-example contributions, and clear accumulated gradients before the next independent update. Inference needs only the forward pass. If dropout is present, training randomly suppresses activations (with compensating scaling in inverted dropout), while evaluation disables this randomness. A plain affine/ReLU MLP has no such mode-dependent layer.

Worked Example: XOR with Two ReLUs​

For x1,x2∈{0,1}x_1,x_2\in\{0,1\}, define

h1=ReLU⁡(x1−x2),h2=ReLU⁡(x2−x1),h_1=\operatorname{ReLU}(x_1-x_2), \qquad h_2=\operatorname{ReLU}(x_2-x_1), f(x1,x2)=h1+h2.f(x_1,x_2)=h_1+h_2.

The equal pairs produce 0 and unequal pairs produce 1. A single linear boundary cannot separate XOR; the hidden layer first represents whether the inputs differ. This proves that one representation exists—not that SGD will find it from arbitrary initialization or generalize under noise.

Modern Nonlinearities​

  • ReLU is cheap and nonsaturating on its positive side, but units may remain inactive. It is still a useful baseline, not universally obsolete.
  • GELU and SiLU/Swish are smooth elementwise functions common in modern Transformers and some ConvNets.
  • GEGLU/SwiGLU use two projected branches and change width, parameters, and kernels; they are not simple drop-in activations.
  • Sigmoid/Tanh remain useful in gates and bounded states despite saturation at large magnitude.

See Activations and Gated Feed-Forward Networks. Fair comparisons control parameter count, FLOPs, training budget, and implementation kernels.

Capacity Is Not Learnability​

Cybenko's theorem states that finite sums of units with a continuous sigmoidal activation can uniformly approximate any continuous function on a compact cube to any positive error tolerance. The required finite width may grow with the function and tolerance; the theorem does not give a fixed width that suffices for every target. It also does not guarantee sufficient training data, an efficient representation, successful optimization, calibrated probabilities, causal validity, or out-of-distribution generalization. “Universal approximator” is not a model-selection result.

Inductive Bias and Counterexamples​

An MLP sees a flat vector. Translation, sequence order, graph structure, and permutation invariance do not appear automatically. CNNs, RNNs, attention, and graph networks encode stronger sharing or computation patterns and can change sample efficiency.

Conversely, on stable tabular features with limited data, a linear model, tree, or small MLP may beat a fashionable architecture. Compare them under the same split and budget.

Minimal Acceptance​

  1. Assert batch, feature, and output shapes.
  2. Overfit a tiny sample to test that the implementation can learn.
  3. Compare against constant, linear, or tree baselines.
  4. Fit learned preprocessing only inside training folds.
  5. Report multiple seeds, failures, and resources—not only the best run.
  6. Test deployment assumptions with Evaluation Under Distribution Shift.

This note focuses on conventional supervised gradient training; it does not treat an MLP as a biological model or survey every activation and optimizer.

Extend the XOR example in TensorFlow Playground: choose a nonlinear dataset, change hidden-layer width or activation, and watch the decision boundary during training. Then increase noise and compare training with test loss; a more flexible boundary can fit noise as well as structure.

Explore connectionsOpen network