Multilayer Perceptron
A hidden layer builds intermediate features before producing an answer. In this tiny network, two neurons measure how the inputs differ. Switch their nonlinearity off to see why adding layers alone is not enough.
Why does XOR need a hidden layer?
XOR means “one input is 1, but not both”. Toggle the inputs, then remove ReLU: the same weights stop solving the problem.
[0, 1] → h = [0, 1] → Output = 1
ReLU keeps positive differences and clips negative ones to zero. Adding the two hidden values detects unequal inputs.
These weights are chosen by hand, not trained. This demonstrates a working representation, not a guarantee that training will find it.Compare all four inputs
| x₁ | x₂ | Target | Output |
|---|---|---|---|
| 0 | 0 | 0 | 0 |
| 0 | 1 | 1 | 1 |
| 1 | 0 | 1 | 1 |
| 1 | 1 | 0 | 0 |
Open the code · NumPy / JAX
The interaction above runs in JavaScript. These Python examples use the same data and equations, starting from the defaults—not your current controls. JAX can compute the loss gradient with jax.grad. JAX autodiff guide
In a new local project, save the example as demo.py. CPU setup: uv init ai-lab → cd ai-lab → uv add numpy jax → uv run demo.py
import jax.numpy as jnp
# NumPy: replace the import above with: import numpy as jnp
x = jnp.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
w1 = jnp.array([[1., -1.], [-1., 1.]])
w2 = jnp.array([1., 1.])
h = jnp.maximum(x @ w1, 0.)
print(h @ w2) # [0., 1., 1., 0.]
print((x @ w1) @ w2) # [0., 0., 0., 0.]: no ReLU
An MLP alternates affine maps with elementwise nonlinearities. For input , hidden width , and output width :
with and . Without , consecutive affine maps collapse into one; depth alone would not enlarge the function class.
Heads, Losses, and Assumptions
The hidden activation and output semantics serve different purposes. Do not select a library loss first and retrofit the question afterward.
What Backpropagation Computes
For one loss , let . The chain rule gives
Autodiff removes hand-coded derivatives, not shape, numerical, or objective assumptions. A zero gradient may indicate an optimum, saturation, a dead ReLU, an incorrect mask, or a disconnected graph.
One Scalar Backward Pass
Take , , , , , target , and . Forward evaluation gives , , , . Backward evaluation gives , , , and . The bias gradients are and . All derivatives use the same forward-pass weights; updating before computing would differentiate a different computation.
Backpropagation computes gradients; an optimizer applies the update. For a mean minibatch loss, average per-example contributions, and clear accumulated gradients before the next independent update. Inference needs only the forward pass. If dropout is present, training randomly suppresses activations (with compensating scaling in inverted dropout), while evaluation disables this randomness. A plain affine/ReLU MLP has no such mode-dependent layer.
Worked Example: XOR with Two ReLUs
For , define
The equal pairs produce 0 and unequal pairs produce 1. A single linear boundary cannot separate XOR; the hidden layer first represents whether the inputs differ. This proves that one representation exists—not that SGD will find it from arbitrary initialization or generalize under noise.
Modern Nonlinearities
- ReLU is cheap and nonsaturating on its positive side, but units may remain inactive. It is still a useful baseline, not universally obsolete.
- GELU and SiLU/Swish are smooth elementwise functions common in modern Transformers and some ConvNets.
- GEGLU/SwiGLU use two projected branches and change width, parameters, and kernels; they are not simple drop-in activations.
- Sigmoid/Tanh remain useful in gates and bounded states despite saturation at large magnitude.
See Activations and Gated Feed-Forward Networks. Fair comparisons control parameter count, FLOPs, training budget, and implementation kernels.
Capacity Is Not Learnability
Cybenko's theorem states that finite sums of units with a continuous sigmoidal activation can uniformly approximate any continuous function on a compact cube to any positive error tolerance. The required finite width may grow with the function and tolerance; the theorem does not give a fixed width that suffices for every target. It also does not guarantee sufficient training data, an efficient representation, successful optimization, calibrated probabilities, causal validity, or out-of-distribution generalization. “Universal approximator” is not a model-selection result.
Inductive Bias and Counterexamples
An MLP sees a flat vector. Translation, sequence order, graph structure, and permutation invariance do not appear automatically. CNNs, RNNs, attention, and graph networks encode stronger sharing or computation patterns and can change sample efficiency.
Conversely, on stable tabular features with limited data, a linear model, tree, or small MLP may beat a fashionable architecture. Compare them under the same split and budget.
Minimal Acceptance
- Assert batch, feature, and output shapes.
- Overfit a tiny sample to test that the implementation can learn.
- Compare against constant, linear, or tree baselines.
- Fit learned preprocessing only inside training folds.
- Report multiple seeds, failures, and resources—not only the best run.
- Test deployment assumptions with Evaluation Under Distribution Shift.
This note focuses on conventional supervised gradient training; it does not treat an MLP as a biological model or survey every activation and optimizer.
Extend the XOR example in TensorFlow Playground: choose a nonlinear dataset, change hidden-layer width or activation, and watch the decision boundary during training. Then increase noise and compare training with test loss; a more flexible boundary can fit noise as well as structure.