Skip to main content

Classification with a Neural Network

A multilayer neural network composes affine transformations with nonlinear activation functions. The nonlinearity lets it represent decision boundaries that a single logistic unit cannot.

One-hidden-layer model​

For an input vector xx, hidden weights W(1)W^{(1)}, hidden bias b(1)b^{(1)}, output weights W(2)W^{(2)}, and output bias b(2)b^{(2)}:

z(1)=W(1)x+b(1),a(1)=ϕ(z(1)),z(2)=W(2)a(1)+b(2),y^=σ(z(2)).\begin{aligned} z^{(1)} &= W^{(1)}x+b^{(1)},\\ a^{(1)} &= \phi\left(z^{(1)}\right),\\ z^{(2)} &= W^{(2)}a^{(1)}+b^{(2)},\\ \hat y &= \sigma\left(z^{(2)}\right). \end{aligned}

Here ϕ\phi may be ReLU, GELU, tanh, or another hidden activation. For binary classification, the sigmoid output y^\hat y is interpreted as a predicted probability and can be trained with binary cross-entropy:

L=−[ylog⁡y^+(1−y)log⁡(1−y^)].L=-\left[y\log \hat y+(1-y)\log(1-\hat y)\right].

Backpropagation​

Backpropagation applies the chain rule from the loss toward earlier layers. With sigmoid plus binary cross-entropy, the output pre-activation gradient simplifies to

δ(2)=∂L∂z(2)=y^−y.\delta^{(2)}=\frac{\partial L}{\partial z^{(2)}}=\hat y-y.

The output-layer gradients are

∂L∂W(2)=δ(2)(a(1))T,∂L∂b(2)=δ(2).\frac{\partial L}{\partial W^{(2)}}=\delta^{(2)}(a^{(1)})^T, \qquad \frac{\partial L}{\partial b^{(2)}}=\delta^{(2)}.

Propagating through the hidden activation gives

δ(1)=(W(2))Tδ(2)⊙ϕ′(z(1)),\delta^{(1)}=(W^{(2)})^T\delta^{(2)}\odot\phi'\left(z^{(1)}\right),

and therefore

∂L∂W(1)=δ(1)xT,∂L∂b(1)=δ(1).\frac{\partial L}{\partial W^{(1)}}=\delta^{(1)}x^T, \qquad \frac{\partial L}{\partial b^{(1)}}=\delta^{(1)}.

The symbol ⊙\odot denotes elementwise multiplication. Shapes matter: writing the computation in matrix form avoids an ambiguous chain of scalar subscripts.

Parameter update​

Plain gradient descent with learning rate α\alpha updates each parameter θ\theta by

θ←θ−α∂L∂θ.\theta\leftarrow\theta-\alpha\frac{\partial L}{\partial\theta}.

Practical training usually operates on mini-batches and may use Adam or another optimizer, but the gradients still come from the same forward graph and chain rule.

In TensorFlow Playground, choose a classification dataset, keep only the two coordinate inputs, and compare a network with no hidden layer against one with a hidden layer. Watch the decision boundary change as training proceeds; the playground illustrates representation and training rather than reproducing this note’s exact calculation.

What this note leaves out​

This small network is enough to show the calculation, not enough to choose a production architecture. Initialization, normalization, regularization, optimizer state, multiclass outputs, numerical stability, and data quality all affect training. The maintained multilayer perceptron note covers the broader architecture.

Shapes and a complete backward calculation​

For dd input features and hh hidden units, use column vectors: x∈Rdx\in\mathbb R^d, W(1)∈Rh×dW^{(1)}\in\mathbb R^{h\times d}, b(1),a(1),δ(1)∈Rhb^{(1)},a^{(1)},\delta^{(1)}\in\mathbb R^h, W(2)∈R1×hW^{(2)}\in\mathbb R^{1\times h}, and scalar b(2),z(2),δ(2)b^{(2)},z^{(2)},\delta^{(2)}. Each weight gradient has the shape of its weight matrix. The outer product δ(1)xT\delta^{(1)}x^T is therefore h×dh\times d.

For a minimal numerical example, let d=h=1d=h=1, x=2x=2, y=1y=1, and use ReLU, ϕ(t)=max⁡(0,t)\phi(t)=\max(0,t). Set W(1)=0.5W^{(1)}=0.5, b(1)=0b^{(1)}=0, W(2)=1W^{(2)}=1, b(2)=0b^{(2)}=0. The forward values are z(1)=a(1)=z(2)=1z^{(1)}=a^{(1)}=z^{(2)}=1, y^≈0.731059\hat y\approx0.731059, and L≈0.313262L\approx0.313262. Since the hidden input is positive, ϕ′=1\phi'=1, giving

δ(2)=δ(1)≈−0.268941,(∂L∂W(1),∂L∂b(1),∂L∂W(2),∂L∂b(2))≈(−0.537883,−0.268941,−0.268941,−0.268941).\delta^{(2)}=\delta^{(1)}\approx-0.268941,\qquad \left(\frac{\partial L}{\partial W^{(1)}},\frac{\partial L}{\partial b^{(1)}},\frac{\partial L}{\partial W^{(2)}},\frac{\partial L}{\partial b^{(2)}}\right) \approx(-0.537883,-0.268941,-0.268941,-0.268941).

With α=0.1\alpha=0.1, updating all four parameters from this same forward pass yields approximately (0.553788,0.026894,1.026894,0.026894)(0.553788,0.026894,1.026894,0.026894) in that order. The next logit is about 1.1918751.191875, and the loss is about 0.2651690.265169. Updating the output weights before computing δ(1)\delta^{(1)} would mix two different parameter states and is not this gradient step.

At ReLU input zero, the ordinary derivative does not exist. PyTorch uses the minimum-norm subgradient, zero, as its backward convention; see Autograd mechanics. Away from such kinks, check an analytic gradient against [L(θ+εej)−L(θ−εej)]/(2ε)[L(\theta+\varepsilon e_j)-L(\theta-\varepsilon e_j)]/(2\varepsilon) on a tiny example. This tests differentiation, not whether training will find a global optimum.

What the optimization loop can establish​

Backpropagation computes derivatives; the optimizer chooses parameter changes. For a batch loss, average the per-example gradients. Mini-batch gradients fluctuate, so every step need not decrease full training loss. Jointly learning hidden and output weights generally makes the objective nonconvex, unlike fixed-feature logistic regression. Different initializations can lead to different stationary regions; a small update may also reflect a tiny learning rate, saturated activations, or inactive ReLUs.

Use finite-value checks, an iteration budget, gradient and loss diagnostics, and held-out validation for early stopping. Validation-based stopping limits overfitting; it is not an optimality certificate. Compute binary cross-entropy from logits using the stable log-loss expression. These distinctions between optimization and generalization are developed in Deep Learning, Chapter 8.

Explore connectionsOpen network