Classification with a Neural Network
A multilayer neural network composes affine transformations with nonlinear activation functions. The nonlinearity lets it represent decision boundaries that a single logistic unit cannot.
One-hidden-layer model
For an input vector , hidden weights , hidden bias , output weights , and output bias :
Here may be ReLU, GELU, tanh, or another hidden activation. For binary classification, the sigmoid output is interpreted as a predicted probability and can be trained with binary cross-entropy:
Backpropagation
Backpropagation applies the chain rule from the loss toward earlier layers. With sigmoid plus binary cross-entropy, the output pre-activation gradient simplifies to
The output-layer gradients are
Propagating through the hidden activation gives
and therefore
The symbol denotes elementwise multiplication. Shapes matter: writing the computation in matrix form avoids an ambiguous chain of scalar subscripts.
Parameter update
Plain gradient descent with learning rate updates each parameter by
Practical training usually operates on mini-batches and may use Adam or another optimizer, but the gradients still come from the same forward graph and chain rule.
In TensorFlow Playground, choose a classification dataset, keep only the two coordinate inputs, and compare a network with no hidden layer against one with a hidden layer. Watch the decision boundary change as training proceeds; the playground illustrates representation and training rather than reproducing this note’s exact calculation.
What this note leaves out
This small network is enough to show the calculation, not enough to choose a production architecture. Initialization, normalization, regularization, optimizer state, multiclass outputs, numerical stability, and data quality all affect training. The maintained multilayer perceptron note covers the broader architecture.
Shapes and a complete backward calculation
For input features and hidden units, use column vectors: , , , , and scalar . Each weight gradient has the shape of its weight matrix. The outer product is therefore .
For a minimal numerical example, let , , , and use ReLU, . Set , , , . The forward values are , , and . Since the hidden input is positive, , giving
With , updating all four parameters from this same forward pass yields approximately in that order. The next logit is about , and the loss is about . Updating the output weights before computing would mix two different parameter states and is not this gradient step.
At ReLU input zero, the ordinary derivative does not exist. PyTorch uses the minimum-norm subgradient, zero, as its backward convention; see Autograd mechanics. Away from such kinks, check an analytic gradient against on a tiny example. This tests differentiation, not whether training will find a global optimum.
What the optimization loop can establish
Backpropagation computes derivatives; the optimizer chooses parameter changes. For a batch loss, average the per-example gradients. Mini-batch gradients fluctuate, so every step need not decrease full training loss. Jointly learning hidden and output weights generally makes the objective nonconvex, unlike fixed-feature logistic regression. Different initializations can lead to different stationary regions; a small update may also reflect a tiny learning rate, saturated activations, or inactive ReLUs.
Use finite-value checks, an iteration budget, gradient and loss diagnostics, and held-out validation for early stopping. Validation-based stopping limits overfitting; it is not an optimality certificate. Compute binary cross-entropy from logits using the stable log-loss expression. These distinctions between optimization and generalization are developed in Deep Learning, Chapter 8.