Softmax Regression
Softmax regression is a linear model that assigns probabilities to mutually exclusive classes, commonly used for single-label classification. It produces one unnormalized score, or logit, per class:
Follow one example with three classes and logits . Exponentiating gives , whose sum is . Dividing each entry by gives probabilities .
Softmax converts the logits into nonnegative values that sum to one:
Implementations subtract the largest logit before exponentiation for numerical stability; this does not change the result because softmax is invariant to adding the same constant to every logit. At inference, argmax can operate directly on logits because softmax preserves their ordering.
Each output combines all four inputs using its own weights and bias. The three output nodes are logits, not probabilities: softmax normalizes them afterward. There is no hidden layer. With the matrix convention above, this example has W of shape 4 × 3 and three bias terms.
Cross-Entropy
For a one-hot target , the per-example cross-entropy is (with denoting the probability assigned to the observed class)
Suppose the second class is correct, so the one-hot target is . Only its probability contributes: . The model favors the first class, but raising the second class’s probability would reduce this loss.
Minimizing this objective corresponds to maximizing the conditional likelihood of the observed class labels under the model.
The information-theory definition of cross-entropy compares a target distribution with a predictive distribution. Here the one-hot is the empirical target for one observed example, and is the prediction. Natural logarithms give loss in nats. This does not imply that the population's conditional label distribution is itself one-hot.
A Stable Loss and Its Gradient
For input features, is , is length , and a batch of rows produces logits. Let be the correct class index, distinct from the one-hot vector . The softmax derivation yields
Compute this log-sum-exp expression or use a fused cross-entropy loss that accepts logits. Taking a logarithm of already rounded probabilities can produce infinity even after stable softmax.
In the same example, and , so the stable loss is . Adding to every logit cancels out of this expression and leaves both loss and probabilities unchanged.
Subtract the target from the prediction to find the direction of correction:
The negative second component calls for raising the correct logit; the positive components call for lowering the others. A gradient-descent step on these logits subtracts a positive multiple of this vector.
To connect the logits to the trainable parameters, differentiate and use :
This is reverse-mode differentiation through the affine logits: the loss gradient is passed backward and scaled by the corresponding input. The matrix formula collects those coordinate derivatives. A mean batch loss averages these per-example gradients.
Boundaries
- Mutually exclusive multiclass: softmax regression commonly models a single class label per example. The softmax function also serves other purposes, such as computing attention weights.
- Multilabel: use independent outputs, commonly sigmoid probabilities with a binary loss, when several labels may be true simultaneously.
- Decision rule:
argmaxchooses a class, but asymmetric costs may require a different rule. - Probability quality: good classification accuracy does not imply calibrated probabilities.
- Linear boundary: nonlinear feature relationships require transformed features or a more expressive model.
Evaluate against class frequencies and a simple baseline, then inspect a confusion matrix and per-class errors. Aggregate accuracy can hide poor behavior on rare or important classes.
Continue with the Multilayer Perceptron to introduce nonlinear hidden representations. See Dive into Deep Learning: Linear Neural Networks for Classification for derivations and maintained implementations.