Skip to main content

Classification with a Logistic Unit

Logistic Regression as a Linear Classifier​

A linear score followed by a sigmoid and trained with log loss is logistic regression, sometimes described as a single logistic neuron. It is not the classical perceptron, which uses a threshold decision rule and a different update algorithm.

Mathematical Formulation​

For inputs x1,…,xnx_1,\ldots,x_n, weights w1,…,wnw_1,\ldots,w_n, and bias bb, the linear score is:

z=∑i=1nwixi+bz = \sum_{i=1}^{n} w_i x_i + b

This equation represents the linear combination of inputs and their respective weights, with the bias term added to account for offsets. For classification, the sigmoid function, σ(z)\sigma(z), is used as the activation function, transforming zz into a probability between 0 and 1:

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

This function outputs a value in the range (0,1)(0, 1), making it suitable for binary classification tasks.

Sigmoid Function​

sigmoid function

The sigmoid function, denoted as σ(z)\sigma(z), plays a crucial role in machine learning, especially in logistic regression and neural networks, due to its ability to map any real-valued number into the (0,1)(0, 1) interval. This property is particularly useful for modeling probabilities.

Definition​

The sigmoid function is defined as:

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

where z∈Rz \in \mathbb{R} is the input to the function.

Properties​

  • Domain and Range: The function maps the domain of all real numbers R\mathbb{R} to the range (0,1)(0, 1).
  • Asymptotes: It has horizontal asymptotes at y=0y = 0 and y=1y = 1, implying that σ(z)\sigma(z) approaches 00 as z→−∞z \to -\infty and 11 as z→∞z \to \infty.
  • Symmetry: The identity σ(−z)=1−σ(z)\sigma(-z)=1-\sigma(z) makes the graph point-symmetric about (0,12)(0,\tfrac12), not about the origin.
  • Sigmoid of Large Positive and Negative Values: For large positive values of zz, σ(z)\sigma(z) approaches 11, and for large negative values, σ(z)\sigma(z) approaches 00.

Derivative of the Sigmoid Function​

The derivative of the sigmoid function is significant in machine learning algorithms, particularly in the optimization process. It can be computed using the chain rule of calculus and exhibits a simple form that is computationally efficient.

Calculation​

Let's denote the derivative of σ(z)\sigma(z) with respect to zz as σ′(z)\sigma'(z). The calculation proceeds as follows:

  1. Start with the definition σ(z)=(1+e−z)−1\sigma(z) = (1 + e^{-z})^{-1}.
  2. Applying the chain rule, we get:
σ′(z)=ddz(1+e−z)−1=−(1+e−z)−2⋅(−e−z)\sigma'(z) = \frac{d}{dz}\left(1 + e^{-z}\right)^{-1} = -(1 + e^{-z})^{-2} \cdot \left(-e^{-z}\right)
  1. Simplifying, we obtain:
σ′(z)=e−z(1+e−z)2\sigma'(z) = \frac{e^{-z}}{(1 + e^{-z})^2}
  1. By adding and subtracting 11 in the numerator and rearranging, we find:
σ′(z)=11+e−z(1−11+e−z)\sigma'(z) = \frac{1}{1 + e^{-z}} \left(1 - \frac{1}{1 + e^{-z}}\right)
  1. Finally, recognizing that the terms in the parenthesis represent σ(z)\sigma(z) and 1−σ(z)1 - \sigma(z) respectively, we arrive at the elegant result:
σ′(z)=σ(z)(1−σ(z))\sigma'(z) = \sigma(z)(1 - \sigma(z))

Gradient Descent for Logistic Regression​

Gradient descent is employed to minimize the error between the predicted and actual classifications. It adjusts the weights and bias to reduce the loss function, calculated using the log loss for classification:

L(y,y^)=−[ylog⁡(y^)+(1−y)log⁡(1−y^)]L(y, \hat{y}) = -[y \log(\hat{y}) + (1 - y) \log(1 - \hat{y})]

where yy is the observed label and y^\hat{y} is the predicted probability from the sigmoid. A class label is obtained only after choosing a decision threshold. The loss function L(y,y^)L(y, \hat{y}) measures how well the predicted probability agrees with the observed label. The optimization's goal is to minimize LL by adjusting the model parameters, specifically the weights (ww) and bias (bb).

To understand how changes in ww and bb affect LL, we calculate the partial derivatives of LL with respect to these parameters. This involves understanding how LL is influenced by y^\hat{y} and, in turn, how y^\hat{y} depends on each parameter.

Chain Rule Application​

The calculation of ∂L∂wi\frac{\partial L}{\partial w_i} and ∂L∂b\frac{\partial L}{\partial b} involves the application of the chain rule of calculus, expressed as:

  • ∂L∂wi=∂L∂y^⋅∂y^∂wi\frac{\partial L}{\partial w_i} = \frac{\partial L}{\partial \hat{y}} \cdot \frac{\partial \hat{y}}{\partial w_i}
  • ∂L∂b=∂L∂y^⋅∂y^∂b\frac{\partial L}{\partial b} = \frac{\partial L}{\partial \hat{y}} \cdot \frac{\partial \hat{y}}{\partial b}

The term ∂L∂y^\frac{\partial L}{\partial \hat{y}} is common across these expressions and is crucial for understanding the gradient's direction and magnitude.

Derivative Calculations​

Derivative of LL with respect to y^\hat{y}​

Given the log loss function, the derivative of LL with respect to y^\hat{y} is calculated as:

∂L∂y^=−yy^+1−y1−y^\frac{\partial L}{\partial \hat{y}} = -\frac{y}{\hat{y}} + \frac{1 - y}{1 - \hat{y}}

This expression represents how the loss function gradient depends on the difference between actual and predicted values.

Derivative of y^\hat{y} with respect to ww and bb​

The predicted probability y^\hat{y} is the sigmoid of the linear score. The derivatives of y^\hat{y} with respect to wiw_i, and bb are informed by the derivative of the sigmoid function:

  • ∂y^∂wi=y^(1−y^)xi\frac{\partial \hat{y}}{\partial w_i} = \hat{y}(1 - \hat{y})x_i
  • ∂y^∂b=y^(1−y^)\frac{\partial \hat{y}}{\partial b} = \hat{y}(1 - \hat{y})

Final Gradient Expressions​

The final expressions for the partial derivatives of the loss function with respect to the parameters are:

  • ∂L∂wi=−(y−y^)xi\frac{\partial L}{\partial w_i} = -(y - \hat{y})x_i
  • ∂L∂b=−(y−y^)\frac{\partial L}{\partial b} = -(y - \hat{y})

These gradients guide the update steps in the gradient descent algorithm, indicating the direction and magnitude by which the parameters should be adjusted to reduce the loss.

Gradient Descent Update Rule​

The gradient descent update rules for the weights and bias are as follows, where α\alpha is the learning rate:

  • wi:=wi−α∂L∂wiw_i := w_i - \alpha \frac{\partial L}{\partial w_i}
  • b:=b−α∂L∂bb := b - \alpha \frac{\partial L}{\partial b}

With an appropriate step size, these updates seek lower loss; finite parameter convergence also depends on the data, as the separable case below shows.

Conclusion​

A linear score, sigmoid probability, and log loss form logistic regression. Gradient descent uses the compact gradients above to fit its weights. Keeping this terminology separate from the classical perceptron avoids mixing two related but distinct algorithms.

One update, and when no finite optimum exists​

For x=(2,−1)x=(2,-1), y=1y=1, and (w1,w2,b)=(0,0,0)(w_1,w_2,b)=(0,0,0), the score is z=0z=0 and the probability is 1/21/2. The gradient is (−1,1/2,−1/2)(-1,1/2,-1/2). With α=0.1\alpha=0.1, update all parameters together to (0.1,−0.05,0.05)(0.1,-0.05,0.05); then z=0.3z=0.3, y^≈0.574443\hat y\approx0.574443, and the loss drops from log⁡2≈0.693147\log2\approx0.693147 to 0.5543550.554355.

For multiple samples, average (y^i−yi)xij(\hat y_i-y_i)x_{ij} and (y^i−yi)(\hat y_i-y_i) at the current parameters before updating. Writing pi=y^ip_i=\hat y_i and including the intercept coordinate in x~i\tilde x_i, the loss Hessian is

H=1N∑ipi(1−pi)x~ix~iT⪰0.H=\frac1N\sum_i p_i(1-p_i)\tilde x_i\tilde x_i^T\succeq0.

Thus logistic regression with fixed features is convex in its parameters. Convexity alone does not guarantee a finite minimizer: on strictly linearly separable data, scaling a separating score sends every correct-class probability toward one and the loss toward zero, while weights diverge. A quadratic penalty on all parameters makes the objective coercive and strictly convex; in practice, whether the intercept is penalized must be specified. A small loss change can therefore indicate growing weights, not convergence to finite optimal parameters.

For threshold 0.50.5, predict class 11 when z≥0z\ge0; the boundary is linear in the input features. Choose a different threshold when decision costs require it, rather than changing the derivative formulas. Evaluate the loss directly from logits with the stable expression in Log Loss; do not take logs of probabilities rounded to 00 or 11.

Explore connectionsOpen network