Skip to main content

Regression with a Linear Unit

A linear unit predicts a real number from an affine combination of inputs. This note derives its squared-loss gradients; it does not use the threshold activation or mistake-driven update of a classical perceptron.

A linear unit for regression​

Training adjusts weights and bias to minimize the chosen loss. This is exactly linear regression, not a biological model of a neuron. The textbook Dive into Deep Learning uses the same linear output and half-squared-loss convention.

Mathematical Representation​

Given inputs x1,x2,...,xnx_1, x_2, ..., x_n with corresponding weights w1,w2,...,wnw_1, w_2, ..., w_n and a bias term bb, the output y^\hat{y} of the linear unit is given by:

y^=∑i=1nwixi+b\hat{y} = \sum_{i=1}^n w_i x_i + b

This output can be used for predictions in linear regression problems, where y^\hat{y} might represent the predicted value of a dependent variable, such as the price of a house.

Loss Function​

A common regression loss is mean squared error (MSE). Here we use half the MSE to simplify derivatives:

L(y,y^)=12N∑i=1N(yi−y^i)2L(y, \hat{y}) = \frac{1}{2N} \sum_{i=1}^{N} (y_i - \hat{y}_i)^2

where yiy_i is the actual value, y^i\hat{y}_i is the predicted value, and NN is the number of samples.

Gradient Descent​

Gradient Descent Algorithm​

To minimize the loss function, gradient descent updates the parameters as follows:

  • wi(new)=wi(old)−α∂L∂wiw_i^{(new)} = w_i^{(old)} - \alpha \frac{\partial L}{\partial w_i}
  • b(new)=b(old)−α∂L∂bb^{(new)} = b^{(old)} - \alpha \frac{\partial L}{\partial b}

Where α\alpha represents the learning rate, a hyperparameter that controls the step size during the optimization process.

Derivatives Calculation​

The update rules rely on the calculation of derivatives of the loss function with respect to each parameter. These derivatives are obtained using the chain rule for differentiation. For a model with a simple quadratic loss function (L=12(y−y^)2L = \frac{1}{2}(y - \hat{y})^2), the derivatives are as follows:

Loss Function Derivative with Respect to Predictions​

dLdy^=−(y−y^)\frac{dL}{d\hat{y}} = - (y - \hat{y})

Partial Derivatives of Predictions​

  • With respect to bias (bb):
dy^db=1\frac{d\hat{y}}{db} = 1
  • With respect to weight (w1w_1):
dy^dw1=x1\frac{d\hat{y}}{dw_1} = x_1
  • With respect to weight (w2w_2):
dy^dw2=x2\frac{d\hat{y}}{dw_2} = x_2

Chain Rule Application​

The chain rule is applied to compute the gradient of the loss function with respect to each parameter:

  • For bias (bb):
dLdb=dLdy^⋅dy^db=−(y−y^)\frac{dL}{db} = \frac{dL}{d\hat{y}} \cdot \frac{d\hat{y}}{db} = - (y - \hat{y})
  • For weight (w1w_1):
dLdw1=dLdy^⋅dy^dw1=−(y−y^)⋅x1\frac{dL}{dw_1} = \frac{dL}{d\hat{y}} \cdot \frac{d\hat{y}}{dw_1} = - (y - \hat{y}) \cdot x_1
  • For weight (w2w_2):
dLdw2=dLdy^⋅dy^dw2=−(y−y^)⋅x2\frac{dL}{dw_2} = \frac{dL}{d\hat{y}} \cdot \frac{d\hat{y}}{dw_2} = - (y - \hat{y}) \cdot x_2

Update Rules​

Integrating the derivatives back into the gradient descent formula, the parameters are updated iteratively:

  • w1(new)=w1(old)−α⋅[−(y−y^)⋅x1]w_1^{(new)} = w_1^{(old)} - \alpha \cdot [ - (y - \hat{y}) \cdot x_1 ]
  • w2(new)=w2(old)−α⋅[−(y−y^)⋅x2]w_2^{(new)} = w_2^{(old)} - \alpha \cdot [ - (y - \hat{y}) \cdot x_2 ]
  • b(new)=b(old)−α⋅[−(y−y^)]b^{(new)} = b^{(old)} - \alpha \cdot [ - (y - \hat{y}) ]

Through repeated application of these updates, gradient descent aims to converge to the optimal values of w1,w2,w_1, w_2, and bb that minimize the loss function, leading to a model with minimized prediction error.

Conclusion​

The chain rule separates the loss derivative from the prediction derivative. The same pattern extends to nonlinear networks, while this model remains affine in its inputs.

From one sample to a batch​

The factor 1/21/2 makes the displayed loss half the MSE; it cancels the 22 from differentiating a square. For sample index ii and feature index jj, let ri=y^i−yir_i=\hat y_i-y_i. Then the full-batch gradients are

∂L∂wj=1N∑irixij,∂L∂b=1N∑iri.\frac{\partial L}{\partial w_j}=\frac1N\sum_i r_i x_{ij},\qquad \frac{\partial L}{\partial b}=\frac1N\sum_i r_i.

The scalar formulas above correspond to N=1N=1. For x=(2,−1)x=(2,-1), y=3y=3, and all parameters initially zero, r=−3r=-3, so (∂L/∂w1,∂L/∂w2,∂L/∂b)=(−6,3,−3)(\partial L/\partial w_1,\partial L/\partial w_2,\partial L/\partial b)=(-6,3,-3). With α=0.1\alpha=0.1, the simultaneous update gives (w1,w2,b)=(0.6,−0.3,0.3)(w_1,w_2,b)=(0.6,-0.3,0.3). The prediction becomes 1.81.8 and the half-squared loss falls from 4.54.5 to 0.720.72.

For a batch, average the gradients computed at the same parameters before updating. With fixed features, this quadratic objective is convex, but unique weights require a full-column-rank design matrix. Suitable step size is still necessary for convergence. A single linear unit cannot express arbitrary nonlinear relationships, and low training loss is not evidence of good predictions on new data.

Explore connectionsOpen network