Regression with a Linear Unit
A linear unit predicts a real number from an affine combination of inputs. This note derives its squared-loss gradients; it does not use the threshold activation or mistake-driven update of a classical perceptron.
A linear unit for regression
Training adjusts weights and bias to minimize the chosen loss. This is exactly linear regression, not a biological model of a neuron. The textbook Dive into Deep Learning uses the same linear output and half-squared-loss convention.
Mathematical Representation
Given inputs with corresponding weights and a bias term , the output of the linear unit is given by:
This output can be used for predictions in linear regression problems, where might represent the predicted value of a dependent variable, such as the price of a house.
Loss Function
A common regression loss is mean squared error (MSE). Here we use half the MSE to simplify derivatives:
where is the actual value, is the predicted value, and is the number of samples.
Gradient Descent
Gradient Descent Algorithm
To minimize the loss function, gradient descent updates the parameters as follows:
Where represents the learning rate, a hyperparameter that controls the step size during the optimization process.
Derivatives Calculation
The update rules rely on the calculation of derivatives of the loss function with respect to each parameter. These derivatives are obtained using the chain rule for differentiation. For a model with a simple quadratic loss function (), the derivatives are as follows:
Loss Function Derivative with Respect to Predictions
Partial Derivatives of Predictions
- With respect to bias ():
- With respect to weight ():
- With respect to weight ():
Chain Rule Application
The chain rule is applied to compute the gradient of the loss function with respect to each parameter:
- For bias ():
- For weight ():
- For weight ():
Update Rules
Integrating the derivatives back into the gradient descent formula, the parameters are updated iteratively:
Through repeated application of these updates, gradient descent aims to converge to the optimal values of and that minimize the loss function, leading to a model with minimized prediction error.
Conclusion
The chain rule separates the loss derivative from the prediction derivative. The same pattern extends to nonlinear networks, while this model remains affine in its inputs.
From one sample to a batch
The factor makes the displayed loss half the MSE; it cancels the from differentiating a square. For sample index and feature index , let . Then the full-batch gradients are
The scalar formulas above correspond to . For , , and all parameters initially zero, , so . With , the simultaneous update gives . The prediction becomes and the half-squared loss falls from to .
For a batch, average the gradients computed at the same parameters before updating. With fixed features, this quadratic objective is convex, but unique weights require a full-column-rank design matrix. Suitable step size is still necessary for convergence. A single linear unit cannot express arbitrary nonlinear relationships, and low training loss is not evidence of good predictions on new data.