Gradient Descent in Two Variables
Gradient descent uses the slopes in both coordinate directions to search for a minimum of . A concrete step shows how the two coordinates change together.
Practical Example
Consider the illustrative room-temperature function
This is a different model from the preceding sauna example, and it is not a calibrated physical temperature model. The bounded room matters: without the domain restriction, the polynomial is unbounded below, so there is no coolest point to find.
First gradient step
Start at with learning rate . To calculate the partial derivatives, let . Then and , giving the gradient
At the starting point, . Subtract times each component:
Both components must be computed from the old point before updating either variable. The new point is still inside the room, and falls from approximately to .
Conceptual Overview
The gradient collects the two partial derivatives. At a point where it is nonzero, it points in the direction of greatest instantaneous increase; its magnitude is that maximal rate per unit distance. Gradient descent moves in the opposite direction, the direction of steepest local decrease. This is a local direction, not a guarantee that any chosen finite step decreases the function or reaches a minimum.
Mathematical Formulation
For the current point , the update is
The learning rate scales both components of the displacement. A large step can overshoot a minimum; a very small one can make convergence slow.
Algorithm for Two Variables
Choose an initial point , compute its gradient, and apply the update repeatedly. For a box-constrained room, keep every new point feasible by projecting the update:
where clips each coordinate into . Monitor the objective decrease and the projected-gradient residual
Set an iteration cap as well. Small changes between iterates alone do not prove a minimum. For the box constraint, checking only the raw gradient is insufficient: a boundary optimum need not have zero gradient. Even a small projected residual is only a stationarity test for this nonconvex problem.
Challenges and Considerations
For this particular model, the minimum can also be found exactly. On , attains its unique maximum at . Maximizing each subtracted term gives the global minimum with . By contrast, starting at gives zero gradient and never reaches that minimum. In general, gradient descent may converge to a local rather than a global minimum; trying different starting points can reduce this risk but does not prove global optimality.
At , the Hessian is . Linearizing the update gives the sufficient local contraction range . Thus is close to the stability limit in but moves slowly in . From the stated start, 1000 steps give approximately ; this empirical run does not prove convergence from every start.
In Distill’s interactive explanation of momentum, compare plain gradient descent with momentum on elongated quadratic contours. Changing the step size helps explain why steep and shallow directions progress differently, as in the Hessian calculation here; the demo uses a separate objective.