How a model learns by following the slope of its own error.
A gradient always points uphill.
The learning rate controls how big each step is.
∇f(x) is the direction of steepest ascent. Descent flips its sign.
Too large and you overshoot the minimum. Too small and training takes forever.