Gradient Descent: training is just rolling downhill

The curve is the model's error at each setting of one dial. Training measures the slope, steps the dial the other way, repeats. Click anywhere on the curve to drop the ball, pick a step size, and watch where “learning” actually lands.

Click the curve to drop the ball anywhere. There are two valleys: the shallow one on the right and the deep one on the left. The ball does not know which is which — it only ever checks the slope under its feet.

position (w)
--
error (loss)
--
steps taken
0
last step size
--

Controls

The wait, what?

Nothing in that animation understood anything. The ball checks one number — the slope under its feet — and steps the other way. That is the entire training loop of modern AI: measure the error, step against the slope, millions of times, on millions of dials at once.

And notice the two ways it fails. Drop the ball on the right slope and it settles happily in the shallow valley — a real answer, just not the best one; the ball has no idea a deeper valley exists. Crank the learning rate and it bounces straight out of the valley or flies off the board entirely.

Where training lands is an accident of where it started and how big a step it takes — not of the model knowing the right answer. Same curve, same rules, different landing every time.