A Software Engineer's Intuition for Gradient Descent, Part 2: The Step Size
In Part 1 we built gradient descent around one update rule: $$\theta := \theta - \alpha \nabla f(\theta)$$ Almost all of that article was about the gradient ∇f(θ), the part that tells us which way is
Sep 24, 202612 min read10


