What does the learning rate control in gradient descent?
- A The size of the step taken towards the loss minimum on each update
- B The number of training examples used
- C The number of layers in the network
- D The regularisation strength
Answer
The size of the step taken towards the loss minimum on each update
Too high and training diverges or oscillates; too low and it converges slowly or stalls in a poor region.





