Advanced NLP Interview Questions #1 - The Learning Rate Warm-Up Trap
The hidden physics of why Transformers can't handle a 1e-3 LR at initialization.
Youโre in a Senior ML Interview at Google DeepMind. The interviewer hands you a marker and sets the trap:
โWe are training a ๐๐ณ๐ข๐ฏ๐ด๐ง๐ฐ๐ณ๐ฎ๐ฆ๐ณ from scratch using ๐๐ฅ๐ข๐ฎ. We set a constant Learning Rate of 1e-3. Predict the first 1000 steps.โ
90% of candidates walk right into the trap.
Most candidates say: โIt converges. ๐๐ฅ๐ข๐ฎ is an adaptive optimizer, it adjusts per-parameter learning rates automatically. 1e-3 is a standard default. It might be noisy at first, but the loss will go down.โ
The reality? Your loss curve doesnโt just oscillate, it explodes. You hit NaNs within 50 steps. You just wasted a cluster run.
Here is the physics of why: ๐๐ณ๐ข๐ฏ๐ด๐ง๐ฐ๐ณ๐ฎ๐ฆ๐ณ๐ด, unlike ๐๐ฆ๐ด๐๐ฆ๐ต๐ด, lack strong inductive biases at initialization.


