AI Interview Prep

AI Interview Prep

Advanced NLP Interview Questions #1 - The Learning Rate Warm-Up Trap

The hidden physics of why Transformers can't handle a 1e-3 LR at initialization.

Hao Hoang's avatar
Hao Hoang
Dec 09, 2025
โˆ™ Paid

Youโ€™re in a Senior ML Interview at Google DeepMind. The interviewer hands you a marker and sets the trap:

โ€œWe are training a ๐˜›๐˜ณ๐˜ข๐˜ฏ๐˜ด๐˜ง๐˜ฐ๐˜ณ๐˜ฎ๐˜ฆ๐˜ณ from scratch using ๐˜ˆ๐˜ฅ๐˜ข๐˜ฎ. We set a constant Learning Rate of 1e-3. Predict the first 1000 steps.โ€

90% of candidates walk right into the trap.

Most candidates say: โ€œIt converges. ๐˜ˆ๐˜ฅ๐˜ข๐˜ฎ is an adaptive optimizer, it adjusts per-parameter learning rates automatically. 1e-3 is a standard default. It might be noisy at first, but the loss will go down.โ€

The reality? Your loss curve doesnโ€™t just oscillate, it explodes. You hit NaNs within 50 steps. You just wasted a cluster run.

Here is the physics of why: ๐˜›๐˜ณ๐˜ข๐˜ฏ๐˜ด๐˜ง๐˜ฐ๐˜ณ๐˜ฎ๐˜ฆ๐˜ณ๐˜ด, unlike ๐˜™๐˜ฆ๐˜ด๐˜•๐˜ฆ๐˜ต๐˜ด, lack strong inductive biases at initialization.


AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
ยฉ 2026 Hao Hoang ยท Privacy โˆ™ Terms โˆ™ Collection notice
Start your SubstackGet the app
Substack is the home for great culture