AI Interview Prep

AI Interview Prep

LLM System Design Interview #7 - Why Your 1B → 70B Training Exploded

Most engineers blame learning rate. Elite engineers blame parameterization - and fix scaling with MuP instead of wasting millions on sweeps.

Hao Hoang's avatar
Hao Hoang
Nov 05, 2025
∙ Paid

You’re in an AI Engineer interview at Google DeepMind and the interviewer asks:

“Your 1B parameter proxy model trains perfectly with a 1.2e-4 learning rate. You scale the model to 70B, and the training immediately explodes. What’s the most 𝘭𝘪𝘬𝘦𝘭𝘺 reason and how do you fix it 𝐰𝐢𝐭𝐡𝐨𝐮𝐭 running a new, expensive hyperparameter sweep?”

Most candid…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture