AI Interview Prep

AI Interview Prep

Advanced NLP Interview Questions #12 – The Optimizer State Memory Trap

Why switching from SGD to Adam can instantly triple your VRAM usage and crash a 7B training run.

Hao Hoang's avatar
Hao Hoang
Dec 18, 2025
∙ Paid

You are in a Senior Machine Learning Engineer interview at Google DeepMind and the interviewer asks:

“You just switched a 7B parameter training run from SGD to Adam to speed up convergence. The model size is identical, but the cluster immediately crashes with a 𝘊𝘜𝘋𝘈 𝘖𝘶𝘵-𝘖𝘧-𝘔𝘦𝘮𝘰𝘳𝘺 (𝘖𝘖𝘔) error. Why?”

🚫 Don’t say:

“Adam is computationally more expensive so it uses more memory,” or “I probably need to lower the batch size.”

That’s a junior guess. It ignores the mechanics of the optimizer.

The reality is that Adam isn’t just an algorithm, it’s a VRAM glutton. The candidates fell into the 𝐎𝐩𝐭𝐢𝐦𝐢𝐳𝐞𝐫 𝐒𝐭𝐚𝐭𝐞 𝐓𝐫𝐚𝐩.

Unlike SGD, which is stateless, Adam maintains two additional scalar states for every single parameter in your network to track the learning trajectory:

- 𝘔𝘰𝘮𝘦𝘯𝘵𝘶𝘮 (𝘍𝘪𝘳𝘴𝘵 𝘔𝘰𝘮𝘦𝘯𝘵)

- 𝘝𝘢𝘳𝘪𝘢𝘯𝘤𝘦 (𝘚𝘦𝘤𝘰𝘯𝘥 𝘔𝘰𝘮𝘦𝘯𝘵)

Here is the “3𝐱 𝐏𝐚𝐫𝐚𝐦𝐞𝐭𝐞𝐫 𝐑𝐮𝐥𝐞” strictly for memory planning:

- 𝘚𝘎𝘋: 𝘙𝘦𝘲𝘶𝘪𝘳𝘦𝘴 𝘔𝘦𝘮𝘰𝘳𝘺 ≈ 𝘞𝘦𝘪𝘨𝘩𝘵𝘴 + 𝘎𝘳𝘢𝘥𝘪𝘦𝘯𝘵𝘴.

- 𝘈𝘥𝘢𝘮: 𝘙𝘦𝘲𝘶𝘪𝘳𝘦𝘴 𝘔𝘦𝘮𝘰𝘳𝘺 ≈ 𝘞𝘦𝘪𝘨𝘩𝘵𝘴 + 𝘎𝘳𝘢𝘥𝘪𝘦𝘯𝘵𝘴 + 𝘔𝘰𝘮𝘦𝘯𝘵𝘶𝘮 + 𝘝𝘢𝘳𝘪𝘢𝘯𝘤𝘦.

For a 7B model in FP16 (2 bytes/param), your weights are ~14GB. But Adam demands an additional ~28GB just to store those optimizer states.

AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Keep reading with a 7-day free trial

Subscribe to AI Interview Prep to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture