AI Interview Prep

AI Interview Prep

LLM System Design Interview #15 - The FLOPs Fallacy

Why two 70B models with identical compute still run at different speeds - and how GQA wins by breaking the memory bottleneck.

Hao Hoang's avatar
Hao Hoang
Nov 12, 2025
∙ Paid

You’re in a Senior AI Engineer interview at Meta and the interviewer asks:

“You’re A/B testing two 70B models - one Multi-Head Attention (MHA), one Grouped Query Attention (GQA). Your colleague argues they’ll have the same inference speed since FLOPs and parameter counts are identical. Is this assumption correct?”

Thanks for reading AI Interview Prep! Sub…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture