You’re in a Senior AI Engineer interview at Meta and the interviewer asks:
“You’re A/B testing two 70B models - one Multi-Head Attention (MHA), one Grouped Query Attention (GQA). Your colleague argues they’ll have the same inference speed since FLOPs and parameter counts are identical. Is this assumption correct?”


