AI Interview Prep

AI Interview Prep

LLM System Design Interview #9 - The Memory Wall

Why attention isn’t too slow - it’s too hungry. How FlashAttention wins by trading O(N²) memory I/O for O(N²) fast compute.

Hao Hoang's avatar
Hao Hoang
Nov 07, 2025
∙ Paid

You’re in a Staff ML Engineer interview at a OpenAI and the interviewer asks:

“Your team needs to support a 10M context window. An engineer says it’s impossible because standard attention is O(N²) compute. Why is that the wrong bottleneck to focus on, and how does FlashAttention actually solve the real problem?”

Thanks for reading AI Interview Prep! Subsc…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture