AI Interview Prep

AI Interview Prep

LLM System Design Interview #14 - The Two Faces of Inference

Why your chatbot’s latency and your batch job’s throughput fight over different GPU limits - and how to balance compute-bound and memory-bound workloads.

Hao Hoang's avatar
Hao Hoang
Nov 11, 2025
∙ Paid

You’re in an ML Engineer interview at Google and the interviewer asks:

“We need to serve our model for two different use cases: a low-latency chatbot that needs a fast 𝐓𝐢𝐦𝐞-𝐭𝐨-𝐅𝐢𝐫𝐬𝐭-𝐓𝐨𝐤𝐞𝐧 (𝐓𝐓𝐅𝐓), and a high-throughput batch summarization job. How do these two workloads stress the GPU differently, and what fundamental tradeoff are you …

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture