You’re in an ML Engineer interview at Google and the interviewer asks:
“We need to serve our model for two different use cases: a low-latency chatbot that needs a fast 𝐓𝐢𝐦𝐞-𝐭𝐨-𝐅𝐢𝐫𝐬𝐭-𝐓𝐨𝐤𝐞𝐧 (𝐓𝐓𝐅𝐓), and a high-throughput batch summarization job. How do these two workloads stress the GPU differently, and what fundamental tradeoff are you …


