AI Interview Prep

AI Interview Prep

LLM System Design Interview #21 - The GRPO Length Trap

How a "good idea" in the DeepSeek objective secretly incentivizes 10000-token failures - and bankrupts your inference budget.

Hao Hoang's avatar
Hao Hoang
Nov 17, 2025
∙ Paid

You’re in a Senior AI Engineer interview at Google DeepMind, and the interviewer asks:

“We’ve implemented the original DeepSeek GRPO paper to train our new math chatbot. On uncertain queries, the Chain-of-Thought (CoT) is suddenly exploding to 10000 tokens. An engineer on the team says this is great, the model is just thinking harder and learning to back…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture