You’re in a Senior AI Engineer interview at Google DeepMind, and the interviewer asks:
“We’ve implemented the original DeepSeek GRPO paper to train our new math chatbot. On uncertain queries, the Chain-of-Thought (CoT) is suddenly exploding to 10000 tokens. An engineer on the team says this is great, the model is just thinking harder and learning to back…


