AI Interview Prep

AI Interview Prep

LLM System Design Interview #16 - The RoPE Misconception That Breaks Training

Why adding RoPE to token embeddings silently destroys positional information — and how DeepMind engineers actually apply it inside every layer.

Hao Hoang's avatar
Hao Hoang
Nov 14, 2025
∙ Paid

You’re in an AI Engineer interview at Google DeepMind and the interviewer asks:

“A new engineer implements RoPE by adding a rotational embedding to the token embeddings at the bottom of the model. The training loss is flat. What fundamental misunderstanding do they have about how and where RoPE is actually applied?”

Don’t say: “The position information is…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture