You’re in an AI Engineer interview at Google DeepMind and the interviewer asks:
“A new engineer implements RoPE by adding a rotational embedding to the token embeddings at the bottom of the model. The training loss is flat. What fundamental misunderstanding do they have about how and where RoPE is actually applied?”
Don’t say: “The position information is…


