AI Interview Prep

AI Interview Prep

Advanced NLP Interview Questions #21 – The PPO vs DPO Implementation Trap

Why replacing PPO with DPO is not a free lunch - and how gradient saturation turns “simpler” into “weaker.”

Hao Hoang's avatar
Hao Hoang
Dec 27, 2025
∙ Paid

You are in a Senior AI Interview at OpenAI. The interviewer sets a trap:

“Our engineers want to rip out 𝘗𝘗𝘖 (𝘗𝘳𝘰𝘹𝘪𝘮𝘢𝘭 𝘗𝘰𝘭𝘪𝘤𝘺 𝘖𝘱𝘵𝘪𝘮𝘪𝘻𝘢𝘵𝘪𝘰𝘯) and replace it with 𝘋𝘗𝘖 (𝘋𝘪𝘳𝘦𝘤𝘵 𝘗𝘳𝘦𝘧𝘦𝘳𝘦𝘯𝘤𝘦 𝘖𝘱𝘵𝘪𝘮𝘪𝘻𝘢𝘵𝘪𝘰𝘯). They argue it’s strictly better because it simplifies the stack. Do we approve the PR?”

90% of candidates walk right into it.

They say “Yes, absolutely. PPO is unstable and requires maintaining a separate 𝐑𝐞𝐰𝐚𝐫𝐝 𝐌𝐨𝐝𝐞𝐥 (𝐑𝐌) and 𝐕𝐚𝐥𝐮𝐞 𝐇𝐞𝐚𝐝. DPO optimizes the same objective analytically without the extra inference overhead. It’s a free lunch.”

They just fell for the “𝐈𝐦𝐩𝐥𝐞𝐦𝐞𝐧𝐭𝐚𝐭𝐢𝐨𝐧 𝐅𝐚𝐥𝐥𝐚𝐜𝐲”. They are optimizing for engineering convenience, not mathematical reality.

AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Keep reading with a 7-day free trial

Subscribe to AI Interview Prep to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture