AI Interview Prep

AI Interview Prep

Advanced NLP Interview Questions #13 – The Knowledge Distillation Trap

Why training on ensemble hard targets throws away the most valuable signal—and how dark knowledge actually transfers intelligence.

Hao Hoang's avatar
Hao Hoang
Dec 19, 2025
∙ Paid

You are in a Senior AI Interview at OpenAI. The interviewer sets a quiet trap:

“We need to distill a massive 10-model ensemble into a single small model for low-latency serving. Why is training the student on the ensemble’s final output tokens a complete waste of compute?”

90% of candidates walk right into it.

Most candidates say: “It’s not a waste. The ensemble is the ‘teacher.’ If the ensemble predicts Cat with 99% confidence, we should treat Cat as the ground truth and train the student to predict it using standard Cross-Entropy Loss.”

Technically correct. Practically useless.

Here is the mental shift: If you only train on the “Hard Target” (the winner), you are throwing away the single most valuable asset the ensemble created.

AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Keep reading with a 7-day free trial

Subscribe to AI Interview Prep to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture