AI Interview Prep

AI Interview Prep

Advanced NLP Interview Questions #2 – The Gradient Shockwave Trap

How a single random layer can silently erase millions of dollars of pretraining, and how to stop it.

Hao Hoang's avatar
Hao Hoang
Dec 09, 2025
∙ Paid

You’re in a Senior ML Interview at NVIDIA. The interviewer sets a trap:

“You attach a new, random linear head to a pre-trained Transformer. Do you unfreeze all layers and start backprop immediately?”

90% of candidates walk right into the trap.

Their answer is: “Of course. End-to-end training allows the backbone to adapt to the new task immediately. If we have the compute, why artificially restrict the model by freezing layers?”

It feels efficient. It works in the tutorials.

-----

𝐓𝐡𝐞 𝐑𝐞𝐚𝐥𝐢𝐭𝐲: They aren’t accounting for 𝐓𝐡𝐞 𝐆𝐫𝐚𝐝𝐢𝐞𝐧𝐭 𝐒𝐡𝐨𝐜𝐤𝐰𝐚𝐯𝐞.

AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture