AI Interview Prep

AI Interview Prep

Generative Vision Interview Questions #16 - The Mid-Noise Paradox

Why 2x-ing your training compute won't fix blurry images, and how switching to a logit-normal curriculum forces your model to fight the structural war where it actually matters.

Hao Hoang's avatar
Hao Hoang
Jun 24, 2026
∙ Paid

You’re in a Senior ML Engineer interview at Midjourney and the interviewer asks:

“Your DiT trains beautifully at 512×512. You bump inference to 1024×1024 and it generates garbage, warped anatomy, repeated limbs, a teddy bear with three faces. Before you touch the VAE or the sampler, where do you look first?”

Don’t say: “The model didn’t see high-res data,…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture