AI Interview Prep

AI Interview Prep

Generative Vision Interview Questions #10 - The Receptive Field Illusion

Why stacking convolutions to build global context quietly turns distant features into mush, and how trading inductive bias for self-attention secures true edge-to-edge coherence.

Hao Hoang's avatar
Hao Hoang
Jun 18, 2026
∙ Paid

You’re in a Senior ML Engineer interview at Midjourney and the interviewer asks:

“Your teammate wants to ship a U-Net for your new high-res image model because ‘convolutions capture both local and global features.’ You disagree. Defend it.”

Don’t say: “Transformers are just better, everyone uses DiT now.”

Too lazy. They just failed the question.

Here’s the…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture