AI Interview Prep

AI Interview Prep

LLM Inference Interview Questions #6 - The Multimodal Perception Trap

Why dropping frontier VLMs onto dense web UIs is a hidden trap, and how foundation labs use cheap synthetic data to turn a prompting party trick into reliable execution.

Hao Hoang's avatar
Hao Hoang
Aug 04, 2026
∙ Paid

You’re in a Senior AI Engineer interview at Google and the interviewer asks:

“Your web agent uses set-of-marks, you screenshot the page, draw numbered boxes on every element, and let the VLM click by number. It works in your demo but in production the model keeps clicking box 41 when it meant box 14, and it ignores half the page. The VLM is state-of-the-…

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture