AI Interview Prep

AI Interview Prep

Machine Learning System Design Interview #14 - The Gallbladder Illusion

Why a 99.2% medical-AI model nearly killed a patient - and how the Black Patch Protocol exposes false understanding.

Hao Hoang's avatar
Hao Hoang
Dec 01, 2025
โˆ™ Paid

You are in a Senior Machine Learning Interview at Google for Health. The interviewer sets a trap:

โ€œOur research team just handed you a gallbladder segmentation model with 99.2% test set accuracy. Is it ready for production?โ€

90% of candidates walk right into the wall.

The candidate looks at the metrics and nods. They talk about verifying the F1 score on the hold-out set, setting up a canary deployment, and maybe checking inference latency on the T4 GPUs.

They assume โ€œ๐‡๐ข๐ ๐ก ๐€๐œ๐œ๐ฎ๐ซ๐š๐œ๐ฒโ€ = โ€œ๐‡๐ข๐ ๐ก ๐”๐ง๐๐ž๐ซ๐ฌ๐ญ๐š๐ง๐๐ข๐ง๐ .โ€

The interviewer stops you. โ€œWe checked all that. We deployed it. And it almost killed a patient.โ€

AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
ยฉ 2026 Hao Hoang ยท Privacy โˆ™ Terms โˆ™ Collection notice
Start your SubstackGet the app
Substack is the home for great culture