Machine Learning System Design Interview #14 - The Gallbladder Illusion
Why a 99.2% medical-AI model nearly killed a patient - and how the Black Patch Protocol exposes false understanding.
You are in a Senior Machine Learning Interview at Google for Health. The interviewer sets a trap:
โOur research team just handed you a gallbladder segmentation model with 99.2% test set accuracy. Is it ready for production?โ
90% of candidates walk right into the wall.
The candidate looks at the metrics and nods. They talk about verifying the F1 score on the hold-out set, setting up a canary deployment, and maybe checking inference latency on the T4 GPUs.
They assume โ๐๐ข๐ ๐ก ๐๐๐๐ฎ๐ซ๐๐๐ฒโ = โ๐๐ข๐ ๐ก ๐๐ง๐๐๐ซ๐ฌ๐ญ๐๐ง๐๐ข๐ง๐ .โ
The interviewer stops you. โWe checked all that. We deployed it. And it almost killed a patient.โ


