LLM System Design Interview #20 - Why Raw User Data Will Always Fail Fine-Tuning
The goal isn’t to model the average user - it’s to model the valuable one.
You’re in a Senior ML Engineer interview at Perplexity and the interviewer asks:
“Your PM wants to fine-tune our new model on a random 1M sample of live user prompts to improve real-world performance. You tell them it’s a terrible idea. Why?”
Most candidates say: “Because the data is noisy, has PII, and needs to be cleaned.”
This is a weak answer. It’s tru…


