RAG Interview Questions #8 - The IVF-PQ Compression Trick
Why blindly scaling HNSW is a scaling nightmare, and how to gracefully trade imperceptible coverage loss for massive infrastructure savings.
You’re in a Staff ML Engineer interview at Google and the interviewer asks:
“Your RAG system hits 99% recall on HNSW in the demo. Now it’s 200M vectors and your RAM bill just triggered a budget review. Do you keep HNSW? Why or why not?”
Don’t say: “HNSW is faster and more accurate, so we keep it and add more RAM.”
That answer just told them you’ve never ru…


