π The RAG Interview (Official Release) + Free Part V
A deep dive into vector index economics, the exact reasoning expected when a retrieval interviewer changes one constraint on you.
Todayβs the day. The RAG Interview is officially live.
Last time I gave away a chapter. This time itβs a whole part, Part V, all three chapters, about 80 pages. No email, no strings.
Hereβs what Part V covers:
Indexing and Vector Search. Why exact search dies at high dimension, derived rather than asserted. HNSW built up from a skip list. IVF, LSH, DiskANN. Product quantization, and the recall/latency/memory triangle youβre choosing an index inside of.
Then the part that decides interviews: index memory arithmetic, end to end.
If you read RAG Interview Questions #8 on IVF-PQ compression, or #11 on tombstones and the zombie hub problem, Part V is where both arguments are worked out in full.
Read it. If itβs useful, the full book picks up from there.
Whatβs in the full book:
β The generator side: what the model memorized, and the five causes of factual error
β Chunking, representation, and whether semantic chunking actually pays on real data
β The full ranking progression: BM25, learning-to-rank, dense retrieval, ColBERT, reranking
β Routing, HyDE, and the agentic loops that decide how many times to retrieve
β Evaluation across three layers, plus trust: credibility, prompt injection, datastore poisoning
β 7 end-to-end design drills, including a system that got worse with no obvious reason why
β 123 practice questions at core, senior, and staff tier, answers held separately
β A formula appendix worth printing out
1,141 pages, 41 chapters, 249 sections. Every number derived in front of you or attributed to a named source. Free updates for life.
RAG25 takes 20% off for the first 200 readers.
If youβve been reading the RAG Interview Questions series this summer, thank you. This is the long version of that whole argument. Reply and let me know what you think. I read every one.
Hao



