, 6 min read
We spent three weeks blaming the prompt. It was the retriever the whole time.
On debugging RAG in production, and why the failure is almost never where you think.
The failure mode looked like a prompting problem: vague answers, confident hallucinations, inconsistent tone. It wasn't. The retriever was surfacing the wrong chunks, and no amount of prompt engineering was going to fix retrieval quality. The fix was boring: better chunking, per-document retrieval settings, and a way to see exactly what got retrieved for a given answer. Debuggability beats cleverness every time.