Skip to content

Retrieval Is Where RAG Fails: Lessons from a Payments Assistant and a Context-Allocation Paper

#retrieval-augmented-generation #rag #hallucination #llm-evaluation #search #llm-engineering

The model wasn't the liar ​

I almost shipped a RAG assistant that invented a payment API. Not a hedged "I'm not sure," but a complete authentication flow with real-looking endpoints, real-looking headers, and a cited source URL. The URL wasn't in the corpus. It wasn't anywhere. The model fabricated the citation for content it also fabricated, with zero hedging.

The reflex is to blame the model. I did too, at first. Swap in a bigger model, tune the prompt, add "only answer from the context provided" in bold. The answers didn't improve, because the failure was upstream. A production RAG checklist published this year cites industry analysis landing on the same number: when RAG fails, the failure is in retrieval roughly 73% of the time, not generation. The LLM gets blamed for a mistake that happened several steps before it saw a token.

This article threads together a debugging story, a production checklist, and a context-allocation paper. The through-line is simple: retrieval is the link that breaks, and most of the fixes are cheap.

Similarity scores lie exactly when you need them ​

My first theory was retrieval confidence. Set a similarity threshold, refuse to answer below it, done. I checked the actual numbers before writing that fix.

CaseTop-1 similarityWhat happened
Correct in-corpus answer0.718Correct answer
Worst fabrication (Interswitch)0.712Fully invented, fake citation
Correct decline (out-of-domain)0.691"Not in my knowledge base"

The worst hallucination had higher retrieval similarity than the cleanest correct decline. There is no threshold that lets the good case through and blocks the bad one. A confidence cutoff would have been a fix that felt right and did nothing.

A new arXiv paper calls this the "diagnostic illusion." Standard relevance proxies, cosine similarity included, fail catastrophically on hard negatives. The paper's replacement is a causal leave-one-out probe: drop each retrieved chunk from the context, measure how the generation shifts, and you get a direct read on what the model actually relies on. Similarity tells you what looks relevant. The probe tells you what the model used.

That distinction matters because the two diverge exactly in the cases that hurt: when a retrieved chunk is topically similar but factually wrong for the question.

Same-domain substitution: the bug that looks like the model ​

The chunks my retrieval pulled back for "Interswitch Quickteller" were real. They were Monnify's quickstart and Paystack's accept-payments guide. Similar topic: authentication, checkout, webhooks. The model wasn't confused about the domain. It never checked whether the retrieved text actually named the provider I asked about, versus a different provider talking about something similar.

That's sneakier than "doesn't know when it doesn't know." It's "knows something adjacent and doesn't notice the adjacency."

My first fix was one paragraph in the system prompt: before answering, check whether the specific provider named in the question is actually named in the context excerpts. Retrieval is similarity-based and will sometimes hand you excerpts from a different provider just because the topic is similar. If it isn't named, say so. Re-ran the five failing prompts: five for five clean. I shipped it.

Then an independent check re-ran the same five prompts fifteen times, three runs each. Ten of fifteen came back clean. Not five of five. Two-thirds.

And the failures weren't random. They split cleanly in two. Interswitch, Paga, OPay: nine for nine, 100% reliable. Kuda and PalmPay: one of three, zero of three. PalmPay fabricated an x-palmpay-signature header and a full HMAC handler on every single run. Kuda got silently rerouted to Paystack's live charge endpoint, with an invented bank code stated as fact, in two of three.

Two things were true, and I'd only checked one. First, the chat call runs at temperature 0.2 with no fixed seed, so the same prompt doesn't reliably produce the same answer. My original five for five was one draw from a distribution, not a property of the fix. Second, the failure was concentrated exactly where retrieval is most ambiguous. Kuda and PalmPay's webhook-verification content is topically near-identical to Paystack's and Monnify's, same HMAC-SHA512 shape, same header pattern. Their retrieved chunks sit in the tightest, most confusable similarity band I measured, cosine 0.654 to 0.676, five chunks within 0.022 of each other. The instruction I wrote asks the model to notice when a retrieved chunk doesn't actually name the asked-about provider. It's least able to notice that exactly when the retrieved chunk is close enough to look plausible.

A soft instruction was never going to close that gap reliably. The thing it fights, retrieval similarity between near-duplicate topics, doesn't go away because you asked nicely.

The fix that worked: stop asking, start checking ​

The corpus only covers four providers. That's a small, enumerable set. So the question "is this provider actually in scope" doesn't need an LLM's judgment at all. I added a deterministic gate ahead of retrieval: a list of about 25 known African fintech and banking brands that are not in the corpus, matched by word boundary against the incoming question. Name one without also naming an in-corpus provider, and the question gets declined before retrieval or generation ever runs. No temperature, no seed, no chance to fabricate. Just a string match.

Same fifteen trials, freshly re-cloned: fifteen for fifteen, every response near-instant, no LLM call at all. Regression checks stayed clean. An in-corpus question still runs the full pipeline untouched, and a comparison question like "how does Kuda compare to Paystack for webhook handling?" correctly falls through to the softer instruction instead of getting blanket-refused.

The gate has a weakness, and I found it when I tested a question I hadn't written: "how do I implement BVN verification for a customer with a Wema Bank account?" was declining, and it shouldn't have. BVN verification is covered. Wema is the customer's bank, not the API being asked about. The gate asks "does the question contain a listed name" when the real question is "which provider's API is this question about."

There's a second trap underneath: two sources of truth. The corpus says what's in scope, the denylist says what's out, and nothing checks they stay disjoint. The day you scrape Kuda into the corpus and forget to remove it from the list, the gate will decline questions about a provider you actually cover. A false decline looks byte-for-byte identical to a correct one, no error message anywhere. The fix is to derive the in-scope provider list from the corpus at index time, since the corpus is already the authority on what's covered, and run a control-pair test on every index build: one in-corpus prompt must pass through to retrieval, one listed-out provider must decline.

The two stranger-written prompts that found the original bug were one-time runs, never wired into a regression suite. That's a gap. Found prompts are gold precisely because you couldn't have written them.

The gate doesn't generalize. A provider I didn't think to enumerate still depends on the soft instruction that measured 100% for three providers and 0 to 33% for two. That's a limitation, not a solved problem, and it's written down in the repo instead of implied away.

Quick Take: Retrieval similarity measures what looks relevant, not what the model actually used, so fix the retrieval, not the prompt.

The checklist: where pipelines quietly break ​

The retrieval checklist's core claim matches my experience: naive RAG, chunk, embed, cosine similarity, stuff into the prompt, was always a prototype. Production retrieval is a chain, and it can break at any link.

Two structural fixes matter most. First, decouple the indexing path from the query path. If re-indexing forces live search offline, you stop iterating on chunking or embedding, and a frozen pipeline is a stale pipeline. Second, chunk so each piece stands alone. Fixed-size splitting cuts sentences mid-thought, tables mid-row, code mid-function. The retrieved chunk looks relevant and is missing the half that mattered. Structure-aware splitting on the document's own boundaries, or semantic chunking that starts a new chunk where meaning shifts, both beat blind fixed-size splitting.

The single most common retrieval mistake is vector-only search. Pure vector search is great at meaning: ask "how do I fix login problems" and it surfaces chunks about authentication even when none use the word "login." But it falls apart on anything exact. An error code like ERR_SSL_PROTOCOL_ERROR, a SKU, a function name. Semantic similarity is meaningless for a serial number. The fix is hybrid search: BM25 plus dense embeddings, fused with reciprocal rank fusion. The consensus across BEIR and MTEB is blunt: hybrid beats either one alone on basically every public benchmark.

Pipeline stageNaive defaultProduction fixPayoff
ChunkingFixed-size, 1000 charsStructure-aware or semanticChunks answer questions on their own
EmbeddingRaw body textPrepend heading and summary contextVectors align with real queries
SearchVector onlyBM25 + dense + RRFHandles exact matches and meaning
RankingFirst-pass orderCross-encoder rerank5 to 15 MRR points, up to 3x nDCG
ContextAs many chunks as fitTop 3 to 8, strongest at edgesAvoids "lost in the middle"

Reranking and context assembly ​

The mistake that cost me the most quality for the least obvious reason: I assumed that if the right chunk was retrieved, the model would use it. Where it lands in the list matters enormously.

Vector search uses a bi-encoder. It encodes the query and each chunk separately and compares vectors. Fast, but it trades away fine-grained relevance. The best chunk often gets retrieved at position 8, buried under seven "pretty relevant" ones. Models demonstrably ignore information stranded in the middle of a long list. The right answer is in the context and the model still misses it.

Reranking fixes this. Retrieve a broad candidate set with hybrid search, top 20 to 50, then run a cross-encoder reranker that scores each query-chunk pair jointly. Keep the top 3 to 8. The impact is large, not marginal: a cross-encoder reranker commonly adds 5 to 15 points of MRR on hard sets, and on some reasoning-heavy benchmarks reranking pushed nDCG@10 from about 0.13 to 0.40. Roughly 3x, just from reordering the same candidates you already retrieved.

The recipe that beats most production deployments: retrieve about 20 via hybrid search, rerank to about 5, send 3 to 5 to the LLM. Reranking 100+ candidates rarely pays. The signal lives at the head.

Context assembly matters too. Order: put the strongest chunks at the very start and end, not buried in the center. Volume: more chunks is not better. Stuffing 30 chunks in "to be safe" dilutes the signal and invites the model to average across noise. Citations: ask the model to cite which chunk supports each claim. That discourages free-floating fabrication and gives you a way to verify the answer against its sources.

Long-context windows are not a substitute for retrieval. Frontier models have million-token windows now, and the reflex is "just dump everything in." Resist it. Dumping the whole corpus is slower, more expensive, and less accurate, because the model still has to find the needle. Use the big window for synthesis across long documents, not as a retrieval replacement.

Agentic RAG, retrieval inside a reasoning loop, is for questions that need multiple hops. It is not a fix for a broken basic pipeline. If your chunking is bad and you have no reranking, an agent will just make bad retrieval calls, repeatedly, more expensively.

Context allocation: monolithic widening is a trap ​

The arXiv paper "The Laws of Context Allocation" formalizes what the checklist hints at. The prevailing strategy of monolithic context widening, adding more and more chunks to a single generation, is an architectural trap penalized by relevance decay. The paper proves this with a deconfounded factorial grid and shows an alternative: allocate compute iteratively across multiple sequential generations. That shift drives portfolio recall gains of 16.7 to 20.5 absolute percentage points, scaling up to 32B models.

The mechanism is a closed-loop submodular scheduler. It uses the causal probe from earlier to decide what evidence the model actually relied on, then decides what to retrieve next. An attribution-steered contrastive decoder overrides what the paper calls "LLM attention inertia": the tendency to keep leaning on the same evidence even when it's exhausted. The system forces fresh evidence integration instead.

This is the formal version of "fewer, better chunks beat more chunks." The checklist says send 3 to 5 chunks that earned their place. The paper shows why: relevance decays as you widen the context, and sequential, feedback-driven retrieval beats a single monolithic pass.

Key numbers73% of RAG failures trace to retrieval, not generation. 0.712 similarity for the worst fabrication, 0.691 for the cleanest correct decline. No threshold separates them. 10 of 15 trials passed after the soft-instruction fix. One clean run was a coin flip. 16.7 to 20.5 absolute percentage points of recall gain from sequential generation over monolithic context widening.

RAG for classification, not just Q&A ​

RAG isn't only for question answering. A recent paper applies it to multi-label classification of environmental mitigation obligations in hydropower licensing documents: 5,860 paragraphs, a structured 135-category taxonomy, and severe label scarcity. 40 of 135 categories have no training examples. 26 have fewer than five.

A supervised BERT-based pipeline works on well-represented categories but achieves F1 of zero on unseen classes, regardless of augmentation strategy. The RAG pipeline conditions classification on retrieved category definitions, which enables zero-shot generalization across the full label space. The hybrid system combines both: BERT detection for its high recall on seen categories, RAG classification for zero-shot coverage of the long tail.

The hybrid reaches Micro F1 of 0.524, beating the BERT-only pipeline at 0.477 and the RAG-only pipeline at 0.416 across every training-support bucket. The practical read: when your label space has a long tail, retrieve the definitions at inference time instead of hoping the model memorized them. Same principle as the payments assistant. Re-inject the drowned-out content at query time rather than hoping it survived pretraining.

The second read path: authorization for replayed citations ​

There's a failure mode that has nothing to do with retrieval quality, and it's easy to miss. My copilot persists the source cards it cites: which documents backed each answer, scores, names. That's table stakes for a trustworthy RAG product. An answer without its evidence is just vibes.

The question that changed how I shipped it: six months from now, a user opens that old conversation and the cards render again. Who authorized them the second time?

The comfortable answer: nobody has to, it's the user's own history. The uncomfortable answer: persistence is not permission. What a turn was allowed to show at write time proves nothing about what it may show at read time. The moment you persist retrieval results, your history endpoint becomes a second read path into the same data your search guards so carefully. Entitlements drift between write and read. The database doesn't know about your entitlement checks. Ship it naively and you've built an unguarded side door into the exact data you spent months gating.

The design that works: re-derive the replay from today's entitlements, not from the fact that the rows exist.

What drifted since the turn was writtenWhat the replay shows now
Nothing, everything still onFull transcript + source cards
Document-search toggle switched offTranscript text replays, source cards withheld
Copilot entitlement switched offHistory answers 404, no sessions, no turns

Two details do quiet work here. Denial answers 404, not 403, so a denied caller learns nothing, not even that the session exists. And denial costs zero reads: when the toggle is off, the session and turn queries are never even awaited.

The finer case is the interesting one. Copilot is on, document-search is off. A