Skip to content

Reading Is Not Using: The Retrieval-Integration Gap in RAG

#retrieval-augmented-generation #knowledge-graphs #multi-vector-embeddings #agentic-search #evaluation

Reading Is Not Using: The Retrieval-Integration Gap in RAG ​

Your retriever returns the right passages. Recall@k looks great. And the model's answer still ignores the one fact that matters. You re-check the pipeline. Retrieval is accurate. The context contains the evidence. The judgment doesn't use it.

Retrieval accuracy is not grounding ​

A paper called "Reading Is Not Using" gives this failure a name: the retrieval-integration gap. The authors ran controlled experiments on LLM analysts processing financial disclosures. They held the focal firm's information fixed and varied only unrelated context, from 2,000 to 128,000 tokens. That range spans a short memo to a full 10-K filing. Direct retrieval stayed accurate the whole way. But the influence of a risk disclosure on the model's investment judgment fell to the experimental noise floor.

The pattern replicated across model families and judgment tasks. It survived when the researchers removed real disclosures from actual 10-K filings. More capable models postponed the gap, but none eliminated it.

The uncomfortable conclusion: retrieval-based evaluation can certify systems whose judgments ignore information they demonstrably retrieved. If your eval measures whether the right passage surfaces, you're measuring reading, not using.

What the gap looks like under the hood ​

Causal memory interventions in the same paper show how information dies. Two mechanisms jointly transmit a disclosure into a judgment: compressed summaries and source-text lookup. When either is missing, the disclosure's influence drops.

Workflow architecture decides whether transmission succeeds. Chunk-and-summarize pipelines evict relevant information somewhere between chunking and the context window. A targeted, structured restatement placed adjacent to the decision restores its influence. Same model, same corpus, same question. Different plumbing, different judgment.

AI analyst performance is therefore jointly determined by model capability and workflow architecture. You can't finetune your way out of a pipeline that discards evidence before the model sees it.

Key numbers

  • 2,000 to 128,000 tokens of added context: retrieval stays accurate, judgment influence hits the noise floor.
  • Up to 0.24 NDCG@10 lost to truncation on medical passages averaging 941 tokens.
  • 92.05% strict accuracy on BrowseComp-Plus with 30.21% lower online inference cost.
  • 14.5 hours on a single RTX 3090 to beat every general-purpose retriever on a medical eval.

Evidence blindness: the agentic version of the same disease ​

The retrieval-integration gap has a sibling in agentic search. The AtlasNav paper studies Direct Corpus Interaction (DCI), where agents work directly against a full corpus instead of a pre-chunked retrieval index. The corpus is fully accessible, yet under a finite interaction budget, reachable evidence can remain unusable.

The authors document three stages of failure. Required evidence fails to surface in the first place. A surfaced supporting document never gets opened. An opened document fails to expose its decisive fragment. They call this progressive silent loss "evidence blindness," and it's measurable: stage-wise evidence realization tracks how much of the required evidence actually reaches the final answer.

The cause is structural. Raw interaction adds little reusable corpus organization. Dynamic-workspace methods reconstruct a query-conditioned interaction space from scratch for every query and trajectory. Either way, useful structure is recovered online, under time pressure, and then thrown away.

Quick Take: Retrieval accuracy and judgment quality are decoupled. You can measure perfect retrieval and still ship a system whose decisions ignore everything it retrieved.

Workflow architecture decides what gets used ​

The financial paper and the AtlasNav paper converge on the same point: how you represent and route information matters as much as whether you retrieved it.

Both pipelines retrieve the same disclosure. The top one compresses it into a summary that loses the signal. The bottom one restates it in a form the model can act on, positioned where the decision happens. Same retriever, different outcome.

Persistent structure beats per-query reconstruction ​

AtlasNav takes the "representation matters" argument to its logical end. Instead of reconstructing corpus structure for each query, it organizes the corpus once into a Corpus Atlas, then lets every query navigate that shared structure adaptively.

On BrowseComp-Plus, AtlasNav hits 92.05% strict accuracy, meaning the final answer has to match exactly, with 30.21% lower recorded online inference cost than the previous best dynamic-workspace system. That cost reduction matters when you're paying per token for agentic loops. Under matched budgets, AtlasNav realizes the complete required evidence earlier, and it approaches the same model's evidence-supplied empirical reference more rapidly. The representation principle transfers to PhantomWiki's controlled 10K-1M scaling and to heterogeneous enterprise knowledge.

ApproachHow structure is createdEvidence realizationReported result
Raw DCI interactionLittle reusable structure; search from scratchRequired evidence often fails to surfaceNo reusable organization
Dynamic workspaceReconstructed per query from query and trajectoryFaster, but structure discarded after each queryPrior best on BrowseComp-Plus
Corpus Atlas (AtlasNav)Organized once, navigated by every queryComplete evidence realized earliest under matched budgets92.05% strict accuracy, 30.21% lower online cost

The pattern matches the financial paper: information becomes usable when the workflow invests in structure up front, rather than hoping the model will find and use evidence on its own.

The retrieval layer itself: token-level matching ​

None of this absolves the retriever. The third piece in this cluster is practical: a Sentence Transformers walkthrough of training multi-vector embedding models, the ColBERT-style late-interaction family. These models keep one small vector per token and score with MaxSim, where every query token finds its best-matching document token. That preserves the fine-grained signals a single dense vector has to average away, at the cost of a bigger index.

The most striking finding is about truncation. Most released retrieval checkpoints were trained on short passages. Classic ColBERT truncates at 180 or 300 tokens; many dense models at 256 or 512. My medical evaluation corpus had passages averaging 941 tokens, and the cost showed up immediately: up to 0.24 NDCG@10. That's more than any difference between model architectures. Your retriever isn't underperforming because of the architecture. It's underperforming because it never reads past the first page.

Finetuning changes the picture, and the starting point matters more than you'd expect. I finetuned six checkpoints with an identical recipe on 25k medical question-passage pairs from MIRIAD. The pre-supervised checkpoints adapted best: mLateOn-unsupervised jumped from 0.9087 to 0.9398 NDCG@10. The finished checkpoints barely moved, or regressed, at every learning rate I tried. A fresh projection on a strong dense backbone reached within 0.03 of the existing checkpoints from nothing but the projection and 25k pairs.

The whole run took 14.5 hours on a single RTX 3090, no cloud GPU needed. A punctuation skiplist, which excludes punctuation tokens from document-side scoring, shrank the index by 9.6% for free in my ablation. And the classic [MASK] query expansion trick? I tested it in four configurations and none made a measurable difference. Don't assume the old recipe transfers.

Symbolic constraints for grounded answers ​

The fourth piece covers what happens when you need guarantees, not just better odds. LLM-based knowledge graph QA usually picks between two strategies: full semantic parsing into SPARQL, which is brittle against complex schemas and incomplete graphs, or free-form LLM reasoning over the graph, which is more robust but offers no formal guarantees.

The CES-PK framework (Constrained Entity Selection under Partial Knowledge) takes a middle path. Let the LLM generate candidate answers, then verify them with lightweight symbolic constraints derived from the question: type constraints, relation constraints, exclusion constraints. No executable logical form required.

The key design choice is three-valued semantics: satisfied, violated, unknown. Real knowledge graphs are incomplete, so rejecting anything not explicitly present would wrongly discard valid answers. The unknown state preserves recall, while filtering violated candidates improves precision. Satisfied constraints then become positive symbolic evidence for ranking what remains. Tested on the Hetionet biomedical knowledge graph, precision improves and recall holds.

This is the verification layer the other three papers are implicitly asking for. If retrieval can be accurate while judgments ignore it, and agents can hold evidence without using it, you need a check that catches the failure. Symbolic constraints won't catch everything, but they catch a specific, common class: answers that contradict the question's own constraints.

What trips people up ​

Evaluating retrieval instead of judgment ​

If your eval measures recall@k or hit rate, you can ship a system whose decisions ignore what it retrieved. The financial paper shows this directly. Measure the downstream decision: does the retrieved evidence change the answer, and is the change correct?

Truncating documents to a checkpoint's training cap ​

A checkpoint that truncates at 300 tokens silently discards two-thirds of a 941-token document before scoring. That costs up to 0.24 NDCG@10, more than any architecture choice. Lift the caps or train your own model with the document length your data actually needs.

Finetuning from the finished checkpoint ​

The natural starting point, a polished general-purpose checkpoint, is the weakest option for domain adaptation. Pre-supervised checkpoints adapt far better, and a fresh projection on a strong backbone is a close runner-up. The finished checkpoint carries general-purpose tuning that your domain training then has to undo.

Treating the corpus as a flat pile of chunks ​

Agents fail in three silent stages: evidence doesn't surface, a surfaced document doesn't get opened, an opened document doesn't expose the decisive fragment. Organize the corpus once into navigable structure instead of reconstructing it per query. The AtlasNav results put a number on it: 30.21% lower online inference cost.

Assuming your knowledge graph is complete ​

If your verification rejects anything not explicitly present, you'll wrongly reject valid answers under open-world conditions. Three-valued semantics, with an explicit unknown state, preserves recall while still filtering the violations you can prove.

One thing to remember ​

All four results point the same direction. Retrieval is only a means. The end is whether information changes judgment, and that depends on the workflow architecture that carries evidence, the corpus representation that makes it navigable, the retrieval layer that preserves its fine-grained signals, and the symbolic checks that verify it. A better retriever helps. It doesn't fix a pipeline that discards evidence after retrieval.

Where this leaves you ​

If you're building a RAG system for high-stakes decisions, in finance, medicine, or law, stop evaluating retrieval in isolation. Add a structured restatement step adjacent to the decision and measure whether the retrieved evidence changes the judgment. That's the only metric that matters.

If you're building agentic search over a large corpus, invest in organizing the corpus once into a persistent structure rather than reconstructing it per query. The AtlasNav results suggest you'll realize complete evidence earlier and cut online inference cost by roughly 30%.

If you're finetuning a retriever for a specialized domain, start from a pre-supervised checkpoint or a fresh projection on a strong backbone, lift document length caps before training, and budget for a single-GPU run: 14.5 hours on an RTX 3090 beat every general-purpose retriever on a medical eval. One thing to watch: evaluation is shifting from retrieval metrics to judgment metrics, and "does the retrieved evidence change the answer" will likely become the standard eval within a year.