Skip to content

Compress the Context, Rewire the Retriever, or Drop Vectors: Three Bets on a Cheaper RAG

#retrieval-augmented-generation #context-compression #semantic-retrieval #reinforcement-learning #inference-cost

The retrieval stack has a context problem ​

RAG has an unspoken efficiency problem. Most systems still chunk documents into 256-512 token blocks, embed them into a vector index, and fetch top-k by cosine similarity. That works on clean, short text. Point it at a 400-page regulatory filing or a medical textbook, and retrieval quality falls apart. The vector store returns chunks that are similar to the query, not chunks that answer it.

I spent a week this quarter building a QA bot over a stack of financial disclosures. Every lookup returned the same executive summary paragraph, no matter what I asked. The answer was in section 4.2. The retriever never got there. This is the gap the PageIndex project describes as "similarity does not equal relevance," and it is the most expensive failure in production RAG. Every wrong chunk you retrieve still gets charged against your LLM budget.

Three recent pieces of work attack this from different angles. TopoCompress compresses long contexts into coherent semantic spans before they hit the LLM, no training required. PAO fine-tunes retrievers with selective RL so they can be adapted without rebuilding a frozen document index. PageIndex drops the vector index entirely and walks a tree of sections with an LLM. All three share a goal: make retrieval cheaper and more reliable on long, structured documents.

TopoCompress: compress before you retrieve ​

Long-context compression is the most direct lever on inference cost. Feed an LLM 50 pages of PDF and you are paying for every token, even the ones that have nothing to do with the query. The problem is what most compressors do to get there. They train on specific models, they depend on the target LLM's attention for scoring, or they break the evidence structure of the document.

TopoCompress, from a paper posted on arxiv this month, takes a different route. It splits the context into semantic spans, scores each span against the query using dense and lexical relevance, then builds a hybrid graph. Nodes are spans. Edges connect spans that are semantically similar or sequentially adjacent. Query-guided scores propagate over that graph, and the algorithm selects a coherent set of spans for the final prompt. The graph step matters because it keeps neighboring spans together, so you don't end up with isolated fragments that mislead the model.

The headline result: TopoCompress matches the strongest baseline in quality while using a 4x smaller compression budget. Put concretely, if your previous compressor kept 8,000 tokens, you now keep 2,000 and get the same answer. That is a 75% cut in prompt cost. It also compresses 1.41x faster than the fastest baseline, which matters when you process a stream of documents rather than one-off queries.

4x smaller budget: TopoCompress matches the strongest baseline while keeping 75% fewer spans in the final prompt. 1.41x faster: about 30% lower compression time than the fastest baseline. 98.7% on FinanceBench: PageIndex's tree-search retrieval misses fewer than 2 questions in 100 on this financial QA benchmark, far ahead of vector RAG. 16.6x cost gap: feeding a 420-page PDF directly costs nearly 17 times more than tree-based retrieval.

The advantage of being training-free and model-agnostic is that compression does not couple your retrieval stage to your generation model. You can compress with one model and generate with another. No alignment or fine-tuning pass is required before the compressor works.

PAO: tune the retriever without touching the index ​

Compression addresses the cost of what you feed the model. A second problem lives upstream: the retriever itself is often bad. Dual-encoder retrievers are trained with contrastive loss, which optimizes for embedding similarity. Rerankers, the stage after them, capture finer relevance preferences. When those two objectives disagree, the retriever surfaces candidates the reranker never wanted.

RL is the standard fix for objective mismatch: use reward-model feedback to adapt the retriever. But the PAO authors found something that breaks in practice. On industrial-scale systems, the document index is frozen. You cannot re-embed billions of documents every week. Standard policy-gradient updates penalize negative samples by pushing them away, and in a frozen high-dimensional space, indiscriminate pushing disrupts the pre-trained semantic manifold. Retrieval quality gets worse before it gets better, sometimes permanently.

PAO's fix is selective. It only applies gradient updates to retrieved items with positive advantages. Nothing gets pushed away. The query embedding gets pulled toward high-reward regions while the global embedding topology stays intact. On a large industrial dataset and on public benchmarks, that beats standard RL and distillation baselines by a significant margin.

The constraint sets the audience for this method. If you have a frozen index measured in the hundreds of millions of embeddings, PAO is the only one of these three approaches that assumes that reality. The others assume you can re-index or add a scoring layer. PAO assumes you cannot.

Quick Take: The three lines of work agree on the diagnosis: similarity search over fixed chunks is the wrong abstraction for long documents. They disagree on the fix: compress the context, tune the retriever, or replace the index.

PageIndex: replace the index with a walking tree ​

The most aggressive bet is PageIndex, an open-source project that removes vector databases and chunking from the pipeline entirely. Instead of embedding chunks, it builds a hierarchical tree index per document, deriving the structure from the document's own layout rather than from an LLM. Retrieval then becomes an agentic search: the chat model walks the tree, deciding which section to open next, the way a human expert navigates a long report.

This is a fundamentally different cost model. Indexing is a one-time cost: roughly $0.001 per page with a basic index model, and about 13 seconds to 4.5 minutes for documents between 9 and 1,098 pages. Every subsequent query reuses the tree and reads only the nodes the reasoning reaches. The result shows up in the project's published benchmark: 98.7% on FinanceBench, with cost numbers that explain why the approach scales.

The chart spells out the practical breakpoint. At 52 pages, handing the model the whole PDF is only twice as expensive. At 420 pages it is almost 17 times as expensive, and at 805 pages it stops working entirely because the document exceeds the context window. The savings matter, but the bigger win is that the tree index answers questions naive long-context approaches cannot.

The README taps into a sentiment I keep seeing in GitHub discussions. People are calling vector RAG "vibe retrieval," and I understand the frustration: you get a plausible answer and no way to verify where it came from. Tree-based retrieval returns explicit section references, which is exactly what auditors and analysts need. On GitHub, the recurring complaint I see is the one I ran into myself: the vector store returns the right-sounding section, not the right section. PageIndex does not fix relevance magically, but it does make the retrieval path auditable. For regulated industries, that alone justifies the indexing cost.

How the three approaches compare ​

These are not interchangeable tools. Each one changes a different layer of the stack.

MethodWhat it changesTraining neededConstraintCost profileBest fit
TopoCompressAdds a compression layer before generationNoneKeeps your existing retriever and indexPer-query scoring cost, then 75% fewer prompt tokensLong contexts where prompt cost dominates
PAORetriever embedding weightsRL with a reward modelDocument index stays frozenOne training run, inference unchanged afterwardTeams that cannot rebuild a large production index
PageIndexReplaces the whole retrieval layerNoneNeeds LLM calls for indexing and chatOne-time indexing cost, then per-query reasoning costSingle long PDFs: legal, finance, medical, technical manuals

One caveat about PageIndex: it shines on documents with real layout structure. For messy web text or chat logs, the tree has less to work with. TopoCompress and PAO are more general because they operate on any sequence of text.

Common pitfalls ​

Positive-advantage-only updates matter more than the RL details. If you try to fine-tune a retriever with vanilla policy gradients on a frozen index, you will watch retrieval quality collapse as the embedding geometry deforms. I made this mistake on a production search stack: the fix was to stop penalizing negatives entirely and only pull queries toward positive items.

Don't over-compress without measuring downstream accuracy. A 4x budget cut is impressive, but on contract review or medical records, a dropped clause is worse than a higher bill. Run TopoCompress against your own question set before adopting it. The paper's results are on HotpotQA and similar benchmarks, not on your data.

Don't treat PageIndex as a drop-in vector replacement. The indexing step is a real one-time cost. If you have a million-document corpus, you need the file-system layer. The single-document path does not scale that far. And the chat model doing the tree walk is your main per-query expense. A weak chat model will cut corners in the search.

Finally, don't ignore the objective mismatch that PAO targets. If your retriever is optimized for contrastive similarity and your reranker for relevance, the retriever will keep surfacing the wrong candidates even if the reranker is perfect. Aligning those objectives stops being optional once your corpus gets large and specialized.

One thing to remember: relevance is a relationship between a question and a document, not a property of a chunk. Any retrieval step that ignores the question is guessing. The reason all three of these approaches reduce inference cost is that they put the question at the center of what gets retrieved, so the LLM spends its tokens where they count.

The bottom line ​

If you are building QA over long, structured documents and cannot fine-tune anything, adopt a structure-aware retrieval approach like PageIndex or TopoCompress, because chunk-then-embed will keep missing the relevant-but-not-similar passages and cost you more in wasted tokens.

If you have a frozen multi-million-document index and want better retrieval without rebuilding embeddings, use PAO's positive-advantage-only RL instead of policy gradients, because indiscriminate negative penalties will silently destroy retrieval quality.

If you are paying large prompt bills for long-context LLM calls, layer a training-free compressor like TopoCompress in front of generation, because a 4x smaller budget means the same answers at 75% lower cost. Watch this space for the next few months. Compression methods are getting faster, and model-agnostic tools are eating into the market share of trained compressors.