Appearance
Retrieval Is Where RAG Wins or Loses: Editable Memory, Late Interaction, and the New Embedding Stack
The retrieval bottleneck
Most RAG pipelines fail before the LLM ever sees the prompt. They retrieve too much, too coarsely, and hand the mess to the model to sort out. The generation side of RAG is mature. The retrieval side is where quality is won or lost.
Four recent projects make the same point from different angles. rEDMRec compresses expensive LLM reasoning into editable memory that a small model can query directly. Sentence Transformers v6.0 adds multi-vector late-interaction models that keep token-level detail instead of averaging it into one vector. LightRAG layers a knowledge graph over vector storage to capture relationships between documents. TinySearch filters web content locally with BM25 and ONNX embeddings, so a 4B model never drowns in scraped pages.
Each one attacks a different layer of the stack.
| Approach | What it stores | Query cost | Index cost | Best for |
|---|---|---|---|---|
| Dense bi-encoder | 1 vector per document | One dot product | Tiny | General similarity at scale |
| Late interaction | ~125 vectors per document | MaxSim over token pairs | 10-40x dense | Multi-requirement queries, exact tokens |
| Graph RAG (LightRAG) | Entities, relations, chunks | Graph traversal + vector search | Heavy, one-time | Cross-document reasoning |
| Editable memory (rEDMRec) | Typed experience channels | Lightweight student retrieval | One-time distillation | Recommendations, personalization |
Key Numbers
- 13.3% max HR@1 improvement over the second-best baseline on ML-1M (rEDMRec)
- 64% average reduction in web tokens entering the context window (TinySearch, 8 research queries)
- 42x the index size of multi-vector vs. dense MiniLM on the same 4,874 passages
- 60 labeled pixels → weighted F1 0.84 for mangrove mapping (OlmoEarth Tiny)
- 77.92 vs 51.59 on MLDR: multi-vector vs. dense on long-document retrieval
Reasoning you can edit: rEDMRec
LLMs are good at recommendation reasoning. Given a user's history and a candidate item, they can extract preferences and explain why one item fits better than another. The problem is that this reasoning is expensive to repeat on every ranking request, and once produced, it's consumed and discarded. Nothing is reused. Nothing can be corrected when the user's tastes drift.
rEDMRec treats reasoning as something to cache, not regenerate. A teacher LLM's reasoning is distilled once into four typed experience channels:
- long-term preference: who the user is over months
- short-term context: what they're doing this week
- item-perception: how the user sees items
- counterfactual hard-negative comparisons: why item A beats item B
A memory controller maintains these channels with Add/Delete/Modify/Keep operations, refined by K-agent debate. A lightweight student LLM ranks candidates purely by retrieving from this memory. The teacher never runs again at inference time, which decouples online cost from reasoning depth.
The results hold up across the board. On ML-1M, Amazon Beauty, and Steam, with ten student backbones, rEDMRec beats zero-shot, few-shot, and RAG on every backbone, and GraphRAG on most. The best gain is 13.3% HR@1 over the second-best baseline on ML-1M. For a recommendation service, that's the difference between the right item surfacing first and getting lost entirely.
The ablations tell a more careful story. Short-term context is the only channel that helps consistently across model sizes. Long-term preference, item-perception, and counterfactual comparisons help at some capacity tiers and hurt at others, and on the strongest students their contribution can reverse. That matches intuition: a capable student can infer long-term patterns from raw history, but it has no way to know what you were thinking last week.
The debate-based memory optimization also cuts bank duplication by 7.4 percentage points while raising HR@1 by up to 0.029 over six optimization epochs. Less redundant memory, better rankings.
When one vector isn't enough: late interaction
A dense embedding model compresses a whole text into one fixed-size vector. Everything the model noticed has to fit in those 384 or 768 numbers, and similarity is one dot product between two summaries. That compression is lossy in a specific way: a rare entity, an exact identifier, or one key clause in a long passage all compete for room in the same vector.
Multi-vector models skip that compression. They run the same transformer, but instead of pooling token embeddings into one vector, they project each token to a small dimension (classically 128) and keep all of them. A 9-token document becomes a 9x128 matrix. Scoring uses MaxSim: for each query token, take its highest similarity against any document token, then sum across the query. It's a soft alignment. Every query token points at the document token that best explains it.
The practical effect shows up on multi-requirement queries. "Green sofa with wooden legs and rounded cushions" lets each requirement find its own evidence. A single-vector model has to blend all four into one point, so a green sofa with the wrong legs ends up sitting close to the one you asked for. Late interaction also preserves exact matches: a product code or surname keeps its own token, where a dense model averaged it in with everything else. And the alignment doesn't have to be lexical. Encode "Where do penguins live?" against "Penguins inhabit Antarctica" and the query token "live" matches "inhabit" at 0.94, with no shared characters.
The cost is index size. Encoding 4,874 Natural Questions passages with LateOn produced 608,414 token vectors, an average of 124.8 per passage. That's about 42x the storage of a MiniLM dense index.
Compression changes the picture. The same 608,414 vectors take 92 MB as a fast-plaid index, since PLAID stores a centroid id plus a quantized residual per vector instead of the raw vector. For reference, a 4096-dimensional dense model like Qwen3-Embedding-8B needs about 80 MB for the same passages. A compressed multi-vector index sits in the same territory as the dense indexes people already run.
The quality gap grows with document length, because more text has to fit in the same fixed vector. On MLDR, a long-document retrieval benchmark, the multilingual multi-vector model scores 77.92 against its dense sibling's 51.59. If your chunks are long, single-vector retrieval is quietly losing information on every query.
Quick Take: If you're retrieving long chunks with a single dense vector, you're leaving retrieval quality on the table, and the gap shows up exactly where you'd expect: multi-requirement queries, long documents, and out-of-domain data.
Graphs on top of vectors: LightRAG
LightRAG takes a different route: put a knowledge graph on top of the vector store. It's positioned as a lightweight alternative to Microsoft GraphRAG, and the key difference is cost. GraphRAG generates community reports and does multi-hop reasoning, which is expensive. LightRAG skips both, which drastically cuts the number of LLM calls during indexing and querying.
The dual-layer architecture keeps entities and relationships in a graph while text chunks live in a vector store. Retrieval pulls from both, so you get detailed facts and abstract concepts in the same response. That matters for vertical domains like legal and finance, where a question spans multiple documents and requires connecting entities across them.
The practical details carry the project. Four chunking strategies are selectable: fixed-length, recursive character, vector semantic, and paragraph semantic. The paragraph strategy aligns chunk boundaries with the document's native structure: headings, paragraphs, tables. That fixes the classic failure where a heading gets separated from its content, or a long table loses its header row when split. The docx parser even detects and corrects section headings automatically, which helps documents with inconsistent outlines.
Role-specific LLM configuration is the detail I'd steal first. LightRAG needs four model roles: EXTRACT for entity-relation extraction, QUERY for writing final answers, KEYWORDS, and VLM for multimodal content. You can assign a fast, cheap model to extraction and a stronger one to query answering. The docs suggest Qwen3-30B-A3B as a reasonable local minimum for extraction, a 30B-parameter MoE model with 3B active parameters that runs on a single high-end GPU.
Storage backends now include Neo4J, MongoDB, PostgreSQL, and OpenSearch, so you can pick one store instead of juggling three. Multimodal parsing comes via MinerU and Docling, handling PDFs, images, tables, and formulas. The knowledge graph connects multimodal content to body text, which matters for operation manuals and academic papers.
Keeping the web out of your context window: TinySearch
I've been running TinySearch for a few weeks with a 9B local model, and it fixed a problem I'd given up on. The old workflow was: scrape the web, stuff 50k tokens of pages into the context window, and hope the model found the five paragraphs that mattered. Giving a small model that mess and expecting it to triage defeats the point. A 4B or 9B model burns its limited capacity on filtering instead of answering.
TinySearch filters before the model sees anything. It searches the web, reads the pages that matter, and selects the useful parts locally. The filtering is BM25 plus local ONNX embeddings. No LLM in the middle. The chunks returned are the original page text with source URLs attached, so the model gets evidence, not summaries.
The v0.6.1 release added bring-your-own-browser support over CDP. Instead of TinySearch launching its own headless Chromium, you point it at a browser you operate, with your own profile, proxy, cookies, and fingerprinting setup. I found that sites which blocked the default setup started working once I pointed it at my normal browser profile. That alone made scraping viable on sites that don't love fresh headless sessions.
The token savings hold up in the benchmark. Across 8 research queries, the same webpages went from 146,878 tokens to 53,426, roughly 64% less web content entering the model context. That's not a universal guarantee: bloated sites saw 80%+ reductions while already-clean pages barely changed. But for a local model with a 32k or 64k context window, cutting the input by two-thirds is the difference between a coherent answer and a model that lost the thread.
The workflow I've settled into is: search, let the model choose useful URLs, scrape those URLs, and feed the model only the relevant evidence. Search and scrape are separate tools now, and search is deliberately cheap: it doesn't even start Chromium or load the embedding model. Related links come back ranked against the query, so the agent decides where to go next instead of the tool crawling half the internet.
What the Community Is Saying: The TinySearch thread captures the local-model crowd's priorities. The author builds around a specific pain: a 4B or 9B model given 50k tokens of scraped pages will burn its capacity on triage instead of answering. People running small local models recognized the pain immediately. The open asks are for feedback on self-hosted browser setups and weird sites where CDP still breaks. That's where the next round of fixes will come from.
Embeddings for a specific domain: OlmoEarth
Domain-specific embeddings are the quiet workhorse of applied RAG, and OlmoEarth shows what that looks like for Earth observation. The platform exports embedding vectors from its open-source foundation models as Cloud-Optimized GeoTIFFs, with one band per embedding dimension. Locations with similar surface characteristics end up with similar vectors; locations that differ land far apart.
Three encoder variants are available:
| Variant | Dimensions | Parameters | What it means in practice |
|---|---|---|---|
| Nano | 128 | 1.4M | Embeds on CPU, tiny COG exports |
| Tiny | 192 | 6.2M | The sweet spot for most tasks |
| Base | 768 | 89M | Best quality, needs a GPU for batch work |
The outputs are quantized to int8, values from -127 to 127, with -128 reserved for nodata. A COG with one band per dimension is easy to share and works with any geospatial tool: QGIS, GDAL, rasterio, or your own scripts.
The few-shot segmentation result is the headline. With 60 labeled pixels, 20 per class, a logistic regression over the Tiny embeddings produced a coherent land-cover map of a coastal mangrove region in Vietnam with weighted F1 of 0.84. The classifier saturates quickly: going from 30 to 300 labels barely changes accuracy, because the embeddings are doing the heavy lifting. One afternoon of labeling, wall-to-wall mapping.
The same embeddings power change detection without any training. Monthly embeddings for September 2023 and September 2024, per-pixel cosine distance, and the Park Fire burn scar in Butte County lights up immediately. Unsupervised PCA exploration reproduces field boundaries in the Netherlands without ever being told what a parcel is.
The lesson for RAG builders: if your domain has a foundation model, frozen embeddings from it will beat generic text embeddings for similarity search, and a linear probe on top will beat prompt engineering for classification. The 60-pixel result is the strongest argument for domain-specific embeddings I've seen this year.
Common Pitfalls
Treating the document length cap as a suggestion is the most common multi-vector mistake. Checkpoints truncate at their training cap. A 662-token passage through LateOn's cap of 300 comes back as 273 vectors, and the rest is gone. If your chunks are longer than the cap, lift it explicitly with encode_document(..., processing_kwargs={"text": {"max_length": 512}}), and expect the index to grow in proportion.
Using a thinking model for extraction in graph RAG is another one. Entity-relation extraction runs on every chunk, and a reasoning model makes it slow and expensive. Use a fast non-thinking model for the EXTRACT role and save the strong model for QUERY.
Filtering the web with an LLM when you're running a small model defeats the purpose. If your filtering step needs an LLM, you've just moved the context-window problem upstream. BM25 plus local embeddings is enough to select relevant chunks, and it's free.
Exposing lightrag-server without authentication is a security hole waiting to happen. It binds to 0.0.0.0 by default and every endpoint is public unless you configure LIGHTRAG_API_KEY or AUTH_ACCOUNTS. The Ollama-compatible routes stay open even then unless you set WHITELIST_PATHS=/health.
Assuming one embedding model generalizes across domains is the quiet failure. The OlmoEarth result is the counterexample: domain-specific embeddings turn 60 labels into a wall-to-wall map. Generic text embeddings would need orders of magnitude more labels. Check your domain for a foundation model before you fine-tune anything.
One thing to remember
Every project in this cluster moves work out of the generation step and into the retrieval step. rEDMRec caches reasoning instead of regenerating it. Multi-vector models defer interaction to scoring time. LightRAG indexes relationships once so queries don't re-derive them. TinySearch filters before the model sees anything. If your RAG pipeline feels expensive or imprecise, the fix is probably upstream of the LLM.
The Bottom Line
If you're building a recommendation or personalization system, adopt the editable-memory pattern from rEDMRec: distill expensive reasoning once into typed channels and let a small student model retrieve from it, because you'll get up to 13.3% better HR@1 on ML-1M while cutting per-request inference cost to a fraction.
If you're retrieving long documents or handling multi-requirement queries, switch to a multi-vector model: the 77.92 vs 51.59 gap on MLDR is the difference between finding the right passage and missing it, and a PLAID-compressed index costs about the same as a 4096-dim dense index.
If you're running local or small models, skip the LLM-in-the-loop filtering and use BM25 plus local embeddings like TinySearch: a 64% reduction in context tokens is