Appearance
The default instinct is to slap RAG on everything
Every LLM failure seems to demand the same fix: "we should add RAG." We've all done it. The model is missing a fact, so we throw a search index at it and hope the retrieved chunks anchor the answer. Sometimes it works. But three recent papers in the medical and NLP space say the reflex is wrong in a specific, reproducible way.
Retrieval is not grounding. Grounding means the model's output is tied to evidence you can inspect. Retrieval just stuffs text into the prompt. The difference is the whole ballgame.
The papers cover three different problems: rare-disease diagnosis, radiology report summarization, and long-tail named entity recognition. They reach a consistent conclusion: static, always-on retrieval is worse than a learned decision about whether to retrieve, and it's worse than fusing structured evidence into the generation process. Let me show you the numbers.
Key numbers: Rare-disease fusion improved Recall@1 by 7.86 points on Phenopacket Store and 20.18 points on RAMEDIS. A fusion model trained on other LLMs still lifted DeepSeek-V4-Flash from 0.1657 to 0.2176 without retraining. In radiology, RAG alone gave no benefit and added hallucination risk. Adaptive retrieval for NER (NE-R1) gained 2.52 points F1 in-domain and 1.18 points zero-shot.
Three papers that test the RAG reflex
Let's set the stage. The first paper, Learning to Fuse LLMs with Ontology Rankers for Rare-Disease Diagnosis, asks whether an LLM can improve a deterministic ontology ranker without giving up its evidence trail. The second, Improving Health Literacy through Lay Summarization of Radiological Reports, compares NER, RAG, and both, for turning clinician-speak into patient-friendly text. The third, NE-R1, adds reinforcement learning to make retrieval happen only when the model's own knowledge isn't enough.
Different domains, same thread: the value of retrieval depends entirely on how you decide when to use it and how you connect it to generation.
First: fusion beats replacement in rare-disease diagnosis
Ontology rankers like Phenomizer have been diagnosing rare diseases for over a decade. They take a patient's phenotypes, match them against disease ontologies, and return a ranked list of candidate diseases. Every candidate comes with an evidence trail: these phenotypes matched, these didn't, here's the supporting literature. The problem is that the ranker is rigid. LLMs are flexible and broad, but their differential diagnoses read like confident guesses. No citations, no trace.
The fusion approach doesn't ask which system wins. It takes both ranked lists, the agreement between them, and the ontology support behind each candidate, and trains a small model to decide how much to trust each system for an individual case.
The results are striking. Across eight open LLMs, fusion improves Phenomizer's Recall@1 by 7.86 percentage points on Phenopacket Store and 20.18 points on RAMEDIS. That means for every 100 rare-disease cases, roughly 8 more correct diagnoses move to the top of the list on Phenopacket Store, and 20 more on RAMEDIS. In a domain where diagnostic delays average years, that's not a metric, it's a life.
The most practical result is the one with DeepSeek-V4-Flash through an API. A fusion model trained only on the other eight LLMs improved DeepSeek's Recall@1 from 0.1657 to 0.2176, a 5.19-point gain, with zero retraining. The fusion model learned a general policy for weighing ranker and LLM outputs, and it transferred across model families. For 90.8% of correct fused diagnoses, the disease retained candidate-level ontology evidence, meaning the output isn't just right, it's inspectable.
They also fixed a documentation problem that would have invalidated earlier comparisons: they removed a test-set leakage pathway where benchmark cases and ontology annotations came from the same publications. That's a hidden trap for anyone building medical evaluation sets, and we'll come back to it in the pitfalls section.
Quick take: Retrieval and generation are not competitors. The winning move is to fuse the ranker's evidence with the LLM's knowledge, and let a learned model decide how much to trust each per case.
Then: RAG alone makes radiology summaries worse
The second paper might be the most uncomfortable one for RAG enthusiasts.
Many patients now paste their radiology reports into ChatGPT and ask what they mean. The paper tries to build something safer: automated lay summaries that preserve clinical meaning without the jargon. They compared four setups: plain generation, generation with NER-extracted findings, generation with RAG context from relevant medical sources, and the combination of NER and RAG.
The findings don't match the default instinct. NER consistently improves readability and overall quality. RAG alone offers no benefit, and it actively introduces hallucinations from irrelevant retrieved terms. When I've debugged this kind of pipeline, the failure mode is familiar: the retriever pulls in anatomy textbook passages that match a keyword but don't describe the patient's actual condition. Those irrelevant terms get woven into the summary as confident assertions.
Combining RAG with NER degrades performance in few-shot settings but improves readability when the model is fine-tuned. The fine-tuned BioBART with NER wins overall. This pattern is important: the NER step is what keeps the generation entity-aware, and the retrieved context only helps if the model is trained to use it.
For patient-facing text, every hallucinated finding is a real clinical risk. A static RAG setup that frequently injects unrelated terms is not an acceptable trade for slightly better fluency.
Then: adaptive retrieval wins for long-tail NER
The third paper starts from a different observation. NER models handle common entity types fine, but long-tail and domain-specific entities break them because the knowledge just isn't in the parameters. RAG can supply the missing knowledge, but it adds cost and noise when the model already knows the answer.
NE-R1 solves that with a "retrieval-on-demand" mechanism. The model first decides whether to retrieve. If the entity looks familiar, it generates from parametric memory. If the entity is rare, it queries an external dictionary or knowledge base. The decision is learned end-to-end with reinforcement learning, using a chain-of-thought reward that balances accuracy against retrieval benefit.
Here's why that matters: the model is essentially learning when to trust itself. That's what the radiology paper was missing, and it's what the fusion paper does implicitly by weighing the ontology ranker. The average F1 gain is 2.52 points across in-domain benchmarks and 1.18 points in zero-shot cross-domain evaluation. Those numbers sound small, but on long-tail entity benchmarks, a 2.5 point F1 gain often corresponds to catching the rare entity types that matter most, like uncommon disease names or obscure protein functions.
The other benefit is efficiency. If the model retrieves for, say, 20% of tokens instead of 100%, you cut the retrieval cost substantially. On a production API, that's real money.
Let me put the three approaches side by side.
| Paper | Retrieval trigger | Evidence handling | Practical result |
|---|---|---|---|
| Rare-disease fusion | Always, but weighted per case | Ontology candidates aligned with patient phenotypes | Recall@1 +7.86 pp (Phenopacket Store), +20.18 pp (RAMEDIS) |
| Radiology lay summaries | Static RAG | Retrieved terms dumped into prompt | No benefit, more hallucinations, best result from NER alone |
| NE-R1 NER | Learned, on-demand | Only retrieved when parametric knowledge is insufficient | F1 +2.52 pp in-domain, +1.18 pp zero-shot |
The difference between the first and second rows is the missing selection step. In radiology, retrieval happened without checking whether retrieved terms were relevant to the specific patient. In the fusion model, the ontology ranker already enforces phenotype matching, so the evidence is aligned by construction. In NE-R1, the model learns to retrieve only when the evidence is likely to matter.
That's the pattern. Effective grounding requires either a structured evidence source that's already aligned with the task, or a learned policy for when to reach outside the model. Blind retrieval is neither.
Here's the decision flow that NE-R1 implements, and that I think every RAG system should consider:
The mermaid diagram shows a loop: the model gets feedback on whether retrieval actually helped, and that's what trains the retrieval-on-demand policy. Without that feedback, retrieval is just a fixed input transform.
What these results mean for production RAG
The transferable lesson isn't about medicine. It's about when retrieval actually earns its keep.
Every retrieval call adds latency and cost. More importantly, it adds a distribution shift. The retrieved chunks come from a different corpus than the model's training data, and the model can latch onto irrelevant snippets as if they were authoritative. The radiology paper shows this in a very concrete way: irrelevant terms caused hallucinated findings, which in a clinical setting is dangerous, not just inaccurate.
The fix isn't better retrieval reranking, at least not as a complete solution. The two successful systems in these papers either use a structured evidence source with built-in matching logic (the ontology ranker) or learn a retrieval policy through reinforcement learning (NE-R1). Both end up doing something that static RAG doesn't: they make the decision to retrieve contingent on the case.
For a generic RAG setup, that means you should be looking at retrieval latches, not just retrieval heads. Can you predict query difficulty? Can you measure retrieval benefit after the fact? Those questions are closer to NE-R1's reinforcement learning setup than to a vector database.
Common Pitfalls
I've seen these failure modes enough times in production systems to want to call them out explicitly.
Don't concatenate retrieved chunks without filtering. In the radiology study, irrelevant terms were the direct cause of hallucinations. If you're retrieving passages that match a keyword but don't bear on the specific question, you're poisoning the model. Filter by entity relevance, or better, align the retrieval with the generation objective first.
Watch out for test-set leakage when building medical benchmarks. The rare-disease team had to explicitly remove a leakage pathway where ontology annotations and benchmark cases came from the same publications. If your evaluation data and your knowledge base share sources, you're measuring memorization, not generalization. Check your provenance before you quote your RAG numbers.
Don't assume RAG helps in few-shot settings. In the few-shot radiology experiments, combining RAG with NER was worse than NER alone. The model doesn't have enough training signal to learn how to ignore irrelevant context. If you're in a low-resource setting, start with entity extraction and skip retrieval until you can fine-tune.
Avoid always-on retrieval. It's not just wasteful, it's actively harmful. NE-R1's gains come from deciding when the parametric memory is enough. For common entities, retrieval adds nothing but latency and cost. Implement a gating mechanism based on uncertainty, or train a policy like NE-R1.
Don't treat evidence as a byproduct. The reason the fusion model is useful is that 90.8% of correct diagnoses retain candidate-level ontology evidence. If your system aims for traceability, design for it from the start. Adding retrieval after the fact won't give you inspectable rationale.
The Bottom Line
If you're building a system where every answer must be traceable to evidence, fuse a structured ranker (or a rules engine) with an LLM instead of asking the LLM alone. The rare-disease fusion results are strong: 7.86 to 20.18 points Recall@1 gains, and the evidence trail survives in the vast majority of correct outputs.
If you're constrained by latency or API cost, make retrieval adaptive. NE-R1's on-demand policy matches or beats always-on RAG on F1 while spending less on retrieval calls. That's the kind of improvement you can feel in your monthly bill.
One thing to watch: the radiology study's static RAG failure isn't a quirk, it's a pattern. Retrieval quality is not just relevance, it's alignment with the generation objective. Within a year, I expect more systems to move toward learned retrieval decisions and evidence-aware fusion, because that's where the measurable gains are.