Appearance
The old contract is broken
For most of the LLM era, a model's job was simple: process whatever context you gave it, in one forward pass, and generate the answer. Agentic AI blows up that contract. The systems that matter now decide what to look at, when to look, and when they've looked enough. Evidence acquisition is part of the reasoning loop.
Three papers released this month converge on the same shift from different directions. OmniSeek turns an Omni-LLM into an active audio-visual reasoning agent. A forecasting agent learns to search under a Brier-score reward. Mimir applies the pattern to long-horizon physical control in irrigation. Around them, the plumbing is maturing in parallel: the MCP Python SDK hit v2, Agent-Reach turned web-platform access into a one-sentence install, and agno standardized the agent-platform runtime.
The pattern connecting all of it is simple: the LLM proposes, and a verification layer disposes.
The single-pass assumption is broken
OmniSeek stops treating an audio-visual sequence as a blob to ingest. Instead of processing the whole thing in a single forward pass, the agent decides dynamically whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence from long contexts. The retrieved raw segments get appended back into the context for the next reasoning step. That's a real efficiency win: the model pays attention to what matters, not to everything.
Cold-starting this behavior took a 170K-trajectory corpus called OmniTraj-170K, built from multi-hop chain-of-thought trajectories with interleaved audio and visual evidence. Supervised training instills multi-turn tool use; two-stage reinforcement learning with verifiable rewards then optimizes the policy.
The clever bit is the Audio-Visual Necessity objective. It rewards successful trajectories whose reasoning depends on both modalities and discourages single-modality shortcuts. That matters more than it sounds. If audio alone can answer most questions, a lazy agent learns to ignore video entirely. The objective exists because the shortcut is the default.
Search is a learned skill
The "Do Your Own Research" paper makes the same point with a cleaner experiment. Prior forecasting systems either freeze the research context before training or deploy agentic search only at test time. Either way, the skill of gathering evidence is never shaped by the reward. This work closes that gap: an agent trained with single-epoch GRPO under a Brier-score reward, acquiring its own context at rollout time through web search, page reading, and financial time series.
The training environment is built from 2,100+ resolved Polymarket questions, with layered leak filtering so the agent only sees information published before each question's cutoff. The model is Qwen3.5-35B-A3B, meaning 3B active parameters. That's a model you can run on a single commodity GPU, not a cluster.
Training changes how the agent touches information. Calibration improves by 30-40%. Search attempts per rollout fall from 3.8 to 2.25. That second number is the tell: the agent searches less because it searches better. Evidence discipline is learned, not prompted.
Evaluated head-to-head against four frontier models, the trained policy beats all of them at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs 0.256, n=265), at roughly 5% of the inference cost. The margin is narrow, but the cost difference is not. And the margin widens on the hardest questions, the ones the human crowd itself hadn't decided. That's exactly where an agent should earn its keep.
Key numbers: the forecasting agent's calibration improves 30-40% and its search attempts drop from 3.8 to 2.25 per rollout after training. OmniSeek cold-starts multi-turn tool use from 170,000 audio-visual reasoning trajectories. Mimir uses about 51% less irrigation water than the historical schedule replay. AutoSynthData's synthetic SFT adds 7.2 percentage points to Pass@1 on EnterpriseOps Gym Hybrid, a 35% relative gain.
Physical environments don't reset
Mimir moves the evidence loop into a regime where mistakes compound. In irrigation, every daily decision alters soil-water dynamics for the rest of the growing season. No reset button. No safe episode rerun. And the agent has to improve from experience without rewriting the physical rules that keep execution safe.
Mimir organizes around two repair timescales. The fast loop turns an LLM output into a proposal, then runs it through a structured physical interface and a deterministic simulator: numerically check, revise, bound, then execute. The slow loop watches for recurrent failure patterns and consolidates them into persistent contextual principles that condition future proposals. The physical model, the evaluator, and the execution constraints stay immutable. The LLM can change its own context. It cannot change physics.
Results justify the architecture. Across multiple sites, crops, and years, Mimir posts the lowest aggregate control cost among the evaluated references and uses about 51% less irrigation water than the historical schedule. Ablations show control cost rises when you remove forward simulation, verified revision, or persistent context. And the counterintuitive result: model-scale studies show no monotonic gain from increasing LLM size. Bigger is not automatically better when the bottleneck is the verification loop, not raw reasoning. If your agent fails a long-horizon task, reaching for a bigger model is the last lever, not the first.
| System | What the agent controls | Training signal | Key result |
|---|---|---|---|
| OmniSeek | Which modality (audio/video) and temporal window to retrieve | 170K trajectories + two-stage RL with AV Necessity objective | Evidence seeking becomes part of the reasoning loop |
| DYOR forecaster | Web search, page reading, time-series queries | Single-epoch GRPO under Brier-score reward | 30-40% calibration gain, 40% fewer searches |
| Mimir | Daily irrigation proposals | Verified revision + persistent contextual principles | 51% less water, lowest control cost in the study |
Quick Take: every system that works treats the LLM as a proposal engine and puts a separate verification layer between the proposal and the execution.
The training data bottleneck
You can't train an agent to work well in your environment if you have no data from that environment. AutoSynthData, from ServiceNow CoreAI, attacks that specific bottleneck. It evaluates a target model and a stronger teacher in the target environment, distills the failures into sanitized capability specification cards, generates new tasks that exercise the weak capability, verifies each one, and post-trains on the survivors. As the model improves, the curriculum shifts toward what it still finds hard.
The task abstraction is the most useful idea here: a task is a system specification, a user prompt, and a verifier. Generated tasks must be feasible (a valid trajectory exists), realistic (someone would plausibly ask it), and difficult for the current model. The verifier needs consistency, soundness, and completeness. Every candidate passes two gates: positive verification (does the intended solution actually succeed?) and negative verification (do mutated wrong outcomes fail?). Failed candidates go to a critic for bounded repair before discard.
On EnterpriseOps Gym Hybrid, 2,000 generated samples (about 18 hours of generation, Gemma-4-26B-A4B-it as target with Qwen3.8-27B as teacher) improved mean Pass@1 by 7.2 percentage points, a 35% relative gain, and lifted verifier success from 63.01% to 68.55%. That closes 59% of the original gap to the reference model. On ITSM, 1,994 samples lifted Pass@1 from 18.77% to 27.18%. The generator never saw the original evaluation tasks; everything came from the specification cards. That's what makes this feel like progress rather than benchmark leakage.
The plumbing gets standardized
The tooling layer is consolidating fast. The MCP Python SDK v2 cuts a server down to two type-hinted functions and a docstring. No JSON Schema to write, no request parsing, no protocol handling. The same package is a full client, and transports cover stdio, Streamable HTTP, and SSE. A URL is a deployable transport; the client treats it that way.
Agent-Reach occupies a different layer. It's a capability layer that picks, installs, diagnoses, and routes web-platform access for agents, rather than a single tool. One line in a prompt installs it. Each platform channel has a primary and a backup backend, and channels are probed at runtime rather than checked for existence. When Bilibili's anti-bot blocked yt-dlp in June 2026, the project had already surveyed alternatives and routed traffic to bili-cli; users noticed nothing. agent-reach doctor tells you exactly which path each platform is currently using, and which ones silently rotted.
What the community is saying: cookie-based platform access is the divisive line. The pain is real, and it's everywhere. Twitter's API requires paid access, Reddit blocks datacenter IPs, and login-walled platforms like Xiaohongshu need a browser session. Some teams refuse to hand browser cookies to any tool, and the Agent-Reach README is blunt about the account-ban risk, recommending dedicated burner accounts because a cookie is full account access and platforms detect scripted API calls. I found this the hard way: my first attempt to give an agent Twitter search access ended with a restricted account in under a week. The repo stores credentials only in ~/.agent-reach/config.yaml with 600 file permissions and never uploads them, but file permissions don't protect you from a platform ban. Read the safety section before you point this at anything you care about.
agno targets teams shipping agents as products: an SDK, a runtime, a web UI, 50+ API endpoints with SSE and websockets, Postgres for sessions and traces, JWT-based RBAC, OpenTelemetry observability, and human-approval gates for sensitive tools. The pitch is owning your agent stack: data, memory, and security posture stay yours. It deploys on any container platform, with starter templates for Railway, AWS, GCP, and Azure.
Production looks boring now
Two more projects show how fast this is maturing. DocsGPT is a self-hosted, MIT-licensed platform that turns documents into agents. Visual workflows, GraphRAG knowledge sync from Drive, SharePoint, and Confluence, answers with citations, guardrails for PII and prompt injection, and an OpenAI-compatible API plus an MCP server for clients like Claude and Cursor. One command installs it, and everything runs in your environment, including Postgres, Redis, and the vector store.
On the learning side, the production-agentic-rag-course makes a point that veteran search engineers will appreciate: build the BM25 keyword layer before you add vectors. Week by week it constructs a research assistant over arXiv, from infrastructure to data pipelines to hybrid RRF retrieval, then adds the agentic layer last. Guardrails, document grading, query rewriting, and adaptive retrieval, all in a LangGraph workflow, exposed through a Telegram bot. The numbers along the way are the practical ones: an 80% prompt reduction for 6x faster responses, and 150-400x query speedups from Redis caching. That's what production looks like: boring infrastructure, observability, and retrieval that doesn't need an LLM to do the heavy lifting.
Common pitfalls
The most common failure I see is freezing the evidence-gathering policy. Bolt web search onto a frozen model and it will search badly, because nothing shapes its search behavior with the outcome. The forecasting paper is the control case: calibration improved 30-40% and search attempts nearly halved only after GRPO training under a Brier-score reward. If you can't train the search policy, at least evaluate it with tool-use traces before you ship.
Multimodal agents take shortcuts, and your reward is the reason. Without an explicit necessity objective, an audio-visual agent learns to answer from whichever modality is cheapest to process. OmniSeek's Audio-Visual Necessity objective blocks that path on purpose. Ask what your reward can be satisfied without: if the hard-evidence route isn't required, the agent won't take it.
Weak verifiers poison synthetic agent data. AutoSynthData accepts a task only after both a positive gate (the intended solution passes) and a negative gate (mutated wrong outcomes fail). Skip the negative gate and you train your agent to produce states that look successful but aren't. A lax verifier is worse than no data, because the model learns to be confidently wrong in exactly the ways the verifier can't see.
In physical domains, never hand the LLM the actuator. Mimir keeps the physical model, evaluator, and execution constraints immutable; the LLM proposes, the numerical layer disposes. The ablation shows what you lose otherwise: control cost rises when forward simulation or verified revision is removed. If your agent touches real equipment, the simulator is the safety boundary, and no model output should be able to override it.
Cookie credentials on shared or main accounts cost me a Twitter account. Any tool that uses a logged-in browser session inherits the full account, and platforms detect scripted behavior. The Agent-Reach docs are explicit about using dedicated burner accounts for exactly this reason, and the project's install flow defaults to checking without modifying your system until you explicitly allow it. Treat a cookie as a root credential and scope it accordingly.
One thing to remember: every system in this cluster, from OmniSeek to Mimir to AutoSynthData, runs the same deep pattern. The model generates proposals, and a separate, verifiable mechanism decides what counts as evidence and what gets executed. The verification loop is the product. Design your verifier before your agent, because every agent eventually learns to game whatever check you put in front of it.
The Bottom Line
If you're building a research or analysis agent, train the search behavior under an outcome reward rather than bolting tools onto a frozen model. The DYOR results show calibration gains of 30-40% and a 40% drop in wasted searches, at 5% of the frontier inference cost.
If you're working in physical or enterprise environments where errors compound, put an immutable verification layer between the LLM and execution, the way Mimir uses its simulator and AutoSynthData uses its task gates. Removing that layer is the single clearest way to make control cost and error rates climb.
If you're shipping agent products, standardize on MCP for tool interfaces and a platform runtime like agno instead of hand-rolled integrations. MCP v2 makes a server roughly two type-hinted functions, and the protocol now covers stdio and HTTP transports. One thing to watch: web-platform channels will keep breaking as anti-bot enforcement tightens, and routing layers like Agent-Reach will become standard infrastructure within six months.