Appearance
The Gap Between Demo Agents and Production Agents
A systematic survey of LLM-based agents for software and systems security, spanning 2023 to 2026, lands on a blunt conclusion: "we have built agents able to act but not yet agents whose authority is bounded or whose behavior is auditable." That sentence should sit in the back of your head if you're shipping agentic software today. It explains the gap between your demo and your production anxiety.
The survey covers architecture, perception, memory, reasoning, action space, orchestration, and self-improvement. Across all of that, the field has converged on one uncomfortable truth. Agents can chain tools, retain state, and revise plans. They cannot reliably tell you why they did something, or stop themselves from doing something harmful. That's not a security niche problem. It's the core problem of every agent system.
75.9% of real user-agent interactions are multi-turn, not single-turn (PersonaForge). 35.8% → 66.0% pass@1 on Countdown from RL-based tool learning (Tool-DAPO). 22–27% success-rate lifts from isolating the decisive error agent (DoCtOR). +4.1% composite score from training on multi-turn simulations (PersonaForge). 75.1% top-1 skill retrieval accuracy with profile-conditioned routing (SkillFeed).
Real Users Don't Ask Single-Turn Questions
PersonaForge opens with a number that should embarrass most agent benchmark suites. The team analyzed 16,000 real sessions and found that 75.9% of user-agent interactions are multi-turn. Yet most training sets and benchmarks assume informationally complete, single-turn queries. You're training a kiosk and deploying a conversational colleague.
They built a simulation framework with a four-dimensional persona space and something called Reverse Deep Construction, which anchors synthetic queries in real user seeds. That produced a 6.3K-record training set and PersonaForge-Bench, a manually annotated 138-task benchmark across 20 professional domains.
On Qwen3.5-27B, training with PersonaForge data improves the composite score by +4.1%. The biggest wins are Task Completion (+6.0%) and Response Quality (+6.8%). A 4% composite gain sounds modest, but in practice it's the difference between an agent that asks clarifying questions and one that guesses your intent. Trained agents also use fewer turns and fewer tool calls per task. That's not token savings for their sake. It means the agent converges faster toward an answer a human will accept.
Tools Are Learned, Not Given
The Countdown task isn't glamorous, but it's a perfect microscope for tool integration. The Tool-DAPO paper starts by dissecting reasoning failures on this arithmetic-heavy task. Calculation errors account for a substantial share of wrong answers. No prompting trick fixes that. The model needs a calculator.
They first teach the model useful tool-use patterns with SFT, then apply on-policy RL (RLOO, RLOO++, GRPO, DAPO) with only final-answer reward. Every RL method I've tested across various domains benefits from tool integration, with roughly 10-point gains across pass@k. But Tool-DAPO is the standout: pass@1 jumps from 35.8% (Tool-SFT) to 66.0%.
That 30-point jump matters. It's a pure RL effect: the model isn't rewarded for calling the calculator, only for getting the final answer right. RL figures out that tool calls lead to correctness. If you're building agents for accounting, data analysis, or anything with arithmetic, don't settle for a model that "thinks" in floating point. Give it a calculator and let RL form the habit.
Quick Take: The field has moved from proving agents can act to figuring out how to bound, audit, and evaluate them. Expect reliability work, not capability demos, to dominate the next year.
When Multi-Agent Systems Fail, One Agent Usually Started It
Multi-agent systems fail a lot. The usual response is to make everyone reflect on the failure. DoCtOR argues that's backwards. Most failures trace back to a single decisive error agent. The other agents were just doing their jobs. Forcing them to reflect contaminates their memory with wrong insights.
DoCtOR (Diagnose-then-Correct PPO-enhanced Reflection) runs three steps: identify the decisive error step and agent via automated failure attribution, use counterfactual reasoning to generate a corrected step, then have only the decisive agent reflect. The results are the strongest in this cluster: 22% improvement over initial success rates on HotPotQA, 26% on ChartQAPro, 27% on Mind2Web, beating Reflexion, Retroformer, and COPPER.
The implication is practical. Next time your multi-agent workflow fails, audit which agent actually led it astray before calling a postmortem. Your reviewer agent probably isn't the culprit. Don't let it learn a lesson it doesn't need.
Context Is the Silent Agent Killer
Context management is where long-horizon agents go to die. Preserving every interaction history creates a continuously growing working context. Current proactive methods only offer search, deletion, and summarization. That's not enough.
ContextPilot systematically expands the toolset with planning, long-term memory, and soft context offloading. More importantly, it changes the RL credit assignment. Instead of rewarding all context edits with the same trajectory-level reward, ContextPilot estimates action-level advantages from branched trajectories that pass through each edit. That's how you learn which actions actually matter.
The result: stronger performance on long-context QA and deep search with a more compact working context. If your agent's context window is ballooning, you don't need a bigger model. You need better context hygiene. A 128K window is a budget, not a goal.
Skills Should Match the Person, Not Just the Prompt
Skill routing is becoming the new retrieval problem. SkillFeed points out a failure mode that's easy to miss: a task-only router can pick a semantically plausible skill that's wrong for the requesting user. Task relevance and skill suitability are not the same thing.
Their benchmark holds the task fixed and changes the user profile. When the profile changes which skill is appropriate, adding profile conditioning yields a 35.1-point retrieval gain. That's massive. The same request from a privacy-conscious researcher and a deadline-driven startup founder should route to different skills. Most router benchmarks wouldn't catch that.
On SkillFeed-Bench, the full framework reaches 75.1% top-1 accuracy, 23.1 points over a pretrained routing baseline. If you're building a skill library for internal agents, user constraints matter as much as task semantics.
Security Agents Need Bounded Authority and Auditable Behavior
The security survey's warning is stark because the stakes are asymmetric. An agent that misplaces a file is annoying. An agent that issues a network request or modifies a firewall rule is a liability. The paper catalogs a field that has produced capable agents but not proven assessment protocols. Datasets don't overlap, metrics are incomparable, and baselines are inconsistent.
What does that mean for you? If your agent touches files, runs code, or calls privileged APIs, you need two non-negotiable properties. First, bounded authority: explicit limits on what the agent can do, enforced by the runtime, not the model's good intentions. Second, auditable behavior: a trace you can inspect after a run to answer "why did this agent take that action?" Without those, you're just running a stochastic script.
Sequential vs. Event-Driven: How to Compose Agents
Concurrency is an architectural choice, not a feature list. The Mozaik hackathon announcement makes the contrast clear. Sequential workflows are easier to reason about when every step depends on the previous one. But they turn into bottlenecks the moment agents need to react independently to new information.
Mozaik's event-driven model puts agents in a shared AgenticEnvironment. Participants emit events, register handlers, and react without waiting for a central scheduler. I've felt this in my own projects. A planning agent that blocks a researcher until it finishes reasoning is a design smell. If your researcher and analyst can work on different sources simultaneously, they should.
| Sequential workflow | Event-driven model |
|---|---|
| Agents follow a fixed order | Agents react to events as they arrive |
| Each stage waits for the previous one | Multiple agents work concurrently |
| Orchestration logic encodes every interaction | Participants define their own reactions |
| Adding an agent may require rewiring the whole chain | New participants just join the environment |
| Long-running inference blocks downstream | Non-blocking inference keeps other activity moving |
| Agents are tightly coupled to the pipeline | Agents stay loosely coupled and reusable |
The point isn't to make every workflow concurrent. Some tasks genuinely need ordered execution. But if you have agents observing the same activity and making independent judgments, event-driven composition is worth a prototype.
What the Community Is Saying
The Apodex AMA on r/LocalLLaMA had the right energy for this moment. The team released open checkpoints in FP8, int4, and NVFP4 variants alongside an open-source harness called FrontierAgent. That's a signal: people are running agent loops on local GPUs, not just cloud APIs.
When I've run tool-calling models at 4-bit precision, function arguments are the first thing to break. A leaderboard won't show that. One malformed tool call kills a 30-step agent run. The community chatter around these releases is increasingly about observability and failure recovery, not token throughput. The Mozaik announcement makes the same move: it frames the architecture around reacting, sharing state, and coordinating, not around which model scores highest on a benchmark.
What Trips People Up
- Treating pass@1 as a proxy for reliability. A model that nails the first attempt on a benchmark can still derail on a 40-step task when a tool returns an unexpected shape.
- Spreading reflection across all agents after a failure. The decisive error is usually one agent's fault. Forcing everyone to reflect contaminates shared context with noise, and you'll spend even more tokens unlearning wrong lessons.
- Letting context grow unbounded. Without proactive offloading or summarization, you'll pay for every ignored token and degrade attention on the ones that matter.
- Ignoring user constraints in skill routing. Matching on task similarity alone gives you semantically plausible but personally inappropriate recommendations. Profile-conditioned retrieval matters most exactly when the profile changes the reference skill.
- Assuming sequential workflows are safe because they're simple. They're not wrong for every task, but they turn concurrency into a rewrite once agents need to react to events independently.
One Thing to Remember: An agent that can't explain why it acted is a liability, not an asset. When you're debugging a multi-agent pipeline, ask "which agent made the decisive error?" before asking every agent to reflect.
The Bottom Line
- If you're building a multi-agent system where failure cost is high, adopt failure attribution like DoCtOR before scaling up. Isolate the decisive error agent and reflect on it alone, or you'll train on contaminated memories and chase phantom regressions.
- If you're stuck evaluating agents on single-turn benchmarks, switch to multi-turn simulation like PersonaForge. A 6.3K-record training set lifts composite scores by 4.1 points, and the task-completion gains translate directly to fewer user round-trips in production.
- If you're hitting context limits, try proactive context management instead of renting a bigger window. ContextPilot's action-level RL keeps accuracy up while shrinking the working context, which means lower cost and faster inference at the same time.