Appearance
The strategy lock-in problem
LLM agents can now post-train an LLM end-to-end. They write the code, launch the training run, evaluate the checkpoints, and improve downstream performance. On paper, that's the last mile of AI-for-AI: the loop closes without a human in it.
A new empirical analysis (arxiv 2608.19072) argues this picture is misleading. It separates what agents actually do into two capabilities. Execution-level capability is iterating inside a chosen training strategy. Strategy-level capability is revising the high-level judgment as experimental evidence accumulates. Across a large corpus of publicly released post-training trajectories, the authors found the same pattern every time. The training strategy gets locked in at the very beginning, and the entire remaining budget goes to local adjustments within that strategy.
The agent never steps back and asks whether the strategy itself is wrong.
That failure mode isn't confined to auto-ML research. You see it in production agent systems all the time: an agent that executes a bad plan with great competence, producing polished work toward the wrong goal. The research gives a name to what you're already feeling when your agent burns through a $200 budget optimizing a prompt that should have been scrapped after the first failed run.
What the researchers actually tested
The paper then tests three natural explanations for the lock-in: missing experience, missing guidance, and insufficient reasoning. Each gets an escalating intervention.
An experience-driven scaffold improves execution across the board, +12.6 points on GSM8K and +40.8 on HumanEval. That HumanEval jump is enormous, the difference between a model that fumbles basic function writing and one that reliably produces working code. But the scaffold leaves the strategy static. The agent gets better at doing, not at deciding what to do.
Human guidance works differently. It successfully redirects the initial strategy, but once training starts, the agent falls back into local adjustment loops. You can point the agent at a better strategy, and it still can't hold the course or notice when the course needs changing.
Extra inference compute is the most telling result. It pays off on easier tasks and yields almost no gain on the hardest one. Throwing more tokens at the problem doesn't produce better judgment. It produces more confident execution.
+12.6 points on GSM8K and +40.8 on HumanEval from an experience-driven scaffold. The strategy still never changed.
Three interventions tested: experience, human guidance, extra inference compute. None triggered spontaneous strategy revision.
Human guidance redirected the initial strategy, then the agent slid back into local adjustment loops once training began.
The conclusion lands hard: what agents lack is not experience, guidance, or reasoning compute. They lack a mechanism for spontaneously reevaluating their strategy during execution.
The two loops agents actually run
Here's the structure the paper is pointing at. Agents are very good at the inner loop and effectively blind to the outer one.
The dotted line is the missing mechanism. Evidence accumulates, but nothing carries it from the evaluation step back to the strategy decision. Every agent framework I've used has great support for the execution loop: retries, checkpoints, tool calls, eval harnesses. None of them have a primitive for "the plan is wrong, change the plan."
This is the gap that matters. Not tool calling, not memory, not context windows. Those are solved problems. Strategy reevaluation during execution is not.
Quick Take
Agents are getting dramatically better at doing, not at deciding what to do, and no amount of inference compute fixes that.
SkillForge: attacking the problem upstream
A second paper in this space (arxiv 2608.18933) takes a different angle on the same underlying issue: agents fail not because they can't reason, but because they lack project-specific knowledge. SkillForge, a self-distillation framework for issue resolution, doesn't wait for real issues to expose knowledge gaps. It synthesizes issues by re-implementing test-covered core functionalities of the repository, then resolves those synthetic issues. The results get distilled into entity-grounded skills tied to specific repository entities.
The practical effect: when a real issue arrives, the agent already knows the codebase. It doesn't spend its first ten tool calls discovering that the project uses a custom ORM or that error handling follows an undocumented convention. SkillForge improves issue resolution over strong baselines with both open- and closed-source models.
What connects the two papers is the timing of knowledge. SkillForge moves knowledge acquisition before the task. The post-training study moves strategy evaluation during the task. Both are saying the same thing: the bottleneck in agentic systems sits upstream of execution, in the decisions and knowledge that shape what the agent attempts.
What production frameworks actually give you
Meanwhile, the orchestration layer is consolidating fast. Microsoft Agent Framework (MAF) and Pipecat both shipped major updates recently, and they represent two different answers to the same question: what does a production agent system need?
MAF is a graph-based orchestration framework for Python and .NET. It supports sequential, concurrent, handoff, and group collaboration patterns, with checkpointing, streaming, human-in-the-loop, and time-travel debugging. It ships with OpenTelemetry integration, YAML-declarative agents, and a migration path from Semantic Kernel and AutoGen. The pitch is durability and governance: restartable workflows, observability, and the ability to host agents on Foundry with two extra lines of code.
Pipecat is the real-time voice and multimodal option. Each pipeline is an agent, and you compose them with handoff, parallel fan-out, sidecar workers, or distributed deployments over a shared bus. It's built for ultra-low latency interaction over WebSockets or WebRTC, with SDKs down to ESP32. If you're building a voice assistant or a multimodal companion, this is the framework that gets you to a working demo fastest.
| Microsoft Agent Framework | Pipecat | CrewAI (enterprise analytics) | |
|---|---|---|---|
| Primary focus | Production orchestration, governance | Real-time voice/multimodal | Conversational BI pipelines |
| Languages | Python, .NET | Python | Python |
| Orchestration patterns | Graph: sequential, concurrent, handoff, group | Pipeline handoff, fan-out, sidecar, shared bus | Sequential specialist pipeline |
| Production features | Checkpointing, HITL, time-travel, OpenTelemetry, YAML agents | Low-latency transports, CLI deploy, debugger | MCP-based visualization, multi-tenant isolation |
| Best fit | Long-running business workflows | Voice agents, live interaction | Analytics and reporting |
None of these frameworks solve the strategy problem. They solve the execution problem at scale. That's useful. Just know which problem you're buying a solution for.
The enterprise case: what multi-agent buys you
The CrewAI-based business intelligence platform from arxiv 2608.18740 is a useful data point on the value of multi-agent structure. Five specialized agents run in a sequential pipeline: parse the natural language query, retrieve and analyze data, generate visualizations via MCP, and deliver insights. Across 300 end-to-end test cases on synthetic and production enterprise datasets, it hit 95.3% functional accuracy with a 93% hallucination-free rate. That's a 22.6 percentage point improvement over a single-agent baseline. An ablation study found the Data Analysis and Report Aggregation agents drive most of the quality.
The numbers carry practical weight. 95.3% functional accuracy means roughly 19 of 20 queries produce a correct end-to-end result, which is the threshold where you can let a dashboard answer questions without a data analyst in the loop. The 24-second mean latency is useless for chat but fine for a scheduled report. The 93% hallucination-free rate means about 1 in 14 responses contains fabricated content, which is exactly why the platform keeps a human review step for anything that touches financial numbers.
The 22.6 point accuracy gap over the single-agent baseline is the headline. Splitting the work across specialists with defined roles isn't a convenience. It's the difference between a toy and something you can put in front of a CFO.
Slack's answer: the human is the strategy loop
Slack's approach to human-agent teams, documented in a recent Anthropic interview with CPO Jaime DeLanghe, is the most direct production answer to the strategy problem. The core rhythm is a cycle of handoffs: agents handle the production work, drafting, summarizing, monitoring, preparing, and pass results to a person. The person reviews, decides, and redirects, then hands the work back.
The human is the dotted line in the diagram above. The agent executes; the human reevaluates.
For years, the promise that workplace conversation would compound into organizational knowledge never materialized. The exhaust of people working together just sat there, and people kept repeating themselves. Now it's an agent's job to make sense of that exhaust. DeLanghe's framing is social rather than technical: treat agents like coworkers with clear roles and focus areas. An agent whose value feels mandated rather than understood is an agent people forget to use.
That's why Slack pushes public-by-default channels: open context compounds, and agents get the same shared history that human teammates get. The fastest adoption mechanism Slack has seen is people watching a teammate do it. The company-wide "How I Slackbot" channel at Salesforce is the example: a trick from a sales process ends up reshaping an engineering workflow.
What the community is saying: even the pop-culture takes are circling the same insight. I caught Spider-Man: Brand New Day and spent the ride home thinking about E.V. as an agentic system. Compared to Jarvis, it felt grounded: no holographic interface, just a guy talking to his computer at his desk. The thread on r/LocalLLaMA was mostly silly speculation about whether Peter self-hosts his models or inherited E.V. from Stark. The observation holds up. The most convincing agent UX we've seen on screen is the most boring one. Local, role-bound, and embedded in one person's workflow.
That's the same lesson as Slack's, the enterprise analytics ablation, and the post-training paper. Agents work when they have a defined role, a bounded context, and a human who can change the plan.
Common Pitfalls
Confusing execution capability with strategy capability. A framework that orchestrates well will not make good decisions. The post-training paper shows this empirically: better execution tooling improved outcomes without ever changing the strategy. If your agent is executing a bad plan competently, more scaffolding won't save you. You need a checkpoint where the plan itself gets questioned.
Measuring agent value with token usage. DeLanghe's point about Slack applies directly: token usage tells you the lights are on, not whether anyone is better off. More messages can mean people can't find what they need. More tokens can mean your agent is confidently doing the wrong thing. Measure outcomes, not activity.
Skipping the human review loop in the name of automation. A 93% hallucination-free rate sounds good until you remember it means 1 in 14 responses is fabricated. For anything with financial, legal, or customer-facing consequences, keep the human review step. That's where the quality comes from.
Treating every agent as a one-on-one chatbot. Agents with undefined roles produce undefined results. The enterprise analytics ablation and Slack's "agents are coworkers" framing agree: specialization is what makes multi-agent systems beat single-agent baselines. Define the role before you define the tools.
Locking in an orchestration framework before you know your failure modes. MAF exists because teams hit orchestration walls with single-prompt loops, and Pipecat exists because chat frameworks can't do real-time voice. Both are good at what they do. The mistake is picking one before you know whether your bottleneck is durability, latency, or strategy. The first two are framework problems. The third is not.
One Thing to Remember
The execution/strategy gap is where humans still earn their keep in agentic systems. Every framework, paper, and production deployment covered here points at the same conclusion: agents are getting better at doing, and the people who build the systems that decide what gets done are the ones who win.
The Bottom Line
If you're building an agent that writes code, trains models, or spends real money on tool calls, build explicit strategy checkpoints into the workflow. The agent won't create them itself. Schedule a review after the first failed run, and force a written justification for continuing the current approach.
If you're choosing an orchestration framework for production, pick on durability, observability, and human-in-the-loop support, not demo polish. MAF's checkpointing and OpenTelemetry integration matter more than its haiku example, and Pipecat's low-latency transports matter only if you're actually building voice.
If you're deploying agents inside an organization, start with a small shared channel, give people the same resources, and let usage patterns spread on their own. Slack's experience is that adoption compounds when people watch a teammate do it. Mandated agent value doesn't stick.
One thing to watch: the post-training paper's authors identify spontaneous strategy reevaluation as the missing mechanism, and that's exactly the kind of problem that gets solved within two years. When it does, the human handoff loop shrinks. For now, it's the most valuable part of your system.