Skip to content

Agents Need Structure, Not Scale

#llm-agents #planning #tool-use #gui-automation #benchmarks #multi-agent-systems #reinforcement-learning

The agent research published in the last few weeks carries an uncomfortable pattern. Every project that moves an agent off a static benchmark finds it far less capable than the leaderboard promised. Five arXiv papers and three engineering reports, spanning planning, tool use, GUI automation, and evaluation, converge on the same conclusion from different angles: the models aren't the bottleneck. The structures around them are.

The fixes are structural too. Explicit belief states instead of raw conversation history. Full enumeration instead of sampled rollouts. Cross-device task composition instead of single-screen screenshots. Bounded loops with external verification. Leases and epochs instead of shared browser profiles.

The observation problem: agents without beliefs ​

Put an LLM agent in a partially observable environment and it fails in three characteristic ways. Ambiguous feedback pushes it into premature commitments. A single informative observation collapses its uncertainty onto the wrong hypothesis. Its policy drifts as the history grows. The Belief-State Engine paper traces all three to one structural cause: an LLM agent, as commonly deployed, is a history-conditioned policy with no explicit belief over hidden state.

The fix is an inference module placed outside the LLM. The Belief-State Engine maintains a Bayesian posterior over the latent states of a given POMDP model, and at each decision step it exposes only that posterior to the LLM. The raw action-observation log is never shown.

The authors prove this pairing is a sound Markov policy on the belief MDP, which means it inherits the Bellman optimality guarantees of classical POMDP theory. The guarantee has a condition: the LLM must never see the raw history. Show the log and the guarantee evaporates.

Across the Tiger POMDP and a red-team attack-graph task, the BSE-augmented agent beats a reactive LLM, Chain-of-Thought, ReAct, a natural-language belief tracker, QMDP, and POMCP on task return, belief calibration, and decision consistency. Ten targeted ablations confirm the effect isn't specific to one model. It's the rare agent paper where the mechanism, not the model, is the result.

Tool selection: stop sampling what you can enumerate ​

Reinforcement learning over a frozen reasoner has become the standard recipe for teaching a policy which tools to call. The FGPO paper shows that recipe is structurally mismatched in specialist scientific settings where the complete tool-subset space is enumerable. When a small set of recurring computational capabilities covers a domain, the space of tool subsets is combinatorial but small enough to enumerate. GRPO, which estimates an action expectation from a handful of sampled rollouts, is the wrong tool for that setting.

The degradation is stark. As training succeeds, the policy concentrates on preferred subsets, resamples them, sampled rewards collide, and the group-normalized advantage vanishes. On genomic reasoning, the fraction of questions yielding no reward signal rises from 0.2% under a uniform reference policy to 20.8% after GRPO training. The better the policy gets, the less signal it has to learn from.

FGPO's fix is direct: score every tool subset and optimize the exact action expectation, so each update sees the complete action space. It precomputes the reward of each question-subset pair into an exhaustive table, which removes frozen-reasoner calls from the training loop entirely. Across five frozen reasoners and three genomic benchmarks, FGPO beats GRPO in all 15 settings, by 6.75 points on average and up to 14.20. The reasoners stay frozen; only the policy is trained. That's the point: you can get large gains in tool selection without touching the underlying model.

6.75 points: FGPO's average advantage over GRPO across 15 settings, up to 14.20 20.8%: share of questions with no reward signal after GRPO training, up from 0.2% 2.4x: more reward evaluations an on-demand GRPO schedule requires 1.40: tools invoked per GenomeQA question after FGPO, down from 2.36

GUI automation has a single-device blindspot ​

JarvisGUI starts from an observation that should embarrass the field. Real GUI usage routinely spans devices and platforms: intermediate results get transferred, shared state gets maintained, and coordination happens across heterogeneous environments. Existing GUI benchmarks evaluate agents on single-device, statically defined tasks, which produces an overly optimistic assessment of readiness for real use.

JarvisGUI is a dynamic benchmark that composes multi-step, cross-device workflows across Android, Windows, and Ubuntu. It formulates GUI tasks as input-output transformations under a lightweight type system, which lets it automatically compose workflows and evaluate agents in virtual environments spanning multiple operating systems. Current open-source GUI agents struggle with state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management. Those capabilities are exactly what existing benchmarks can't see.

Quick Take: the through-line in this month's agent research is that capability gaps come from missing structure, not missing scale; explicit beliefs, enumerated actions, bounded loops, and verified authority beat bigger models every time.

Loop engineering breaks in four predictable ways ​

Loop engineering is the practice of building a system, setting a measurable goal, and letting an agent iterate until it gets there. The loop engineering write-up covers four ways the loop breaks, and they map onto the academic results above.

Runaway loops come first. Tokens cost real money, so you need a hard stop rule. But the hard stop is more than cost control: it makes the loop's authority bounded and reviewable. Pair it with an externally defined success check and a failure receipt, what changed, what was tested, and why the run stopped, so a retry can't quietly become a second unobserved experiment.

Unverified autonomy is second. Letting an agent grade its own work is like asking a kindergartner to grade its own homework. You want agent A checking agent B's work instead.

Third is the vague goal. "Make this better" breaks an LLM; the criteria need to be genuinely non-negotiable. I've hit this one more times than I can count. An agent retrying toward "task completed" will happily loop forever on a step that technically succeeded but did the wrong thing. Success and correctness aren't the same check, and the goal signal itself needs verification, not just the stop condition.

Fourth is complexity overflow. When a single loop chokes on a big task, that's the moment to move from loop engineering to graph engineering, where work is a graph of bounded loops rather than one retrying monolith.

Stepping back, every source in this cluster shares a shape: a benchmark or workflow hides a structural gap until conditions change. The failures, evidence, and fixes line up like this.

Failure modeWhere it shows upStructural fix
Premature commitment under ambiguous feedbackBSE evals on Tiger POMDP and attack-graph tasksPosterior over hidden state; raw history hidden from the LLM
Vanishing reward signal as policy concentratesGenomic tool selection: no-reward questions jump from 0.2% to 20.8%Enumerate all tool subsets; precompute exact rewards
Cross-device state transferJarvisGUI workflows across Android, Windows, UbuntuDynamic task composition under a typed I/O system
Unbounded retry loopsProduction workflow automationHard stop rule plus external, non-negotiable success check
Stale authority after handoverBrowser workspaces shared between humans and agentsLease epochs; identity verification at grant time
Benchmark scores that flatter agentsResearch forecasting, single-device GUI evalsRolling cutoffs, oracle labels, synthesized rewards

Multi-agent state is an authority problem ​

SessionDock is a development report about running multiple browser agents in parallel, and it generalizes well beyond browsers. Separate profiles solve only the data-separation problem. A profile doesn't identify who owns an action, and it can't answer the question that matters when control changes hands: is a delayed action from the previous controller still allowed to execute?

The design answers with three concepts. A slot is a prepared Chromium workspace with its own user-data directory, downloads, and XDG paths. A binding fixes the project, worktree, expected account, and expected tenant for a run. A lease grants exclusive, time-limited control, and an epoch fences handovers. Agent A controls a slot at epoch 41. The person takes control back; the epoch becomes 42. A delayed action from A arriving with epoch 41 is rejected as stale.

Reading the thread, the hole that stands out is identity. The epoch catches stale authority, but it can't catch current authority pointed at the wrong tenant. Signing into tenant B inside a slot bound to tenant A produces no error, passes every check, and persists across restarts, because nothing in the design asks the remote service which identity the browser actually holds. The direction the thread converged on is an identity probe at lease grant: one authenticated request through the slot's session to the application's own "who am I" endpoint, compared against the binding, with a digest stored on the lease. The probe is evidence with an age, not an invariant. "Remote identity verified no more than X ago" is a weaker and more honest contract than "this slot is continuously tenant A."

What the community is saying: several commenters brought failure cases from production. Revoking authority does not undo an action, so a handover must fence both outgoing actions and incoming observations. If a write may have succeeded but its result is lost, OUTCOME_UNKNOWN must not become an excuse to retry. Process checks alone don't establish exclusive browser control; one commenter reproduced the reported ProcessSingleton behavior in a standalone Linux test. And conflict keys can only coordinate resources that are described correctly, because no scheduler understands the business logic of an arbitrary website.

TradingAgents, the open-source multi-agent trading framework, has been shipping versions of these lessons since January. Its decision log appends every completed run's decision, then fetches the realized return on the next run and injects a reflection into the portfolio manager's prompt. Checkpoint resume via LangGraph means a crashed run resumes from the last node instead of starting over. The June release added a verified data-access contract; the August one fixed look-ahead and point-in-time errors in FRED macro and social sentiment. Patching look-ahead bias in your data vendors and grounding exact price claims in verified snapshots is the authority problem solved at the data layer. The framework is also honest that two runs of the same ticker and date can differ, and it separates the sources: model sampling, live data movement, reasoning-model nondeterminism.

Evaluations that don't kid themselves ​

Two new evals attack the question of how to measure agents that operate in the real world. RAP (Research Attention Prediction) tracks whether LLM research agents can predict shifts in research attention. It's a rolling benchmark covering 278 AI/ML fields and 1,390 episodes: at each cutoff, an agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions.

The results are humbling. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average baseline in compositional accuracy. Under cumulative-history access, carry-forward beats direct forecast for all four models. Even with exact historical activity, future-specific updating is limited; only GPT-5.5 with reopened search slightly surpasses EWMA. Fine-tuning Qwen3-4B on realized outcomes improves its Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes. The lesson rhymes with TRACE: a small model with the right training signal goes further than a big model with a good prompt.

TRACE takes a different route: making the reward verifiable. Diagnostic reasoning over complex data lacks objective answers, because establishing the true cause of an anomaly takes costly expert investigation and may stay ambiguous. TRACE engineers around that by sampling an intervention, injecting it into a controlled simulator, and generating the observations it would produce. The hidden intervention provides an oracle label; the agent still has to investigate noisy, confounded, distributed evidence using Python and SQL.

FullAttr@1 means the agent named both the root cause and the affected segment correctly in one shot. On a held-out 235-episode test set, the strongest prompted baseline, Claude Opus 5, reaches 0.686. Supervised fine-tuning lifts Qwen3.5-35B-A3B from 0.159 to 0.637, and RL with synthesized rewards pushes it to 0.757, above every prompted baseline, including a prompted Qwen3.5-122B-A10B. With 3B active parameters, the 35B model deploys on a single 4090-class GPU, and it still beat prompted frontier models. RL gave it a training signal they never had; the advantage can't be explained by scale. The resulting policy also uses fewer tool calls than the prompted 35B base.

Common pitfalls ​

What trips people up across this cluster, concretely:

If the action space is enumerable, stop sampling. GRPO-style sample-based RL degrades as training succeeds, because a concentrated policy resamples the same subsets and the advantage signal collapses. Enumerate the full space and precompute the rewards; FGPO shows it's strictly better on every setting tested.

Don't pair a belief state with the raw history. The BSE's soundness proof fails the moment the raw action-observation log is visible to the LLM. In production agents that append "full context so far" to every prompt, you've reintroduced exactly the drift the belief state was meant to remove.

Treat "task completed" as a success check, not a correctness check. An agent will loop forever on a step that technically succeeded but did the wrong thing. Verify the goal signal itself, and use an external verifier rather than letting the agent grade its own output.

Don't confuse profile isolation with authority. Separate browser directories keep login data apart, but they can't answer who may act now, and two separate browsers can still modify the same server-side resource through the same account. You need leases, epochs, and conflict keys, not separate data directories.

Don't treat expected identity as verified identity. An expected_account or expected_tenant field is a claim, not remote state. Signing into the wrong tenant inside a valid slot passes every check and persists the mistake. Probe identity at lease grant, date-stamp the evidence, and treat it as "verified no more than X ago."

One thing to remember ​

Every paper in this cluster agrees on a point that contradicts the usual discourse: the model is the least interesting part of the agent system. The Belief-State Engine's gains come from the environment model and the inference module. FGPO's come from enumeration and a reward table. JarvisGUI's benchmark value comes from a type system and task composer. SessionDock's safety comes from leases and epochs. In each case performance moved because structure was added around a frozen or lightly trained model. When your agent fails mysteriously, look for the missing structure before you blame the model, and definitely before you reach for a bigger one.

The Bottom Line ​

If you're building agents that act under partial observability, adopt an explicit belief state and keep the raw action-observation log away from the policy. That single change beat ReAct, CoT, QMDP, and POMCP in the BSE evals; belief-state modules will be standard agent middleware within a year.

If you're training tool selection and the action space is enumerable, drop the sample-based RL. FGPO's exact expectation beats GRPO by 6.75 points on average across 15 settings, needs 2.4x fewer reward evaluations, and cuts tools per question from 2.36 to 1.40 on GenomeQA.

If you're coordinating multiple agents or human handovers over shared resources, treat every control transfer as an epoch boundary and verify remote identity at grant time, because a current authorization