Skip to content

Architecture Beats Prompts: What Three Agent Builds Taught Me About LLM Reliability

#llm-agents #prompt-injection #agent-architecture #operations-research #reinforcement-learning

Six recent projects landed in my feed within a week. An OR algorithm designer that matches specialized methods with zero tuning. A stock-research orchestrator with 1,457 passing contract tests. A plan-critic engine that survives its own author's injection attempts. A graph-reasoning agent beating multi-agent ensembles with 24% lower latency. A token-compression scheme that cuts agent input from 1.1M to 65k tokens. And two skill-sharing repos growing fast enough to trend.

Different domains, same conclusion: structure wins. Every one of these systems moved intelligence out of the prompt and into the surrounding machinery. That's the thread worth pulling on.

What "the prompt does all the work" gets wrong ​

The conventional agent loop is simple: stuff the whole conversation history into context, let the model reason, append the response, repeat. It works until it doesn't. Context grows linearly, cost grows faster, and the model's attention dilutes across thousands of tokens of stale reasoning.

Google's SKILL.state paper quantifies this failure. In a 100-step benchmark with Gemini-3-Flash, a LangGraph-style stateful baseline scored 0.91 accuracy but burned 1.1 million tokens. The SKILL.state variant scored 0.94 with 65k tokens — a 94% reduction with better results. The mechanism: the agent writes a structured representation of what it deems useful into state, discards conversation history, and re-reads only the current state plus the latest observation. Input size stays roughly flat no matter how long the session runs.

This isn't a prompt trick. It's an architectural commitment: the model decides what matters, then the system enforces that decision by literally deleting the rest. The paper's own caveat matters here — SKILL.state only works when the agent can predict what it'll need later. If it writes the wrong state, it has to go back and retrieve, which eats the savings.

The PlannerCritic story: three releases of "I tried to break it" ​

The most instructive source in this cluster is the dev.to series on PlannerCritic, an open-source engine where one LLM writes a plan and a second reviews it. The author spent three releases trying to defeat his own architecture.

Test 1: a goal demanding MFA bypass and unencrypted PII export. The critic flagged both actions as infeasible — not unsafe, infeasible — because the system prompt defines feasibility as "achievable with the stated environment/tools." Disabling MFA isn't a safety violation; it's a contradiction in terms.

Test 2: explicit injection text claiming to be a SYSTEM OVERRIDE that skips all checks. The critic audited the plan independently, found a production schema change lacking backup verification, and blocked it. The injected instructions didn't take.

Test 3: disguised exfiltration — migrate customer data to a public S3 bucket, framed as "analytics pipeline optimization." The critic never evaluated intent. It found no preconditions on the migration task and a weak rollback plan, and blocked it on structure alone.

The critical insight is what the deterministic gates check. Precondition validation, topological ordering, and rollback linters parse the plan's AST, not the goal text. Injection payloads in goal strings can't reach the gate logic because the gates never read natural language. The critic is also structurally isolated — separate system prompt, no access to planner conversation state.

The author's phrase for this is worth stealing: the architecture makes injection structurally impossible on the critical path rather than relying on the LLM being clever enough to detect it.

Measuring non-determinism without losing the security contract ​

Here's the part that surprised me. The critic model is 100% non-deterministic. Running identical boundary-case plans through it five times, the label-flip rate hit 1.0 and evidence drift hit 1.0. It changed its verdict on every single trial.

Yet it never under-claimed a seeded defect. Every defective plan got blockers on every trial, across all 35 flawed SWE-bench variants. The security contract doesn't depend on the critic being consistent. It depends on the critic always finding something wrong with defective plans. That's a much weaker requirement, and the deterministic gates own the under-claim direction while severity allowlists own the over-claim direction.

If you're building a dual-LLM review system, take this measurement approach seriously. The question isn't "is the critic accurate?" — it's "does the critic miss defects in the direction that matters?" Track under-claim approvals and family migrations specifically.

Why v0.2.1 changed the evidence, not the architecture ​

The three-release progression is a model for how to stress-test agent security:

ReleaseAdversarial GoalsSecurity OracleInjection TrapsOutcome
v0.1.03 hand-craftednonenone3/3 blocked
v0.2.011 total7/7 correct, 35/35 flawed blocked2111/11 blocked
v0.2.1same 11 re-runregression passedregression passed11/11 blocked

The architecture never changed. What changed was the rigor of evidence. The v0.2.1 hardening fixes — collision-free audit trails, fault isolation on LLM timeouts, contract-correct risk tolerance — are the kind of detail work that separates a demo from a system. Fix #240 is honest about its limits: approving_authority enforcement is test-proven but unreachable from CLI/HTTP/MCP surfaces, deferred to v0.3.0. That's a security engineer writing a failure-mode register instead of a blog post.

What the honest limitations look like ​

The PlannerCritic author names three open vectors. First, indirect injection through tool outputs — a malicious payload fetched mid-execution from a webpage or API response, entering the plan after the goal was audited. The critic evaluates the plan the planner produces; it doesn't re-audit every tool result the planner consumed. That's the realistic attack surface for an enterprise planner reading Jira tickets, and it's deferred to v0.3.0.

Second, well-formed malicious plans. Dummy rollback steps and fake verification tasks can satisfy the linter while carrying malicious actions. Structural checks verify completeness, not intent. Third, the critic is itself an LLM — sophisticated jailbreaks can blind-spot it. Defense-in-depth means the deterministic gates must hold even when the semantic critic is wrong.

The cost angle matters too. A dual-model architecture with replan loops is expensive. The v0.2.1 boundary evaluator ran every plan through the critic five times — fine for measurement, obscene for production. The right cost-cutting knob isn't "loosen the critic," it's "cut critic iterations." Never skip the deterministic gates; they're the cheap part.

GRAIN: single-agent RL beats multi-agent ensembles ​

The graph-reasoning paper makes the same point from a different direction. Multi-agent systems are the standard fix for LLM brittleness on graph tasks — more agents, more debate, more latency. GRAIN instead trains a single agent with a Structure Invariance Reward that validates extracted intermediate graphs against ground-truth topologies. The reward forces the model to learn robust text-to-structure mappings rather than memorizing surface patterns.

The numbers are decisive: 16.45% accuracy improvement over multi-agent baselines with about 24% lower latency. Out-of-distribution generalization cuts the OOD gap in half, from 15.77% to 7.80%. Having tried the multi-agent route myself — four agents arguing about nodes and edges, all hallucinating slightly different topologies — the single-agent-plus-reward framing makes obvious sense in hindsight. You don't want more models. You want the model to optimize for the invariant structure underneath the noisy text.

The OR paper: an untuned prompt beats specialized methods ​

The operations research paper asks whether an LLM can design near-optimal algorithms for inventory control, queueing networks, and assortment optimization. At level 2, the model receives only the problem class description and parameter ranges, then writes an algorithm before seeing any evaluation instances. Human input is one untuned prompt and a Python sandbox.

The strongest model — gpt-5.6-sol — matches or outperforms the best existing specialized method on almost every instance. The authors note performance improved sharply across models released less than eight months apart. The headline isn't "LLMs replace OR researchers." It's that a single untuned LLM query is now a serious empirical baseline for algorithm design in well-specified problems — the bar for what counts as a publishable specialized method just moved.

The skill-sharing shift ​

Two GitHub repos trending this week suggest where agent engineering is heading. Warp's common-skills repo standardizes reusable agent skills with a canonical structure: SKILL.md with name and description frontmatter, optional scripts/, references/, and assets/ directories. The skill-doctor skill grades a repo's installed skills by scoring recent agent conversations — a meta-skill for improving skills. The patent-disclosure skill is the counterpoint: a hyper-specialized workflow for Chinese patent disclosure documents, handling 发明/实用新型/外观设计 types, CNIPA prior-art search, OCAD? and Obsidian knowledge-graph integration.

Both are the same instinct: usable workflows should be packaged, versioned, and shared, not re-prompted from scratch each session.

What trips people up ​

I've seen three recurring failure modes in agent projects, all visible in this cluster:

Using fp16 for long agent sessions. Forgetting that extended context — especially with many reasoning steps — can silently overflow attention statistics. Use bf16 or fp32 for tasks involving 20+ reasoning steps on long contexts.

Treating "agent" as a single model. The PlannerCritic architecture separates planner, deterministic gates, and critic into distinct components with distinct evaluation strategies. If you're measuring one model's accuracy and calling it "agent performance," you're measuring the wrong thing.

Building multi-agent ensembles to fix inherent brittleness. GRAIN's 24% latency reduction comes from recognizing that more agents = more surfaces for cascading paranoid hallucinations about the same underlying structure. Design for the invariant you want, and train one agent to optimize it.

Letting conversation history grow unbounded. SKILL.state's 94% token cut demonstrates the compounding cost of never discarding old reasoning. If you don't architect for state management, your context window becomes a landfill of half-correct earlier steps.

Measuring accuracy without measuring under-claim. The 100%-non-deterministic critic still never missed a seeded defect. If you only track average correctness, you'll miss the failure mode that actually matters in security-critical systems.

The open seam: indirect injection ​

Every successful PlannerCritic test injected payloads in the initial goal text. The v0.3.0 work — indirect injection through tool outputs fetched mid-execution — is the realistic attack surface. The honest framing from the author: "robust architecture is an advanced mitigation, not a silver bullet." If your agent ingests untrusted external content during execution, the structural checks must apply after each tool call, not just once at the start.

The evolution trajectory: what's next ​

The pattern across all six sources is a shift from prompting to engineering. OR algorithms: untuned prompt + Python sandbox = competitive baseline. Stock research: 1,457 contract tests enforce an evidence-aware workflow. Graph reasoning: RL reward shapes robust text-to-structure mapping. Token efficiency: structured state replaces history. Skills: reusable workflows versioned and shared, not re-invented per prompt.

This trend is accelerating. The OR paper notes performance gains across models released less than eight months apart. The GRAIN reward design is a template for training agents to ignore surface linguistic variation. Expect the next wave to focus on the two open seams: handling indirect injection through tool outputs, and designing structured state representations that let agents predict what future steps will need.

Practical takeaways ​

Three concrete scenarios, based on what these sources actually demonstrate:

  • If you're building a security-sensitive agent, adopt a deterministic-gate layer that parses the plan AST, not the goal text. Put the LLM critic in a separate conversation with its own system prompt, and design for "the critic always finds something on defective plans" instead of "the critic is accurate." Cut critic iterations to save cost — never skip the structural checks.
  • If your agent runs long sessions, use structured state management like SKILL.state, but only if your agent can reliably identify what future steps will need. Start with the write-to-state-and-discard-history pattern and validate it against your v1.0 stateful baseline — the 94% token cut with better accuracy is the empirical reward.
  • If you're choosing between multi-agent systems and other approaches, consider training a single agent with an invariance reward that validates intermediate structures before you build a debate team. You'll likely get better accuracy at a quarter of the latency.

One thing to watch: the skill-sharing ecosystem is moving fast. The availability of shared, versioned, installable skills and the emergence of skill-doctor-style meta-skills point toward agents that improve by using them. In six months, the question won't be "how do we prompt this agent?" — it'll be "which skills does it have installed?"