Appearance
Which harness actually works
A new agent harness ships roughly every week. Claude Code, OpenCode, DeepAgents, Pi, TrueForge, plus whatever landed on Hacker News this morning. Most wrap the same loop: a model, tool calling, and context management. The real differences show up in token efficiency and model portability, not the loop itself.
I've cycled through most of the open ones this quarter. Claude Code is the most mature by a wide margin, but the token burn gets painful on long tasks that loop through a few dozen tool calls. DeepAgents is an interesting middle ground if you want structured orchestration on the LangGraph stack. TrueForge caught my attention for a different reason: it separates the model from the runtime. Swapping from Opus to GLM or a local model is a config change, not a migration.
That separation matters more than the orchestration features. If the runtime is model-neutral, you can chase the best price-to-quality point as models ship without rewriting your agent loop. And that's where the numbers get interesting.
The benchmark that cuts through the noise
I ran a comparison last week that I hadn't seen anyone do properly. Fourteen cross-system tasks with three MCP servers behind them. A CRM, an issue tracker, and a doc store. Identical model, identical prompts, identical tasks. Only the harness changed, and in one run the model under TrueForge.
| Harness + model | Tasks solved (of 14) | Cost per run | Tokens per run |
|---|---|---|---|
| Claude Managed Agents + Opus 4.8 | 11 | $11.80 | 10.0M |
| TrueForge + Opus 4.8 | 11 | $8.60 | 3.7M |
| TrueForge + GLM-5.2 | 11.7 avg | $3.00 | 3.8M |
The clean comparison is the first two rows. Both harnesses solved 11 of 14 tasks with Opus, but TrueForge used 63% fewer tokens and cost about 30% less per run. Tool call counts explain part of it: TrueForge averaged 19 calls per task against 32 for the managed setup. Fewer calls means less context pressure and fewer chances for a call to fail mid-task.
Key numbers: 63% fewer tokens per run, same model, different harness. 19 tool calls per task vs 32. TrueForge + GLM-5.2 at $3.00 per run, solving slightly more tasks on average than the managed Opus setup at $11.80.
Those numbers have real weight. At a thousand agent tasks a week, the gap between the two Opus setups is over $3,000. The token gap matters beyond cost too. Ten million tokens per run means the context window fills up, forcing compaction, slowing the loop down, and eventually losing detail. Three to four million keeps the whole task in view.
The caveats: TrueForge has real gaps. No first-class tracing or eval tooling yet. No code-execution sandbox, so you have to plug one in. And its context compaction is intentionally lossy, which is dangerous on long unstructured tasks. It is not a feature-for-feature replacement for a managed platform. What it proves is that an open, model-neutral runtime can already be competitive where it counts.
The debate on r/LocalLLaMA keeps circling whether the harness is the right layer to optimize. After running this benchmark myself, I'm convinced it isn't.
Model neutrality changes the economics
The third row of that table is the one I keep coming back to. GLM-5.2 through an open harness solved more tasks on average than Opus through the managed platform, at $3.00 per run. Opus is a great model. But if the harness wastes tokens, the model's quality is subsidizing the runtime's inefficiency.
The model sets the ceiling on intelligence. The harness decides how much of that intelligence you actually pay for.
Quick Take: A harness can swing token usage by 60% or more with the same model. Model neutrality is what lets you capture that saving without a migration.
The stack is settling into four distinct layers, each solving one problem. The harness runs the agent loop. Skills package instructions and scripts. MCP servers expose tool surfaces, and gateways route models and control spend.
Compare harnesses all you want, but you're missing three of the four layers. Gateways handle the money. MCP handles the tools. Skills handle the capabilities. The harness is the most replaceable piece, and it's the one getting the most attention.
Gateways and routers control the spend
Two gateway projects landed this week, solving different problems.
Experiential is an open source gateway that sits in front of hosted, BYOK, and local models behind one OpenAI-compatible API. You control which users and agents can use which models, and how much they can spend. The local wizard sets a $50 command budget by default, which is a spending cap you get without building billing infrastructure. It also ingests OpenTelemetry traces from your agent traffic, fits a custom router against them, and can fine-tune a model you own on the collected traces. Production traffic becomes training data without a separate pipeline.
codex-lb is narrower. It pools multiple ChatGPT accounts behind an OpenAI-compatible endpoint, tracks per-account tokens and cost with 28-day trends, and exposes per-key rate limits. Point Codex CLI, OpenCode, or any OpenAI client at it, and the pool handles failover. If you've ever burned an entire day's rate limit mid-task, the appeal is immediate.
Chart the benchmark costs and the shape of the problem is obvious.
Nobody wins that chart. The lesson is that cost per task is a property of the whole stack, and the gateway is where you get visibility into it.
Skills and MCP are how capabilities ship
OpenAI's skills repo got deprecated recently, with the project page pointing at the Plugins repository instead. That's a signal: the skill format is mature enough to fold into a plugin system. Skills are folders of instructions, scripts, and resources that agents can discover and use. Codex already auto-installs skills from .system, and the pitch is write once, use everywhere.
The capability catalog is growing fast enough that community collections struggle to keep up. Hands-On-AI-Engineering now catalogs 50-plus projects, from multi-agent research assistants with shared memory to agentic form fillers and offline medical RAG systems, all assembled from the same handful of layers.
MCP is the tool surface underneath that. blender-mcp is the best example I've seen of turning a desktop app into an agent surface. A Blender addon runs a socket server, and a Python MCP server speaks the protocol to Claude or any MCP client. You get scene inspection, object and material manipulation, arbitrary Python execution, and asset pipelines for Poly Haven, Sketchfab, and Poly Pizza's roughly 10,600 low-poly models.
| Layer | What it does | Examples from this week |
|---|---|---|
| Harness | Agent loop, context, tool execution | Claude Code, TrueForge, OpenCode, DeepAgents |
| Skill | Reusable instructions and scripts | OpenAI skills, browser-use skill |
| MCP server | App and data surface for agents | blender-mcp, GitHub MCP, Trivago MCP |
| Gateway | Model routing, budgets, key pooling | Experiential, codex-lb |
Two details from blender-mcp matter before you wire it up. First, safe mode. BLENDER_MCP_SAFE_MODE=1 validates every script before it runs and blocks file reads, network access, and anything that keeps running after the script ends. Normal modeling work still passes. I'd turn it on even for personal use, because the default is "the AI can run arbitrary Python in your 3D editor."
Second, licensing. Roughly 69% of the Poly Pizza catalog is CC-BY, which requires crediting the creator. The addon writes a formatted credit line onto each imported object as a custom property, so attribution survives in the .blend file. Filter with licence="CC0" if you'd rather skip credit handling.
Setup pain clusters in predictable places. GUI-launched MCP clients don't inherit your shell PATH, so a bare uvx command fails with spawn uvx ENOENT even though it works in a terminal. The fix is an absolute path, or a cmd /c wrapper on Windows. On machines with conda or pyenv auto-activated, uv can grab the wrong interpreter; pinning Python 3.11 with UV_PYTHON_PREFERENCE=only-managed avoids the conflict. If a failed install keeps replaying, uv cache clean blender-mcp && uvx --refresh blender-mcp clears it.
Browser automation became a skill
browser-use now ships as a skill you can paste into Claude Code, Codex, Cursor, or Hermes. The agent installs it, connects to your browser, and takes instructions. One-off tasks go through the agent. Repeatable automation uses the Python library. That split is the right call: agents handle the "do this once" work, code handles the scheduled and parallel workloads.
The project claims the top spot on the Odysseys leaderboard at 87.4% across 200 long-horizon web tasks, ahead of computer-use agents from OpenAI, Anthropic, Google, and Microsoft. Leaderboards deserve skepticism, but that's still a strong data point for a library that runs on your laptop.
The honest constraint is production. Chrome eats memory, and running many parallel browser agents gets expensive fast. The maintained path for scale is their cloud API, with proxy rotation, CAPTCHA handling, and stealth fingerprinting. Self-hosting browser agents at scale is still a real engineering project, not a config option.
Common Pitfalls
Don't let uv pick a system Python for MCP servers. On conda or pyenv machines the server can fail to install or import. Pin Python 3.11 and set
UV_PYTHON_PREFERENCE=only-managed.Run one MCP server instance per tool. blender-mcp explicitly warns against running the server in both Cursor and Claude Desktop. Port conflicts show up as random timeouts on the first command, which are miserable to debug.
Don't blame the model before measuring the harness. The same Opus model cost about 30% more per run through one harness than another on that benchmark. Compare harnesses first, then optimize the model.
Never trust lossy context compaction on long tasks. TrueForge compacts context to trim tokens, and the compaction is intentionally lossy. Fine for a 20-step task. Dangerous for multi-hour jobs, where details silently vanish and the agent starts deciding on a summary it shouldn't trust.
Don't mistake a gateway for an agent. Experiential and codex-lb route models and enforce budgets. They don't make your agent smarter. Pointing a harness at a gateway without per-key limits just moves the overspend problem one layer down.
One Thing to Remember
The harness is the least interesting layer in the stack. Token efficiency and model portability are where the money is, and capabilities are moving into skills and MCP servers. When you evaluate a new agent tool, ask which layer it actually occupies. Not whether it beats Claude Code in a demo video.
The Bottom Line
If you're running agent workloads at scale, benchmark your harness before you blame your model. Harness choice swung token usage by 63% on the same model in my test, and that's a cost and correctness issue, not a tuning nicety.
If you're cost-constrained, keep the model behind a gateway and treat it as swappable. The open harness plus GLM-5.2 ran at $3.00 per task with a comparable solve rate, about 75% under the managed Opus setup.
If you're building capabilities, ship them as skills and MCP servers. The skill format is folding into plugin systems, and MCP is becoming the universal tool surface. Anything you build against one harness should port to the next one.