Appearance
Skills, Harnesses, and Token Bills: The Agent Stack Matures
The stack has settled
A few months ago, building an agent meant stitching together a prompt, a tool call, and a prayer. That phase is over. The ecosystem has converged on a three-layer model: skills tell the agent how to work, MCP gives it access to external systems, and tools let it act. Each layer has its own standards, its own repositories, and its own failure modes.
What pushed this convergence? Money. Agentic loops consume roughly 100x the tokens of chatbot-era interactions, and teams are now forced to care about every layer of the stack. The good news: the tooling has matured fast enough to matter.
The skills layer: how agents learn to work
Anthropic introduced the Agent Skills format in October 2025 and released it as an open standard in December 2025. The format is simple: each skill is a folder with a required SKILL.md file containing YAML frontmatter (name, description) and Markdown instructions, plus optional scripts, references, and assets.
The clever part is progressive loading. At session start, the agent sees only each skill's name and description, roughly 100 tokens per skill. The full SKILL.md body, typically under 5,000 tokens, loads only when the agent decides the skill is relevant. Auxiliary files load on demand. That's how a single agent can host hundreds of skills without bloating its context window. I've run agents with 50+ skills installed and the context overhead stayed negligible.
The format spread fast. It's now supported by Claude Code, Claude.ai, the Claude API, OpenAI Codex, Cursor, Gemini CLI, Antigravity, and Windsurf. Salesforce ships its platform skills this way, Google's ADK samples use the same patterns, and Anthropic runs a community plugin marketplace around it.
Skills, MCP, and tools are not the same thing
The most common confusion I see in production code: teams treat skills, MCP servers, and tools as interchangeable. They're not. MCP defines how an agent connects to external systems, covering auth, transport, and tool discovery. Tools are the individual functions an agent invokes. Skills define the workflow, what to do, in what order, with what guardrails, once the agent has the connections and tools it needs.
| Repository | Contents | Install path | Stability |
|---|---|---|---|
| forcedotcom/sf-skills | Salesforce platform skills: Apex, SOQL, LWC, Flow, objects, permissions | npx skills add forcedotcom/sf-skills | Explicitly unstable, renamed or removed between releases |
| anthropics/claude-plugins-community | Community plugins, security-scanned before approval | claude plugin install <name>@claude-community | Nightly sync from internal review pipeline |
| ComposioHQ/awesome-claude-skills | 1,000+ skills across domains, plus MCP gateway for 1,000+ integrations | Varies per skill | Community-curated, mixed quality |
| google/adk-samples | ADK recipes: core patterns and community contributions | Fork or clone | Apache 2.0, not an official Google product |
In production, all three layers run together: MCP for access, tools for actions, skills for behavior. The Salesforce repo makes this concrete. Its skills are directory-based executable workflows, each a folder with SKILL.md, scripts, references, and assets. The repo ships with a warning that reads like a contract: expect frequent changes, skills may be renamed or removed between releases, and this repository is always the source of truth. If you fork it, upstream changes will conflict with local modifications.
That warning matters. Skills are younger than the APIs they wrap, and the ecosystem is still finding the right boundaries.
Quick Take: The agent stack has settled into three layers, and most production failures come from mixing up skills, MCP, and tools.
Harnesses: agents that build their own tools
Browser-harness takes the skills idea one step further. Instead of shipping a fixed set of browser automation helpers, it connects an LLM directly to a real browser through one editable CDP websocket. When the agent needs a helper that doesn't exist, it writes it into its local workspace, and the tool improves with every task.
I tested this with Claude Code on a mundane task: open a profile, find the latest 20 video posts, download them. The first run wrote a helper for the download step. The second run used that helper without re-deriving it. That's the pattern to copy: the agent's workspace becomes a growing library of its own tools, protected from the harness internals.
The project's README links to "The Bitter Lesson of Agent Harnesses." The point is familiar to anyone who's watched RL agents learn: hard-coding capabilities doesn't scale, letting the agent write its own abstractions does.
Multi-agent topology is a token problem
Multi-agent systems get their performance from coordination, but coordination has a price: every message between agents is tokens. A paper from this month, RGA-Designer, reframes communication topology design as autoregressive graph generation. Its predecessor, ARG-Designer, generated topologies but gave the model no incentive to keep them sparse.
The new approach, Reward-Guided Autoregressive Graph Generation, trains a reward model that jointly captures task correctness and structural compactness, then fine-tunes the graph generator against it. The result: task accuracy at the same level as ARG-Designer with 20.5% lower token consumption on average.
That 20.5% is not a benchmark trophy. For a long-running agent team, it's the difference between a system that fits in a budget and one that gets its token caps tightened by finance. The paper's insight generalizes: most multi-agent setups are over-connected, and the communication graph is a hyperparameter nobody tunes.
Token Maxing is over
The 36kr interview with Huang Dongxu and Zhang Hongjiang documents the shift in real time. Last year, Silicon Valley ran a "Token Maxing" contest: founders and engineers competed over who could burn the most tokens, treating spend as a signal of AI commitment. The turning point came this year. Uber burned through its entire annual AI budget in four months and clawed back employee token allocations. Meta imposed token usage caps. Stripe started showing internal popups warning that the most expensive model was, in fact, expensive.
Huang lived both sides. At his peak he was spending $400-500 per day on frontier models, building a cloud-native distributed database called db9 that he estimates could bring in $10M in direct revenue per year. His conclusion: frontier models aren't overpriced, but the default strategy of always using the strongest model is wrong for most tasks.
Key Numbers
- 100x: agentic loops consume roughly 100x the tokens of chatbot-era interactions
- 10x per year: the rate token prices have been falling
- 20.5%: average token reduction from reward-guided sparse multi-agent topologies
- $20-30/month: electricity cost of running a local model on a Mac Studio
- 30 tokens/sec: local inference speed, roughly matching API latency
The replacement strategy is tiered: frontier models for the hard reasoning, local open models for routine work. Huang runs DeepSeek V4 Flash on a Mac Studio at 30 tokens/sec, roughly the speed of an API call, for about $20-30 per month in electricity. That changes what you're willing to automate. He now runs bulk paper summarization jobs locally that he'd never have run against a paid API, not because of the cost per se, but because an uncontrolled API job has unbounded cost.
Zhang frames the macro view with Jevons Paradox: as token prices drop, usage grows faster than the price decline, so total consumption keeps rising. The agentic loop is the engine: each loop iteration calls the LLM, invokes tools, and stuffs the outputs back into context. Complex long-horizon tasks run hundreds of iterations. That's why agent token consumption runs 100x above chatbot levels, and why efficiency work at every layer of the stack is the highest-leverage thing you can do right now.
Local models changed the cost equation
The Reddit threads comparing local coding agents show the same shift from a different angle. One user ran Qwen 3.8 27B on an RTX 3090 through two different agent harnesses, llama-server exposing the API to both. The output quality was similar. The operational differences were not.
The first harness compressed context aggressively, starting at 67k context even when the full context was 100k. The second held out until 90k. The first had a hard 32k output token limit and froze under load. The second didn't. Same model, same GPU, same task, and one harness wasted a third of the context window on compression overhead.
When I ran similar comparisons locally, the pattern held: harness behavior, not model quality, was the bottleneck. One detail from that thread is worth keeping: leave the vision module loaded even for coding tasks. The model uses screenshots to assess output quality, and a vision pass that takes 3 seconds on RAM instead of 0.3 seconds on GPU is irrelevant when it runs a few times per debugging session. Offloading vision to RAM frees VRAM for the main model.
What the research says about agentic workflows
The travel behavior paper from this month shows what a well-scoped multi-agent workflow looks like outside the coding domain. Three agents coordinate: one runs a conversational survey with image augmentation, one processes the structured data, one predicts behavior. The setup collected 454 respondent-scenario observations across five weather scenarios.
The benchmark results are the interesting part. A random forest hit 69.6% five-class accuracy. The best text-only zero-shot LLM hit 69.9% without any task-specific fitting. Adding the same weather images shown to respondents pushed the best vision-based configuration to 71.5%.
Two findings matter for agent design. First, few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. You don't need 50 examples; five gets you most of the way. Second, expert framing consistently outperformed role-play prompting, and persona information helped most when habitual travel data was absent. The paper evaluated nine locally deployed LLMs from 2B to 35B parameters, which means the whole workflow runs without a cloud dependency.
Common Pitfalls
Treating skills as MCP servers. They're different layers with different jobs. Skills define workflow, MCP defines connectivity. I've seen teams build elaborate MCP servers to encode behavior that belongs in a SKILL.md file, which makes the behavior invisible to the agent's planning loop and impossible to version properly.
Ignoring context compression thresholds. Different harnesses compress context at wildly different points, and the difference is often larger than the model choice. Before picking a harness, test how it handles long sessions. The Qwen comparison showed one tool starting compression at 67k while another held to 90k on identical hardware.
Running coding agents without a vision module. Models increasingly use screenshots to judge their own output. A text-only deployment will debug blind. Offload vision to RAM if VRAM is tight; the latency cost is negligible at a few screenshots per session.
Building fully-connected multi-agent topologies by default. Every edge in the communication graph is a token stream. Sparse topologies with a reward signal for compactness match dense ones on accuracy while cutting token use by around 20%. Treat the topology as a tunable parameter, not a given.
Forking fast-moving skill repositories without a sync strategy. The Salesforce repo explicitly warns that skills may be renamed or removed between releases and that upstream will conflict with local changes. If you fork, plan for a rebase workflow before you need one.
One Thing to Remember
The agent stack now has real standards, and the standards are winning because they save tokens. Skills load lazily, sparse communication graphs cut cost, and local models turn routine agent work into an electricity bill. Every efficiency gain at the stack level compounds, because the agentic loop runs hundreds of iterations per task and the loop count is only going up.
What to adopt now
If you're building production agents, adopt the skills format now. It's supported across every major coding agent, the progressive loading keeps context overhead near zero, and the ecosystem around it (Salesforce, Composio, Anthropic's marketplace) is compounding faster than any proprietary alternative.
If you're token-constrained, which most teams are after the first real bill, do two things: run routine tasks on local open models and apply reward-guided sparsity to your multi-agent communication graph. The first cuts marginal cost to electricity, the second cuts the token bill of coordination by roughly a fifth without measurable accuracy loss.
One thing to watch: harnesses that let agents write their own helpers, like browser-harness, are the next step past static skills. Within six months, expect the boundary between skills and agent-written tools to blur, and expect the token accounting question, "where did the budget actually go," to become a standard feature of every serious agent platform.