Skip to content

Agents Are Finally Getting Infrastructure

#ai-agents #agent-frameworks #mcp #web-extraction #multi-agent-simulation #claude-code #autogpt

Agents Are Finally Getting Infrastructure ​

The job hunt bot that knows what you want ​

Debbie O'Brien is four months into a new role after two layoffs in a row. When she needed to search again, she didn't open LinkedIn. She opened Grok Bot, created a bot named "job hunt," and told it to look her up online and find her next role.

The bot came back with a picture of her career: Palma-based teacher builder, Playwright community at Microsoft, a short applied AI DevRel stint at Block, now platform engineer on Zephyr's AI platform. It had also found a tweet from March where she called Anthropic's developer education lead her ideal role, except for the 25% office requirement. Remote work is a hard constraint with her kids. The bot remembered both.

That's the moment agentic workflows stop being a demo. The bot runs daily searches, filters out roles she's already seen, flags the ones closest to her profile, and offers an opinion on which exception is worth testing. She can ask it to prep for a specific role, and it pulls the full job description, walks her through it, and sets up interview prep. All from a chat interface, on mobile, in the background.

The comments on her post picked up on something important. One reader noted that the happy path was never the expensive part; the boundary behavior, stale state, and partial failure paths are what decide whether an agent design holds up in production. When I've run agents like this myself, that matches my experience exactly. The impressive demo is the easy 20%. The daily grind of remembering context, skipping stale results, and handling a job board that changed overnight is where the real engineering lives.

Five layers, one stack ​

That job hunt bot looks like a single product, but it's actually the top of a stack that has quietly consolidated over the past year. Agent work has split into distinct layers, each with its own tools, its own failure modes, and its own economics.

Each layer in this cluster has a representative project that's either trending or already huge. The pattern across all of them is the same: the layer is getting boring, which is exactly what you want from infrastructure.

LayerWhat it doesRepresentative project
PlatformTurns a description into a running agent, handles scheduling and costAutoGPT, Grok Bot
HarnessRuns the agent, manages tools, context, and setupDeepSeek Harness, Claude Code
ConfigurationDefines agent behavior via agents, commands, hooks, skillsbuildwithclaude, claude-code-templates
Data and toolsFeeds agents web data and connects them to external systemsCrawl4AI, MCP servers
SimulationTests scenarios in a multi-agent sandbox before you commitMiroFish

Harnesses: the unopinionated lesson ​

The harness layer is where I've seen the most frustration historically, and the most improvement recently. A harness is the runtime an agent lives in: it manages model calls, tool execution, the context window, and the setup experience.

A recent Reddit thread about DeepSeek Harness captures the shift. The poster's complaint history is telling: they'd tried other harnesses and found setup frustrating. With DeepSeek Harness, setup was zero. Nothing to configure. The killer detail: they got it to integrate with SimpleX, an encrypted messaging protocol, by simply asking it to.

I had the same experience when I tested it. No waiting on a PR to merge. No one telling me to RTFM. No googling for community plugins. I just described what I wanted, and the harness figured out the integration. That's the difference between a harness that gets out of your way and one that fights you.

The design choice that makes this work is being unopinionated. DeepSeek Harness doesn't force you into a particular agent shape. It molds to how you want to behave. The poster's words: they'd procrastinated on writing their own opinionated harness, and the open-source one did a better job than they would have.

Quick Take: Setup friction in agent harnesses is now a choice, not a requirement. If your harness fights you on day one, that's a design failure, not a skill issue.

The configuration layer is exploding ​

Above the harness sits the configuration layer: the agents, commands, hooks, and skills that define what your agent actually does. This is where the ecosystem has exploded.

Two projects dominate this layer right now. buildwithclaude is a plugin marketplace for Claude Code with a curated collection: 117 agents, 175 commands, 28 hooks, 26 skills, and 51 bundled plugins. It also indexes the broader ecosystem: 20,000+ community plugins, 4,500+ MCP servers, and 1,100+ marketplaces.

claude-code-templates (aitmpl.com) takes a different angle. It's a collection of 100+ ready-to-use agents, commands, settings, hooks, and MCPs, installable with a single npx command. It aggregates from multiple sources: 139 scientific skills from K-Dense, 21 official Anthropic skills, 48 agents from wshobson, and workflow skills from Jesse Obra's superpowers.

Key numbers: 117 curated agents, 175 slash commands, 28 hooks, 26 skills in one marketplace. 4,500+ MCP servers indexed. 20,000+ community plugins. 1,100+ marketplaces. The bottleneck is no longer availability, it's discovery.

The practical implication of 4,500+ MCP servers: whatever tool you need, there's probably a ready-made connector. Database, API, Slack, GitHub. You don't write integration code anymore. You search a marketplace and install.

The two projects differ in philosophy. buildwithclaude is a marketplace you add to Claude Code via a plugin command. claude-code-templates is a CLI that installs components directly. One is a store, the other is a package manager.

buildwithclaudeclaude-code-templates
ModelMarketplace via /plugin marketplace addCLI installer via npx claude-code-templates
Curated contents117 agents, 175 commands, 28 hooks, 26 skills100+ agents, commands, settings, hooks, MCPs
Ecosystem index20k+ plugins, 4,500+ MCP servers, 1,100+ marketplacesAggregates from 10+ named sources
ExtrasWeb UI at buildwithclaude.comAnalytics, mobile chat view, health check
SponsorshipNoneBright Data (web data skills and MCPs)

The warning sign in both: 175 commands is a lot of surface area. Installing everything is how you end up with conflicting agents and ambiguous slash commands. More on that in the pitfalls section.

Platforms: AutoGPT wants your whole pipeline ​

The platform layer sits above the harness. AutoGPT is the clearest example. 185,000+ GitHub stars makes it one of the most battle-tested agent projects in the open-source world. Andrej Karpathy called AutoGPTs "the next frontier of prompt engineering."

The current AutoGPT is a full platform, not a single agent. AutoPilot turns a plain-English description into a working agent. The visual builder lets you drag, connect, and branch blocks for exact control. The marketplace lets you start from proven agents and customize them. Agents run on demand, on a schedule, or from a trigger, and connect to 45+ platforms including Gmail, Slack, Notion, Jira, Salesforce, and Stripe.

The 45+ connected platforms matter. Your agent can read your calendar, draft replies in your CRM, and file issues in your tracker without you writing a single integration. The scheduling and trigger system is what separates this from a chat bot: an agent that runs every morning at 6am and reports back is qualitatively different from one you have to prompt.

AutoGPT's real decision point is hosting. The managed platform is paid, with usage-based agent runs. Self-hosting is free under the Polyform Shield license, but you bring your own infrastructure and model keys.

AutoGPT PlatformSelf-hosted
CostPaid plan plus agent usageNo license fee; pay your own infra and models
SetupManagedDocker and configuration required
Model accessBuilt inBring your own API keys
OperationsManaged by AutoGPTManaged by you
Data controlHosted by AutoGPTRuns on your infrastructure
License constraintn/aPolyform Shield: free for personal and internal use, can't resell as a competing hosted service

The license detail deserves a second read. Polyform Shield means you can use AutoGPT freely inside your company, but you can't wrap it and sell it as a hosted service. If you're building a commercial product on top of agents, that constraint shapes your architecture.

Web extraction: the layer everyone forgets ​

Every agent above needs data. Most of that data lives on the web, and most of it is a mess. This is the layer everyone forgets until their agent starts hallucinating because it was fed garbage.

Crawl4AI exists because its creator needed web-to-Markdown conversion in 2023, and the "open source" option wanted an account, an API token, and $16, and still under-delivered. He built Crawl4AI in days. It went viral. It's now the most-starred crawler on GitHub with 50,000+ stars.

The core idea is simple: turn web pages into clean, LLM-ready Markdown. Headings, tables, code blocks, and citation hints preserved. Navigation noise removed. There's a BM25-based filter that keeps only the content relevant to a query, and a pruning filter that drops low-value sections entirely.

Two features stand out for agent work. Prefetch mode gives 5-10x faster URL discovery, which means deep crawls that used to take an hour finish in minutes. And the Docker deployment now ships with MCP integration, so Claude Code and similar tools can query it directly.

The security history is also instructive. The v0.9.0 release made authentication default-on because the Docker API had critical vulnerabilities: RCE, SSRF, auth bypass, file write, XSS, and a hardcoded JWT secret. That's a reminder that agent infrastructure is server infrastructure. If you self-host it, you own the security posture.

Multi-agent simulation: rehearse before you commit ​

The newest layer is simulation. MiroFish is a multi-agent prediction engine that builds a parallel digital world from seed information: breaking news, policy drafts, financial signals. Thousands of agents with independent personalities, long-term memory, and behavioral logic interact and evolve. You inject variables from a "God's-eye view" and watch the trajectories.

The pipeline is straightforward: extract seed entities, build a graph with GraphRAG, generate personas, run the simulation on a dual-platform engine, then generate a report. The simulation engine is powered by OASIS from CAMEL-AI, which is a nice example of the stack consolidating: even a prediction engine builds on someone else's agent infrastructure.

Use cases range from serious to playful. Policy teams can test public relations moves at zero risk. Novelists can ask what happens next in a story: MiroFish's demo includes predicting the lost ending of "Dream of the Red Chamber" from the first 80 chapters. Financial and political prediction examples are in the works.

I'm skeptical of the prediction claims, and you should be too. Multi-agent simulation is compelling as a rehearsal tool, but treating simulated social evolution as ground truth is a leap. What's genuinely useful is the test bench: before you deploy an agent into production, you can run it against a simulated environment and watch for emergent failure modes.

Common Pitfalls ​

I've been running agents and watching this ecosystem for a while. Here's where people trip up.

  1. Confusing the harness with the agent. The harness is the runtime. Your agent's behavior lives in the configuration layer: the agents, commands, and skills you install. When an agent underperforms, audit those definitions first, not the harness. Blaming the runtime for a prompt problem wastes a week.

  2. Feeding agents garbage data. An agent is only as good as its extraction layer. If you're scraping raw HTML or pasting noisy web content into a context window, no prompt engineering fixes it. Use an LLM-ready extraction pipeline like Crawl4AI's fit-Markdown or a proper MCP connector. The cost of bad data compounds: every downstream step inherits it.

  3. Installing everything. The marketplaces make it too easy. all-agents@buildwithclaude gives you 117 agents that overlap and conflict. Commands with ambiguous names shadow each other. Start with 3-5 components that map to your actual workflow, then add as you hit real needs.

  4. Self-hosting with default security. Crawl4AI's Docker API shipped with RCE and SSRF vulnerabilities until the maintainers made auth default-on. If you self-host any agent infrastructure, bind to loopback, require tokens, and treat request bodies as untrusted. The default posture of "it's just an internal tool" is how you get pwned.

  5. Assuming stateless agents hold up. Long-running agents forget context, re-read stale pages, and lose track of what they've already reported. The job hunt bot works because it skips roles it already showed you. That requires explicit state management: resume state, memory stores, on-state-change callbacks. If your agent doesn't have these, it will fail on day three, not day one.

One thing to remember ​

The agent stack is young but it's converging fast. Every layer now has a boring, reliable default: AutoGPT for platforms, DeepSeek Harness for runtimes, a marketplace for configuration, Crawl4AI for data, MiroFish for simulation. The job hunt bot works because each layer got boring enough to ignore. That's the definition of infrastructure maturing.

The Bottom Line ​

Three takeaways, depending on where you sit.

If you're building a personal agent for job hunting, research, or daily monitoring, skip the platform entirely. Use a harness like DeepSeek Harness or Grok Bot, give it your public profile, and let the configuration layer do the work. You don't need to write code; setup is minutes, not days.

If you're building agent infrastructure for a team, start with the data and tools layer. Wire up Crawl4AI for web extraction and MCP servers for tool access before you touch the platform. That's where production agents fail, and it's the layer that's cheapest to get right early.

If you're choosing between managed and self-hosted agents, let the license and data constraints decide. AutoGPT's Polyform Shield means you can't resell it as a hosted service, so a commercial product needs its own platform layer. If data control matters more than convenience, self-host. If you want agents running today, pay for the managed platform.

One thing to watch: the simulation layer is the least proven and moving fastest. Expect the harness and configuration layers to consolidate within six months, and expect multi-agent simulation to either produce prediction value or fade into demo territory. Either outcome will be visible from the GitHub trending page.