Appearance
The Agent Stack Is No Longer About the Model
The easy part is over
The LLM is the easy part of an agent. Writing the loop is easy too, roughly 80 lines of Python once you know the pattern. The hard part is everything that wraps that loop: memory that survives an editor swap, procedures that survive a version bump, logs that show causality instead of just order, and permissions that don't turn your host machine into the blast radius.
That shift is the real story in agentic engineering. The model matters less every quarter. The runtime around it matters more. The past month of releases shows exactly where the effort is concentrating: memory you own, skills that encode hard-won details, execution trees instead of flat logs, and benchmarks that refuse to be gamed.
Memory has two jobs and they are not the same
I ran the cuchly test myself after reading the writeup. Teach a fact to your assistant in a fresh session. Then ask from a different editor. Then from a different machine. Then ask a colleague. The shape of the four numbers tells you which product you are holding.
A vendor's memory is bound to the vendor's runtime. It makes one assistant continuous with itself. That is genuinely valuable, and it is what most people mean when they ask for memory. The test showed it clearly: 5/5 recalls in the same tool, 0/5 everywhere else. That is session memory. It answers the question "what did I just say?"
Project memory is bound to the repository instead. It spans editors, machines, people, and model upgrades. When someone leaves your team, session memory leaves with them. Project memory does not, because it was never theirs. It answers the question "what does this project know?"
The deeper insight comes from the comment thread on that post: a symptom and its fix share almost no vocabulary. "deploy hangs at Build image" and "runner disk was full; docker prune" sit in different embedding neighborhoods. Semantic closeness and causal connection are different relations, and a vector index measures only one of them. The author who built that store added explicit edges (caused-by, fixed-by, superseded-by) beside the similarity path. That is the only reason the store earns its keep. funes from Hugging Face makes the same bet. It parses agent traces into turns, embeds locally, fuses BM25 with vectors, reranks with a cross-encoder, and reweights by recency. It also keeps the raw text. No fact is distilled at write time. A summary flattens the exact finding you needed three weeks later.
One thing I found in the comments: a similarity search will happily return a generic hub node as the nearest neighbor to everything. Several people reported that general lessons like "cache stampede" or "retry idempotency" showed up in half of all recalls. The fix is not to remove them. It is to demote any node that is a neighbor to everything.
Quick Take: If your memory dies when your editor does, you don't have project memory. You have a session cache.
Skills are turning prompts into procedures
Skills are folders with a SKILL.md: name, description, markdown instructions. Anthropic's repo shows the template, and the Agent Skills spec is open at agentskills.io. Simple enough. The power shows in what people are packaging on top of it.
publishing-kit packages the entire multi-platform publishing lifecycle as a Claude Code skill. Write one markdown file, and it builds the dev.to, AWS Builder Center, Medium, and LinkedIn versions, checks them, and posts where an API exists. The interesting part is how many hard-won details it encodes. dev.to renders markdown with hard breaks on, so a source wrapped at 95 columns arrives with a stray break in every paragraph. The author measured it: 47 of 62 paragraphs in a published article carried a stray break. Medium drops data URI images on paste. 0 of 4 survived; from real URLs, 4 of 4 survived.
You do not need to rediscover those facts. The skill already knows them. That is the entire point of the format.
The humanizer skill goes further. It embeds 35 patterns from Wikipedia's "Signs of AI writing" page. It makes a first pass, checks the draft against the patterns and the original claims, then rewrites anything that still sounds artificial. It won't invent a name, a date, or a number. The end result is a first-person piece of prose that reads like it came from a human, because the patterns that machine text hits are the exact things it removes.
The academic-research-skills suite shows the ceiling. It runs a 10-stage pipeline with integrity gates, citation verification, and multi-agent peer review. A real run caught 15 fabricated references and 3 statistical errors before the paper ever reached a reviewer. That is not a toy. That is a procedure encoded well.
| Layer | What it holds | Key tool this month | The insight that changes how you build |
|---|---|---|---|
| Memory | Cross-session, cross-tool knowledge | funes, cachly | Similarity search cannot connect symptom to cause; store explicit causal edges. |
| Procedures | Repeatable instructions plus scripts | Anthropic skills, publishing-kit, humanizer | Skills encode measured quirks like 47/62 hard-wrapped paragraphs, so you stop rediscovering them. |
| Observability | Execution paths and ownership | AgentInspect, Pier | Trees show the path, not just the answer; expose failures that recovery hides. |
| Authority | Permissions at the OS level | ShrekOS | Move the trust boundary out of the agent; let the OS grant capabilities on request. |
| Benchmarks | Long-horizon, verifiable tasks | DeepSWE, CAD 1000 Hours | Use program-based verifiers in isolated environments instead of LLM judges. |
Debug the path, not the answer
Flat logs make me invent causality. The AgentInspect writeup makes the strongest case I have seen for execution trees as the default debugging projection for agent runs.
A support agent run: plan starts, inventory request fails with a 503, retries and succeeds. As a tree, ownership is explicit. The indentation is not decoration. It tells you which higher-level operation owned the failure. With flat logs, you match IDs to infer that same structure, and you often guess wrong.
The tree also shows failure that success hides. A fallback that saves the run is visible. A retry that cost 250ms of latency is visible. Without it, the final answer looks identical and the dysfunction disappears into the word "success."
When I examined my own traces, I found a fifth shape missing from the four canonical ones. Sibling calls to the same tool with paraphrased inputs and no failure between them. Not a retry, nothing failed. Not a fallback, nothing was replaced. The agent recalled, got a plausible answer, then recalled again with a reworded query. The tree turned it into a countable thing instead of the vague feeling that the agent was chatty. We built a deterministic check out of it: same tool, siblings under one parent, input similarity above a threshold, zero failures in between.
One caveat from the comments deserves wide circulation. An empty retry shape is not evidence that no retry happened. A step wrapper around a client that retries internally records one node whether the upstream saw one request or three. The tree becomes the artifact people trust, so trace metadata needs an explicit field: attempts_observed versus attempts_unknown.
The next generation of benchmarks refuses to be gamed
DeepSWE brings 113 tasks drawn from active open-source repositories across TypeScript, Go, Python, JavaScript, and Rust. 113 tasks is small enough to run in a day but varied enough that memorization is useless. It uses program-based verifiers instead of LLM judges. The reference patch is held out. Grading happens in a pristine container via the Pier framework, which forks Harbor to support CLI agents in air-gapped tasks. Pier adds per-agent network allowlists, so agents get only the access they need.
On the computer-use side, CAD 1000 Hours gives us 1,021.64 hours of recorded workflows across 10 CAD, BIM, and structural-analysis applications. That is 256.6 GB of screen recordings, synchronized mouse and keyboard events, and frame-level narration. 82,350 downloads last month tell you the demand. That volume of authentic interaction data is what a model needs to click through SOLIDWORKS at inference instead of seeing an unfamiliar UI for the first time.
The measured numbers back up why this data matters. A chart of what these tools actually cost and produce:
The zero represents a failure to arrive. On the harder task, compaction flattened the finding that mattered, so the question could never be answered.
The authority problem: who decides
ShrekOS names the tension that every serious agent builder eventually hits. An AI agent should not have to be trusted with the whole machine to be useful on the machine. Every time you make an agent more useful, you drift toward one of two bad places. Restrict until useless, or trust until dangerous.
Containers, namespaces, cgroups, egress filtering, AppArmor, immutable base images. Each solves a piece of the puzzle. None of them gives you a coherent operating model for a workload that decides at runtime which capabilities it needs. The project is composition, not invention. The author planned to install ffmpeg through a naive apt install on the host, then realized the permanent mutation touched the package database, pulled un-audited dependencies, and ran install scripts as a normal operation. Multiply that by every tool an agent decides it needs over months, and the host becomes an unpredictable pile of state.
The fix is to move the trust decision out of the agent entirely. The agent is a client. It can be trusted to make a request. The operating system decides whether to grant it, and under what constraints.
Common Pitfalls
Indexing traces on arrival instead of on resolution. If you index the thrashing, you train your memory on dead ends. A session that ends without a resolution still writes its last state. The worth-while unit is the symptom, cause, and fix triple, and that triple only exists once the incident closes. Write at close time, looking backward.
Letting a green retry gate lull you. A max-calls check counts instrumented calls. If the client retries below your wrapper, the gate stays green while the dependency hit the endpoint three times. Record attempts_observed versus attempts_unknown, and fail closed when evidence is incomplete.
Treating a clean execution tree as proof the answer is correct. Trees show the path, not the content. A required retrieval step can return irrelevant documents. A model call can produce unsupported claims. Use semantic evaluators and human review for content quality.
Running the agent as root so it can install ffmpeg. Give the agent a capability request path instead of a password. An OS-level policy file defines what an agent may touch before it starts. The boundary is a hard limit, not a suggestion the model has promised to respect.
Trusting embeddings for causal questions. Neural search has no native concept of "far from everything." It always returns an argmax. Build a refusal path: below some confidence, the correct answer is "no known cause," not the nearest neighbor.
One thing to remember
Memory, skills, execution trees, and OS-level permissions are all answers to the same question: how much of the messy decision-making can you move out of the model and into deterministic, versionable infrastructure? Every piece of this stack does exactly that.
The Bottom Line
If you are building agents that act on your actual machine, adopt an OS-level permission layer or a strict sandbox. Agent-side promises not to use curl dangerously do not hold up under prompt injection. Give the agent a request path, not the host.
If your agents work across a team, skip vendor-locked session memory and wire in a project memory that spans editors and models. The knowledge must survive the person who learned it and the model that wrote it. Test it with the four-boundary script before you trust it.
One thing to watch: the Agent Skills standard is consolidating around agentskills.io, and Anthropic's reference repo is open. Within six months, expect skills to absorb most of the custom MCP glue people write for deterministic procedures. Memory and permissions will move down the stack into the operating layer. Procedural knowledge will move into versioned artifacts you can diff and audit.