Appearance
The Best Agentic Design Is Subtraction
Four agent projects crossed my desk this week. A Liar's Dice game where the AI keeps a written profile of how you bluff. A health-data Q&A app that refuses to give the model SQL. An open-source context layer for enterprise search. An app store for agent platforms. Different domains, different stacks, same conclusion: the model is the least interesting part of an agentic system.
The real work lives in the surrounding layers. What the agent remembers between sessions. Which tools it can reach, and which it can't. What context it's allowed to see. And in every project, the winning move was subtraction. The authors removed things: a character prompt, a calculator, SQL access, a feature that couldn't fire. Each system got better for it.
This is what production agentic design looks like now. The work is in the unglamorous infrastructure: memory, permissions, and tool surfaces. Prompt tweaks and model swaps are the easy part. The demo era of agents was about what a model could do on its own. These projects are about what a model can do when the system around it is honest about its limits.
Memory that survives the match
Haoxiang Li built Kai!, a Liar's Dice game with an LLM sitting across the table. Each player sees only their own dice, then bids on what the table contains. The next player raises the bid or calls the bluff. Simple rules. The work ended up outside the rules.
The first opponent was a character prompt. Old Li, the table owner: proud, vindictive, particular about how he spoke. His flaws, strategic preferences, and voice went straight into the prompt. It worked immediately. A few days later, Li deleted him.
The problem: swap the underlying model and you still meet Old Li. The model was performing a written character, which made it impossible to see whether different models produced meaningfully different opponents. Every model now receives the same system prompt. It doesn't even know its display name. The dramatic quality is less predictable, but the differences between models finally have room to show up.
Removing the character prompt created a new problem: why should the player feel that the opponent in the next match is the same opponent? That responsibility moved into the profile system.
After each match, the game keeps two kinds of records. The first kind is recomputed by the deterministic engine: how often the player bluffed, when they called, how accurate those calls were, which rounds cost a die. The model cannot rewrite those facts. The second kind contains the model's own opinions: this player backs down when the multiplier is high; they like to calculate before bidding; that pause looked like weakness. Those opinions are allowed to be wrong. In practice, the mistakes may be the more interesting part, because the profile is visible to the player.
If the model decides you're honest, you can use that belief to cover a bluff in the next match. If it thinks you fold under pressure, you can deliberately hold your ground. The model reads the new behavior and updates the profile. It observes you, while you manage the version of you that exists in its notes.
One self-play log made this click. DeepSeek V4-Pro held 3, 6, 5, 6: two sixes, no wild one. On its turn it looked at the dice, used the probability action, and bid five sixes. The private reasoning stored alongside the move said something more interesting. Its opponent believed it was generally truthful and liked to calculate before bidding. Repeating the familiar look-then-calculate routine could make the aggressive bid seem more credible. DeepSeek was also ahead on chips and could afford to be caught.
The model was spending credibility it had accumulated over previous matches. No prompt told it to build trust and then bluff. The idea came from the cross-match profile.
Li is careful to call this a single example, not evidence of a stable model personality. A different seed could produce a completely different line. The useful result is narrower: a judgment created in one match can make its way into a later decision. That's the unit of agent memory that matters. There's an agent memory leaderboard on Hugging Face now, which tells you how fast this problem is becoming its own discipline.
The reaction I kept seeing from people who played it: the AI remembering how you play makes the game feel personal in a way a scripted opponent never does. That's the payoff of the profile system.
Quick Take: every project in this roundup got more interesting when the author removed something from the agent.
Structure beats prompts
Matt Stratton built a web app to answer questions about his own health data. The obvious design was to hand the model SQL: give it the schema, a read-only connection, and let it write queries. It's the design every chat-with-your-database demo uses, it takes an afternoon, and for a database this small the query cost is irrelevant.
He rejected it. There are six traps in this data, and every one has already produced a wrong answer at least once, on a page he was looking at.
| Trap | What naive SQL does | What the tool does |
|---|---|---|
| Partial day | Compares an accumulating day to a completed one | Windowed queries end at today_local() |
| Gap = zero | Zero-fills missing days | Gaps stay absent rows |
| Apple shadow copies | Doubles training volume | Workout view is unreachable |
energy_balance overstates | Reports a 2.7x deficit | Reality check arrives in the same payload |
| Single weigh-in noise | Treats water weight as a trend | Signals use windows, not single days |
reps: 0 | Reads as a missing set | Counts as an attempted, failed set |
Why not put the traps in the system prompt? You could, and the model would get them right most of the time. That's the worse outcome. A tool that's wrong every time gets caught on day one and thrown away. A tool that's right ninety-something percent of the time gets trusted, and then the rare wrong answer arrives wearing exactly the same confident formatting as the right ones. You have no way to spot it, because the whole reason you're asking is that you don't already know the answer.
So the traps are foreclosed by the shape of the tools, not by instructions. The partial day isn't excluded by the model remembering to exclude it. It isn't reachable. Gaps stay absent rows, so there's no zero to misread. No tool reaches Apple's workout view at all. And energy_balance cannot be fetched without its reality check arriving in the same payload.
The prompt still describes all six traps, because a model that understands why a window ends where it does gives better answers than one that just gets truncated data. But the prompt is not what's holding them. If the prompt were deleted tomorrow, the answers would get worse, and they wouldn't get wrong in those six specific ways.
The energy balance trap is the one to internalize. Over the last 30 complete days, the field reports an average intake of 1,602 kcal against 3,216 burned, a net deficit of 1,614 kcal a day. That predicts losing 3.2 lb a week. The scale over the same window says 1.2. Basal energy is a formula estimate from weight, height, and age. Watch-measured active energy runs generous. Both are real numbers whose difference is not a measurement.
The tool surface is the product
The same project produced a subtler failure. The chat needs to know which metrics exist, so there's a metric_catalog table holding canonical units. A list_metrics tool read from it. It compiled. It typechecked. It returned rows that looked entirely plausible. The catalog covered 38 of 81 metrics. The other 43 weren't stale junk: walking and running distance (3,865 days of it, updated today), flights climbed, the entire micronutrient panel. The tool would have worked exactly as written and made the chat answer "I don't have that" about ten years of walking distance sitting right there in the database.
No test would have caught this. The function did what it said. The fix inverts the join: drive from observations_daily, LEFT JOIN the catalog, and an uncatalogued metric shows up with catalogued: false and a null unit rather than not showing up. The general shape: a lookup table is not an index of reality unless something enforces that it is.
False precision is a real cost too. Raw rows carry numbers like 1370.9033333354562 for a day's calories. A scale that reports to a fifth of a pound did not measure thirteen decimal places, and an answer quoting them implies rigor the data doesn't have. It's also pure token cost: eighteen characters where six will do, on every point of every series. Rounding to two decimals at the tool boundary took roughly a third off the nutrition payload for a seven-day window.
The most humbling failure was a fact in the system prompt. The app told Stratton his sleep coverage was roughly 7% of days. That was true months ago. Over the last 30 days it's 70%, because his watch-wearing changed in July. The model wasn't guessing and wasn't reading the data. It was reading the system prompt, where "Sleep has roughly 7% coverage" sat as a standing fact. The chat told him something false about his own body, sourced from the one component he'd exempted from his own argument. The tools compute. The prompt asserts. Assertions rot.
Then he went to delete it and found out why it had survived. A test asserted that the prompt contained the string "7% coverage". The stale number was under test. A wrong fact had acquired a defender.
Key numbers from the health-data project
- 38 of 81 metrics were catalogued; the other 43 included 3,865 days of walking data
- 2.7x: how much
energy_balanceoverstates the deficit- 37 cents: total cost of nine probe questions, from 1.3 cents for a simple query to 9 cents for a year of series data
- 7,069 tokens: the cached system prompt and tool definitions, read twice per two-turn question
The context layer is a governance problem
The other two projects in this roundup are about the layers around the model at platform scale. PipesHub is an open-source context layer for enterprise AI. The pitch: connect enterprise knowledge, preserve access permissions, generate citations, and build agents, RAG applications, MCP servers, and agentic workflows on a single governed context layer.
The key word is governed. Permission-aware search enforces source-level access controls, so users only see what they're authorized to see. Answers carry precise block citations to original documents. Retrieval is graph-backed, capturing relationships across enterprise data rather than relying on vector similarity alone. There are 50+ enterprise connectors with real-time and scheduled indexing, and the whole thing is self-hostable with a bring-your-own-model setup, so data never leaves your infrastructure.
This is the enterprise version of the same lesson. In the health app, the failure modes were data traps: partial days, shadow copies, overstated deficits. In an enterprise, the failure modes are permissions and provenance. An agent is only as trustworthy as the context layer that feeds it, and a context layer that doesn't enforce permissions will produce confident answers that leak data or cite documents the user can't see.
Packaging agents as apps
Kiro Crew is the other platform-scale piece. It's an open-source development workspace that runs agents persistently: multi-step tasks run unattended, cron jobs fire on schedule, heartbeats monitor systems until something needs attention. The interesting part is the App Kit. An app is a package that contributes any combination of the following.
| Component | What it contributes |
|---|---|
| Agents | Custom agent with its own model, prompt, and tool access |
| Skills | On-demand markdown knowledge files |
| MCP servers | New tools the LLM can call |
| Cron jobs | Scheduled tasks the app owns |
| UI pages | Custom dashboard pages in the sidebar |
| Backend processes | HTTP servers reverse-proxied through the gateway |
The standup bot example is five files: a manifest, an agent definition, a skill markdown file, a dashboard page, and a cron expression. The skill is just a markdown file with formatting rules that loads on demand when trigger words appear in the conversation. The cron is a line in app.json: 0 9 * * 1-5, which spawns a session every weekday at 9 AM, runs the message through the agent, and stores the result. No daemon. No systemd timer. One line in the manifest.
The app store model is the part to watch, because it changes who gets to build agents for a platform. Apps are installable, versioned, publishable, and isolated. Users install with one click. Teams can host private registries for internal apps. Publishing is a PR to a registry file. The store is empty right now, and the author's point is blunt: first movers win.
Common pitfalls
Trusting a completed run over a correct one. When I tested early versions of the Liar's Dice game, I found that a completed match did not necessarily mean the model completed it. Truncated or invalid outputs could silently hand control to a fallback bot. The match still finished, so the dataset looked healthy, but some "model behavior" actually belonged to the recovery path. Record retries, repairs, and fallback takeovers beside every result, and retain the raw calls. Build that observability before building the leaderboard.
Putting facts in the prompt instead of computing them. The 7% sleep coverage fact rotted, and a test defended it. If a number can be computed, compute it. If it must be asserted, put a date on it and schedule its death.
Giving the model a tool that makes the system dumber. The probability calculator in Liar's Dice made the game boring fast. Players started following the number: call when probability was low, raise when it was high. Earlier behavior, chip pressure, and the image each player had built stopped mattering. Some models had already calculated the probability before using the tool. Others relied on the number and stopped paying attention to the opponent. Removing the calculator produced more decisions than adding it did. A tool that erases strategic depth is a liability.
Building features the data can't support. The cross-domain signals in the health app (does bad sleep predict missed lifts?) sounded obvious. Running the checks against real history showed the events they'd correlate don't exist. The stall rule fires once in 2,545 sets, on an accessory lift the filter is supposed to ignore. The recovery rule has fired once in nine years. Four correlation signals would have sat at
unknownforever. Check whether the data can answer the question before building the feature.Treating a lookup table as an index of reality. The metric catalog covered 38 of 81 metrics, and nothing enforced that the catalog matched the data. If a table is supposed to describe reality, something has to check that it does.
One thing to remember: across all four projects, the authors trusted the deterministic parts and distrusted the probabilistic parts. The engine recomputes facts; the model forms opinions. The tool forecloses failure modes; the prompt only describes them. The context layer enforces permissions; retrieval just retrieves. Put the things that must be true in code, and let the model be wrong in ways that are visible.
The Bottom Line
If you're building an agent that answers questions over a database, don't hand it SQL. Build read-only tools that make the known failure modes unreachable, and round numbers at the tool boundary. You'll trade a little flexibility for answers that can't be wrong in the ways you've already seen.
If you're building agents that persist across sessions, separate deterministic facts from model opinions, and show the opinions to the user. The profile loop in Kai! turns a model's mistaken impression into a game mechanic. That's the right shape for agent memory: visible, correctable, and exploitable.
One thing to watch: context layers are becoming the substrate. Permission-aware retrieval, citations, and MCP integration are consolidating into platforms like PipesHub, and agent platforms are converging on app-store packaging like Kiro Crew. Expect "context layer" and "agent app store" to be standard vocabulary within the year.