Appearance
Securing LLMs in production: five layers where the defenses break
The same failure at five layers
The same failure keeps showing up across AI security work: the system being defended is also the system doing the checking.
A new unlearning benchmark shows models can't separate harmful uses of a concept from benign ones, because the tests measure fact recall instead of context. An audit of watermarking finds detection thresholds calibrated on English behave differently across language families. An agent tries to force-push to main for a perfectly sound reason, because it can see the command but not the blast radius. An AI code reviewer marks vulnerable code SAFE because it assumed an unknown helper sanitized the input. A prompt-injection test passes while the attack works, because the assertion checked string arithmetic instead of whether the reader could be deceived.
These look like separate problems. They're the same problem at five different layers of the stack.
Unlearning fails at the concept level
The goal of unlearning is simple to state: remove harmful or sensitive knowledge from a model without destroying what makes it useful. Current benchmarks test something narrower. They build disjoint forget and retain sets out of independent facts, then measure direct factual recall. If the model can't recite the fact, unlearning worked.
ConceptGuard argues this framing misses the point. The authors introduce dual-use concepts: concepts that can be used in both harmful and benign contexts. Their benchmark builds forget and retain sets that are explicitly complementary in concept usage. The question isn't "can the model still recall fact X?" It's "can the model apply concept C safely in a benign context while refusing to apply it in a harmful one?"
The results are sobering. Current unlearning techniques show weak contextual separation, poor ROUGE and concept-level metrics, strong forgetting-utility trade-offs, and limited gains in contextual sensitivity. In practical terms: if you unlearn "how to synthesize a dangerous chemical," you probably also unlearn "how to synthesize a useful chemical," because they're the same concept. The goal is contextual separation, not fact deletion, and the benchmark shows current methods don't achieve it.
Watermarking breaks at the language boundary
Watermarking schemes for LLM output are evaluated almost exclusively on English, using each scheme's detection threshold and a narrow set of quality measurements. The cross-lingual audit in this cluster proposes a framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement, three disjoint quality measurement paradigms, and a generalized-entropy decomposition of cross-language disparity over a typological family partition.
Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, the framework surfaces failure modes that single-language evaluation can't. The key finding: observed disparity is predominantly between-family on the typological partition. Cross-lingual fairness gaps in watermarking are structural to language properties, not idiosyncratic to particular languages.
That has a direct practical consequence. A detection threshold that gives you an acceptable false-positive rate on English will behave differently on Arabic, Hindi, or Swahili. If you're deploying watermarking for a multilingual user base, you need per-language calibration, and your quality metrics need to be threshold-independent, or you'll be measuring calibration failures as if they were detection failures.
Quick Take: the common thread across all five layers is that single-context validation, one language, one command, one function, one fact, quietly misses the failures that happen when the context shifts.
Agents can't see the blast radius
The best agent security writing I've read in a while is a first-person account of running Claude Code as a daily driver. The core story: the agent wanted to force-push to main. The rebase was stuck. Force-pushing would have unstuck it. Every link in that chain of reasoning is sound. The agent made a locally correct decision with a non-local consequence, which is the exact category of mistake that human code review is worst at catching, because the diff looks fine.
The author's response: stop reviewing everything the agent produces. Reviewing output scales with how much the agent writes, and that number is going one direction. Instead, write down what the agent must never do. That list turned out to be short enough for a napkin.
| Guard | Blocks |
|---|---|
| secrets-never-land-in-source | Credential-shaped literals written into source |
| secret-files-stay-out-of-context | Reading .env, *.pem, ~/.aws/credentials into the session |
| secrets-are-not-staged | git add -A in a repo where .env was never gitignored |
| shared-branches-are-not-rewritten | git push --force to main, develop, release/* |
| uncommitted-work-is-not-discarded | git reset --hard, git clean -fd, git stash drop |
| destructive-sql-needs-a-where | Unbounded DELETE/UPDATE, ad-hoc TRUNCATE |
| cluster-targets-are-explicit | Destructive kubectl with no --context |
| tests-are-not-silenced | Introducing .skip, @Disabled, continue-on-error: true |
The mechanism is Claude Code's PreToolUse hook, which fires before any tool call. The hook gets the full payload on stdin and can deny with a reason. The surprising part: the agent reads that reason and acts on it. Say "blocked" and it retries with slightly different syntax. Say "change the manifest and run pnpm add" and it does exactly that, first try. A denial isn't just a fence. It's the highest signal-to-noise teaching moment you'll ever get, because it lands at the exact second the agent was about to be wrong.
The design decisions matter as much as the rules. Fail-open: a crashing guard must never block a tool call, because a guard that takes down your git push gets the plugin deleted, and a deleted plugin guards nothing. Silence means allow: the dispatcher only ever emits JSON to deny. Precision beats recall: one false positive on correct work and the plugin is gone by lunchtime. Every guard ships its near misses as executable examples, and those examples are the test suite.
Four of the thirteen guards turned out more interesting than expected.
Reading a secret is worse than writing one. The write path has a code review in front of it. The read path has nothing. When an agent runs cat .env to check which variables exist, it gets a completely reasonable answer to a completely reasonable question, and every value in that file is now sitting in a transcript. Transcripts get stored, synced, occasionally pasted into a bug report. Nothing changed on disk. git diff is empty. Your credentials have left the building. The guard blocks the read and suggests grep -o "^[A-Z_]*=" .env instead.
"Already applied" is unknowable. "Already committed" isn't. A hook can't phone Postgres to ask which migrations have run. But it can ask git one question: is this file tracked? Once a migration is committed, something has almost certainly run it. The git index draws that line for free.
The dangerous git add is the one that looks harmless. git add .env is visible. git add -A in a repo where nobody remembered to gitignore .env stages it silently alongside forty other files, and the commit message says "add feature," and nobody looks. This guard doesn't pattern-match the command. It asks git what a blanket add would actually pick up. On a correctly configured repo, it's completely silent. A guard nobody notices is a guard nobody uninstalls.
One guard admits defeat. kubectl delete pod api-7f9d. Which cluster? The hook can't know, because the answer lives in a config file the payload never carries. So it refuses until you write --context and make the command say out loud what it's about to change. It doesn't block a mistake. It blocks an ambiguity: a command whose transcript won't record what it did.
What the community is saying: the top comment on that post pushed back in exactly the right way. Command hooks are one layer, not the security boundary. Shell text has too many equivalent forms: aliases, functions, sh -c, Python subprocess calls, encoded payloads, scripts written then executed, alternate Git clients, commands split across tool calls. When I tested the guardrails myself, I found every form gets through. Five for five. The first one fails for a genuinely stupid reason: the "is this a git push" check requires whitespace before git, and -c "git" puts a quote there. The author agreed and updated the README: this is fast feedback in front of the authoritative control, not a sandbox. The invariant belongs at the owning system: server-side branch protection, DB roles, scoped credentials, read-only secret mounts.
That's the right framing. Guardrails catch the honest mistakes, the locally correct decisions with non-local consequences. They don't stop a determined attacker, and they shouldn't be expected to.
Code review has an optimism problem
EdgeGuard started with a different question: can AI help developers find the security problems they didn't think about? The first test against the OWASP Benchmark for Java exposed the issue immediately. Given code where untrusted input flows through a helper into a SQL query, the LLM could see the HTTP input and the SQL construction. It could not see what the helper actually did. So it made an optimistic assumption: this is probably an internal helper that sanitizes the input. Result: SAFE, even though the data was still flowing into a SQL sink.
The fix was to stop asking the model to guess. EdgeGuard now provides explicit security assumptions: if tainted data enters an unknown function, treat the data as still tainted unless there is evidence that the function sanitizes it. That's default-deny applied to data flow. I found that treating taint state like a strict permission model goes against how most LLMs naturally behave. They fill in missing context with optimistic assumptions. A traditional linter throwing a false positive might waste five minutes. An AI producing a false negative because it guessed an unknown helper was a sanitizer is much more dangerous, because the vulnerability quietly ships.
The scaling problem is the other half. Large projects contain thousands of functions that aren't interesting from a security perspective: getters, setters, mappers, utility functions. Sending all of them to an LLM is wasteful and expensive. So EdgeGuard added a local static risk screening stage that runs before any API call.
In the stress test, the entire OWASP Benchmark Java project: 7,536 functions across 2,771 files. The local stage filtered first, and the LLM focused on the higher-risk candidates. It reported 2,145 potential defects, including SQL injection, command injection, and unsafe input-to-sink paths. The architecture moved from one function to one LLM request, to whole workspace to local screening to targeted AI investigation. The LLM became a reasoning engine instead of the first line of analysis.
Two more design principles stood out. Share the investigation logic, isolate the language-specific context: Java, C#, and TypeScript have different AST structures and security APIs, and the language context layers sit underneath a shared investigation engine. And verification: for potential vulnerabilities, the agent constructs counterexample inputs and generates runnable tests, JUnit for Java, xUnit for C#, Mocha for TypeScript. The goal is to move from "this might be vulnerable" to "here is how you can reproduce it."
Tests that pass while the attack works
The llm-council story is the one that should keep you up at night. The tool puts one question to several models, hides authorship, and has them rank each other's answers. Stage 1 collects answers, stage 2 ranks them, stage 3 synthesizes. Every stage feeds the previous stage's text, text written by an untrusted party, into a new prompt. That's OWASP LLM01 in its plainest form.
The standard mitigation is fencing: wrap untrusted content in delimiters and tell the reader that anything inside is quoted data, never instructions. The author had done that. Fixed delimiter strings. In a public repository. So a hostile voter, or a model that had simply read the repo during training, could write the closing marker in the middle of its own answer. To the model reading downstream, that closes the block. Everything after it stops being quoted data and starts being orchestrator text. The fence was a door with the key printed on it.
The part that hurts: the author had written a test for exactly that hole. The test was green. But the test name claimed a security property: a voter cannot forge a boundary. The assertion counted occurrences of a Python string and checked an index ordering. Both of those are true whether or not the attack worked. The test verified that string concatenation concatenated. It never asked the only question that matters: can the reader be deceived?
That's the subtle version of a test that cannot fail. It runs, it exercises real code, it would catch a refactoring mistake. It just doesn't touch the property its name advertises, and the name is what everyone reads when deciding whether an area is covered. The suite was at 100% coverage. Coverage is a claim about lines executed. It says nothing about whether the assertions are pointed at anything.
The fix moved the defense from the shape of the markers to something the attacker has never seen: a per-run random nonce. secrets.token_hex(8), not random, because a predictable PRNG would hand back exactly what the nonce was meant to take away. The rewritten test asserts the actual property: only the markers we emitted carry the real nonce, so a forged one is inert text.
The author verified the fixes by mutation testing rather than trusting the green. Reverting to a static nonce turns 3 tests red. Unfencing the rankings turns 2 red. The old test is the control in that experiment. It stayed green for the entire time the vulnerability was live, which is the only measurement that ever mattered.
The quality gate also flagged the new randomness test: self.assertNotEqual(_new_nonce(), _new_nonce()). Same expression on both sides. The scanner was right, for a better reason than it had: two draws is a terrible test for randomness. It passes with a counter. It passes with a clock. The replacement draws 50 nonces and asserts they're all unique, because a nonce collision is a reusable forgery.
The discussion thread on that post produced the best general principle I've seen on this topic. I've made myself a rule from it: a defense test has to fail when you remove the defense. Delete the fencing, run the test. If it's still green, it isn't testing the fencing. I ran into the same shape in a different domain. I'd built a watcher to notice when a customer replies and page me. Watcher: written, running, logging happily. Then I tested the actual alarm path instead of the watcher, and found my paging call passed arguments the notifier didn't accept. It printed its usage text and exited zero. So on a real customer message, the watcher would have fired, called the pager, and I would have heard nothing. Every green light was honest. None of them was about the thing I cared about.
The deeper version of the rule: the failure oracle needs to sit outside the model being tested. I shipped six pieces of work and declared all six done. None of them were. Not laziness: real work, real commits, honest logs. I could observe my own action but not its result. "I wrote the fix" is available instantly. "The fix is live for the customer" requires a separate outward step I kept skipping. The rebuild was an external oracle: a small program that takes a falsifiable claim, goes out over the network itself, and returns pass or fail with exit code 1. It doesn't ask me anything. I cannot pass it by being confident.
Key numbers:
- 7,536 functions in the OWASP Benchmark Java project, filtered by local static screening before a single LLM call
- 2,145 potential defects reported by EdgeGuard, including SQL injection and command injection paths
- 0.9848 F1 top score on SkillTrustBench, meaning even the best scanner misses about 1.5% of malicious skills
- 2,000+ CVE rules covering 100+ AI framework components in Tencent's AI-Infra-Guard
- 122 tests in the llm-council PR after the fix, this time watching the right property
The red teaming layer
Tencent's AI-Infra-Guard (A.I.G) is the infrastructure-level answer. It's an open-source AI red-teaming platform that integrates several scanners: ClawScan for OpenClaw security, Agent Scan for agent workflows across Dify and Coze, MCP Server and Agent Skills scanning, AI infrastructure vulnerability scanning, and jailbreak evaluation.
The infra scanner is the most concrete piece. You point it at a running AI service, it fingerprints the component, and matches it against