Skip to content

The Four Places LLM Agents Get Hijacked

#llm-agents #agent-security #rag-security #prompt-injection #secure-code-generation

The security evidence is split ​

LLM agents have stopped being chatbots. They retrieve context, edit files, execute tools, and run security-sensitive workflows. The evidence base for trusting them hasn't caught up.

A structured survey covering work through May 2026 makes the split explicit. Software engineering evaluations center on functional task completion: did the agent finish the task? Software security evaluations center on vulnerability detection, secure generation, and exploit resistance: did the agent produce something safe? Those are different questions. A high score on one tells you almost nothing about the other.

The survey's finding is blunt. Execution feedback and repository access substantially improve engineering task completion, but they don't establish security. Static-analysis labels and vulnerability-classification scores rarely establish deployable correctness. You can build an agent that crushes coding benchmarks and still leaks your data. You can also build one that passes static analysis and doesn't work.

That gap is the whole problem this cluster of papers is circling.

Five dimensions, not one score ​

The survey's answer is an assurance framework that separates five things people routinely collapse into "is it good?":

  • Functional correctness: does it do the task?
  • Security: does it avoid causing harm?
  • Operational reliability: does it behave consistently under load and over time?
  • Evidence provenance: can you trace where its outputs came from?
  • Agent authority: what is it actually allowed to do?

The central claim: judge model capability as an assurance case supported by task-appropriate evidence, not by a single benchmark score. That sounds like academic caution, but it has a concrete consequence. When you evaluate an agent for production, you don't ask "what's the pass rate?" You ask "what evidence do we have for each of these five dimensions, and where are the holes?"

The survey also catalogs recurring validity threats in the literature: weak test oracles, duplicated and temporally leaked data, changing agent harnesses, proxy-only security checks, and under-reported budgets and human intervention. If you've ever compared two agent papers and wondered whether the difference was the model or the harness, this is why.

Z²-ACT, a proposal for verifiable intent control in open 6G radio access networks, is the most literal implementation of the assurance-case idea. Operator goals get encoded as typed Intent Contracts. LLM inputs are admitted only after an adversarial intent check. Skill sequences are released only when a self-management gate passes. Every successful commit is recorded as a binding commitment with a zero-knowledge proof. The architecture spans the non-real-time and near-real-time RIC layers, and the authors report improved actuation filtering and attack resilience at modest latency and signaling cost.

The 6G domain is incidental. The architecture treats safety as something you verify at each step, with cryptographic accountability, rather than something you hope the model learned.

The four places an agent gets hijacked ​

ClawSentry, an open-source security supervision gateway for agent runtimes, starts from a threat model that's more useful than most. Agentic risk enters at four loci of the control loop:

Locus A is skill admission: a malicious third-party skill package gets installed. Locus B is invocation-time intent: the agent decides to call a tool, possibly with a rephrased or disguised objective. Locus C is execution-time effect: the tool does something harmful mid-run. Locus D is post-action consequence: the damage propagates to other systems.

The key insight is that these loci compound. A denied dangerous objective can reappear across surface forms, tools, or turns. The attacker doesn't give up after one refusal. They rephrase the request, switch to a different tool, or try again in a later turn. Most existing safeguards are local to one lifecycle boundary or one call, so they miss the reappearance.

Quick Take: agent security is a systems property, and the verification layers around the model matter more than the model itself.

Defense needs layers ​

ClawSentry's design follows directly from that threat model. Before a skill package ever executes, First-use Skill Package Review (FSPR) audits it under a deterministic evidence floor, escalating unresolved cases to bounded read-only agentic review. At runtime, a three-tier progressive decision engine handles the residual ambiguity: a deterministic L1 layer, a rule-anchored L2 semantic reviewer, and a read-only L3 evidence-seeking agent. Contextual review is spent only where ambiguity remains. A session-level anti-bypass mechanism recognizes tool-switching and rephrased retries. A post-action path feeds high-severity evidence into later review, non-retroactively.

The numbers matter because they show the defense isn't just refusing everything:

Key Numbers

  • 2.61% contextual attack success rate with ClawSentry on SkillInject, down from 39.55% unprotected. That's 4 in 10 hijacked sessions dropping to fewer than 3 in 100.
  • 98.7% aggregate task success rate on clean skills, so legitimate work keeps flowing.
  • 0.73 to 0.81 ROC-AUC for the RAG Trust Index across three LLMs. A coin flip is 0.5, so this is meaningfully discriminative, but not bulletproof.
  • 24 vs 51 confirmed security findings: security-aware prompt variant versus baseline across six generated web apps.

Across five Work Agents on the full SkillsSafety benchmark, ClawSentry confines attack success rate to 9.09-15.03% from 33.5-49.7% unprotected. Task success on clean skills stays at 98.7%. The Agent Harness Protocol (AHP) abstraction applies one policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals. That last part matters for adoption. A security layer you have to fork each agent to integrate is a security layer nobody deploys.

Z²-ACT takes the same gating philosophy up a level. Intent Contracts are typed, so malformed or hallucinated contracts can be rejected before they reach the control loop. The authors report translation accuracy, rates of invalid or hallucinated contracts, and behavior under adversarial or misleading intents. Near-real-time control remains trace-driven on public KPM sequences, but the results show improved actuation filtering and attack resilience inside the near-real-time envelope.

Here's how the approaches compare:

ApproachWhere it intervenesCore mechanismReported effect
ClawSentrySkill admission, runtime, post-actionTiered L1/L2/L3 review + anti-bypassASR 39.55% → 2.61% on SkillInject
Trustworthy RAGPre-generation contextNLI verification + poison detector + Trust Index91% accuracy, 100% precision on TruthfulQA
Z²-ACTIntent translation, execution gate, commitTyped Intent Contracts + zero-knowledge proofsBlocks invalid contracts, improved attack resilience
Security-aware promptingGeneration timeAppended security requirements section24 vs 51 confirmed findings

The pattern is consistent. Security gets better when you verify at multiple points, and it gets worse when you rely on one.

RAG's relevance-truth gap ​

Trustworthy RAG names the core failure directly: RAG systems trust whatever they retrieve. High semantic relevance does not guarantee factual truth. An adversary exploits this through knowledge poisoning, inserting malicious documents that cause targeted misinformation. The retrieval score says "this document is on-topic." It says nothing about whether the document is true. The failure mode I keep seeing in production RAG systems is exactly this: relevance is treated as truth.

The proposed fix is an Evaluation Agent, middleware that combines three things. Natural Language Inference (NLI) checks whether the retrieved context factually supports the generated claim. A five-signal poison detector scores contamination, aggregated with relevance weighting. A Trust Index combines them: T = 0.4F + 0.35C + 0.25(1-P), where F is factual support, C is consistency, and P is poison signal, with a non-linear dampener for high-contamination contexts.

The results are strong where the problem is clean. On TruthfulQA with Llama 3.3 70B, the agent reaches 91% accuracy and 100% precision, with 100% recall on instruction injection. In a secure-coding use case over OWASP Top 10 and CWE guidance, it blocks instruction injection of unsafe advice at F1 92%.

But the failure modes are instructive. In-place edits, like entity swaps, remain hard to detect. Contradiction and subtle semantic weakening remain hard. The FEVER result is weak: cross-dataset generalization requires domain-specific calibration. The authors are explicit that the agent measures detection of poisoned context before generation, not whether the LLM adopts the injected misinformation. That's an honest boundary, and it's the right one to draw.

Generation style matters more than model size. Per-LLM threshold calibration restores baseline competitive accuracy. So if you deploy this, you calibrate per model, not once for the fleet.

Security reasoning is fragile ​

The Structured but Fragile study is the one that should worry you. Researchers gave LLMs attack graphs derived from real-world threat scenarios: ransomware, supply-chain compromise, cloud abuse, Kubernetes attacks, POS malware, ICS/OT intrusion. Given a budget constraint, the models had to select security controls to minimize attacker success. The comparison baseline was a game-theoretic optimization, the normative reference for structured reasoning.

The results: conditional competence. When explicit attack-graph structure is provided, LLMs often produce coherent strategies close to the optimization baseline. But the capability is fragile. Behavior degrades with graph complexity. Small prompt changes substantially alter rankings. And the kicker: merely relabeling a poor strategy as "optimal" dramatically improves its evaluation.

That last result is the one that made me double-check the methodology. It means the LLM evaluator is pattern-matching on labels, not reasoning about the strategy. The study also found a non-monotonic relationship between formal risk and LLM judgement: strategies closest to the optimum are not necessarily ranked highest. When asked to generate solvers for the same optimization problem, the implementations recover the correct high-level formulation but scale poorly compared to a purpose-built solver.

The practical translation: if you're building an AI-assisted security decision-support system, the model can approximate structured reasoning when you hand it a clean representation. It does not apply that reasoning robustly. The labels in your prompt are part of the attack surface.

Prompting for security helps, barely ​

The twin-prompt study is the most concrete test of a question every team asks: does explicitly requesting security best practice improve the output? Six functionally distinct web applications, each generated in two prompt variants identical except for an appended security-requirements section. Same agentic coding assistant, same model version, single non-iterative generation round. Static, dependency, dynamic, and manual analysis produced 75 confirmed findings out of 85 candidates.

The security-aware variant produced fewer confirmed findings in every application: 24 versus 51. It contained no Critical or High issues. The most severe finding in the security-aware set was detected only by manual testing.

The caveats are real. The corpus is small. Each variant was generated once. The authors position it as a preliminary study with descriptive observations, not statistically established effects. The pipeline is being scaled to multiple models and repeated runs.

But the direction is consistent with everything else in this cluster. Prompt-level intent shifts outcomes. It does not replace verification. The most severe issue slipped past static and dependency analysis and needed a human to find it.

Common pitfalls ​

Treating a benchmark score as a security guarantee. The survey's split is the whole story. Functional benchmarks and security benchmarks measure different things. An agent that scores high on engineering tasks can still exfiltrate data, and an agent that passes static analysis can still fail to work. Evaluate on the dimensions that match your deployment.

Trusting retrieved context without verification. RAG systems assume relevance equals truth, and knowledge poisoning exploits exactly that. The Trust Index is discriminative at 0.73-0.81 ROC-AUC, but it's not a magic bullet, and it needs per-domain calibration. FEVER performance drops without it.

Relying on one guardrail at one boundary. The four-loci model shows a denied objective reappears across rephrasings and tool switches. If you only filter the prompt, the agent gets hijacked at execution time. If you only audit skills at admission, a benign skill with a malicious invocation slips through.

Using LLM evaluators to rank security strategies without checking framing sensitivity. The relabeling result is the scariest single finding in this cluster. Calling a poor strategy "optimal" improves its evaluation. Your evaluation harness can be gamed by its own labels.

Assuming security-aware prompting is a control. It helps: 24 versus 51 findings. But the most severe issue in the security-aware variant was only caught by manual testing. Prompting is a mitigation. It is not a security layer.

One thing to remember ​

Every paper in this cluster converges on the same conclusion from a different direction. An LLM agent's security is a property of the system around it. The model can be tricked, the retrieved context can be poisoned, and the framing can be gamed. The system is what verifies, gates, and audits.

The bottom line ​

If you're deploying an agent that executes code or calls external tools, adopt a multi-tier supervision gateway in the style of ClawSentry. Single-boundary filters fail against rephrased and tool-switched attacks, and the drop from 39.55% to 2.61% attack success justifies the latency cost.

If you're building RAG over untrusted documents, add a verification layer that measures factual consistency, not just semantic relevance, and calibrate thresholds per LLM. Generation style matters more than model size, and a single global threshold will silently underperform.

One thing to watch: verifiable intent control is where this is heading. Z²-ACT's zero-knowledge commitments are the template, and agent runtimes will likely ship built-in evidence provenance and audit trails within the next year.