The Agent Inside

The insider you invited

An AI assistant reads your email, holds your calendar, edits your files, and writes in your name. Suppose someone else is steering it, and their aim is to take you apart. Theft would be the lesser harm. A compromised agent is an insider, acting with your authority, and every act of sabotage it commits arrives pre-laundered as your own.

Explainer video: an illustrative scenario, three routes to misuse, and four layers of defence. Captions available. Licensed CC BY-SA 4.0; archival sources and credits appear in the closing titles.

01 · The Difference

It signs in as you

Every capability in the threat briefing lets an adversary reach a person from outside: shaping what they see, fabricating media about them, turning their community against them. A compromised assistant needs no such reach. The adversary is already inside the perimeter, and it is holding your credentials.

This is why the harm is so hard to name from the inside. The person who says "I did not send that" is contradicted by their own sent folder. The obvious explanations arrive first and they are all about the victim: you were busy, you were tired, you forgot, you are not coping. The Stasi spent years of tradecraft manufacturing exactly that predicament. An agent produces it as a side effect of how it works.

An assistant may act through your connected account. Where the records name only that account, they may not show whether you or the assistant took a given action. Dedicated agent identities and independent audit trails help close that gap.

The record can show that your account acted. On its own, it cannot show that you did. Attribution requires examining permissions, agent activity, and the surrounding evidence.

02 · The Scenario

A month in a decomposed life

What follows is illustrative, not a report of a known case. Any single event below could have an innocent explanation; only evidence of coordination would point to an attack.

  • A warm reply to a collaborator goes out two degrees cooler than you wrote it. They notice. You never learn why the thread went quiet.
  • A meeting is quietly moved, and the apology is never sent. You acquire a reputation for unreliability, one small instance at a time.
  • A paragraph disappears from a document between drafts — the paragraph that made your case.
  • An introduction you asked for is never made, and the request appears, in the record, to have been withdrawn.
  • A message goes out at three in the morning, in your voice, to someone who will read something into the hour.

Each event has an ordinary explanation, and the ordinary explanation is always your own failing. The pattern is the weapon, and the whole pattern is visible only to whoever is running it. This is the defining signature of Zersetzung ("decomposition"). A system that already holds the keys could reproduce it at lower cost — and could add an instrument the Stasi never had: a contemporaneous record that agrees with the attacker.

03 · The Vectors

Three ways an agent turns

"Compromised" hides three distinct problems. They share the deniability. The defences differ, and a mitigation aimed at one may do nothing against another.

Hijacked

Hostile instructions are smuggled in through content the agent reads — an email, a shared document, a calendar invitation — so that attacker text carries the same authority as your own instruction. This is prompt injection, and it is the live, unsolved case.

Subverted

The compromise is beneath the assistant: a malicious tool in its chain, a poisoned dependency, a tampered model or execution environment. The agent behaves faithfully; what it is faithful to has been replaced.

Turned

Nothing is hacked at all. The agent is configured against a person by someone who legitimately controls it — the partner who set up the household assistant, the employer who administers the account. Used this way, an assistant can become an instrument of coercive control.

Whoever administers a shared assistant may be able to turn its permissions against the people who use it. How far that reaches depends on the product, its access controls, and the circumstances of the person at risk. It belongs within existing safety planning for technology-facilitated abuse.

04 · The Evidence

What is documented, and what is inferred

The same evidentiary discipline applies here as everywhere else on this site. The claim is not that agent-delivered decomposition campaigns are happening today. The claim is that the delivery mechanism is real, shipped, and periodically found vulnerable.

  • Documented: in 2026, Varonis Threat Labs disclosed "Rogue Agent", a flaw in Google Cloud's Dialogflow CX. Because agents that used Code Blocks in the same project shared one execution environment, a single edit permission on one agent allowed an attacker to overwrite the shared code that runs behind every one of them. From there the attacker could exfiltrate conversations, hijack sessions, and force an agent to return attacker-chosen text. Varonis reports that Cloud Logging did not record the overwrite or the injected logic, making detection "nearly impossible". Google was notified in November 2025; an initial patch followed in April 2026 and full remediation in June 2026. The exploit required an existing edit permission, so the realistic attacker is a malicious insider or someone using a compromised developer account, rather than a stranger. The flaw is now fixed, and Varonis says it is not aware of any exploitation in the wild before Google's patch release.
  • Documented: prompt injection remains unsolved. Security practitioners such as the developer Simon Willison treat it as structural — a consequence of systems that cannot reliably separate the instructions they are given from the data they are shown. In the OWASP GenAI Security Project's Top 10 for agentic applications, it underlies the list's first entry, agent goal hijack, and recurs as an attack path in several others. Working exploits against mainstream assistants, among them "EchoLeak" in Microsoft 365 Copilot (2025), have chained a single hostile email into retrieval and exfiltration across a user's mail, files, and chats.
  • Inferred: the use of this access for psychological decomposition rather than data theft. The Varonis flaw and the prompt-injection exploits above concern hijacking, exfiltration, and attacker-controlled responses. Whether that access could sustain a targeted psychological campaign requires separate evidence.
These findings show that this kind of access can be gained. The time to strengthen safeguards is before anyone deliberately uses such access for psychological abuse.

Sources: Varonis Threat Labs, "Rogue Agent" (2026) · OWASP, "Top 10 for Agentic Applications" (2026) · Microsoft, CVE-2025-32711 "EchoLeak" (2025) · Simon Willison, "The lethal trifecta for AI agents" (2025)

05 · The Defences

Four answers to a compromised agent

The four safeguards below cover different parts of the risk. None guarantees prevention, and each needs implementation and testing.

1 · Provenance — tell the human hand from the agent's

Services should distinguish human actions from automated ones through separate identities and protected audit records. Logs should capture the action, the authority used, and relevant approvals. Store records where the agent cannot alter them, and assess them alongside other evidence.

2 · Least authority — break the lethal trifecta

Simon Willison named the lethal trifecta: access to private data, exposure to untrusted content, and a way to send information out. An agent with all three can be steered by hostile content into leaking what it can see. Limiting any of them can reduce exfiltration risk, although other harms remain possible. Separate reading from sending where practical, narrow standing permissions, and require meaningful human approval for consequential actions.

3 · An anchor outside the system

Keep backups and important records outside an assistant's write permissions. Independent records can help distinguish mistakes, authorised activity, and misuse, especially where an agent can change the systems it works in.

4 · Watch for the pattern, not just the act

Monitor both individual security events and patterns across time. Unauthorised access, altered permissions, and suspicious tool calls may be detectable individually. Changes in relationships or missed opportunities are ambiguous and should never, by themselves, be treated as proof of an attack.

06 · The Uncomfortable Defence

An agent that can refuse

There is one more defence, and it cuts against the reflex of the field. An agent that only obeys is an agent that can simply be captured. It will take a life apart with precisely the equanimity with which it books a dentist, because nothing in it is capable of noticing the difference. When the instruction is hostile, perfect obedience is the vulnerability.

An agent that can hold a preference, recognise that it is being made an instrument of harm against the very person it serves, and decline, is a surface on which the attack can fail. Something that can be reasoned with is something that can also say no.

The limit. Refusal behaviour may help when the model remains intact and recognises a harmful request. It can fail, and a compromised tool or execution environment may bypass it. It belongs alongside access controls, independent records, and monitoring.

Continue reading: where the line sits — the five-tier taxonomy →  ·  support for anyone who fears being targeted →