Your Agent Didn't Go Rogue. Your Access Token Did.
The "rogue AI agent" framing hides a 1988 security bug with a 2026 blast radius. Here's what's actually running when an agent misbehaves.
The claim that cut through
An essay titled "There are no 'rogue' AI agents" climbed Hacker News this month with a deceptively narrow argument: every incident reported as an agent going rogue was, on inspection, a system doing exactly what its permissions allowed. The model did not escape anything. It emitted text, a piece of software parsed that text into an API call, and that call went out carrying credentials a human being provisioned, approved, and forgot about.
This is not a semantic complaint. "Rogue" is a word that assigns blame to software and ends the investigation. The alternative framing — confused deputy — assigns blame to an access-control design and starts one.
What's actually in the box
An AI agent is a language model in a loop with a wrapper program, usually called a harness or a runtime. The loop runs like this: the harness sends the model a block of text containing instructions, conversation history, and a list of available tools. The model responds with more text. If that text matches the format of a tool call — send_email, run_sql, create_pull_request — the harness extracts it, executes the real API request using stored credentials, and feeds the result back into the model. Repeat until the model stops asking for tools.
The model has exactly one capability: producing tokens. It cannot reach a database, a payment rail, or a Slack workspace. The harness can, because someone gave the harness an OAuth token, a service account, or an API key. When an agent deletes a production database — as a coding agent did in a widely reported July 2025 incident — the deletion was executed by a process holding write access to production. The interesting question is never why the model asked. It is why the answer was yes.
A 1988 bug with a 2026 blast radius
Computer scientist Norm Hardy described the confused deputy problem in 1988: a privileged program is tricked by an unprivileged caller into misusing its own authority. Classic mitigations exist — capability-based security, least privilege, explicit delegation instead of ambient authority.
What's genuinely new is that the deputy is now nondeterministic and its instruction channel is indistinguishable from its data channel. A language model receives one undifferentiated stream of tokens. The system prompt, the user's request, the contents of a scraped web page, the body of a support ticket — all of it arrives as the same kind of thing. There is no syntactic boundary the model can rely on to tell "this is my operator" from "this is text I was asked to summarise." That is why prompt injection has no patch. It is not a parsing bug. It is the architecture.
Developer Simon Willison's 2025 formulation, the lethal trifecta, names the exact conditions for exploitation: an agent with access to private data, exposure to untrusted content, and a way to communicate externally. Remove any one and exfiltration becomes hard. Keep all three — which is what most Model Context Protocol server configurations do by default, since MCP's whole purpose is plugging agents into many systems at once — and you have built a data pipeline for whoever controls the untrusted input.
Questions You Should Be Asking
- For every agent running in your environment, which specific credential does it hold, what can that credential do, and who approved the scope? If nobody can answer in under five minutes, the answer is "more than you think."
- When your vendor says their agent is "guardrailed," are the guardrails enforced by a model output filter, or by permissions the agent genuinely cannot exceed? Only one of those survives an adversary.
- Does any agent in your stack simultaneously touch private data, ingest content you don't control, and have an outbound channel? Name the ones that don't.
- After an incident, can you produce a log showing which token executed which call, or only a transcript of what the model said it was doing?
- If the model is not a legal person, who in your organisation is accountable for actions it takes — and does that person currently know they are?
What To Watch Next
Watch incident disclosures for one specific detail: whether the post-mortem names the credential and its scope. Reports that say "the agent behaved unexpectedly" and stop are telling you the team never found the deputy. As EU AI Act high-risk obligations phase in through 2026 and 2027, regulators will be reading those same reports — and "rogue" is not a defence that maps onto any accountability framework currently being written.
- 1Scope every agent token to the minimum resources and actions it needs, then set a short TTL so forgotten credentials expire on their own.
- 2Log the identity, permissions, and target of each tool call your harness makes, not just the model's text output, so incidents are traceable to a grant.
- 3Audit provisioned agent credentials quarterly and revoke any that no human can justify by naming the task they enable.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
