Excessive Agency Is Now the Number 3 LLM Risk
Image generated by AI

Excessive Agency Is Now the Number 3 LLM Risk

OWASP published a new Top 10 for LLM Applications in early August 2026, replacing the 2025 list, and Excessive Agency jumped from sixth place to third.

The ranking this time weighs 75% expert vote and 25% real incident data, drawn from 6,639 logged incidents, so the move is not one committee’s hunch. Prompt Injection and Sensitive Information Disclosure kept first and second place. Unbounded Consumption rose four places. Misinformation rose two. System Prompt Leakage got renamed to Hidden Context Exposure and widened to cover more than just the prompt text. Excessive Agency is the one that should worry anyone running an agent in production, because it stopped being mostly about chatbots.

What excessive agency looks like in practice

A chatbot with excessive agency is annoying. It might say something wrong, confidently. An agent with excessive agency has a tool it did not need for the task in front of it, and that tool is the entire vulnerability. Not the model’s output. The model’s reach.

Say you built a support agent that reads a ticket, checks an order, and issues a refund. Give it a database connection with write access to every table, and the risk is not “the model might hallucinate an order number.” It is “the model has a write connection to your billing tables and a plausible reason to use it.” The model does not need to be compromised for this to go wrong. It needs to be wrong once, with the permissions to act on being wrong.

The same shape shows up outside support bots. A coding agent wired with general shell access, so it can run a test suite, also technically has every other command that shell offers, if a prompt or a bad tool response convinces it to reach for one. Nobody added that capability on purpose. It came bundled with “give the agent a terminal,” because scoping a terminal down to only the commands a test run actually needs is more work than handing over the whole shell and moving on.

The three levers

OWASP frames excessive agency as three separate problems, and they need three separate fixes.

Excessive functionality. The agent carries tools it does not use for its actual job. A general-purpose database tool bolted onto every agent, because it was easier to build once and reuse everywhere, means the summarization agent can technically also delete rows, even though nothing about summarizing a document ever calls for that. Scope the tool set to the task, not to what was convenient to wire up once and never revisit.

Excessive permissions. The agent has more access than the tool call needs. A read-only reporting agent connected through a role that also has write access, because nobody got around to making a second database role, and the existing one was already sitting in the environment variables. This is the cheapest fix on the list, and the one people skip most, because it is invisible until something goes wrong.

Excessive autonomy. The agent can act without anyone checking first. Sending an email, refunding a charge, merging a pull request, deleting a file. All fine for a person to do on their own judgment. All worth a pause when an agent does them on a model’s judgment instead, because a person who makes a bad call usually also notices they made it. A model rarely does.

Scoping this for real

Three habits, roughly in order of how little they cost to add:

  • A tool allowlist per task, not per agent. If today’s job is “answer a billing question,” the agent gets read access to billing and nothing else, even if the same process handles a different task tomorrow with a different allowlist.
  • A read-only database role wherever the task is read-only. Most agent tool wiring defaults to whatever credential is already sitting in the environment, and that is usually the write one, because it was already there.
  • A human approval gate on anything irreversible. Refunds, deletions, merges, emails to a customer. The gate does not need a person clicking a button every time. A threshold works too, auto-approve refunds under 20 EUR, hold anything above it. Something just has to sit between the model deciding and the action happening.
  • A log of which tool the agent actually called, kept separately from the model’s own transcript. If the allowlist ever needs tightening after the fact, this is the record that tells you whether it was already too wide, not the agent’s own account of what it did.

None of this is hard to build. It is easy to skip, because the demo works fine with the wide-open write role, and nobody notices until the agent is wrong about something that matters. The scoping work is boring in exactly the way that good infrastructure work usually is.

Where this ends up in your repo

The permissions your agent has are usually visible before anything runs: a tool definition, a role grant, a config file sitting in the repo. Guard scans code inside the JetBrains IDE and can flag an agent tool wired to more access than the task in front of it needs. It reads what the tool definitions and permissions look like in the repo, though, not what the agent actually did at runtime. It catches the scoping mistake before anything runs on it. It will not tell you about the one time the agent used the access it had to do something you did not expect, because that is a runtime event, not something sitting in the repo to read. That part is still on you, and it is still worth catching the mistake before it exists rather than after it shows up in an incident report.