Skip to main content
skilder

engineering

Least privilege for AI agents: the role is the permission boundary

Agents inherit whatever their tools can reach. Least privilege for AI agents means scoping tools per role, not per user, and treating every tool description as untrusted input.

Author: Nicolas Corod
  • #governance
  • #mcp
  • #ai-agents
  • #tool-use
  • #roles
skilder mascot pointing at the article title on a paper background

“Which employees should be allowed to use AI agents” is the wrong question. Every employee already runs one, in Claude, ChatGPT or Copilot, and most of them have wired it to something. The question that decides whether the deployment survives an audit is narrower: what can the agent reach, and who decided that.

The uncomfortable answer for most teams is that nobody decided. The agent reaches whatever the connected tools can reach, and the tools were connected with the credentials that happened to be at hand.

An agent holds the union of its tools’ permissions

Classic access control assumes a human in the loop who reads a screen and chooses an action. Grant the human a permission set, and the permission set bounds what happens. That model was written down in 1975 by Saltzer and Schroeder as the principle of least privilege: every program and every user operates with the smallest set of privileges the job needs. NIST SP 800-53 carries it forward as control AC-6, and every serious security review checks for it.

An agent breaks the assumption in one specific way. It does not choose an action from a screen. It reads text, from a prompt, a document, a web page or the result of a previous tool call, and decides what to call next. Its effective permission set is not the user’s position in the org chart. It is the union of every scope on every connector it can see, exercised by a planner that takes instructions from untrusted text.

Connect the CRM with an admin token, the mailbox with full send rights and the file share with the user’s own credentials, and the agent now holds a capability no single human in the company was ever granted: read any customer record, and mail it anywhere, in one turn, without a screen to pause on.

The failure has a name: excessive agency

OWASP lists this as LLM06:2025 Excessive Agency in its Top 10 for LLM applications. The entry splits it into three parts worth keeping apart, because each has a different owner:

  • Excessive functionality: the agent can call tools the task never needs. A summarisation agent with a delete endpoint in reach.
  • Excessive permissions: the tools it needs run with more rights than the task needs. Read-only work done over a read-write connection.
  • Excessive autonomy: high-impact actions run without a human confirmation step.

Simon Willison gave the attack shape a name in June 2025: the lethal trifecta. An agent that can read private data, is exposed to untrusted content, and can communicate externally is exploitable by construction. Remove any one of the three legs and the attack falls over. Least privilege is the discipline of removing legs on purpose, per task, instead of hoping the prompt holds.

MCP moved the problem, it did not remove it

The Model Context Protocol (MCP) made tools portable. One server, any client. That is the good news. The specification’s own security best practices page is candid about what portability costs:

  • Confused deputy. An MCP server that proxies a third-party API on behalf of many users can be tricked into acting with one user’s authorisation on another user’s behalf. The server holds the authority; the agent supplies the intent; the intent came from text.
  • Token passthrough. A server that accepts a token issued for another service and forwards it upstream skips every control that service would have applied. The spec forbids it. Plenty of quick integrations do it anyway, because it works on the first try.

The tools section adds the line that matters most for anyone wiring agents to production systems: tool descriptions come from the server and are untrusted unless the server is trusted. A tool description is model input. Whoever controls the description controls part of the plan. That is the same content problem discussed in how MCP and skills compose, read from the security side: every connector dumped into context is one more surface an attacker can write to.

Per-user permissions do not bound an agent

The reflex is to reuse the identity stack. Single sign-on, group membership, a personal token per employee. It feels like least privilege because it is least privilege, for a human.

It fails for an agent on three counts.

The scope is the person, not the task. A finance analyst’s identity can read the general ledger and post journal entries. An agent drafting a variance commentary needs the first and must never have the second. Personal credentials cannot express that split without creating a second identity per task, which nobody maintains.

The grant widens at runtime. Agents discover tools. A client that lists every server the user has ever authorised hands the planner every scope at once, whether or not the current task needs it. The context bloat that hurts accuracy is the same list that hurts security.

Nobody can answer “why did it do that”. When the token is personal, the audit trail says the person did it. The person did not. A prompt injected in a supplier PDF did, and the log cannot tell the difference. That gap is what turns shadow AI from a data-leak worry into an accountability problem, and it is why a community SKILL.md nobody reviewed is a bigger exposure than it looks.

The role is the permission boundary

The unit that fits an agent is not the user and not the tool. It is the role: the work one seat on the org chart actually does, expressed as a bundle of skills (the instructions) and scoped tools (the connectors), with the permissions attached to the bundle rather than to whoever invokes it. A role is configured from a blueprint against a team’s own processes; it is never a ready-made package.

Four properties make a role a real boundary rather than a naming convention. The skilder team has since tested this design against flat tool injection and multi-agent orchestration on 13 scenarios and six models; what the paper measures is the enforcement half of this argument.

Granted centrally, and it does not widen at runtime. Someone with authority decides which servers the role can reach and with which scopes. The agent performing the role sees those tools and no others. Discovery happens inside the boundary, never across it.

Scoped per task, not per person. The variance-commentary role gets ledger read. A separate close-the-books role gets journal write, with a confirmation step on every post. Two roles, two grants, one person allowed to invoke both. The identity stack still authenticates the person; it no longer defines the reach.

Backed by a curated registry, not an open directory. The servers a role can name come from an internal list of connectors that were reviewed, credentialed and scoped once. Adding a server to that list is the authorisation step. A role cannot reach a server that is not on it, which closes the token-passthrough shortcut by construction.

Every call attributed to a person, through the role. The log records who invoked the role, which tool the agent called, with which arguments, and what came back. When a tool result carried an injected instruction, the trace shows the turn where the plan changed. That is the difference between an incident report and a guess.

This is how skilder is built: roles bundle skills and scoped MCP tools, are granted through role-based access control, and are served on demand to whichever agent the employee already uses. The execution trace is kept, attributable and exportable, as described on the trust and security page.

A checklist for the next connector you wire

Least privilege for agents is a habit, not a product. The habit looks like this.

  1. Name the role before the tool. Write one sentence describing the work. If the sentence needs two different write permissions, it is two roles.
  2. Grant read before write, and write with confirmation. Any tool that changes state outside the conversation gets a human step until the role has a track record.
  3. Never pass a personal token to a server. The server gets its own credential, scoped to what the role needs. If the upstream API cannot scope that narrowly, the server wraps it and exposes only the narrow operations.
  4. Treat tool descriptions and tool results as input. Review a server’s descriptions before it enters the registry. Log results. Assume a document can talk to the planner, because it can.
  5. Break the trifecta on purpose. For a role that reads private data and ingests untrusted content, remove the external channel. For a role that must send outward, keep it away from the sensitive read. One of the three legs goes, every time.
  6. Log to the person, through the role. If the trace cannot say who invoked which role and what the agent called, the deployment is not ready for production data.

The teams that get this right do not ship slower. They ship the first role slowly and every role after it fast, because the grant is made once and reused, instead of renegotiated per person and per tool.

Takeaway: stop asking who may use the agent, and start asking what each role may reach. Then make that grant central, scoped, and impossible to widen from inside a prompt.

To see how roles, scoped tools and the execution trace fit together, start with the skilder documentation.

Frequently asked questions

What does least privilege mean for an AI agent?

The agent can only reach the tools, data and actions the task in front of it needs, and nothing else. In practice that means scoping connectors per role rather than per user, granting read before write, and refusing any grant that the agent could widen on its own during a run.

Why is a personal API token the wrong way to connect an agent?

A personal token carries every permission the person holds across every system it touches. The agent then reasons over all of it. One injected instruction in a document or a tool result is enough to turn that reach into an action nobody asked for. The MCP specification calls this the confused deputy problem and warns against token passthrough for the same reason.

Does least privilege slow down agent adoption?

It slows down the first connection and speeds up every one after it. Once a role exists with its scoped tools, the next team reuses it instead of negotiating a new set of credentials. The cost moves from every user to one central grant.

Related articles