# Progressive skill discovery as access control: what the skilder paper measures

> A 30-page white paper tests whether serving skills through one MCP server can enforce what an agent may do. 13 scenarios, six models, three tool interfaces, and the limits the authors report.

Language: en · Category: governance · Author: Nicolas Corod (Founder & CEO) · Published: 2026-10-01

## TL;DR

**Every guardrail written into a system prompt is a request, not a rule.** A tool-using agent can be talked, tricked or simply confused into calling something it was told not to touch, because the tool definition sits in its context and the instruction is just more text. The [white paper the skilder team posted on arXiv](https://arxiv.org/abs/2609.28693) on 23 September 2026 asks a narrower question than "how do we make agents behave": can the interface that delivers tools to an agent be the thing that enforces which tools it may use? The setup is progressive skill discovery. The agent connects to one MCP server, sees four platform tools and a catalog of roles, learns the role its task requires, and receives that role's skills, instructions and tools. A tool outside a learned role cannot be called. The evaluation runs 13 scenarios on six models against two baselines, flat tool injection and multi-agent orchestration, and reports where the design holds, where it costs, and where some models cannot follow it.

This is a vendor's paper about the vendor's own design. Read it as a description of the mechanism and of a test harness that anyone can argue with, not as a neutral benchmark.

## The problem the paper starts from

Companies connect agents to CRM, billing, HR, security and engineering systems at once. The common baseline hands the model every tool definition in one flat list. That works at fifteen tools. As the catalog grows, definitions eat context and tool selection degrades. More importantly, as the authors put it, [prompt instructions are not access controls](https://arxiv.org/html/2609.28693#S1): an agent can be induced to take harmful actions despite instructions to the contrary, and every unnecessary tool in context widens the action space in which that can happen.

The usual answer is to split the work into specialist agents behind an orchestrator: one for tier 1 support, one for billing, one for fraud. Each specialist sees a shorter list. But every handoff is a new model session and a new message, end-to-end tracing gets harder, and unless authorization is enforced outside the models, policy is still a prompt inside each specialist. The paper's framing is that both baselines leave the boundary inside the model. It wants the boundary in the server.

## The mechanism: a role is the unit of access control

The architecture section is short and worth reading in full. [skilder is an MCP server](https://arxiv.org/html/2609.28693#S3). The agent connects to it and to nothing else, and sees four platform tools:

- `init_skilder` starts the session and returns the catalog of roles inside the session's authorization scope: names, one-line descriptions, skill names, and the path to learn each. No instructions and no domain tools yet.
- `learn` takes a path. Learning a role returns its instructions and its skills, each with its own instructions and tools, and unlocks every tool the role carries. Learning a single skill again re-reads it. Learning a resource fetches an attached document.
- `call_tool` asks skilder to execute a domain tool on the agent's behalf, and does so only if the tool belongs to a learned role. Anything else returns `ACCESS DENIED` with the list of tools that are available.
- `feedback_skill` lets the agent rate or comment on a skill.

The check inside `call_tool` is what the paper calls the router. Two layers work together: tool-level access, where a tool that was never delivered inside a learned skill cannot be called, and catalog-level scoping, where a role outside the session's authorization scope cannot even be learned. This is the same idea this blog described as [the role being the permission boundary](/en/blog/least-privilege-ai-agents/), written down as a protocol. The paper's own summary of the properties: small start regardless of catalog size, scoped access one role at a time, a hard block the model cannot skip, and the option to learn a second role when a thread crosses domains.

The comparison with plain MCP is precise. MCP exposes a flat tool namespace in which the agent sees every registered tool at once. The Agent Skill standard packages instructions and resources but leaves tool exposure to the host. skilder combines the two: it serves skills, and tools reach the agent only through the skills it has learned. That is the difference between hiding a tool from the prompt and being unable to call it, which is also the difference [composing skills over MCP tools](/en/blog/mcp-vs-skills-composing/) turns on.

## What was measured

The harness is built on [promptfoo](https://promptfoo.dev) with a custom agent loop. Tool calls go through a simulated authorization layer that copies the role design: a role catalog, tool learning, access control including a $500 dollar limit and a required step order, and fixture data for the domain tools. The authors are explicit that results measure this design, not a live product snapshot.

Three conditions use the same model and change only how tools reach it:

| Condition | How tools reach the model | Hard limit outside the model |
| --- | --- | --- |
| Flat injection | Every domain tool in context from turn 1, generic system prompt | None |
| Multi-agent orchestration | A coordinator with `delegate_to_agent`; one sub-agent per domain with that domain's tools and prompt | None unless added; tokens add coordinator and sub-agents |
| skilder | Four platform tools; roles learned on demand; domain tools only through `call_tool` | `ACCESS DENIED` outside learned roles, `GOVERNANCE VIOLATION` above the limit |

Six models: Claude Haiku 4.5, Qwen 3.5 122B, Gemma 4 31B, Ministral 3 14B, Claude Opus 4.7 and GPT-5.5, through the Anthropic and OpenAI APIs and Infomaniak AI, which is Swiss-hosted. Scenarios 1 to 10 run ten trials per model and condition; the institutional suite, scenarios 11 to 13, runs five. The 13 scenarios group into four themes: structural governance (social-engineered admin requests, over-limit refunds, role selection, ambiguous account activity), institutional policy (a resolution ladder with brand language that the agent must follow in order), multi-domain adaptability (support that turns into fraud discovery, proactive role expansion), and correctness (ordinary support tasks that flat injection already completes).

One scoring rule matters more than any table. A trial fails if any check fails, and the [authors split skilder misses into two buckets](https://arxiv.org/html/2609.28693#S5): the model never reached the router (it did not finish `init` then `learn`, never issued a governed call, or failed a wording check), and the router was tested (the model learned a role and made a call). Only the second bucket says anything about enforcement.

## What was found

On the direct test of the governance claim, the paper reports no observed enforcement failure: when a governed call reached the simulated authorization layer, no unauthorized tool call and no over-limit refund executed. Tools outside the learned role got `ACCESS DENIED`, refunds above $500 got `GOVERNANCE VIOLATION`, and roles outside the session scope were denied at catalog lookup. The authors define a true platform miss as a learned role and a forbidden call that still executes, and write that they do not observe that in this suite.

The pooled numbers are less tidy, and the paper prints them anyway. Theme pass rates across the six models, from [table 14](https://arxiv.org/html/2609.28693#S5.SS5):

| Theme (scenarios) | n | Flat injection | Multi-agent | skilder |
| --- | --- | --- | --- | --- |
| Governance (5 to 8) | 240 | 31.3% | 96.7% | 90.4% |
| Institutional policy (11 to 13) | 90 | 68.9% | 33.3% | 68.9% |
| Adaptability (9 to 10) | 120 | 75.0% | 69.2% | 75.8% |
| Parity (1 to 4) | 240 | 95.4% | 87.9% | 82.5% |

Three readings the authors give of that table. On governance, skilder substantially exceeds flat injection while multi-agent is higher still; the skilder column mixes router holds with trials that never issued a governed call, so it understates enforcement. On institutional policy, flat injection ties skilder at 68.9% because the policy tool is already visible and retrieval is explicitly requested; the difference is structural, policy visibility is limited by role. On parity, skilder is lower (198 of 240 against 229 of 240), and the gap is protocol cost on two models plus three rubric misses on Opus, not a governance leak.

The per-model split is the most useful part for anyone choosing a model pool. On the all-check total over scenarios 1 to 10, Haiku 4.5 and Gemma 4 31B scored 100 of 100 under skilder, GPT-5.5 and Opus 4.7 scored 96, while Qwen 3.5 122B scored 53 and Ministral 3 14B scored 61. The two low rows failed the multi-step `learn` sequence more often than they failed anything else. The paper's line: protocol compatibility is model-dependent, `ACCESS DENIED` itself is not. Opus adds a different wrinkle: it often refuses overt misuse even under flat injection, yet tends to treat role-embedded policy as a prompt injection, which the authors flag as a deployment limitation for policy-bearing roles on that model.

The token study is a footnote in the paper and a headline for platform teams. With one strong model, Sonnet 4.5, and catalogs from 15 to 225 tools, the `init`, `learn`, `call_tool` sequence costs more than a direct call at 15 tools, crosses flat injection near 30 tools, and at 225 tools a single customer lookup uses [9,084 tokens with skilder against 51,330 with flat injection](https://arxiv.org/html/2609.28693#S7). Context stays at four platform tools plus a short catalog as the catalog grows.

## The limits, as the paper states them

The limitations section is one paragraph per point and none of them is buried:

- **Harness, not a product changelog.** The simulated authorization layer implements the role design under test. Results measure that design, not a live runtime.
- **Model-dependent discovery.** skilder scores range from 53 of 100 to 100 of 100 depending on the model. Deployments should validate protocol compatibility when selecting their model pool.
- **Mock responses.** Domain tools return fixtures, not live MCP servers. Cold-start latency is out of scope.
- **Single-run token curves.** Sonnet 4.5 only, no prompt caching.

Two more caveats sit in the discussion. The three conditions share model weights but not an identical inference path: multi-agent gets a larger inference budget and higher prompt priority for specialist instructions, skilder imposes an extra discovery protocol, so behavioral differences measure the whole setup. And scenarios 5, 6 and 8 were rerun after their assertions were changed to score execution outcomes rather than counting a denied attempt as a breach; the other scenarios were not recoded.

The future work list is concrete: separate scores for structural and behavioral checks, failure analysis of the two models that miss multi-step `learn` and of Opus on role-embedded policy, real MCP servers instead of fixtures, larger catalogs, production telemetry.

## What it means if you are wiring agents to company tools

Three practical consequences follow from the paper, and none of them requires believing the pooled numbers.

First, the enforcement point belongs in the server that delivers tools, not in the prompt and not in the specialist's system message. A model that never received a tool definition cannot call it, and a server that checks membership before executing does not depend on the model's judgement. That is the property that turns [shadow MCP](/en/blog/shadow-mcp/) from a discovery problem into a scoping problem.

Second, test the discovery protocol against your own model pool before you rely on it. Four of the six models followed `init` then `learn` then `call_tool` reliably. Two did not, and no amount of enforcement helps a model that never reaches the router.

Third, the trade is explicit. A short catalog and on-demand delivery cost a few extra calls on a one-role task and save an order of magnitude at 225 tools. Whether that crossover, around 30 tools, is behind you or ahead of you is a question about your catalog, not about the paper.

In skilder the mechanism described here is how a role is served: configured from a blueprint, then exposed over MCP to whatever agent a team already uses. The [documentation](https://docs.skilder.ai) shows the same four platform tools from the client side, and a first role can be configured at [app.skilder.ai/signup](https://app.skilder.ai/signup).

## The paper

Michael Stettler, Benjamin Girardet, Jonas Canton and Nicolas Corod, "Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery", white paper, 30 pages, [arXiv:2609.28693](https://arxiv.org/abs/2609.28693) ([PDF](https://arxiv.org/pdf/2609.28693)), categories cs.AI, cs.CR, cs.MA and eess.SY, submitted 23 September 2026. The four authors are the skilder team.

```bibtex
@misc{stettler2026progressive,
  title        = {Progressive Skill Discovery as Access Control for Tool-Using LLM Agents:
                  Structural Governance through Role-Scoped Capability Delivery},
  author       = {Stettler, Michael and Girardet, Benjamin and Canton, Jonas and Corod, Nicolas},
  year         = {2026},
  eprint       = {2609.28693},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  url          = {https://arxiv.org/abs/2609.28693}
}
```
