Skip to main content
skilder

product

You can't manage what you can't measure: the problem of valuing agent skills

Everyone is building agent skills, but almost no one can say which ones are worth keeping. Why skill value resists measurement, and what an honest answer looks like.

Author: Nicolas Corod
  • #skills
  • #ai-agents
  • #observability
  • #ai-roi
skilder mascot next to the article title on a paper background

Every team building with AI agents is quietly accumulating the same thing: skills. A skill to draft a contract clause. A skill to reconcile invoices against a ledger. A skill to turn a messy CSV into a board-ready chart. Since Anthropic introduced Agent Skills in October 2025 as folders of instructions, scripts and resources that an agent loads when a task calls for them, organizations have been producing them at a remarkable pace. (If the concept is new to you, here is why skills make agents smarter.)

Then someone in a planning meeting asks the obvious question: which of these skills are actually worth keeping? And the room goes quiet.

Most teams can tell you how many tokens an agent burned last month. They can tell you the latency of a tool call down to the millisecond. But ask which skill generated value, for whom, and how much, and the honest answer today is that nobody really knows. The instrumentation built for AI systems measures the engine, not the work it does. The teams that close this gap first will invest in their skill libraries with confidence instead of guesswork.

Why the question is suddenly urgent

For a while, skill sprawl didn’t matter much. You had a handful of capabilities, you knew them all by name, and “is this useful?” was a judgment call you could make over coffee.

That era is ending. Skill libraries are growing into the dozens and hundreds. They are shared across teams, embedded in several agents, versioned, forked and deprecated. Many organizations now bundle them into roles: a role is the work one seat on the org chart actually does, packaged as a set of skills, the tools they need and the permissions that go with them. Each skill carries a cost: the time to build and maintain it, the tokens it consumes at runtime, the review effort to keep it safe and current.

Cost, notably, is the easy part to measure. The benefit side is where everything falls apart, and that asymmetry quietly distorts every decision. You can prove what a skill costs; you can only hand-wave at what it returns.

Why skills resist measurement

The trouble isn’t that teams are lazy about metrics. Agent skills have structural properties that defeat the measurement habits brought over from traditional software.

A skill doesn’t cleanly “execute”

In conventional software, a function call is a crisp event. It starts, it runs, it returns, and you can log all of it. A skill is fuzzier. The model decides, in context, whether and how to draw on a skill. The same request might invoke a skill once, twice, partially or not at all. The boundary between “the model did this” and “the skill did this” is genuinely blurry: the skill shapes the model’s behaviour rather than running as an isolated unit. So even the most basic metric, did this skill run and what did it produce, is surprisingly hard to capture reliably.

Nobody knows who consumed it

A skill rarely serves one person. It gets embedded in an agent, that agent serves a workflow, that workflow serves many users across several teams. When value finally shows up (a deal closed faster, a report that didn’t need rework), there is no clean lineage tracing it back to the specific skill that helped. Ask “who actually benefits from this skill, and how often?” and you are usually reconstructing the answer from fragments.

Output is not outcome, and there’s no counterfactual

Suppose you could count executions perfectly. You still wouldn’t have measured value, because value is the outcome a skill produced: time saved, errors avoided, quality improved. Connecting an invocation to a business result means tracing a long, confounded causal chain, and at the end of it sits the hardest wall of all: the counterfactual. To know what a skill added, you would need to know what would have happened without it. Agents are non-deterministic and expensive to run twice, so the clean A/B comparison that would settle the question almost never gets done.

Asking people doesn’t fix this. In a randomized controlled trial run by METR in 2025, 16 experienced open-source developers expected AI tools to speed them up by 24%. With the tools, they actually took 19% longer, and afterwards they still believed they had been 20% faster. Self-reported time savings are weak evidence.

Credit is shared, and division is unsolved

Real tasks compose skills. A single completed job might lean on a retrieval skill, a reasoning skill and a formatting skill in sequence. When the outcome is good, how do you divide the credit? A skill can be necessary without being sufficient; its marginal contribution depends on what else is in the chain. Outside carefully controlled experiments, most teams resolve this by not resolving it: they credit the whole pipeline and move on. Treating the catalogue as a graph of skills with explicit dependencies at least makes that chain visible.

The tooling was built for the model, not the skill

Underneath all of this is an observability stack designed around model calls and tool calls. The OpenTelemetry semantic conventions for generative AI, the closest thing to an industry standard, define spans, metrics and events for model clients and MCP, with token usage and latency as the headline metrics. They were not designed to follow a skill across invocations, consumers, versions and outcomes. Until skills become first-class citizens in telemetry, the data simply isn’t there to answer the value question.

The proxy-metric trap

Faced with all this, teams reach for the metrics they can get: invocation counts, adoption rates, “this skill was used 4,000 times last quarter.” It feels like progress. It is also a trap.

Usage is not value. A heavily invoked skill might be doing trivial work, while a skill used twice a month might be the one preventing a compliance incident. Optimizing for usage rewards the busy over the important, and pushes teams to deprecate exactly the wrong things. It also misses what skills are for: capturing know-how rather than knowledge, which is precisely the part no counter sees.

What honest measurement looks like

None of this is unsolvable; it’s unsolved. A new layer is forming above raw model-and-token telemetry, whose job is to make skills measurable. Call it skill observability. Wherever it lives, it has to deliver a few capabilities:

  • Skills as observable entities. Instrumentation at the skill boundary, so an invocation is a real, loggable event with its own identity.
  • Consumer lineage. Every invocation tagged with the agent, person, team, workflow and task that triggered it.
  • Outcome linkage. A way to connect invocations to downstream signals: human review, task success, time to completion.
  • Counterfactual capability. Running a workflow with and without a given skill, at least in evaluation, so marginal value can be estimated rather than assumed.
  • Attribution for composition. A principled way to divide credit when several skills cooperate on one outcome.
  • A shared definition of value. The same units (time saved, errors avoided, cost to serve) for every skill, so they can be compared.

One discipline matters more than the list: call an estimate an estimate. Even the most careful large-scale work does. When Anthropic analysed 100,000 real Claude conversations in November 2025, it estimated an average time reduction of about 80% on tasks, then spelled out that the figure does not count time spent checking or refining the output outside the chat, and that it is not a prediction. That is the right posture for skill value too: time saved and its worth are derived from real usage, task by task, and reported as estimates, not as measured savings.

This is the posture skilder takes. Usage is tracked per role, so an organization can see which capabilities are used, by whom and how often, alongside an estimate of the time saved and what that time is worth. Those figures are estimates derived from real usage, and they are labelled that way.

The bottom line

Right now, most organizations are flying their skill libraries blind. They see the fuel gauge (cost) with perfect clarity and almost nothing about where the trip is taking them. That is tolerable with five skills. With five hundred, many of them living in personal folders where SKILL.md feeds shadow AI, it is a real liability.

The first step isn’t a better dashboard; it is deciding that skill value is something you design to measure, starting at the moment a skill is invoked. If you want a quick read on where your organization stands, the 10-question AI check-up takes two minutes and asks for nothing. The teams that build this discipline early won’t just trim their libraries more wisely; they will have a defensible estimate of what their agents are actually worth.

Frequently asked questions

How do you measure the value of an agent skill?

Start by making the skill an observable entity: log each invocation at the skill boundary, tag it with who triggered it and for which task, then link it to an outcome signal (task success, human review, time to completion). From that usage you can derive an estimate of time saved and what that time is worth. It stays an estimate, because the counterfactual (the same task done without the skill) is rarely observed directly.

Is usage a good metric for agent skills?

Usage tells you a skill is being called, not that it adds value. A heavily used skill can do trivial work, while a rarely used one can prevent a costly error. Usage is a useful input for an estimate, never the estimate itself.

Can you measure AI time savings precisely?

Not precisely. Controlled studies show people misjudge their own speed-up, and large-scale analyses of real conversations present their figures as estimates with explicit caveats. The defensible approach is to estimate time saved from real usage, task by task, and to say clearly that it is an estimate.

Related articles