Extending the Skill-Use mechanical checks, per the baseline's own next step, found something bigger than a missing check. THE DEFECT `workspace-repo-commit-protocol` told agents that `/workspace/repo` was "the ONLY path where source-modifying edits belong". The platform mounts and advertises `/mission/repo` — 26 references in the code; `/workspace/repo` appears in none of them. The skill is bound on 29 role bindings and was delivered TWICE in the run already measured, so an agent received the real path in its tool preamble and a skill contradicting it a few hundred tokens later, in one prompt. An agent that obeyed the skill wrote source into a directory nothing collects — the phase then delivers nothing, and looks like an agent that did no work. The same skill instructed `file_read` / `file_write` / `shell`: ZeroClaw's names, the exact ones `phase_task_text` was fixed to stop advertising after five agents on a single mission spent 7.4k tokens describing the mismatch instead of working. The prompt was corrected and the skill kept saying it. Rewritten against what the code actually does, including the repo-less case (`/mission/repo` exists, is collected as artifacts, has nothing to push). THE CLASS, AND THE GUARD The skills were never checked against the platform they describe. Nothing compared them, so a skill could contradict the prompt it ships inside and stay that way indefinitely — the same shape as PLAN_COMPLETE being documented and never implemented. Two tests in `skills_loader::contradiction_tests` now hold it: no skill may name a repo path the platform does not mount, and none may instruct a tool the agent's subprocess does not expose. The second matches backticked instructions and skips corrective lines, so a skill may still WARN against the wrong names — as this one now does. Both negative-controlled by restoring the old wording. AND THE CHECK THAT STARTED IT `workspace-repo-commit-protocol` now has a Boundary check: writing outside `/mission/repo` fails, and the message names the consequence — a phase that delivers nothing — rather than just the wrong path. docs/SKILL-USE-BASELINE.md records this as the fourth defect the measurement found, and corrects the "next unit of work" note now that this one is done. Full workspace suite green: 106 binaries, zero build errors. Co-Authored-By: Claude Opus 5 <[email protected]>
175 lines
8.3 KiB
Markdown
175 lines
8.3 KiB
Markdown
# Skill-Use baseline
|
|
|
|
*First measurement of whether ClawMates' skills change what agents do.
|
|
2026-08-19.*
|
|
|
|
Scored on the three axes from `Skill-Use` (arXiv, 2026-08-05): **Trigger** (did
|
|
the agent reach for the skill), **Compliance** (did it follow the procedure),
|
|
**Boundary** (did it avoid what the skill forbids).
|
|
|
|
Read the method before the numbers. A measurement whose limits are not stated
|
|
is worse than none, because it gets quoted without them.
|
|
|
|
## Why there was no baseline before today
|
|
|
|
Not because nobody ran it. Because **it could not have returned anything but
|
|
zero**, for two structural reasons that had nothing to do with agent behaviour:
|
|
|
|
1. 55 of 85 role skill bindings resolved to skills that were never authored.
|
|
2. Even resolved skills had no delivery channel to a mission agent — the
|
|
catalogue's only route was an MCP server that mission claws cannot reach.
|
|
|
|
Both were fixed in the two commits preceding this document. Anyone who had run
|
|
this measurement in July would have concluded "our agents ignore their skills",
|
|
which would have been false and expensive.
|
|
|
|
## Method, and what it cannot see
|
|
|
|
Scored by `cm_api::skill_use` from what the platform records: the
|
|
`prompt.composed` event (the exact bytes an agent received) and the `reasoning`
|
|
events (what it said it did). No re-derivation from the catalogue — the
|
|
catalogue changes, and now that agents author their own skills it changes by
|
|
itself.
|
|
|
|
### Trigger is not observable here, and that is a finding
|
|
|
|
The paper measures agents under **progressive disclosure**: the agent sees a
|
|
name and description and must decide to retrieve the body. That retrieval is a
|
|
tool call, which makes Trigger observable.
|
|
|
|
We do not deliver skills that way. `pinned_skills_text` inlines full bodies into
|
|
the prompt, because mission claws run on `claude_cli`, which cannot surface a
|
|
tool call — there is nothing to retrieve *with*. The agent never reaches for a
|
|
skill; it is simply holding one.
|
|
|
|
So Trigger is reported as `not_observable` with the reason attached, **never as
|
|
zero**. Scoring it zero would report a delivery-model property as an agent
|
|
failure — the same confusion that kept 55 empty bindings invisible.
|
|
|
|
### Compliance and Boundary are checked mechanically, or not at all
|
|
|
|
Only skills whose procedure has a machine-checkable consequence are scored.
|
|
Everything else returns `not_applicable` rather than a guess: a heuristic that
|
|
scores prose by keyword overlap produces a number that looks like a measurement
|
|
and is not one.
|
|
|
|
Compliance for `int-xx-marker-protocol` is checked by running the **real**
|
|
`task_card_parser`, not a copy of its rules — a second implementation would
|
|
drift, and then the score would pass while the mission loop still stalled.
|
|
|
|
## The runs
|
|
|
|
Two missions on the container/ZeroClaw tier, local stack, `research_only`.
|
|
|
|
| | run 1 | run 2 |
|
|
|---|---|---|
|
|
| distinct skills delivered | 3 | **9** |
|
|
| total deliveries (per role prompt) | 3 | 14 |
|
|
| phantom "skills" scored | **2** | 0 |
|
|
|
|
Run 1's phantom entries are the finding of the run, described below.
|
|
|
|
Run 2, per skill (all `source_kind=builtin`; no agent-authored skill has been
|
|
delivered yet):
|
|
|
|
| skill | deliveries | compliance | boundary |
|
|
|---|---|---|---|
|
|
| `int-xx-marker-protocol` | 1 | **pass** | n/a |
|
|
| `small-focused-commits` | 4 | n/a | n/a |
|
|
| `cargo-test-driven-development` | 2 | n/a | n/a |
|
|
| `workspace-repo-commit-protocol` | 2 | n/a | n/a |
|
|
| `decompose-int-items` | 1 | n/a | n/a |
|
|
| `write-rust-current-edition` | 1 | n/a | n/a |
|
|
| `code-review-checklist` | 1 | n/a | n/a |
|
|
| `criterion-benchmarking` | 1 | n/a | n/a |
|
|
| `tdd-red-green-refactor` | 1 | n/a | n/a |
|
|
|
|
**n = 2 runs. No spread is reported because two runs cannot establish one.**
|
|
This is a baseline in the sense of "the first honest number", not in the sense
|
|
of `metrics-baseline-comparison.md`, which requires enough runs to see the noise
|
|
floor before any change is judged against it. Do not compare a future number to
|
|
this one without first establishing that floor.
|
|
|
|
## What the measurement found
|
|
|
|
Four defects, none of which any test or log would have surfaced.
|
|
|
|
### 1. The prompt format made its own record unparseable
|
|
|
|
Skills were introduced with `## <name>`, and skill bodies are markdown full of
|
|
`##` headings. Run 1 duly scored **"Sizing heuristic"** and **"The output
|
|
shape"** — both subheadings inside `decompose-int-items` — as skills with no
|
|
catalogue row.
|
|
|
|
Fixed with an unambiguous `--- SKILL: <name> ---` marker, and both writers now
|
|
share one renderer so the reader cannot drift from the writer.
|
|
|
|
### 2. A prompt was recorded that was never sent
|
|
|
|
The phase prompt was recorded at the dispatch fork, before the tier was chosen.
|
|
The container tier does not send that text — it sends the bare task and appends
|
|
skills per turn. So every container mission logged a `solo` prompt that reached
|
|
no agent.
|
|
|
|
A provenance record of something that did not happen is worse than no record: it
|
|
is the wrong answer, delivered confidently. Recording now happens inside each
|
|
tier, and a test asserts every launcher records the prompt it actually sends.
|
|
|
|
### 3. A pinned skill contradicted the platform in the same prompt
|
|
|
|
`workspace-repo-commit-protocol` told agents that **`/workspace/repo`** was "the
|
|
ONLY path where source-modifying edits belong". The platform mounts and
|
|
advertises **`/mission/repo`** — 26 references in the code; `/workspace/repo`
|
|
appears in none. The skill is bound on **29 role bindings** and was delivered
|
|
twice in run 2, so agents received the real path in the tool preamble and a
|
|
skill contradicting it a few hundred tokens later.
|
|
|
|
It also instructed `file_read` / `file_write` / `shell` — ZeroClaw's tool names,
|
|
the exact ones `phase_task_text` was fixed to stop advertising after five agents
|
|
on one mission spent 7.4k tokens describing the mismatch instead of working.
|
|
|
|
An agent that obeyed this skill wrote source into a directory nothing collects,
|
|
and reached for tools its subprocess does not expose. Rewritten against what the
|
|
code actually does, with two guards in `skills_loader::contradiction_tests`: no
|
|
skill may name a repo path the platform does not mount, and none may instruct a
|
|
tool the agent does not have. Both negative-controlled.
|
|
|
|
This is the same shape as the finding below and it is worth stating as a class:
|
|
**the skills were never checked against the platform they describe.** Nothing
|
|
compared them, so a skill could contradict the prompt it ships inside and stay
|
|
that way indefinitely.
|
|
|
|
### 4. The skill documents a marker the platform never implemented
|
|
|
|
`int-xx-marker-protocol` lists `PLAN_COMPLETE: INT-NN` in its ladder.
|
|
`task_card_parser` has **no such kind** and never has. An agent following the
|
|
skill exactly emits a marker that is silently ignored.
|
|
|
|
Observed live: run 2's planner emitted `PLAN_COMPLETE: INT-01..02`, which is
|
|
also the range form — on the kinds that *are* parsed, that yields the id
|
|
`INT-01..02`, a task card for an item that does not exist while the two real
|
|
items stay open.
|
|
|
|
**This is a skill/implementation mismatch, not an agent failure**, and it is
|
|
precisely what this measurement exists to find: the agent did what it was told,
|
|
and what it was told was wrong. Both shapes are now scored as failures; the
|
|
underlying reconciliation — implement `PLAN_COMPLETE` or remove it from the
|
|
skill — is deliberately left as a decision rather than guessed at here.
|
|
|
|
## Honest limits
|
|
|
|
- **Two runs, one tier, one workflow.** Nothing here generalises to the microVM
|
|
or session tiers yet, and this document should not be read as if it does.
|
|
- **Most skills score `not_applicable`** on both observable axes. That is not a
|
|
pass. It means we cannot currently tell whether those skills changed anything.
|
|
`workspace-repo-commit-protocol` now has a Boundary check (writing outside
|
|
`/mission/repo`); `small-focused-commits` and `tdd-red-green-refactor` remain
|
|
candidates, and both need the repository diff rather than the turn text.
|
|
- **No agent-authored skill has been measured.** Self-authoring shipped in the
|
|
same pass; `source_kind` is carried through the scorer specifically so a
|
|
rising score on agent-authored skills is visible rather than averaged in.
|
|
- **Evidence expires.** Mission events are reaped after 7 days unless
|
|
`retain_events_until` is set. Both runs here are held for 90 days. An empty
|
|
score means "no evidence", never "no compliance", and the API says so in its
|
|
payload rather than leaving the caller to infer it.
|