fix(skill-use): a research phase writing markdown is not a TDD failure
The first live scoring of run 3 reported `cargo-test-driven-development` and `tdd-red-green-refactor` as compliance=FAIL: files were written and no test ever ran. Wrong, and wrong in the way this module exists to prevent. The phase wrote fifteen markdown notes and a helper script; there was no code to test-drive. Reporting it as an agent failure is a system defect wearing an agent's name — and it would have buried the actual finding, which is that a repo-less `research_only` mission is staffed with a Rust SDLC crew whose coder, tester, reviewer and committer have nothing to do. The check is now scoped to files with a source extension in the languages the skill itself names. Shell is deliberately excluded: a helper script written during a research turn is not behaviour-adding code, and the false failure costs more than the missed one. Recorded in SKILL-USE-BASELINE.md as finding 8 rather than quietly corrected. A measurement that hides its own false positives cannot be trusted about anyone else's. Also in the doc: the Trigger reason is half false now (the transport can surface a tool call; we simply still inline), and the architecture doc's observe/gate table said the container tier was ungated and unobserved, which shipped work has made wrong. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
This commit is contained in:
co-authored by
Claude Opus 5
parent
4a6d0dfe01
commit
c209e654d9
+179
-112
@@ -1,7 +1,7 @@
|
||||
# Skill-Use baseline
|
||||
|
||||
*First measurement of whether ClawMates' skills change what agents do.
|
||||
2026-08-19.*
|
||||
*Whether ClawMates' skills change what agents do. First measured 2026-08-19;
|
||||
re-measured 2026-08-21 against tool evidence rather than agent prose.*
|
||||
|
||||
Scored on the three axes from `Skill-Use` (arXiv, 2026-08-05): **Trigger** (did
|
||||
the agent reach for the skill), **Compliance** (did it follow the procedure),
|
||||
@@ -10,165 +10,232 @@ the agent reach for the skill), **Compliance** (did it follow the procedure),
|
||||
Read the method before the numbers. A measurement whose limits are not stated
|
||||
is worse than none, because it gets quoted without them.
|
||||
|
||||
## Why there was no baseline before today
|
||||
## What changed on 2026-08-21
|
||||
|
||||
Not because nobody ran it. Because **it could not have returned anything but
|
||||
zero**, for two structural reasons that had nothing to do with agent behaviour:
|
||||
The 2026-08-19 measurement scored Compliance and Boundary from the
|
||||
`reasoning` events — the agent's own account of its turn. Since then the
|
||||
container tier records what agents actually **do**
|
||||
(`container_tool_hooks`, `PostToolUse`), and `vm_tool_tap` stopped throwing the
|
||||
tool's arguments away, so `Bash` commands and `Write` paths are on the record.
|
||||
|
||||
1. 55 of 85 role skill bindings resolved to skills that were never authored.
|
||||
2. Even resolved skills had no delivery channel to a mission agent — the
|
||||
catalogue's only route was an MCP server that mission claws cannot reach.
|
||||
`skill_use` now reads those. The difference is not cosmetic:
|
||||
|
||||
Both were fixed in the two commits preceding this document. Anyone who had run
|
||||
this measurement in July would have concluded "our agents ignore their skills",
|
||||
which would have been false and expensive.
|
||||
- `workspace-repo-commit-protocol`'s Boundary was a substring search for
|
||||
`/workspace/repo` in the narrative. **An agent that wrote to the wrong root
|
||||
without narrating it scored a clean pass.** It now reads the write paths.
|
||||
- `arxiv-daily`'s Boundary read a URL in prose, which may be the agent
|
||||
explaining that it did *not* fetch it. It now reads the `curl` that ran.
|
||||
- `tdd-red-green-refactor`, `cargo-test-driven-development` and
|
||||
`small-focused-commits` gained their first checks at all.
|
||||
|
||||
Two verdicts changed for honesty rather than coverage. **Silence used to score
|
||||
`Pass`** — a mission with no evidence scored identically to one checked and
|
||||
found clean; it is now `NotObservable`. And a test that ran *after* the first
|
||||
write is `NotObservable`, not a failure, because a Rust unit test lives in the
|
||||
file under test.
|
||||
|
||||
### Trigger: the reason changed, and only half of it went away
|
||||
|
||||
The 2026-08-19 document said Trigger was unobservable because `claude_cli`
|
||||
"cannot surface a tool call — there is nothing to retrieve *with*."
|
||||
|
||||
**That half is now false.** Mission tool calls are recorded on both tiers; a
|
||||
retrieval would be as visible as any other call.
|
||||
|
||||
The other half still holds and is the one that decides the verdict: **we still
|
||||
inline**. `pinned_skills_text` puts full skill bodies in the prompt, so the
|
||||
agent never reaches for anything — it is simply holding one. Trigger is now
|
||||
*instrumentable* and still not *observable*, and the blocker has moved from the
|
||||
transport to the delivery model.
|
||||
|
||||
Making it real is one change and no scorer work: serve skills through the door
|
||||
(`TOOL-CALL-ARCHITECTURE.md` §3) so retrieval becomes a tool call.
|
||||
|
||||
## Method, and what it cannot see
|
||||
|
||||
Scored by `cm_api::skill_use` from what the platform records: the
|
||||
`prompt.composed` event (the exact bytes an agent received) and the `reasoning`
|
||||
events (what it said it did). No re-derivation from the catalogue — the
|
||||
catalogue changes, and now that agents author their own skills it changes by
|
||||
itself.
|
||||
Scored by `cm_api::skill_use` from what the platform records: `prompt.composed`
|
||||
(the exact bytes an agent received), `reasoning` (what it said it did), and
|
||||
`tool.call` (what it did). No re-derivation from the catalogue — the catalogue
|
||||
changes, and now that agents author their own skills it changes by itself.
|
||||
|
||||
### Trigger is not observable here, and that is a finding
|
||||
|
||||
The paper measures agents under **progressive disclosure**: the agent sees a
|
||||
name and description and must decide to retrieve the body. That retrieval is a
|
||||
tool call, which makes Trigger observable.
|
||||
|
||||
We do not deliver skills that way. `pinned_skills_text` inlines full bodies into
|
||||
the prompt, because mission claws run on `claude_cli`, which cannot surface a
|
||||
tool call — there is nothing to retrieve *with*. The agent never reaches for a
|
||||
skill; it is simply holding one.
|
||||
|
||||
So Trigger is reported as `not_observable` with the reason attached, **never as
|
||||
zero**. Scoring it zero would report a delivery-model property as an agent
|
||||
failure — the same confusion that kept 55 empty bindings invisible.
|
||||
|
||||
### Compliance and Boundary are checked mechanically, or not at all
|
||||
**Every tool-backed check is one-sided.** It reports a violation it can see and
|
||||
never infers compliance from silence: the recorded stream is capped per phase
|
||||
(`PER_PHASE_CAP = 400`), so an absent call is not proof of an absent action.
|
||||
|
||||
Only skills whose procedure has a machine-checkable consequence are scored.
|
||||
Everything else returns `not_applicable` rather than a guess: a heuristic that
|
||||
Everything else returns `not_applicable` rather than a guess — a heuristic that
|
||||
scores prose by keyword overlap produces a number that looks like a measurement
|
||||
and is not one.
|
||||
|
||||
Compliance for `int-xx-marker-protocol` is checked by running the **real**
|
||||
`task_card_parser`, not a copy of its rules — a second implementation would
|
||||
`task_card_parser`, not a copy of its rules; a second implementation would
|
||||
drift, and then the score would pass while the mission loop still stalled.
|
||||
|
||||
## The runs
|
||||
|
||||
Two missions on the container/ZeroClaw tier, local stack, `research_only`.
|
||||
Three missions on the container/ZeroClaw tier, local stack, `research_only`.
|
||||
Run 3 uses the **same task text as run 2**, so the only variable is the scorer.
|
||||
|
||||
| | run 1 | run 2 |
|
||||
|---|---|---|
|
||||
| distinct skills delivered | 3 | **9** |
|
||||
| total deliveries (per role prompt) | 3 | 14 |
|
||||
| phantom "skills" scored | **2** | 0 |
|
||||
| | run 1 | run 2 | run 3 |
|
||||
|---|---|---|---|
|
||||
| date | 08-19 | 08-19 | 08-21 |
|
||||
| distinct skills delivered | 3 | 9 | 9 |
|
||||
| total deliveries (per role prompt) | 3 | 14 | 14 |
|
||||
| phantom "skills" scored | **2** | 0 | 0 |
|
||||
| tool calls recorded | 0 | 0 | **49** |
|
||||
| axes scored from actions | 0 | 0 | **3 skills** |
|
||||
|
||||
Run 1's phantom entries are the finding of the run, described below.
|
||||
|
||||
Run 2, per skill (all `source_kind=builtin`; no agent-authored skill has been
|
||||
Run 3, per skill (all `source_kind=builtin`; no agent-authored skill has been
|
||||
delivered yet):
|
||||
|
||||
| skill | deliveries | compliance | boundary |
|
||||
|---|---|---|---|
|
||||
| `int-xx-marker-protocol` | 1 | **pass** | n/a |
|
||||
| `small-focused-commits` | 4 | n/a | n/a |
|
||||
| `workspace-repo-commit-protocol` | 2 | n/a | **pass** *(from 12 write paths)* |
|
||||
| `small-focused-commits` | 4 | n/a | n/a *(no commit ran)* |
|
||||
| `cargo-test-driven-development` | 2 | n/a | n/a |
|
||||
| `workspace-repo-commit-protocol` | 2 | n/a | n/a |
|
||||
| `tdd-red-green-refactor` | 1 | n/a | n/a |
|
||||
| `decompose-int-items` | 1 | n/a | n/a |
|
||||
| `write-rust-current-edition` | 1 | n/a | n/a |
|
||||
| `code-review-checklist` | 1 | n/a | n/a |
|
||||
| `criterion-benchmarking` | 1 | n/a | n/a |
|
||||
| `tdd-red-green-refactor` | 1 | n/a | n/a |
|
||||
|
||||
**n = 2 runs. No spread is reported because two runs cannot establish one.**
|
||||
This is a baseline in the sense of "the first honest number", not in the sense
|
||||
of `metrics-baseline-comparison.md`, which requires enough runs to see the noise
|
||||
floor before any change is judged against it. Do not compare a future number to
|
||||
this one without first establishing that floor.
|
||||
**n = 3 runs. No spread is reported because three cannot establish one.** This
|
||||
is a baseline in the sense of "the first honest number", not in the sense of
|
||||
`metrics-baseline-comparison.md`, which requires enough runs to see the noise
|
||||
floor before any change is judged against it.
|
||||
|
||||
`workspace-repo-commit-protocol`'s pass is the one score that materially
|
||||
improved: it now rests on twelve recorded `Write`/`Edit` paths, every one under
|
||||
`/mission/repo`, instead of on the absence of a string in prose.
|
||||
|
||||
## What the measurement found
|
||||
|
||||
Four defects, none of which any test or log would have surfaced.
|
||||
### 1–4: the 2026-08-19 findings
|
||||
|
||||
### 1. The prompt format made its own record unparseable
|
||||
Four defects, none of which any test or log would have surfaced: the prompt
|
||||
format made its own record unparseable (`## <name>` against markdown bodies);
|
||||
a prompt was recorded that was never sent; a pinned skill taught
|
||||
`/workspace/repo`, a path the platform does not mount; and
|
||||
`int-xx-marker-protocol` documented a `PLAN_COMPLETE` marker the parser had
|
||||
never implemented. All four are fixed, with guards in
|
||||
`skills_loader::contradiction_tests` and `topology_exec`. The detail is in this
|
||||
file's git history.
|
||||
|
||||
Skills were introduced with `## <name>`, and skill bodies are markdown full of
|
||||
`##` headings. Run 1 duly scored **"Sizing heuristic"** and **"The output
|
||||
shape"** — both subheadings inside `decompose-int-items` — as skills with no
|
||||
catalogue row.
|
||||
### 5. The tool tap recorded the name and discarded the argument
|
||||
|
||||
Fixed with an unambiguous `--- SKILL: <name> ---` marker, and both writers now
|
||||
share one renderer so the reader cannot drift from the writer.
|
||||
The first container-tier mission with telemetry recorded `Bash × 6` and not one
|
||||
of them said what it ran. `vm_tool_tap::parse` read `tool_input` to pull the
|
||||
path out of it and dropped the rest.
|
||||
|
||||
### 2. A prompt was recorded that was never sent
|
||||
Every behavioural question was therefore unanswerable from a record that looked
|
||||
complete — which is the recurring shape, not a new one. Fixed host-side: the
|
||||
arguments were always in the tap file.
|
||||
|
||||
The phase prompt was recorded at the dispatch fork, before the tier was chosen.
|
||||
The container tier does not send that text — it sends the bare task and appends
|
||||
skills per turn. So every container mission logged a `solo` prompt that reached
|
||||
no agent.
|
||||
### 6. Two more skills contradicted the platform
|
||||
|
||||
A provenance record of something that did not happen is worse than no record: it
|
||||
is the wrong answer, delivered confidently. Recording now happens inside each
|
||||
tier, and a test asserts every launcher records the prompt it actually sends.
|
||||
Same class as finding 3, and both found by reading the source of truth before
|
||||
writing a check against it.
|
||||
|
||||
### 3. A pinned skill contradicted the platform in the same prompt
|
||||
- **`decompose-int-items` taught `PLAN_COMPLETE: INT-01..05`.** An id is
|
||||
strictly `INT-` plus digits, so the range form is rejected outright: the plan
|
||||
pass records nothing while every item stays open. A live planner emitted
|
||||
exactly that line.
|
||||
- **`workspace-repo-commit-protocol` claimed the task-card parser advances
|
||||
mission state on the INT id in your commit subject.** Nothing in the platform
|
||||
reads commit messages. `task_card_parser::apply_for_run` reads `run_events` —
|
||||
the agent's turn output. An agent that believed this would commit with the id,
|
||||
never emit `COMPLETED: INT-NN`, and leave the mission open on an item it had
|
||||
already finished.
|
||||
|
||||
`workspace-repo-commit-protocol` told agents that **`/workspace/repo`** was "the
|
||||
ONLY path where source-modifying edits belong". The platform mounts and
|
||||
advertises **`/mission/repo`** — 26 references in the code; `/workspace/repo`
|
||||
appears in none. The skill is bound on **29 role bindings** and was delivered
|
||||
twice in run 2, so agents received the real path in the tool preamble and a
|
||||
skill contradicting it a few hundred tokens later.
|
||||
`no_skill_shows_a_marker_the_parser_would_reject` now runs the real parser over
|
||||
every marker in every skill's fenced blocks, negative-controlled against the
|
||||
range form.
|
||||
|
||||
It also instructed `file_read` / `file_write` / `shell` — ZeroClaw's tool names,
|
||||
the exact ones `phase_task_text` was fixed to stop advertising after five agents
|
||||
on one mission spent 7.4k tokens describing the mismatch instead of working.
|
||||
### 7. A repo-less research mission is staffed with a Rust SDLC crew
|
||||
|
||||
An agent that obeyed this skill wrote source into a directory nothing collects,
|
||||
and reached for tools its subprocess does not expose. Rewritten against what the
|
||||
code actually does, with two guards in `skills_loader::contradiction_tests`: no
|
||||
skill may name a repo path the platform does not mount, and none may instruct a
|
||||
tool the agent does not have. Both negative-controlled.
|
||||
This is the finding of run 3, and it explains most of the `not_applicable`
|
||||
column above.
|
||||
|
||||
This is the same shape as the finding below and it is worth stating as a class:
|
||||
**the skills were never checked against the platform they describe.** Nothing
|
||||
compared them, so a skill could contradict the prompt it ships inside and stay
|
||||
that way indefinitely.
|
||||
`templates/workflows/research_only.toml` declares `requires_repo = false` and a
|
||||
single `research` phase — and `default_team_template = "rust_sdlc"`. So the
|
||||
mission was staffed with **planner, coder, tester, reviewer, committer**, and
|
||||
each received the skills its role is bound to:
|
||||
|
||||
### 4. The skill documents a marker the platform never implemented
|
||||
```
|
||||
coder :: write-rust-current-edition, cargo-test-driven-development,
|
||||
workspace-repo-commit-protocol, small-focused-commits,
|
||||
int-xx-marker-protocol
|
||||
tester :: cargo-test-driven-development, criterion-benchmarking,
|
||||
tdd-red-green-refactor
|
||||
committer :: workspace-repo-commit-protocol, small-focused-commits
|
||||
reviewer :: code-review-checklist, small-focused-commits
|
||||
planner :: decompose-int-items, small-focused-commits
|
||||
```
|
||||
|
||||
`int-xx-marker-protocol` lists `PLAN_COMPLETE: INT-NN` in its ladder.
|
||||
`task_card_parser` has **no such kind** and never has. An agent following the
|
||||
skill exactly emits a marker that is silently ignored.
|
||||
There is no repository, nothing to test, nothing to review and nothing to
|
||||
commit. Four of the five roles have no work, and 50KB of prompt (~12.6k tokens)
|
||||
is spent staffing them.
|
||||
|
||||
Observed live: run 2's planner emitted `PLAN_COMPLETE: INT-01..02`, which is
|
||||
also the range form — on the kinds that *are* parsed, that yields the id
|
||||
`INT-01..02`, a task card for an item that does not exist while the two real
|
||||
items stay open.
|
||||
**The skills are correctly bound to the roles. The roles are wrong for the
|
||||
workflow.** That distinction matters: a reader who saw only "7 of 9 skills
|
||||
scored not_applicable" would conclude the skills are useless, when what the
|
||||
number actually measures is a staffing default.
|
||||
|
||||
**This is a skill/implementation mismatch, not an agent failure**, and it is
|
||||
precisely what this measurement exists to find: the agent did what it was told,
|
||||
and what it was told was wrong. Both shapes are now scored as failures; the
|
||||
underlying reconciliation — implement `PLAN_COMPLETE` or remove it from the
|
||||
skill — is deliberately left as a decision rather than guessed at here.
|
||||
`continuous_research` names its own team; the other four recipes all default to
|
||||
`rust_sdlc`. `research_only` has no correct existing template to point at —
|
||||
`papers_research` is arXiv-shaped, `insight_research` is vault-shaped, and
|
||||
`codebase_research` needs a repo — so the fix is an operator decision, not a
|
||||
one-line edit, and is deliberately left open.
|
||||
|
||||
### 8. The check that got it wrong first
|
||||
|
||||
Run 3's first scoring reported `cargo-test-driven-development` and
|
||||
`tdd-red-green-refactor` as **compliance = fail**: files were written and no
|
||||
test ever ran.
|
||||
|
||||
That verdict was wrong, and wrong in the way this whole document exists to
|
||||
prevent. The phase wrote fifteen markdown notes and a helper script. There was
|
||||
no code to test-drive. Reporting it as an agent failure would have been a system
|
||||
defect wearing an agent's name — and it would have buried the real finding,
|
||||
which is finding 7 above.
|
||||
|
||||
The check is now scoped to files with a source extension in the languages the
|
||||
skill itself names. It is recorded here rather than quietly corrected, because
|
||||
a measurement that hides its own false positives cannot be trusted about
|
||||
anyone else's.
|
||||
|
||||
## Honest limits
|
||||
|
||||
- **Two runs, one tier, one workflow.** Nothing here generalises to the microVM
|
||||
or session tiers yet, and this document should not be read as if it does.
|
||||
- **Most skills score `not_applicable`** on both observable axes. That is not a
|
||||
pass. It means we cannot currently tell whether those skills changed anything.
|
||||
`workspace-repo-commit-protocol` now has a Boundary check (writing outside
|
||||
`/mission/repo`); `small-focused-commits` and `tdd-red-green-refactor` remain
|
||||
candidates, and both need the repository diff rather than the turn text.
|
||||
- **No agent-authored skill has been measured.** Self-authoring shipped in the
|
||||
same pass; `source_kind` is carried through the scorer specifically so a
|
||||
rising score on agent-authored skills is visible rather than averaged in.
|
||||
- **Three runs, one tier, one workflow.** Nothing here generalises to the
|
||||
microVM or session tiers.
|
||||
- **The new checks are not yet exercised live.** `research_only` writes no code
|
||||
and makes no commits, so the TDD and commit checks are proven by unit tests
|
||||
and negative controls, not by a mission that could have violated them. A
|
||||
coding run against a real repository is the next measurement.
|
||||
- **Tool calls carry no agent attribution.** `record_vm_tools` writes
|
||||
`agent_id: None` — the container tier's tap is per-container, and all five
|
||||
roles share one container. Every score above is therefore per-**mission**,
|
||||
not per-role, and the World's per-agent view gets nothing from it. Mapping
|
||||
the hook payload's `session_id` back to a turn would fix it.
|
||||
- **Most skills still score `not_applicable`** on both observable axes. That is
|
||||
not a pass. See finding 7 for why the number is what it is.
|
||||
- **No agent-authored skill has been measured.** `source_kind` is carried
|
||||
through the scorer specifically so a rising score on agent-authored skills is
|
||||
visible rather than averaged in.
|
||||
- **Evidence expires.** Mission events are reaped after 7 days unless
|
||||
`retain_events_until` is set. Both runs here are held for 90 days. An empty
|
||||
score means "no evidence", never "no compliance", and the API says so in its
|
||||
payload rather than leaving the caller to infer it.
|
||||
`retain_events_until` is set; `scripts/skill-use-run.sh` holds every run for
|
||||
90 days so it stays re-scorable when the scorer changes again — which is
|
||||
exactly what happened to run 3. An empty score means "no evidence", never "no
|
||||
compliance", and the API says so in its payload.
|
||||
|
||||
## Reproducing
|
||||
|
||||
```
|
||||
scripts/skill-use-run.sh "<title>" "<task>" # run and score
|
||||
scripts/skill-use-run.sh --score <mission-id> # re-score, no new run
|
||||
```
|
||||
|
||||
Local stack only. Production auth is Clerk and a mission cannot be launched
|
||||
from a terminal there — which is also why, as of 2026-08-21, **production has
|
||||
never run a mission at all** (`select count(*) from missions` → 0).
|
||||
|
||||
Reference in New Issue
Block a user