Files
clawmates/docs/SKILL-USE-BASELINE.md
T
Omar SobhandClaude Opus 5 c209e654d9 fix(skill-use): a research phase writing markdown is not a TDD failure
The first live scoring of run 3 reported `cargo-test-driven-development`
and `tdd-red-green-refactor` as compliance=FAIL: files were written and no
test ever ran.

Wrong, and wrong in the way this module exists to prevent. The phase wrote
fifteen markdown notes and a helper script; there was no code to
test-drive. Reporting it as an agent failure is a system defect wearing an
agent's name — and it would have buried the actual finding, which is that
a repo-less `research_only` mission is staffed with a Rust SDLC crew whose
coder, tester, reviewer and committer have nothing to do.

The check is now scoped to files with a source extension in the languages
the skill itself names. Shell is deliberately excluded: a helper script
written during a research turn is not behaviour-adding code, and the false
failure costs more than the missed one.

Recorded in SKILL-USE-BASELINE.md as finding 8 rather than quietly
corrected. A measurement that hides its own false positives cannot be
trusted about anyone else's.

Also in the doc: the Trigger reason is half false now (the transport can
surface a tool call; we simply still inline), and the architecture doc's
observe/gate table said the container tier was ungated and unobserved,
which shipped work has made wrong.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
2026-08-21 08:37:15 -07:00

242 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Skill-Use baseline
*Whether ClawMates' skills change what agents do. First measured 2026-08-19;
re-measured 2026-08-21 against tool evidence rather than agent prose.*
Scored on the three axes from `Skill-Use` (arXiv, 2026-08-05): **Trigger** (did
the agent reach for the skill), **Compliance** (did it follow the procedure),
**Boundary** (did it avoid what the skill forbids).
Read the method before the numbers. A measurement whose limits are not stated
is worse than none, because it gets quoted without them.
## What changed on 2026-08-21
The 2026-08-19 measurement scored Compliance and Boundary from the
`reasoning` events — the agent's own account of its turn. Since then the
container tier records what agents actually **do**
(`container_tool_hooks`, `PostToolUse`), and `vm_tool_tap` stopped throwing the
tool's arguments away, so `Bash` commands and `Write` paths are on the record.
`skill_use` now reads those. The difference is not cosmetic:
- `workspace-repo-commit-protocol`'s Boundary was a substring search for
`/workspace/repo` in the narrative. **An agent that wrote to the wrong root
without narrating it scored a clean pass.** It now reads the write paths.
- `arxiv-daily`'s Boundary read a URL in prose, which may be the agent
explaining that it did *not* fetch it. It now reads the `curl` that ran.
- `tdd-red-green-refactor`, `cargo-test-driven-development` and
`small-focused-commits` gained their first checks at all.
Two verdicts changed for honesty rather than coverage. **Silence used to score
`Pass`** — a mission with no evidence scored identically to one checked and
found clean; it is now `NotObservable`. And a test that ran *after* the first
write is `NotObservable`, not a failure, because a Rust unit test lives in the
file under test.
### Trigger: the reason changed, and only half of it went away
The 2026-08-19 document said Trigger was unobservable because `claude_cli`
"cannot surface a tool call — there is nothing to retrieve *with*."
**That half is now false.** Mission tool calls are recorded on both tiers; a
retrieval would be as visible as any other call.
The other half still holds and is the one that decides the verdict: **we still
inline**. `pinned_skills_text` puts full skill bodies in the prompt, so the
agent never reaches for anything — it is simply holding one. Trigger is now
*instrumentable* and still not *observable*, and the blocker has moved from the
transport to the delivery model.
Making it real is one change and no scorer work: serve skills through the door
(`TOOL-CALL-ARCHITECTURE.md` §3) so retrieval becomes a tool call.
## Method, and what it cannot see
Scored by `cm_api::skill_use` from what the platform records: `prompt.composed`
(the exact bytes an agent received), `reasoning` (what it said it did), and
`tool.call` (what it did). No re-derivation from the catalogue — the catalogue
changes, and now that agents author their own skills it changes by itself.
**Every tool-backed check is one-sided.** It reports a violation it can see and
never infers compliance from silence: the recorded stream is capped per phase
(`PER_PHASE_CAP = 400`), so an absent call is not proof of an absent action.
Only skills whose procedure has a machine-checkable consequence are scored.
Everything else returns `not_applicable` rather than a guess — a heuristic that
scores prose by keyword overlap produces a number that looks like a measurement
and is not one.
Compliance for `int-xx-marker-protocol` is checked by running the **real**
`task_card_parser`, not a copy of its rules; a second implementation would
drift, and then the score would pass while the mission loop still stalled.
## The runs
Three missions on the container/ZeroClaw tier, local stack, `research_only`.
Run 3 uses the **same task text as run 2**, so the only variable is the scorer.
| | run 1 | run 2 | run 3 |
|---|---|---|---|
| date | 08-19 | 08-19 | 08-21 |
| distinct skills delivered | 3 | 9 | 9 |
| total deliveries (per role prompt) | 3 | 14 | 14 |
| phantom "skills" scored | **2** | 0 | 0 |
| tool calls recorded | 0 | 0 | **49** |
| axes scored from actions | 0 | 0 | **3 skills** |
Run 3, per skill (all `source_kind=builtin`; no agent-authored skill has been
delivered yet):
| skill | deliveries | compliance | boundary |
|---|---|---|---|
| `int-xx-marker-protocol` | 1 | **pass** | n/a |
| `workspace-repo-commit-protocol` | 2 | n/a | **pass** *(from 12 write paths)* |
| `small-focused-commits` | 4 | n/a | n/a *(no commit ran)* |
| `cargo-test-driven-development` | 2 | n/a | n/a |
| `tdd-red-green-refactor` | 1 | n/a | n/a |
| `decompose-int-items` | 1 | n/a | n/a |
| `write-rust-current-edition` | 1 | n/a | n/a |
| `code-review-checklist` | 1 | n/a | n/a |
| `criterion-benchmarking` | 1 | n/a | n/a |
**n = 3 runs. No spread is reported because three cannot establish one.** This
is a baseline in the sense of "the first honest number", not in the sense of
`metrics-baseline-comparison.md`, which requires enough runs to see the noise
floor before any change is judged against it.
`workspace-repo-commit-protocol`'s pass is the one score that materially
improved: it now rests on twelve recorded `Write`/`Edit` paths, every one under
`/mission/repo`, instead of on the absence of a string in prose.
## What the measurement found
### 1–4: the 2026-08-19 findings
Four defects, none of which any test or log would have surfaced: the prompt
format made its own record unparseable (`## <name>` against markdown bodies);
a prompt was recorded that was never sent; a pinned skill taught
`/workspace/repo`, a path the platform does not mount; and
`int-xx-marker-protocol` documented a `PLAN_COMPLETE` marker the parser had
never implemented. All four are fixed, with guards in
`skills_loader::contradiction_tests` and `topology_exec`. The detail is in this
file's git history.
### 5. The tool tap recorded the name and discarded the argument
The first container-tier mission with telemetry recorded `Bash × 6` and not one
of them said what it ran. `vm_tool_tap::parse` read `tool_input` to pull the
path out of it and dropped the rest.
Every behavioural question was therefore unanswerable from a record that looked
complete — which is the recurring shape, not a new one. Fixed host-side: the
arguments were always in the tap file.
### 6. Two more skills contradicted the platform
Same class as finding 3, and both found by reading the source of truth before
writing a check against it.
- **`decompose-int-items` taught `PLAN_COMPLETE: INT-01..05`.** An id is
strictly `INT-` plus digits, so the range form is rejected outright: the plan
pass records nothing while every item stays open. A live planner emitted
exactly that line.
- **`workspace-repo-commit-protocol` claimed the task-card parser advances
mission state on the INT id in your commit subject.** Nothing in the platform
reads commit messages. `task_card_parser::apply_for_run` reads `run_events` —
the agent's turn output. An agent that believed this would commit with the id,
never emit `COMPLETED: INT-NN`, and leave the mission open on an item it had
already finished.
`no_skill_shows_a_marker_the_parser_would_reject` now runs the real parser over
every marker in every skill's fenced blocks, negative-controlled against the
range form.
### 7. A repo-less research mission is staffed with a Rust SDLC crew
This is the finding of run 3, and it explains most of the `not_applicable`
column above.
`templates/workflows/research_only.toml` declares `requires_repo = false` and a
single `research` phase — and `default_team_template = "rust_sdlc"`. So the
mission was staffed with **planner, coder, tester, reviewer, committer**, and
each received the skills its role is bound to:
```
coder :: write-rust-current-edition, cargo-test-driven-development,
workspace-repo-commit-protocol, small-focused-commits,
int-xx-marker-protocol
tester :: cargo-test-driven-development, criterion-benchmarking,
tdd-red-green-refactor
committer :: workspace-repo-commit-protocol, small-focused-commits
reviewer :: code-review-checklist, small-focused-commits
planner :: decompose-int-items, small-focused-commits
```
There is no repository, nothing to test, nothing to review and nothing to
commit. Four of the five roles have no work, and 50KB of prompt (~12.6k tokens)
is spent staffing them.
**The skills are correctly bound to the roles. The roles are wrong for the
workflow.** That distinction matters: a reader who saw only "7 of 9 skills
scored not_applicable" would conclude the skills are useless, when what the
number actually measures is a staffing default.
`continuous_research` names its own team; the other four recipes all default to
`rust_sdlc`. `research_only` has no correct existing template to point at —
`papers_research` is arXiv-shaped, `insight_research` is vault-shaped, and
`codebase_research` needs a repo — so the fix is an operator decision, not a
one-line edit, and is deliberately left open.
### 8. The check that got it wrong first
Run 3's first scoring reported `cargo-test-driven-development` and
`tdd-red-green-refactor` as **compliance = fail**: files were written and no
test ever ran.
That verdict was wrong, and wrong in the way this whole document exists to
prevent. The phase wrote fifteen markdown notes and a helper script. There was
no code to test-drive. Reporting it as an agent failure would have been a system
defect wearing an agent's name — and it would have buried the real finding,
which is finding 7 above.
The check is now scoped to files with a source extension in the languages the
skill itself names. It is recorded here rather than quietly corrected, because
a measurement that hides its own false positives cannot be trusted about
anyone else's.
## Honest limits
- **Three runs, one tier, one workflow.** Nothing here generalises to the
microVM or session tiers.
- **The new checks are not yet exercised live.** `research_only` writes no code
and makes no commits, so the TDD and commit checks are proven by unit tests
and negative controls, not by a mission that could have violated them. A
coding run against a real repository is the next measurement.
- **Tool calls carry no agent attribution.** `record_vm_tools` writes
`agent_id: None` — the container tier's tap is per-container, and all five
roles share one container. Every score above is therefore per-**mission**,
not per-role, and the World's per-agent view gets nothing from it. Mapping
the hook payload's `session_id` back to a turn would fix it.
- **Most skills still score `not_applicable`** on both observable axes. That is
not a pass. See finding 7 for why the number is what it is.
- **No agent-authored skill has been measured.** `source_kind` is carried
through the scorer specifically so a rising score on agent-authored skills is
visible rather than averaged in.
- **Evidence expires.** Mission events are reaped after 7 days unless
`retain_events_until` is set; `scripts/skill-use-run.sh` holds every run for
90 days so it stays re-scorable when the scorer changes again — which is
exactly what happened to run 3. An empty score means "no evidence", never "no
compliance", and the API says so in its payload.
## Reproducing
```
scripts/skill-use-run.sh "<title>" "<task>" # run and score
scripts/skill-use-run.sh --score <mission-id> # re-score, no new run
```
Local stack only. Production auth is Clerk and a mission cannot be launched
from a terminal there — which is also why, as of 2026-08-21, **production has
never run a mission at all** (`select count(*) from missions` → 0).