# Skill-Use baseline *Whether ClawMates' skills change what agents do. First measured 2026-08-19; re-measured 2026-08-21 against tool evidence rather than agent prose.* Scored on the three axes from `Skill-Use` (arXiv, 2026-08-05): **Trigger** (did the agent reach for the skill), **Compliance** (did it follow the procedure), **Boundary** (did it avoid what the skill forbids). Read the method before the numbers. A measurement whose limits are not stated is worse than none, because it gets quoted without them. ## What changed on 2026-08-21 The 2026-08-19 measurement scored Compliance and Boundary from the `reasoning` events — the agent's own account of its turn. Since then the container tier records what agents actually **do** (`container_tool_hooks`, `PostToolUse`), and `vm_tool_tap` stopped throwing the tool's arguments away, so `Bash` commands and `Write` paths are on the record. `skill_use` now reads those. The difference is not cosmetic: - `workspace-repo-commit-protocol`'s Boundary was a substring search for `/workspace/repo` in the narrative. **An agent that wrote to the wrong root without narrating it scored a clean pass.** It now reads the write paths. - `arxiv-daily`'s Boundary read a URL in prose, which may be the agent explaining that it did *not* fetch it. It now reads the `curl` that ran. - `tdd-red-green-refactor`, `cargo-test-driven-development` and `small-focused-commits` gained their first checks at all. Two verdicts changed for honesty rather than coverage. **Silence used to score `Pass`** — a mission with no evidence scored identically to one checked and found clean; it is now `NotObservable`. And a test that ran *after* the first write is `NotObservable`, not a failure, because a Rust unit test lives in the file under test. ### Trigger: the reason changed, and only half of it went away The 2026-08-19 document said Trigger was unobservable because `claude_cli` "cannot surface a tool call — there is nothing to retrieve *with*." **That half is now false.** Mission tool calls are recorded on both tiers; a retrieval would be as visible as any other call. The other half still holds and is the one that decides the verdict: **we still inline**. `pinned_skills_text` puts full skill bodies in the prompt, so the agent never reaches for anything — it is simply holding one. Trigger is now *instrumentable* and still not *observable*, and the blocker has moved from the transport to the delivery model. Making it real is one change and no scorer work: serve skills through the door (`TOOL-CALL-ARCHITECTURE.md` §3) so retrieval becomes a tool call. ## Method, and what it cannot see Scored by `cm_api::skill_use` from what the platform records: `prompt.composed` (the exact bytes an agent received), `reasoning` (what it said it did), and `tool.call` (what it did). No re-derivation from the catalogue — the catalogue changes, and now that agents author their own skills it changes by itself. **Every tool-backed check is one-sided.** It reports a violation it can see and never infers compliance from silence: the recorded stream is capped per phase (`PER_PHASE_CAP = 400`), so an absent call is not proof of an absent action. Only skills whose procedure has a machine-checkable consequence are scored. Everything else returns `not_applicable` rather than a guess — a heuristic that scores prose by keyword overlap produces a number that looks like a measurement and is not one. Compliance for `int-xx-marker-protocol` is checked by running the **real** `task_card_parser`, not a copy of its rules; a second implementation would drift, and then the score would pass while the mission loop still stalled. ## The runs Three missions on the container/ZeroClaw tier, local stack, `research_only`. Run 3 uses the **same task text as run 2**, so the only variable is the scorer. | | run 1 | run 2 | run 3 | |---|---|---|---| | date | 08-19 | 08-19 | 08-21 | | distinct skills delivered | 3 | 9 | 9 | | total deliveries (per role prompt) | 3 | 14 | 14 | | phantom "skills" scored | **2** | 0 | 0 | | tool calls recorded | 0 | 0 | **49** | | axes scored from actions | 0 | 0 | **3 skills** | Run 3, per skill (all `source_kind=builtin`; no agent-authored skill has been delivered yet): | skill | deliveries | compliance | boundary | |---|---|---|---| | `int-xx-marker-protocol` | 1 | **pass** | n/a | | `workspace-repo-commit-protocol` | 2 | n/a | **pass** *(from 12 write paths)* | | `small-focused-commits` | 4 | n/a | n/a *(no commit ran)* | | `cargo-test-driven-development` | 2 | n/a | n/a | | `tdd-red-green-refactor` | 1 | n/a | n/a | | `decompose-int-items` | 1 | n/a | n/a | | `write-rust-current-edition` | 1 | n/a | n/a | | `code-review-checklist` | 1 | n/a | n/a | | `criterion-benchmarking` | 1 | n/a | n/a | **n = 3 runs. No spread is reported because three cannot establish one.** This is a baseline in the sense of "the first honest number", not in the sense of `metrics-baseline-comparison.md`, which requires enough runs to see the noise floor before any change is judged against it. `workspace-repo-commit-protocol`'s pass is the one score that materially improved: it now rests on twelve recorded `Write`/`Edit` paths, every one under `/mission/repo`, instead of on the absence of a string in prose. ## What the measurement found ### 1–4: the 2026-08-19 findings Four defects, none of which any test or log would have surfaced: the prompt format made its own record unparseable (`## ` against markdown bodies); a prompt was recorded that was never sent; a pinned skill taught `/workspace/repo`, a path the platform does not mount; and `int-xx-marker-protocol` documented a `PLAN_COMPLETE` marker the parser had never implemented. All four are fixed, with guards in `skills_loader::contradiction_tests` and `topology_exec`. The detail is in this file's git history. ### 5. The tool tap recorded the name and discarded the argument The first container-tier mission with telemetry recorded `Bash × 6` and not one of them said what it ran. `vm_tool_tap::parse` read `tool_input` to pull the path out of it and dropped the rest. Every behavioural question was therefore unanswerable from a record that looked complete — which is the recurring shape, not a new one. Fixed host-side: the arguments were always in the tap file. ### 6. Two more skills contradicted the platform Same class as finding 3, and both found by reading the source of truth before writing a check against it. - **`decompose-int-items` taught `PLAN_COMPLETE: INT-01..05`.** An id is strictly `INT-` plus digits, so the range form is rejected outright: the plan pass records nothing while every item stays open. A live planner emitted exactly that line. - **`workspace-repo-commit-protocol` claimed the task-card parser advances mission state on the INT id in your commit subject.** Nothing in the platform reads commit messages. `task_card_parser::apply_for_run` reads `run_events` — the agent's turn output. An agent that believed this would commit with the id, never emit `COMPLETED: INT-NN`, and leave the mission open on an item it had already finished. `no_skill_shows_a_marker_the_parser_would_reject` now runs the real parser over every marker in every skill's fenced blocks, negative-controlled against the range form. ### 7. A repo-less research mission is staffed with a Rust SDLC crew This is the finding of run 3, and it explains most of the `not_applicable` column above. `templates/workflows/research_only.toml` declares `requires_repo = false` and a single `research` phase — and `default_team_template = "rust_sdlc"`. So the mission was staffed with **planner, coder, tester, reviewer, committer**, and each received the skills its role is bound to: ``` coder :: write-rust-current-edition, cargo-test-driven-development, workspace-repo-commit-protocol, small-focused-commits, int-xx-marker-protocol tester :: cargo-test-driven-development, criterion-benchmarking, tdd-red-green-refactor committer :: workspace-repo-commit-protocol, small-focused-commits reviewer :: code-review-checklist, small-focused-commits planner :: decompose-int-items, small-focused-commits ``` There is no repository, nothing to test, nothing to review and nothing to commit. Four of the five roles have no work, and 50KB of prompt (~12.6k tokens) is spent staffing them. **The skills are correctly bound to the roles. The roles are wrong for the workflow.** That distinction matters: a reader who saw only "7 of 9 skills scored not_applicable" would conclude the skills are useless, when what the number actually measures is a staffing default. `continuous_research` names its own team; the other four recipes all default to `rust_sdlc`. `research_only` has no correct existing template to point at — `papers_research` is arXiv-shaped, `insight_research` is vault-shaped, and `codebase_research` needs a repo — so the fix is an operator decision, not a one-line edit, and is deliberately left open. ### 8. The check that got it wrong first Run 3's first scoring reported `cargo-test-driven-development` and `tdd-red-green-refactor` as **compliance = fail**: files were written and no test ever ran. That verdict was wrong, and wrong in the way this whole document exists to prevent. The phase wrote fifteen markdown notes and a helper script. There was no code to test-drive. Reporting it as an agent failure would have been a system defect wearing an agent's name — and it would have buried the real finding, which is finding 7 above. The check is now scoped to files with a source extension in the languages the skill itself names. It is recorded here rather than quietly corrected, because a measurement that hides its own false positives cannot be trusted about anyone else's. ## Honest limits - **Three runs, one tier, one workflow.** Nothing here generalises to the microVM or session tiers. - **The new checks are not yet exercised live.** `research_only` writes no code and makes no commits, so the TDD and commit checks are proven by unit tests and negative controls, not by a mission that could have violated them. A coding run against a real repository is the next measurement. - **Tool calls carry no agent attribution.** `record_vm_tools` writes `agent_id: None` — the container tier's tap is per-container, and all five roles share one container. Every score above is therefore per-**mission**, not per-role, and the World's per-agent view gets nothing from it. Mapping the hook payload's `session_id` back to a turn would fix it. - **Most skills still score `not_applicable`** on both observable axes. That is not a pass. See finding 7 for why the number is what it is. - **No agent-authored skill has been measured.** `source_kind` is carried through the scorer specifically so a rising score on agent-authored skills is visible rather than averaged in. - **Evidence expires.** Mission events are reaped after 7 days unless `retain_events_until` is set; `scripts/skill-use-run.sh` holds every run for 90 days so it stays re-scorable when the scorer changes again — which is exactly what happened to run 3. An empty score means "no evidence", never "no compliance", and the API says so in its payload. ## Reproducing ``` scripts/skill-use-run.sh "" "<task>" # run and score scripts/skill-use-run.sh --score <mission-id> # re-score, no new run ``` Local stack only. Production auth is Clerk and a mission cannot be launched from a terminal there — which is also why, as of 2026-08-21, **production has never run a mission at all** (`select count(*) from missions` → 0).