# Where this left off — 2026-08-21 (second pass) Read `CAPABILITY-REVIEW.md` for the system picture, `TOOL-CALL-ARCHITECTURE.md` for how mission tools actually work, and `SKILL-USE-BASELINE.md` for the measurement — which now scores behaviour rather than the agent's own account of it. ## State of the tree Local suite green: **107 test binaries, 792 tests** (`cargo test --workspace`), and the workspace builds with `--all-targets`. Four measurement missions ran on the local stack (`scripts/skill-use-run.sh`); all four are held 90 days and re-scorable with `--score `. Seven commits on `main` this pass, **not pushed** — a push to `main` auto-deploys to gw-04, and the container-tier work already deployed is unexercised there (see below). ## The premise of the last handoff's item 1 was wrong It said the container-tier gate and tap were "proven locally and unproven in prod", and told you to watch for the first production mission. **Production has never run a mission.** ``` gw-04$ select count(*) from missions; -> 0 gw-04$ select count(*) from mission_events; -> 0 ``` Prod is armed correctly — server restarted with `CLAWMATES_RUNTIME_IMAGE=clawmates-runtime:hooks`, image present. There is simply nothing to watch. Prod auth is Clerk, so a mission cannot be launched from a terminal; someone has to click. Generalise the lesson: before debugging why a deployed thing shows no evidence, check whether anything ran. ### A leaked mission container nothing can reap `cm-runtime-mission-019ff5b157ce77028f308ebd3dc92748` has been `Up` since **2026-08-12** on `clawmates-runtime:sync`, with no `missions` row behind it. `mission_runtime`'s terminal sweeper selects `FROM missions WHERE status IN (…)`, and `teardown_container(mission_id)` is only ever called with an id from that query. Nothing enumerates Docker for `cm-runtime-mission-*` containers with no matching row, so a container whose row is gone is invisible to every reaper. Same shape as the earlier agent-container reap drift, different table. It holds disk on the host whose disk exhaustion broke CI runs 503–506. **Not fixed — your call**, and the fix is small: one sweep that lists containers by name prefix and reaps those with no row. ## What shipped this pass ### The tool tap kept the name and discarded the argument The container tier's first measured mission recorded `Bash × 6` and not one of them said what it ran. `vm_tool_tap::parse` read `tool_input` to pull the path out of it and dropped the rest, so every behavioural question about a phase was unanswerable from a record that looked complete. `Observed.input` now keeps it, bounded: file bodies become a byte count, other long strings truncate with a marker. **Host-side only, no image rebuild** — the arguments were always in the tap file. `tool.call` also gained `detail.path` (the World's SSE reads it and had been getting null on every container-tier call) and `file.touch` gained `detail.abs`. ### Skill-Use is scored from actions `skill_use::Evidence` carries `tool.call` rows alongside the narrative, and every check prefers them. `workspace-repo-commit-protocol`'s boundary was a substring search for `/workspace/repo` in prose — an agent that wrote to the wrong root **without narrating it scored a clean pass**. Two verdicts changed for honesty: silence is `NotObservable` rather than `Pass`, and a test that ran after the first write is undecidable rather than a failure. **Trigger is still `NotObservable`, and half of its old reason is now wrong.** "`claude_cli` cannot surface a tool call" is false. What still holds is that we **inline** skill bodies, so there is no retrieval to observe. The blocker moved from the transport to the delivery model, and the door (§3) closes it with no scorer change at all. ### Two more skills contradicted the platform Both found by reading the source of truth before writing a check against it — which is the only reason they were found. - `decompose-int-items` taught `PLAN_COMPLETE: INT-01..05`. Ids are strictly `INT-`, so the range form is rejected and the plan pass records nothing while every item stays open. - `workspace-repo-commit-protocol` claimed the task-card parser advances mission state on the INT id in your commit subject. **Nothing in the platform reads commit messages** — `apply_for_run` reads `run_events`, the turn output. `no_skill_shows_a_marker_the_parser_would_reject` guards the class, running the real parser over every marker in every skill's fenced blocks. ## Next, in order 1. **`research_only` staffs a Rust SDLC crew, and it is your decision.** The recipe declares `requires_repo = false`, one `research` phase, and `default_team_template = "rust_sdlc"` — so a repo-less markdown mission gets planner, coder, tester, reviewer and committer, four of whom have nothing to do, at ~50KB of prompt. This is why most skills score `not_applicable`: the skills are correctly bound to the roles, and the **roles are wrong for the workflow**. No existing template fits — `papers_research` is arXiv-shaped, `insight_research` vault-shaped, `codebase_research` needs a repo — so this is a new template or a recipe change, not a one-liner. `benchmark`, `refactor`, `research_and_code` and `security_hardening` all default to `rust_sdlc` too; check whether each is right. 2. **The TDD check cannot confirm red-first, and that is structural.** Run 4 (`research_and_code`, real repo) edited `src/lib.rs` once — implementation *and* `#[cfg(test)] mod tests` in the same write — then ran `cargo test` five times. In Rust the unit test lives in the file under test, so that ordering is what following the skill precisely looks like from outside. The check detects "wrote source, never ran a test" and nothing more. If you want red-first, it needs the diff (did the test exist before the impl?), not the tool order. 3. **Attribute tool calls to agents.** `record_vm_tools` writes `agent_id: None`, because the container tap is per-container and all roles share one. Every Skill-Use score is therefore per-**mission**, not per-role, and the World's per-agent view gets nothing from the container tier. The hook payload carries `session_id`; mapping it back to a turn is the fix. 4. **Deploy the door** (`TOOL-CALL-ARCHITECTURE.md` §3). Config, not code: `/zeroclaw-data/clawmates-mcp.json` plus a door-shaped provider alias. It is now the single change that makes **Trigger** a real measurement, and the precondition for skills moving from inlined bodies to progressive disclosure — which would also cut the prompt cost in item 1. 5. **Fold the microVM tier onto `container_tool_hooks`.** It has a gate and a tap by a different route (`vm_tool_tap` installs into the guest, `microvm_executor` drains inside the turn). Two mechanisms for one job is how they drift — and the argument-discarding bug above lived in the shared parser precisely because nobody looked at it from the container side. The fleet has been offline for over a week, so this cannot be tested today. 6. **Pull upstream's egress policy** — `0db7d999a feat(plugins): add shared egress policy foundation (#9137)`. We are ~220 commits behind; this is the one item worth taking, and it is defence for a problem we have not solved. ## Open decisions that are yours - **Push.** Seven commits are local. Pushing `main` triggers CI → auto-deploy to gw-04. - **The orphan container** above. - **Self-authoring scope.** Agents apply their own `skill_candidate` items with no human click (`CLAWMATES_SKILL_SELF_AUTHORING=0` restores the gate). `identity_refinement` and `brain_consolidation` still wait for a human, because they change what an agent IS rather than adding a procedure it can consult. ## Deliberately not done - **The mission executor swap.** Blockers are structural: `cm-runtime`'s `files` tool rejects absolute paths by construction, `shell` runs in a per-agent sandbox with no mission mount, `ToolContext` carries no path or VM handle, and approvals key on `(session_id, message_id)`. - **A tap for the direct-session tier.** Dormant — `CLAWMATES_MISSION_EXECUTOR` is unset in production, so it never runs. - **`cm-brain` offline tests** — 6 of 9 need live `clawbrainhub.com`. - **Graph memory / `clawhdf5-agent`** — in the workspace manifest, used by no crate. ## Operational facts that cost time to learn - The Gitea **actions-log API returns 403** for the token in `deploy/compose/.env`. A token with the `actions` scope remains the highest-value thing to obtain. - **gw-04 uses legacy `docker-compose`**, not the v2 plugin. - **Do not build images by hand on gw-04 while CI may run** — same 150G volume, and the frontend image build is what loses. - The server reaches Docker through a **socket proxy** (`DOCKER_HOST`). Use `container_exec::connect()`, never `connect_with_local_defaults()`. - Prod auth is **Clerk**; the bootstrap password in `deploy/compose/.env` works only against the local stack. - Rebuilding the local server image is a **full Rust compile inside Docker** (~8 min); the layer cache does not preserve `target/`. Budget for it before any measurement that needs new server code. - macOS has no `timeout(1)`. ## The recurring shape, now seven times over **A claim in a comment or a doc, believed and never checked.** Every significant finding this pass came from reading the source of truth — the parser, the recipe, the production table — rather than the text describing it. The two new skill contradictions were found *while writing checks against those skills*, which is the cheapest place to catch them and the reason to always read first. The corollary the measurement itself demonstrated: **its own first verdict was wrong**, and scoring a research phase as a TDD failure would have buried the real finding (item 1). A check that reports a system defect as an agent defect is worse than no check.