107 test binaries, 792 tests, zero failures across the workspace — run, not estimated from the cm-api figure. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
9.8 KiB
Where this left off — 2026-08-21 (second pass)
Read CAPABILITY-REVIEW.md for the system picture,
TOOL-CALL-ARCHITECTURE.md for how mission tools actually work, and
SKILL-USE-BASELINE.md for the measurement — which now scores behaviour rather
than the agent's own account of it.
State of the tree
Local suite green: 107 test binaries, 792 tests (cargo test --workspace), and the workspace builds with --all-targets.
Four measurement missions ran on the local stack (scripts/skill-use-run.sh);
all four are held 90 days and re-scorable with --score <id>.
Seven commits on main this pass, not pushed — a push to main auto-deploys
to gw-04, and the container-tier work already deployed is unexercised there
(see below).
The premise of the last handoff's item 1 was wrong
It said the container-tier gate and tap were "proven locally and unproven in prod", and told you to watch for the first production mission.
Production has never run a mission.
gw-04$ select count(*) from missions; -> 0
gw-04$ select count(*) from mission_events; -> 0
Prod is armed correctly — server restarted with
CLAWMATES_RUNTIME_IMAGE=clawmates-runtime:hooks, image present. There is
simply nothing to watch. Prod auth is Clerk, so a mission cannot be launched
from a terminal; someone has to click. Generalise the lesson: before debugging
why a deployed thing shows no evidence, check whether anything ran.
A leaked mission container nothing can reap
cm-runtime-mission-019ff5b157ce77028f308ebd3dc92748 has been Up since
2026-08-12 on clawmates-runtime:sync, with no missions row behind it.
mission_runtime's terminal sweeper selects FROM missions WHERE status IN (…), and teardown_container(mission_id) is only ever called with an id from
that query. Nothing enumerates Docker for cm-runtime-mission-* containers with
no matching row, so a container whose row is gone is invisible to every reaper.
Same shape as the earlier agent-container reap drift, different table.
It holds disk on the host whose disk exhaustion broke CI runs 503–506. Not fixed — your call, and the fix is small: one sweep that lists containers by name prefix and reaps those with no row.
What shipped this pass
The tool tap kept the name and discarded the argument
The container tier's first measured mission recorded Bash × 6 and not one of
them said what it ran. vm_tool_tap::parse read tool_input to pull the path
out of it and dropped the rest, so every behavioural question about a phase was
unanswerable from a record that looked complete.
Observed.input now keeps it, bounded: file bodies become a byte count, other
long strings truncate with a marker. Host-side only, no image rebuild — the
arguments were always in the tap file. tool.call also gained detail.path
(the World's SSE reads it and had been getting null on every container-tier
call) and file.touch gained detail.abs.
Skill-Use is scored from actions
skill_use::Evidence carries tool.call rows alongside the narrative, and
every check prefers them. workspace-repo-commit-protocol's boundary was a
substring search for /workspace/repo in prose — an agent that wrote to the
wrong root without narrating it scored a clean pass. Two verdicts changed
for honesty: silence is NotObservable rather than Pass, and a test that ran
after the first write is undecidable rather than a failure.
Trigger is still NotObservable, and half of its old reason is now wrong.
"claude_cli cannot surface a tool call" is false. What still holds is that we
inline skill bodies, so there is no retrieval to observe. The blocker moved
from the transport to the delivery model, and the door (§3) closes it with no
scorer change at all.
Two more skills contradicted the platform
Both found by reading the source of truth before writing a check against it — which is the only reason they were found.
decompose-int-itemstaughtPLAN_COMPLETE: INT-01..05. Ids are strictlyINT-<digits>, so the range form is rejected and the plan pass records nothing while every item stays open.workspace-repo-commit-protocolclaimed the task-card parser advances mission state on the INT id in your commit subject. Nothing in the platform reads commit messages —apply_for_runreadsrun_events, the turn output.
no_skill_shows_a_marker_the_parser_would_reject guards the class, running the
real parser over every marker in every skill's fenced blocks.
Next, in order
-
research_onlystaffs a Rust SDLC crew, and it is your decision. The recipe declaresrequires_repo = false, oneresearchphase, anddefault_team_template = "rust_sdlc"— so a repo-less markdown mission gets planner, coder, tester, reviewer and committer, four of whom have nothing to do, at ~50KB of prompt. This is why most skills scorenot_applicable: the skills are correctly bound to the roles, and the roles are wrong for the workflow. No existing template fits —papers_researchis arXiv-shaped,insight_researchvault-shaped,codebase_researchneeds a repo — so this is a new template or a recipe change, not a one-liner.benchmark,refactor,research_and_codeandsecurity_hardeningall default torust_sdlctoo; check whether each is right. -
The TDD check cannot confirm red-first, and that is structural. Run 4 (
research_and_code, real repo) editedsrc/lib.rsonce — implementation and#[cfg(test)] mod testsin the same write — then rancargo testfive times. In Rust the unit test lives in the file under test, so that ordering is what following the skill precisely looks like from outside. The check detects "wrote source, never ran a test" and nothing more. If you want red-first, it needs the diff (did the test exist before the impl?), not the tool order. -
Attribute tool calls to agents.
record_vm_toolswritesagent_id: None, because the container tap is per-container and all roles share one. Every Skill-Use score is therefore per-mission, not per-role, and the World's per-agent view gets nothing from the container tier. The hook payload carriessession_id; mapping it back to a turn is the fix. -
Deploy the door (
TOOL-CALL-ARCHITECTURE.md§3). Config, not code:/zeroclaw-data/clawmates-mcp.jsonplus a door-shaped provider alias. It is now the single change that makes Trigger a real measurement, and the precondition for skills moving from inlined bodies to progressive disclosure — which would also cut the prompt cost in item 1. -
Fold the microVM tier onto
container_tool_hooks. It has a gate and a tap by a different route (vm_tool_tapinstalls into the guest,microvm_executordrains inside the turn). Two mechanisms for one job is how they drift — and the argument-discarding bug above lived in the shared parser precisely because nobody looked at it from the container side. The fleet has been offline for over a week, so this cannot be tested today. -
Pull upstream's egress policy —
0db7d999a feat(plugins): add shared egress policy foundation (#9137). We are ~220 commits behind; this is the one item worth taking, and it is defence for a problem we have not solved.
Open decisions that are yours
- Push. Seven commits are local. Pushing
maintriggers CI → auto-deploy to gw-04. - The orphan container above.
- Self-authoring scope. Agents apply their own
skill_candidateitems with no human click (CLAWMATES_SKILL_SELF_AUTHORING=0restores the gate).identity_refinementandbrain_consolidationstill wait for a human, because they change what an agent IS rather than adding a procedure it can consult.
Deliberately not done
- The mission executor swap. Blockers are structural:
cm-runtime'sfilestool rejects absolute paths by construction,shellruns in a per-agent sandbox with no mission mount,ToolContextcarries no path or VM handle, and approvals key on(session_id, message_id). - A tap for the direct-session tier. Dormant —
CLAWMATES_MISSION_EXECUTORis unset in production, so it never runs. cm-brainoffline tests — 6 of 9 need liveclawbrainhub.com.- Graph memory /
clawhdf5-agent— in the workspace manifest, used by no crate.
Operational facts that cost time to learn
- The Gitea actions-log API returns 403 for the token in
deploy/compose/.env. A token with theactionsscope remains the highest-value thing to obtain. - gw-04 uses legacy
docker-compose, not the v2 plugin. - Do not build images by hand on gw-04 while CI may run — same 150G volume, and the frontend image build is what loses.
- The server reaches Docker through a socket proxy (
DOCKER_HOST). Usecontainer_exec::connect(), neverconnect_with_local_defaults(). - Prod auth is Clerk; the bootstrap password in
deploy/compose/.envworks only against the local stack. - Rebuilding the local server image is a full Rust compile inside Docker
(~8 min); the layer cache does not preserve
target/. Budget for it before any measurement that needs new server code. - macOS has no
timeout(1).
The recurring shape, now seven times over
A claim in a comment or a doc, believed and never checked. Every significant finding this pass came from reading the source of truth — the parser, the recipe, the production table — rather than the text describing it. The two new skill contradictions were found while writing checks against those skills, which is the cheapest place to catch them and the reason to always read first.
The corollary the measurement itself demonstrated: its own first verdict was wrong, and scoring a research phase as a TDD failure would have buried the real finding (item 1). A check that reports a system defect as an agent defect is worse than no check.