7bcf7865f07f7c95e4e1ce3c6bcfe56b918571c7
887
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7bcf7865f0 |
docs: prod ran a mission, and the whole chain held
`ClawHDF5` and `JEPA Research` were launched from the UI. The first completed,
and everything shipped over the previous two passes engaged correctly on its
first execution anywhere outside the local stack:
skills door installed — api_origin() derived the host from the server's
own container id, which had never run where it was not tested
staffing Topic Research, 3 roles / 4 deliveries, not rust_sdlc's 5 / 14
drain 92 tool calls
attribution 92 of 92, across a phase with TWO passes and six turns — the
case attribute_sessions had never met, and it attributes
nothing at all unless the counts match exactly
boundary all 8 Write/Edit paths under /mission/repo
arm inline, 0 retrievals; prod leaves the env unset
judge pass 0 met=false "zero URLs — grep -c http returns 0"
pass 1 met=true "57 http references"
The judge line is the one worth rereading: the loop converged on the exact
mechanically-checked defect it named, and pass 0 would otherwise have shipped
a report whose every claim was unattributed while reporting `completed`.
Two traps recorded rather than smoothed over:
- The drain selects phases `IN ('completed','failed')`, so a phase on its
second pass shows zero tool calls and reads as broken while being correct.
- I reused a diagnostic query with no `WHERE mission_id`. That was fine while
prod held one mission and silently wrong the moment a second launched — it
compared one mission's tap against two missions' events. The production
drain query is correctly scoped; the diagnostic was not.
Unexplained: one agent called the `Agent` tool 4 times. Mission agents are
spawning subagents and nothing in our design accounts for it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
|
||
|
|
accae7fa94 |
docs: the delivery A/B, and the first retrieval nobody asked for
Runs 9 and 10: identical task text, one server process, and a task that never
mentions skills, MCP or retrieval. Run 8 demonstrated the instrument, but its
retrieval was instructed by the task — it showed the pipe worked, not that an
agent would judge relevance.
Under `index`, two of four skills were fetched, and attribution is the part
that matters:
Solveig (lead_researcher) -> web-search-triage
Olamide (report_writer) -> scientific-writing-conventions
Each agent reached for the skill bound to its OWN role and neither reached for
another's. An agent that fetched all four would have shown only that it could.
The regression the A/B existed to catch did not appear: 34% fewer tokens, 59
tool calls against 89, both arms passed the independent judge, and the
deliverables came out slightly larger rather than thinner.
Two readings the data does not support, recorded because the first draft of
this section made one of them:
- Every `tool.call` in a phase carries the DRAIN timestamp, not the call time.
All 59 rows of run 10 read `12:48:12`. Ordering by that column said the
report writer had fetched both skills; `agent_id` says otherwise.
- The prompt saving is 15-43%, not an order of magnitude. Skill bodies are a
minority of a turn prompt. Progressive disclosure is worth doing for Trigger,
not for context economy.
`workspace-repo-commit-protocol` scores Trigger=FAIL beside boundary=pass: it
behaved correctly without reading the rule. That verdict is left standing and
argued with in the text rather than tuned away.
n=1 per arm. A signal, not a rate.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
|
||
|
|
22eeaa6f15 |
feat(auth): the door that can delegate no longer needs a person's session
`/mcp` — `email_send`, `slack_post`, `delegate` — authenticated with `authenticate`, which accepts only `full`. Nothing hands it a token today, so this cost nothing yet; the moment something did, the only credential that worked would have been an owner's session, held by an agent runtime. `SCOPE_AGENT_DOOR` is that credential's narrow form. `full` still works, so the UI and every human caller are unaffected, and the route now names what it accepts rather than accepting everything by default. The test that matters is not that each scope opens its own route: it is that holding one grants nothing the other has. Both tokens live where an agent can read them. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> |
||
|
|
f52cff3e04 |
feat(skill-use): progressive disclosure, as an arm and not a switch
Trigger — did the agent reach for the skill when it applied? — cannot be measured while every body is inlined into the prompt. Nothing was reached for. `skill_use` has been reporting `NotObservable` for that reason, and it was right to. The skills door made retrieval possible; this makes it a delivery arm. `index` sends each pinned skill's name, description, `when_to_use` and the uri that returns its body, and the agent fetches what it judges relevant. `inline` is unchanged and stays the default. An A/B rather than a switch, because `index` can only cost Compliance: under `inline` the procedure sits in front of the model whether or not it noticed it applied. Trading a measured axis for an unmeasured regression in another is not an improvement, so both arms stay runnable and the arm is recorded on the mission row. Three things the mechanism refuses to do: - `index` without a door falls back to `inline`. An index names bodies and says how to fetch them; with no `clawmates_skills` server reachable that is a list of dead ends, and it fails as an agent ignoring its skills rather than as a missing config. `install_skills_door` now returns whether it installed, because the caller needs the answer and not just the log line. - The scorer reads the arm off the recorded PROMPT, not off the mission row. The row says what the mission is configured to do now; the score is being computed against a turn that ran then. - Under `index`, a skill that was offered and never read is a Fail, not the inline arm's `NotObservable` — but only where the skill had a checkable consequence in that phase. Reusing the inline text would have said "this skill was inlined into the prompt" about a skill whose body was never sent, and scoring a real miss as a structural blind spot is the failure this measurement already made once. The arm is per mission (`config.skill_delivery`), not only per deployment. Both arms run against one server process; restarting between them would put a confound in the comparison that the numbers would not show. 829 tests, 108 binaries, green. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> |
||
|
|
b58f0347e6 |
fix(testkit): stop leaking a database per test
`test_pool` creates a database per test and nothing ever dropped it.
Invisible on the testcontainer path — the container dies with the process
and takes them with it. But `CM_TEST_DATABASE_URL` points at a SHARED
server that outlives the run, and that is the path CI uses and the path
`.cargo/config.toml` sets for local development. So on both, every
database ever created is still there, growing with every `cargo test`.
Measured before writing the fix: **3,546 databases, 38 GB** on one
developer machine. After: 391 and 4.3 GB — the remainder being today's,
still inside the window. The docker volume went 42.3 GB to 5.7 GB.
Age comes from the NAME, not the catalogue. Postgres records no creation
time for a database, but the names are `test_<uuid-v7>` and UUIDv7 puts
its millisecond timestamp in the first 48 bits — the same property
`mission_runtime::container_name` already relies on.
Three things the tests pin down:
- a database created just now must read as NEW, or the reaper deletes
one a parallel test binary is still using;
- only names we minted are reapable — `test_scratch` and `clawmates`
survive;
- the window outlasts any test run.
`WITH (FORCE)` because a single leftover session pins a database and the
drop otherwise silently does nothing. Best-effort throughout: a test must
never fail because housekeeping could not run.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
|
||
|
|
72eda8b3d2 |
docs: handoff reflects the pushed state
22 commits pushed, CI green, deployed. Items 1-3 of the previous list are done: staffing, attribution, and the door. Trigger is measured and red-first turned out observable from run outputs rather than from the diff — the previous list was wrong about that, and the skill says why. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
7525be3791 |
test(orphans): the destructive sweep test is opt-in
CI mounts /var/run/docker.sock into the test container and the runner is gw04 — the host that runs production missions. So `cargo test --workspace` there has full access to the production docker daemon, and this test REMOVES containers. `adopt_existing` protects everything already present, but it cannot protect a mission container created in the seconds between that call and the sweep. On a laptop that race is nothing; on gw04 it is somebody's mission. So the destructive case now requires `CM_TEST_ORPHAN_SWEEP=1` and CI simply does not run it. The read-only probes still run everywhere — they create fixtures and inspect them, and never sweep. This is the second time this test's blast radius has bitten: it reaped two real local mission containers on its first run, and this would have been the same mistake with production's daemon. The sweep is not the problem — a sweep is global by nature — the harness around it is. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
5220f3bfea |
feat(skill-use): red-first is observable from what the RUNS reported
The open item said this needed the repository diff rather than tool
order. That was wrong, and the skill says why: "Commit the RED-to-GREEN
pair as one commit." The failing test and its fix land together by
instruction, so the diff and the commit history are as blind as the tool
ordering already was — in Rust one `Edit` adds the implementation and its
`#[cfg(test)] mod tests` in the same call.
The only remaining witness is what each test run itself printed, and the
tap was throwing it away. Claude Code's PostToolUse payload carries
`tool_response` — verified against the real binary, keys
stdout/stderr/interrupted, plus `duration_ms` and `tool_use_id`.
So `Observed.response` now keeps it, for COMMANDS only: a `Read`'s
response is the file it just read and a `Write`'s restates its own
argument — both already knowable, both large, and storing them would
double the biggest write path in the system for nothing.
`bounded_response` keeps the **end** of the output, which is the opposite
of `bounded_input` and deliberately so. An argument's meaning is its verb,
at the start. A command's meaning is its verdict, at the end: `cargo test`
prints hundreds of lines and then `test result: ok` or `FAILED`. A
head-biased truncation would keep the noise and discard the only thing
being stored for — negative-controlled with a 400-line fixture.
`red_before_green` now falls through to the run outcomes:
failing run, then a passing one → Pass, red then green observed
every run failed → Fail, the loop ends on green
every run passed → NotObservable, and the reason says
why: a test that never failed is
equally what a correct implementation
written first looks like
no outputs recorded → NotObservable (pre-capture missions)
Read from the runner's verdict line, not an exit code — the payload
carries none.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
|
||
|
|
b47ae7fa6b |
feat(skill-use): Trigger is observable — score it
The door made retrieval possible; this makes it *measured*. A skill that
arrives by retrieval leaves a recorded tool call, and until now the scorer
ignored it entirely — so the one axis the whole door was built for stayed
`NotObservable` even on a mission where three agents demonstrably reached
through it.
Taken from what run 8 actually recorded, not from the shape I imagined:
ReadMcpResourceTool {"uri":"skill:global/workspace-repo-commit-protocol",
"server":"clawmates_skills"}
`retrieved_skills` reads those URIs through `mcp_skills::parse_uri` — the
function that WROTE them — rather than a second matcher, because two
implementations of one format drift and the drift shows up as a skill
silently scoring nothing.
Trigger is now `Pass` for a skill the agent reached for, and
`NotObservable` for one that was inlined — with a reason that names the
fix rather than the transport: being handed a skill is not failing to
reach for one.
`score` also had to stop reading only the prompt. A skill retrieved and
never inlined is invisible to `skills_in_prompt`, and under progressive
disclosure that is EVERY skill — so the scorer would have reported zero
for the delivery model this axis exists to measure.
Listing the catalogue is browsing; reading a body is the reach. Only
`ReadMcpResourceTool` counts.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
|
||
|
|
4b160c5a1b |
test(orphans): prove the sweep against real containers, both directions
`sweep_orphans` force-removes containers and had never run against a daemon — only its pure decision logic was covered. The two Docker-touching seams are exactly the ones worth exercising for real: what it can see, and whether a checkout holds work no remote has. Three fixtures, three outcomes, one sweep: - unpushed commits, no remote ref → SURVIVES - every commit on a remote ref → reaped - inside the grace window → survives anyway Negative-controlled: making `unpushed_commits` return `None` for a dirty checkout fails with "the probe said a checkout with an unpushed commit holds nothing — this is the exact answer that destroys work". Two real hazards the test surfaced, neither of them in the sweep: 1. **The tests raced each other.** The sweep is global — it reaps every orphaned mission container on the daemon, including fixtures another test in this file just started. A `FIXTURES` mutex serialises them. Found the honest way: the reap test deleted the listing test's fixture and the listing test reported a container it could not see. 2. **The test destroyed real local state.** The sweep asks the DATABASE whether a container is known, and `test_pool()` knows nothing — so on a developer machine it classified the live stack's mission containers as orphans and reaped two of them on the first run. `adopt_existing` now gives every pre-existing mission container a row before sweeping, which makes the test safe AND covers the one case the other assertions missed: a container the platform still knows about is never touched. Negative-controlled both ways with a bystander container: without adoption REAPED, with adoption SURVIVED. Skips cleanly with no Docker, so a runner without one reports "not run" rather than failing — the placeholder-as-result shape `scripts/verify-mission-delivery.sh` was written to avoid. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
42e014976d |
docs: confirm the door from inside a mission, and correct a count I took from a transcript
Run 8, all three agents, per-agent attributed: Pedro / Ebele / Ahmad — ListMcpResourcesTool, ReadMcpResourceTool The agent's own report: 53 resources from `clawmates_skills`, and `skill:global/workspace-repo-commit-protocol` read back as `# Mission repo + commit protocol`. So the wiring works end to end, not just the mechanism. And a correction to the commit before this one. It recorded "58 MCP resources" as a measurement. That number was the model's paraphrase in a probe transcript, not an observation. `resources/list` returns 53 and `select count(*) from skills` is 53. Noted in the doc rather than quietly changed, because it is the same error this project keeps making — a model's self-report treated as evidence — and I made it in the very document arguing for measuring things. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
02d5f5a8c9 |
docs: the door is deployed, and what it does not buy
Proven against the real binary in the runtime container — connect, list (58 resources) and read (`# Mission repo + commit protocol`, the correct first heading). That probe is a two-minute loop; I reached for the ten-minute rebuild-and-run-a-mission one first, and it would have found the container-name bug sooner. No `--allowedTools` change was needed. Recorded because the guess would have been wrong in an expensive way: with no config read on the daemon, "adding" the MCP tools meant overwriting the seed's `tools` list and stripping Write and Bash from every mission agent — to solve a problem that does not exist. The §3 claim that this was "config, not code" is corrected in place: it needed a credential narrow enough to leave in a container an untrusted agent reads, and the measured proof that the credential IS narrow (same token: 58 skills from /mcp/skills, 401 from /api/missions). And what it does not buy, stated plainly: Trigger is still unmeasured, because delivery still inlines. The door makes retrieval possible; making Trigger real means switching to progressive disclosure, which could regress Compliance and so wants an A/B rather than a flip. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
3aeee070b8 |
fix(missions): the door read a field that is not set yet
`install_skills_door` took the container name from `mission.runtime_container_name`, and `on_launch` loads the mission at the top — before `ensure_container` runs and binds that field. So it was always `None`, and the early return had no log, so the door simply never installed and said nothing about it. Verified against a live mission: no log line, no file in the container. That is the same shape as the three hook bugs before it, which is a poor excuse for repeating it. The name is derived from the mission id (`container_name`) instead, guarded on `mission_gateway` being Some — which is exactly the signal that `ensure_container` ran and that this mission has its own container rather than the shared runtime. Every remaining early return now logs. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
73f5d71c55 |
feat(missions): install the skills door, with a credential it is safe to leave
The capability has been built and undeployed since `88eef99d4`: `claude_cli` accepts `mcp_config` and passes `--mcp-config --strict-mcp-config`, so Claude Code's own MCP client can reach our skills server. What was missing was the config document and, underneath it, a credential that could be left in a container an untrusted agent reads. Now both halves happen together — the document goes in, and the daemon is told to pass it — because doing one without the other leaves a door installed and unreachable, which looks exactly like a door nobody walked through. That is the same shape as the hooks that shipped installed and inert three bugs running. The API origin defaults to our own `HOSTNAME` rather than a container name. Mission containers share `clawmates_core` with the server, and the server's name differs between deployments (`clawmates-server-1` locally, `clawmates_server_1` on gw-04); docker's embedded DNS resolves a container id on a user-defined network, so this is self-configuring. Measured from a sibling container: both the id and the name return 200. `--allowedTools` is deliberately NOT touched. The provider passes it only when `tools` is set and the seed already sets it — without it `claude -p` stops mid-turn asking for write permission. Whether MCP tools also need naming there is undocumented in anything we control, and the daemon exposes no config read to merge into the list safely; overwriting it would take `Write` and `Bash` from every mission agent, and that failure would look like agents that stopped working rather than a config that was replaced. So the question gets answered by running a mission with the door installed. Guessing is how the last three defects in this file got in. Every failure degrades to "no door", never to a failed launch. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
2668191e30 |
feat(auth): a credential narrow enough to hand to an agent
`docs/TOOL-CALL-ARCHITECTURE.md` §3 calls deploying the MCP door "config, not code". It is not, and the reason is authentication. `/mcp/skills` authenticates with `AuthService::authenticate`, which returns a full `AuthedUser` carrying the user's role. There is no narrower credential in the system. So pointing a mission container at the door means writing a bearer token into a file inside that container — and mission agents run arbitrary `Bash` with egress and no read gate, which is this platform's own documented security posture. An owner-scoped token there turns "the agent runs commands in a sandbox" into "the agent drives the whole ClawMates API as the owner". Checked before building this rather than assumed: no such credential is in a mission container today. The runtime's config.toml has no `[mcp.servers]` block and no bearer, so the door would have been a NEW exposure, not an existing one. So: `auth_sessions.scope`, defaulting to `full`. `authenticate` now delegates to `authenticate_scoped(token, SCOPE_FULL)`, which means **every existing caller rejects a narrow token** and a route must opt in by naming the scope it accepts. `/mcp/skills` is the only opt-in. Fail closed on purpose. The likely mistake here is adding a scope and forgetting to wire its check; this way that mistake grants nothing rather than granting everything. `mint_scoped` refuses to mint a `full` token — a caller reaching for it wants a narrow credential, and handing back a full one because an argument was wrong is exactly the failure the column exists to prevent, and it would be invisible because the token would work. The test that matters is not that the door accepts the token, it is that nothing else does. Negative-controlled: removing the scope comparison fails `a_scoped_token_is_refused_by_every_unscoped_caller`. `.sqlx` regenerated — `authenticate` is a compile-checked query and CI builds with SQLX_OFFLINE=true. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
8591585e60 |
feat(missions): attribute a phase's tool calls to the agent that made them
`record_vm_tools` wrote `agent_id: None` for every call. The container tap is per-CONTAINER and every role in a phase shares one, so a phase arrived as one undifferentiated stream: every Skill-Use score was per-mission rather than per-role, and the World's per-agent view got nothing from this tier. One `claude -p` invocation is one turn is one agent, and Claude Code stamps each invocation with a `session_id` the tap was discarding. So the distinct sessions, in order of first appearance, are the phase's turns in the order they ran — and `prompt.composed` already records the agent of each turn in that same order, written by the tier as it sends each turn, so it IS the running order rather than a reconstruction of it. **It attributes nothing rather than guessing.** Only when the counts match exactly. A phase whose sessions and turns differ has something this correlation does not model — a retry, a turn that called no tool, two genuinely concurrent agents — and a plausible-looking wrong attribution is worse than none here: it puts one agent's `git push` on another agent's record, and a person later reasons from that. One call missing a session id refuses the whole batch, because a hole shifts every later session onto the wrong turn. The microVM call sites pass no turn agents and so keep today's behaviour exactly. Resolving a graph node to an agent uuid is the fix there, it cannot be tested while the fleet is offline, and guessing would put one node's actions on another node's record. Also restores the `#[cfg(test)]` gate on `repo_less_text_tests`, which my own insertion had taken — those tests would have compiled into release builds. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
3f26dfeaca |
docs: suite is 796 tests across 107 binaries after this pass
Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
19c4de36e4 |
docs: the staffing fix, measured
Run 5 is run 3's task against the new staffing: 5 roles → 3, 14 skill deliveries → 4, 50KB of prompt → 24KB, and 1 of 9 delivered skills applicable → 4 of 4. The agents produced exactly the structure the new team's task specifies — questions.md, evidence.md, REPORT.md — with zero writes outside /mission/repo. The baseline says plainly that the SCORES barely moved, because they did: run 5 is one `pass` and three `not_applicable`. What changed is what `not_applicable` means — "no machine-checkable consequence" rather than "this skill had nothing to do with this phase". Halving the prompt is real but incidental. The finding is that the denominator was wrong: seven of run 3's nine skills were never applicable, so any ratio over them measured staffing, not skill use. Handoff item 1 is closed and the orphan-container section now records what was actually in it. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
6af1149e45 |
feat(missions): reap orphaned runtime containers — unless they hold work
`sweep_once` selects `FROM missions`, and `teardown_container` is only ever called with an id from that query. So a container whose row is gone is invisible to every reaper: nothing enumerates docker, nothing errors, and the only symptom is disk. Found on gw-04 today — `cm-runtime-mission-019ff5b1…`, Up nine days, 2.5G, against a `missions` table with zero rows. `list_mission_containers` is the piece that never existed: without it "which containers exist" is a question the platform cannot ask, and a container the database has forgotten is not merely unreaped, it is unseeable. **The sweep refuses to reap work that exists nowhere else.** That container's checkout held ten commits on a branch that had never been pushed — +3451/-30 across 30 files, eighteen INT items including AES-256-GCM, Ed25519 signing and HNSW batch insert. A reaper that deleted on sight would have destroyed all of it silently, as its designed behaviour. `unpushed_commits` asks the checkout (`git rev-list --all --not --remotes`) and leaves the container alone, loudly, every tick, when the answer is not zero. Every failure path returns `SomeOrUnknown`: a container we cannot question is not a container we may delete. Same for one docker will not date — including a future `Created` from clock skew, which would otherwise underflow into an age past any grace period. Grace is 24h, long on purpose. The row-driven sweep already handles everything the platform knows about, so anything reaching this path is already unexpected. The container above was handled by hand first: bundled, verified, branch pushed to git.redclaw.dev, confirmed on the remote at the branch tip, then removed. 59G free, up from 57G. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
ceec0423ad |
feat(teams): staff research phases with a research team
`research_only` is repo-less, one research phase, "produce a markdown
artifact" — and it defaulted to `rust_sdlc`. So it was staffed with a
planner, a coder, a tester, a reviewer and a committer, four of whom had
nothing to do, each carrying the code-and-commit skills its role is bound
to. Measured 2026-08-21: 9 distinct skills across 5 role prompts, ~50KB,
one applicable. That is what "most skills score not_applicable" in the
Skill-Use baseline has been measuring all along — the skills were
correctly bound to their roles; the roles were wrong for the workflow.
None of the three existing research templates fit, so this adds
`topic_research`: frame the brief into answerable questions, gather
evidence with the URL and the quoted passage, check every claim against
its source, write the report. Three roles, four skills, each checked
against its own `when_to_use` before binding — and two obvious candidates
deliberately NOT bound, because `executive-summary-writing` tells the
writer to discard any item not tied to a named project and
`signal-to-noise-ranking` scores relevance the same way. On a standalone
topic report that discards the deliverable.
`default_phase_teams` lets a recipe staff each phase PURPOSE separately,
resolved into `config.phase_teams` at create. A multi-phase recipe does
not have one job: `research_and_code`'s research phase spends a paragraph
of `task` telling its team not to change source files, because
`rust_sdlc` gave that phase a coder and a committer and they did what
coders do — mission 01a00c57 shipped both INT items during RESEARCH and
the coding phase then delivered +0/-0. Prose was the only lever
available; staffing is the actual one.
Also fixed in the three existing research templates, all verified rather
than inferred:
- `papers_research` bound `arxiv-daily` to its DOMAIN SCOUT. That
skill's entire content is "Do not search arXiv yourself — the harvest
already ran", and its `when_to_use` names Continuous Research
missions, which are the only ones the platform writes a harvest
manifest for. The role whose job is searching was bound a skill
forbidding it.
- Its PAPER READER was told to "fetch the PDF, extract text". The
runtime image has no pdftotext, no mutool and no pypdf — checked in
the container. Every paper would have hit the `[read: abstract only]`
fallback, which reads identically to the fallback working as designed.
- `insight_research` cross-referenced "our repos'" history. A mission
binds ONE repo (`missions.repo_id`).
- `codebase_research` wrote to "the Obsidian vault"; no vault is
mounted, and both it and `papers_research` were committing in "PRs",
which the platform does not open.
And `research_only` itself had neither `task` nor `done_when` — the same
defect `benchmark`, `security_hardening` and `research_and_code` were each
fixed for, and it was left out. A phase with no `done_when` is never
judged. It also still asked for `pdf`, a format nothing generates.
Two new guards, both negative-controlled: every team a recipe names must
exist (a typo currently only logs, and the mission is staffed by the
fallback crew looking deliberate), and every `default_phase_teams` key
must be a purpose `purposes_for` actually emits.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
|
||
|
|
4f4ce34203 |
fix(teams): the wrong repo path was in the TEAM templates too
The `/workspace/repo` guard was written on 2026-08-19 against `skills/`
only. The same wrong path had been sitting in four team templates the
whole time, and nothing looked.
`rust_sdlc` is the default team for five of the six workflow recipes. Its
CODER was told "your working directory is /workspace/repo. All edits
happen there." Its COMMITTER was told to `cd /workspace/repo`. The
platform mounts /mission/repo — `stamp_workspace_paths` pins it there.
Same for the frontend, three.js and mobile coders.
The guards now walk ONE corpus — skills, team templates and workflow
recipes together — because the rule is a property of what an agent is
TOLD, not of which file it was written in. A guard covering one corpus
and not the other reads exactly like a guard covering the problem.
Negative-controlled: widening it failed on all four templates before they
were fixed.
Two more defects in the same committer prompt, both found by reading it:
- `git push` unconditionally, while the `workspace-repo-commit-protocol`
skill bound to that same role says push only when the task says to,
because most missions deliver by diffing the checkout. The role prompt
and its own skill contradicted each other in one prompt.
- `git commit -m "<INT-NN> <title>\n\n<rationale>"` — inside a
double-quoted shell string `\n` is a literal backslash-n, so the
"paragraph" was never on its own line.
And the committer now says what advances the mission loop: the marker in
the turn output, not the id in the commit subject.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
|
||
|
|
6f2b0a8f43 |
docs: record the verified suite numbers in the handoff
107 test binaries, 792 tests, zero failures across the workspace — run, not estimated from the cm-api figure. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
9560aaec41 |
test(skill-use): the coding run, and the parsing bug it found
Run 4 (`research_and_code`, real repo) is the first mission that could
have violated the TDD and commit checks. It exercised both, and found a
bug in one.
Claude Code writes a multi-line commit message as a heredoc inside a
command substitution:
git commit -m "$(cat <<'EOF'
INT-01 Add slugify function to src/lib.rs
…
EOF
)"
`commit_subjects` read the first line of the `-m` value, which is the
heredoc OPENER. Every commit check was scoring `$(cat <<'EOF'` — a string
the agent never wrote. It reported no violation only because that string
is not one of the never-merge messages, which is luck rather than a check.
Regression test built from the exact command in `mission_events`.
The TDD verdict came back `not_observable`, which is the honest answer and
also a real limit worth stating: the agents edited `src/lib.rs` once —
implementation and `#[cfg(test)] mod tests` in the same write — then ran
`cargo test` five times. In Rust the unit test lives in the file under
test, so that ordering is exactly what following the skill precisely looks
like from outside. The check detects "wrote source, never ran a test" and
cannot confirm red-first. Confirming it needs the diff, not the tool order.
Every one of run 4's 33 tool calls stayed inside /mission/repo.
Handoff and baseline updated: production has never run a mission (both
tables empty), a mission container has leaked since 2026-08-12 that no
reaper can see, and `research_only` staffs a five-role Rust SDLC crew on a
repo-less markdown mission — which is what "most skills score
not_applicable" has been measuring all along.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
|
||
|
|
c209e654d9 |
fix(skill-use): a research phase writing markdown is not a TDD failure
The first live scoring of run 3 reported `cargo-test-driven-development` and `tdd-red-green-refactor` as compliance=FAIL: files were written and no test ever ran. Wrong, and wrong in the way this module exists to prevent. The phase wrote fifteen markdown notes and a helper script; there was no code to test-drive. Reporting it as an agent failure is a system defect wearing an agent's name — and it would have buried the actual finding, which is that a repo-less `research_only` mission is staffed with a Rust SDLC crew whose coder, tester, reviewer and committer have nothing to do. The check is now scoped to files with a source extension in the languages the skill itself names. Shell is deliberately excluded: a helper script written during a research turn is not behaviour-adding code, and the false failure costs more than the missed one. Recorded in SKILL-USE-BASELINE.md as finding 8 rather than quietly corrected. A measurement that hides its own false positives cannot be trusted about anyone else's. Also in the doc: the Trigger reason is half false now (the transport can surface a tool call; we simply still inline), and the architecture doc's observe/gate table said the container tier was ungated and unobserved, which shipped work has made wrong. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
4a6d0dfe01 |
test(skill-use): keep the harness that runs the measurement
The first baseline was produced by a throwaway script that no longer exists, so the second measurement could not be run the same way as the first — which is most of what makes two numbers comparable. Local stack only, because production auth is Clerk and a mission cannot be launched from a terminal there. `--score <id>` re-scores a finished run without spending another one, and every run is held for 90 days so it stays re-scorable when the scorer changes again. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
d0b657a24b |
fix(skills): two more skills that contradicted the platform
Same class as the `/workspace/repo` path and the ZeroClaw tool names: the skills were written alongside the platform and never compared to it again. Both found by reading the source of truth before writing a check against it. 1. `decompose-int-items` showed `PLAN_COMPLETE: INT-01..05`. An id is strictly `INT-<digits>`, so the range form is rejected outright — the plan pass records nothing while every item stays open. A live planner emitted exactly that line. Now one id per line. 2. `workspace-repo-commit-protocol` said the task-card parser advances mission state on the INT id in the commit subject. Nothing in the platform reads commit messages; the parser reads `run_events` — the agent's turn output. An agent that believed it could commit with the id and never emit `COMPLETED: INT-NN`, leaving the mission open on an item it had finished. The convention is kept, the mechanism corrected. `no_skill_shows_a_marker_the_parser_would_reject` runs the real parser over every marker in every skill's fenced blocks, negative-controlled against the range form. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
1a6fdfc0e6 |
feat(skill-use): score Compliance and Boundary from actions, not narrative
The scorer read the concatenated `reasoning` text — the agent's own account of its turn, written by the thing being measured and silent about anything it did not think worth mentioning. `Evidence` now carries the recorded tool calls alongside that text and every check prefers them. What that changes, concretely: - `workspace-repo-commit-protocol` Boundary was a substring search for `/workspace/repo` in the narrative. An agent that wrote to the wrong root without narrating it scored a clean pass. It now reads the `Write` and `Edit` paths, and gained the skill's other hard prohibition — force-push — which leaves no trace anywhere else once it succeeds. - `arxiv-daily` Boundary reads the `curl` that ran rather than a URL in prose, which may be the agent explaining that it did NOT fetch it. - `tdd-red-green-refactor` and `cargo-test-driven-development` gain their first Compliance check: files written with no test command anywhere cannot have been red-green under any reading of the loop. - `small-focused-commits` gains a Boundary check on the exact subjects the skill names as never-merge, read out of `git commit -m`. Two verdicts changed for honesty rather than coverage. Silence used to score `Pass`: a mission with no evidence scored identically to one checked and found clean. It is now `NotObservable`. And a test that ran AFTER the first write is `NotObservable`, not a failure — a Rust unit test lives in the file under test, so that ordering is what following the skill most precisely looks like from here. Every tool-backed check is one-sided: it reports a violation it can see and never infers compliance from silence, because the recorded stream is capped per phase. The negative controls earned their keep — they caught `-f` inside a commit message scoring as a force-push, and `git commit -am` yielding no subject at all. Trigger stays `NotObservable`, and half its stated reason is now wrong. "`claude_cli` cannot surface a tool call" is false; we simply still inline. The blocker moved from the transport to the delivery model, and the module says so. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
8cb38d1320 |
feat(missions): keep the tool's arguments, not just its name
The container tier's first measured mission recorded `Bash × 6` and not one of them said what it ran. Every behavioural question about the phase — did it run the tests, did it commit, did it call an API a skill forbids — was unanswerable from a record that looked complete. `vm_tool_tap::parse` already read `tool_input` to pull the path out of it, then dropped the rest on the floor. It now keeps it, bounded: file bodies (`content`, `new_string`, `old_string`, `edits`) become a byte count, and any other over-long string is truncated with a marker saying so. Bounded rather than whitelisted, because a whitelist silently loses the one argument that matters the first time a tool grows a field. `file.touch` keeps the absolute path in `detail.abs` alongside the repo-relative `target`. Normalising is what the map needs and exactly what destroys "did this write land outside the checkout". `tool.call` also gains `detail.path`, which the World's SSE has been reading and getting a null from on every container-tier call. `mission_events::tool_evidence_for_mission` is the reader — the counterpart to `narrative_for_mission`, and the reason it exists: the narrative is what an agent SAID it did. Host-side only. No image rebuild: the arguments were always in the tap file, the first parse threw them away. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
0b4d91889a |
docs: hand-off refresh — container-tier work shipped, stale guidance corrected
TOOL-CALL-ARCHITECTURE.md said "switch claude_cli to stream-json" as the cheapest fix. That was wrong and is now marked so, with what actually happened: zero tool.call events with the parser working perfectly, because TurnEvent::ToolCall only fires for tools ZeroClaw itself executes. Hooks sidestep that entirely, and the doc now leads with the resolution rather than the theory. A fresh session is pointed at this file, so leaving the wrong recommendation on top would have sent it down the same path. NEXT-SESSION.md: state header, and the ordered list rewritten — items 1-3 are done or superseded. "Give the direct-session tier a tap" is dropped with its reason: that tier is dormant (CLAWMATES_MISSION_EXECUTOR unset), and checking before building saved the work. New top item is watching the first production mission, since the gate and tap are proven locally and unproven in prod. Added an operational section for the things that cost the most time: the 403 actions-log API, gw-04's legacy docker-compose, the socket proxy, disk contention between manual builds and CI, and Clerk-only prod auth. Also flagged that SKILL-USE-BASELINE.md's Trigger column is now stale in a good way — tool calls are observable on the container tier, so Trigger can be scored from behaviour instead of prose. That is the highest-value follow-up. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
5a11fae0d6 |
docs: container-tier gate and telemetry shipped; CI failures were disk
Records the verified result (10 tool.call, 4 file.touch on a real mission), how hooks succeed where stream-json could not, the production state and its rollback, and the three same-shaped bugs the live test found. Also records that CI's build failures were disk pressure from my own manual runtime builds on gw-04 — not code — and that a docs-only commit was the first casualty, which made it look like a regression. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
f6e6037aa0 |
ci: make the build job readable too, and reclaim the disk that broke it
Four runs failed at `build` with nothing readable — the actions-log API returns 403 for our token, so "failure" was the whole message. The first casualty was a DOCS-ONLY commit, which made it look like a code regression and cost a cycle chasing one. It was disk. I had been building runtime images on gw-04 while CI ran on the same host; the frontend image build lost the race. Reproduced afterwards with space free and it builds clean, and `docker builder prune` reclaimed 34GB (22G free → 57G). The build job now writes its breadcrumb and a `df -h` snapshot to /tmp/ci-logs on the runner host, and records which services actually got pushed. That last one matters: the failing runs had built and pushed `server` and then aborted on `frontend`, so the registry held a partial set and `:latest` never moved — which presented as "the deploy did not happen" three steps later, nowhere near the cause. Operational note for the next person, me included: building images by hand on gw-04 competes with CI for disk on the same 150G volume. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
e84413d437 |
fix(missions): the container tier now records its tool calls — verified live
Ran it end to end on a real mission. First time the container tier has ever
been observable:
tool.call 10 Bash 6, Read 3, Write 1
file.touch 4 research/tapproof.md
reasoning 5
prompt.composed 5
Three defects found by running it, each of which left every other link
looking correct:
1. The settings document pointed PostToolUse at {TAP_DIR}/tap.sh while the
installer wrote {HOOK_DIR}/tap.sh. Claude Code does not complain about a
hook command that does not exist — it records nothing. Asserting the
script "mentions tap.sh" had passed; the PATHS have to be compared, and
a test now does that for every hook the document names.
2. The mission container runs CLAWMATES_RUNTIME_IMAGE, not the shared
runtime container I had swapped. It was still on an image whose daemon
schema has no `settings` field, so set_claude_cli_settings returned
404 path_not_found — which the error message said plainly, and which is
the only reason this was quick to spot.
3. The sweep used connect_with_local_defaults(). The server reaches Docker
through a socket proxy (DOCKER_HOST), so that connector fails there — and
my code returned Ok(()) on the error, silently. The tap filled up, the
query matched rows, and nothing ran. Now uses container_exec::connect and
logs the failure; a test pins the choice.
All three are the same shape as the bug they were chasing: installed,
inert, indistinguishable from working. The tests added for each compare the
two ends rather than asserting a string appears somewhere.
Full workspace suite green: 107 binaries, 412 lib tests.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
cd59e4798d |
feat(missions): collect the container tier's tool calls
The hooks from the previous commit write a tap file that nothing reads — which is the same shape as the gate that is installed and inert: everything looks wired and no evidence ever appears. The microVM tier records its tools from inside the loop watching the VM. A container turn is driven asynchronously by topology_worker, so there is no such loop and something has to come and collect the file. `drain_finished_container_phases` does, on the same tick as the benchmark baseline and the security scan, reusing `record_vm_tools` so container tool calls land as the same TOOL_CALL / FILE_TOUCH events the World already renders. One shape, two tiers. Idempotent by TRUNCATION, not a marker or a cursor column: `drain` clears the file it read, so a second pass finds nothing. Read-then-clear happens in one exec, and only for phases that have FINISHED — the agent is no longer appending, so the gap between read and clear cannot lose an event. A cursor would have needed a migration and a column that means nothing to anyone else. Two tests exist because the failure is silent either way: the drain must clear what it read (otherwise every tick re-records the same calls and a phase's early files end up weighted by how long the sweep ran), and the tick must actually call the sweep (otherwise the hooks write a file nobody collects). Full workspace suite green: 107 binaries. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
b89606fcf1 |
feat(missions): gate and observe tools on the container tier
The container tier is the one that actually runs missions in production, and it had neither a tool gate nor tool telemetry. The microVM tier has had both since yesterday; the tier that matters had neither. Both gaps have one cause. `claude_cli` runs claude as a subprocess, claude runs its tools inside that subprocess, and those calls never pass through ZeroClaw's executor — the only thing that emits TurnEvent::ToolCall and therefore the only thing the gateway turns into a frame ClawMates can see. Recovering the calls from the CLI's stream-json output did not help: a real mission produced zero tool.call events with the parser working perfectly. The transport was never the problem. Hooks are the way in, and they are proven. Claude Code reads hooks.PreToolUse / PostToolUse from the document given to `--settings` and honours them under `-p` — measured yesterday against the real binary, where the gate blocked a Bash call, recorded the payload, and got its refusal reason back to the model. So the same hook scripts the microVM tier uses are now written into the mission's container, and the provider is pointed at the settings document (`--settings` added to claude_cli in the fork, be9c34b1c). Composed in ONE script for one document: two writers of one settings.json is a silent clobber, and the microVM tier already learned that expensively. Installed on BOTH container paths — created and reused. A hook that exists only on first creation quietly disappears after a server redeploy, and the container outlives the server process. Everything degrades to "no hooks", never to a failed mission: a phase that runs unobserved still delivers; one that fails to start because telemetry could not be installed delivers nothing. Four tests, including two that exist because the halves are inert alone: the installer and the provider prop must both be wired (hooks nobody reads, or a document nobody wrote), and nothing may be written under /mission/repo, where it would arrive as part of the agent's delivered diff. Full workspace suite green: 107 binaries. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
930c7e0b67 |
docs: the PreToolUse gate is verified end to end
Ran it against the real claude binary with the real settings document and the real hook script. Both halves. It blocks: asked to `curl -X POST`, the agent attempted the Bash call, the hook fired FROM --settings, the call was refused, and denied.jsonl recorded the payload with hook_event_name PreToolUse and the exact command. The agent relayed the reason accurately — the text from vm_tool_gate::RULES reached the model, which is the point of writing reasons rather than bare refusals. It allows: `echo` and a harmless `rm -rf ./scratch-nonexistent` both ran and denied.jsonl stayed empty. A gate that blocked everything would have passed the first test; this is the half that rules that out — and two of this gate's four bugs produced exactly that failure. So the last unproven link in the chain is closed, and the gate is real in production rather than plausibly real. One finding worth keeping: asked to `git push --force`, the model refused on its OWN before ever calling Bash, so the hook never fired and the test was inconclusive. A gate test must use a command the model will actually attempt. The model's judgement is not the gate, and testing against something it already refuses measures nothing. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
0be932fd83 |
test(gate): a fixture emitter for the live PreToolUse check, and what it proved
Tried to close the last open question — does the PreToolUse gate actually
fire in a guest — and got most of the way.
Established:
- the generated script blocks and allows correctly under DASH, not just
macOS sh: force-push and `cd /tmp && rm -rf /` return 2, while
`grep -rn 'rm -rf /' docs/` and ordinary work return 0
- without node it allows and writes the `inert` marker, so a gate that
cannot parse is distinguishable from one that matched nothing
- `claude` in the runtime image supports `--settings` (SETTINGS-OK)
- PreToolUse DOES fire under `claude -p` in this image — measured by an
earlier session and recorded in vm_stop_gate.rs:36
Unproven, and now precisely scoped: whether Claude Code honours a
PreToolUse hook supplied via `--settings <path>` specifically, with a real
agent turn. The live attempt hit the weekly subscription rate limit, and
`claude doctor` does not report hooks, so there is no non-LLM confirmation
available.
`emit_guest_assets` (ignored by default) writes the real hook script and the
real settings document to /tmp so the check can be run against the actual
binary in one docker command — no microVM, no fleet. The exact command is in
docs/NEXT-SESSION.md.
Worth stating plainly: if that link is broken, the gate is inert in
production and looks exactly like a gate that found nothing — which is the
failure mode this whole session has been about.
Full workspace suite green: 107 binaries.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
afb1e29bf3 |
docs: streamjson2 built but deliberately not deployed
The corrected runtime image exists on gw-04 and stays there. It delivers no observability until TurnEvent::ToolCall can be emitted for observed calls, so deploying it alone would be a provider output-format change carrying risk for no benefit. Production stays on the known-good :v084. The harmful v1 image was deleted from both hosts so it cannot be redeployed by accident. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
536adddd0f |
docs: stream-json tested live — it does not deliver observability, and v1 was harmful
Deployed the amd64 build to gw-04 and drove a real mission. The agent used Bash and the standard tools; no tool.call events appeared, and the gateway's unmatched-frame histogram still showed only session_start. The reason is structural: TurnEvent::ToolCall is emitted from tool_execution.rs, only for tools ZeroClaw itself runs. Claude Code runs its tools in its own subprocess, so the event never fires. A provider that knows about the calls changes nothing by itself. The first version was also harmful — it returned the observed calls as tool_calls, so the loop tried to execute Claude Code's tool names and fed "Unknown tool: Bash" back to the model. Fixed in the fork; both runtimes rolled back to the known-good image in the meantime. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
ac4fa0b8f7 |
docs: CI green and deployed — record the verified production state
Run 498 passed and deployed. Confirmed on gw-04: 53 skills, 11 templates, zero unresolved bindings, self-authoring announced ENABLED, the new gateway_preflight answering, and migration 0080 applied. Also records that run 497 was cancelled by the concurrency guard rather than failing, and that the stream-json runtime image is still NOT shipped by this pipeline. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
689a5e14a3 |
docs: CI root cause was an apostrophe, not any of the three theories
Records both real causes (run 490 stomped by an overlapping run; 491-496 killed by an apostrophe closing a single-quoted sh -c block), the guard that now catches the second class locally, and what to check when the in-flight run settles — including that a successful build is the FIRST time these commits reach production. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
8f988739ec |
fix(ci): an apostrophe in a comment killed six runs
Runs 491 through 496 failed on one character.
The Rust step is a `docker run … sh -c '…'`. A comment inside that
single-quoted block read `cm-api's vm_tool_gate`, and the apostrophe closed
the quote. Bash died with "unexpected EOF while looking for matching quote"
BEFORE running anything — which is why no log ever appeared, why the
breadcrumb showed the step entered and produced nothing, and why three
separate theories were floated to explain an empty failure.
I introduced it in the commit that installed nodejs, so the fix for run 490
broke every run after it.
Run 490 itself was the stomping: it overlapped run 491, which began by
removing the shared `cm-ci-pg` container out from under it. That is fixed
too, and was a real defect — it was simply not the cause of 491+.
`bash -n` answers this in milliseconds and nothing was running it: a
workflow is not compiled, not linted, and its only feedback is a red build
with a log this deployment cannot read. `tests/workflow_shell_syntax.rs`
now extracts every `run:` block and syntax-checks it, so the failure shows
up before the push rather than six runs later. Gitea's `${{ … }}` is
replaced with a placeholder first — the point is to check OUR quoting, not
to evaluate their templating. Negative control: restoring the apostrophe
fails the test with the file and line.
The block also carries a standing NO APOSTROPHES warning, because the next
person to write a comment there will not be thinking about quoting.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
d23f30e929 |
docs: record the CI investigation honestly, including what is still unknown
Establishes what is verified (the code passes on the runner host, with cargo's real exit code), what is narrowed (493/494 die inside the Rust step before cargo starts; 495 died before step 1), the three theories that were wrong, and the cheapest next experiment. Also records the two things that made this expensive: the actions-log API returns 403 for our token, and my first reproduction piped cargo into `tail` and reported tail's exit code. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
c02dbe2266 |
ci: capture the Rust step's own output, not just cargo's
The breadcrumb narrowed run 494 to the Rust step, and rust.log did not exist — so cargo never started. Whatever failed (apt-get, git config, or docker itself) wrote to the job log, which the actions-log API will not give us. The docker run's stdout and stderr now land on the host too, and the step exits with docker's status. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
24393819bd |
ci: breadcrumb which step dies
The host-log change proved the job never reaches `cargo test` — rust.log is absent while /tmp/ci-logs exists. But TWO steps create that directory, so "the directory exists" does not say how far the job got, and that ambiguity cost a debugging cycle on its own. Each step now overwrites /tmp/ci-logs/STEP on entry, so the last value names the step that died. The postgres step also runs under `set -x`. Verified manually on gw-04 in the meantime: the postgres step's exact commands succeed there (STEP_RC=0), as does the whole Rust container command with cargo's real exit code, as do the three frontend commands. So the failure is something the job does that running its steps by hand does not reproduce — which is precisely what a breadcrumb answers and guessing does not. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
a864f2ccc7 |
ci: make a failed run readable, and stop laundering cargo's exit code
Three runs failed and I debugged all three blind: Gitea's actions-log API
returns 403 for the token we have, so the only evidence was the word
"failure". I twice inferred a cause from that and was twice wrong — first
node, then dash — and a third theory (two runs stomping each other) was
right about a real defect but not about these failures.
Worse, my own reproduction lied. It ran `cargo test ... | tail -80`, so the
reported exit code was TAIL's. A green pipeline over a red suite is exactly
the trap this repo already documents, and I walked into it while hunting a
red build.
- every step writes its full output to /tmp/ci-logs on the RUNNER HOST,
which outlives the container, so a failure can be read afterwards
- the Rust step captures cargo's status in a variable and exits with it,
with the grep and tail in between — no pipe anywhere near the status
- the frontend step runs npm ci / typecheck / vitest separately, keeps
each status, prints all three tails, and fails if any is non-zero.
Previously a `set -e` abort meant later steps produced no output at all
What is now known, verified on the runner host itself with cargo's real
exit code: `cargo test --workspace` PASSES on gw-04 in the CI container
against a CI-shaped Postgres (CARGO_RC=0), and `npm ci`, `typecheck` and
`vitest` all pass there too. So the failing step is not one of those, and
the next run will say which it is instead of leaving it to be guessed.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
72ba4ba523 |
fix(ci): two runs stomped each other, and the logs blamed the tests
Runs 490 and 491 both failed `test`. Neither failure was in the code.
Runs 490 and 491 started 16 minutes apart and a full suite takes longer
than that, so they overlapped. The first thing a run does is
`docker rm -fv cm-ci-pg` — a name every run shared — so the newer run
deleted the older run's database mid-suite. Both failed, and the failures
read as test failures.
Verified before changing anything: the exact CI command, on gw-04, against
the same warm cargo volumes and a Postgres started exactly as CI starts it,
passes on
|
||
|
|
128b423205 |
fix(ci): the tool gate needs node, and an inert gate must say so
The first push of the PreToolUse gate failed CI, and the reason is a property of the gate worth fixing rather than a CI quirk. The hook parses its JSON payload with `node` — no jq in the runtime image, and node is guaranteed there because Claude Code is a node program. CI runs `cargo test --workspace` inside `rust:1.96-slim`, which has no node. The extraction returned nothing, the gate allowed everything, and the two "blocks" tests failed. That is correct behaviour with a dangerous appearance. A gate that cannot read its input must not block the phase — failing closed on a parse error denies every tool call, which is what an earlier `case`-syntax bug did. But allowing silently makes an INERT gate indistinguishable from one that simply matched nothing, which is this codebase's recurring defect exactly. So the gate now records `inert` when node is absent, still allowing, and a test pins both halves: exit 0, and the marker written. The host can check for that file rather than infer a working gate from an absence of denials. CI installs nodejs so the shell tests exercise the gate instead of its inert path. Verified in a rust:1.96-slim container: without node the force-push payload returns 0, with node it returns 2. Also confirmed the generated script behaves under dash — Linux /bin/sh — not only under macOS sh. An earlier apparent dash failure was invalid JSON in the probe command, not the gate. Full workspace suite green: 106 binaries. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
b653dbfe72 |
docs: hand-off note for the next session
Records the state of the tree, the one step not taken (the stream-json runtime image is built and never deployed, so no mission has confirmed tool.call rows end to end), the ordered next steps, the decisions that are the operator's, and what was deliberately left undone with reasons. Also records the two corrections made this session — "missions can't call tools" was wrong, and raw test counts are a bad coverage metric — because both were confidently stated here before being checked. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
547b5d9987 |
feat(missions): a pre-execution gate on mission tool calls
The second half of the tool-call research. Until now a mission agent's
Bash call was gated by nothing, anywhere.
WHY THERE WAS NO GATE
vm_tool_tap is a PostToolUse hook: it fires after the tool has already run
and exit-0s unconditionally, because a non-zero PostToolUse talks back to
the model. It is telemetry and says so. GatePolicy — the §15 door — has one
enforcement site, the chat loop, and its approvals key on
(session_id, message_id), which no mission phase can produce. Meanwhile the
solo tiers run `claude -p --permission-mode acceptEdits` with Read, Edit,
Write and Bash pre-approved.
PreToolUse fires under `claude -p` in this image — measured by vm_stop_gate,
which also proved the exit-2-plus-stderr contract — and had zero callers.
This is that hook.
WHAT IT IS, AND IS NOT
A deterministic policy gate: a short deny list of actions with no legitimate
form inside a mission, blocked before they run, with the reason handed back
so the model can choose differently.
It is NOT the §15 human approval gate, and the module says so. A hook blocks
the agent's process while it runs and a human decision takes minutes to
hours; waiting inside the hook would wedge the turn. This closes the gap
between nothing and something.
The deny list is short on purpose. A gate that blocks legitimate work is
worse than none: the agent cannot ask a human, so it either works around the
block — doing something stranger than what was denied — or burns the turn.
FOUR BUGS THE TESTS FOUND, IN ORDER
1. Substring matching denied `grep -rn 'rm -rf /' docs/`. Searching for a
string is not running it. Now rules anchor to the start of a shell
segment, with a separate flag-style match that exempts text tools.
2. The `case` patterns were unquoted, so a needle containing a space made
the whole script a SYNTAX ERROR — which as a PreToolUse hook exits
non-zero and denies EVERY call. Every text assertion passed while the
script was in that state; only running it under a real `sh` found it.
3. The hook receives JSON, not a command, so "starts with" could never
match — `case` saw `{"tool_name":"bash",...` every time. Now extracts
tool_name and tool_input.command with `node` (no jq in the image; node is
guaranteed because Claude Code is a node program).
4. `IFS='\n'` in POSIX sh sets IFS to backslash and the letter n, not a
newline. Nothing split, so only commands with no separator were ever
tested and `cd /tmp && rm -rf /` sailed through. Now a literal newline.
Every failure path allows. A gate that fails closed on a parse error blocks
the whole phase, which is exactly what bug 2 did.
Wired through vm_tool_tap::guest_settings, still the single writer of the
guest settings document — a third hook makes the clobber it prevents more
likely, not less, and a test asserts all three survive one document and that
PreToolUse points at the gate's own script rather than the tap's.
Honest limit, stated in the module: a determined agent defeats any
string-matching gate. This is aimed at accidents and obvious cases; the real
isolation is the container and microVM boundary.
Full workspace suite green: 106 binaries, zero build errors.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
ea0b989b3f |
docs(research): missions DO call tools — the claim was wrong, and the truth is worse
Deep research into "missions can't call tools at all", which I wrote and which is false. docs/TOOL-CALL-ARCHITECTURE.md has the full findings. WHAT IS ACTUALLY TRUE Three of the four mission paths end in `claude -p` with Claude Code's own toolset and permissions PRE-ACCEPTED: solo microVM Read Edit Write Bash Agent --permission-mode acceptEdits composed microVM same, per node same direct session Read Edit Write Bash acceptEdits So the position is not "no tools". It is: mission agents run Bash and Write with permissions pre-accepted, and nothing in this platform can gate them. That is a stronger finding than the one it replaces — "can't call tools" sounds like a missing feature; "calls tools freely, ungated, and mostly unobserved" is a security posture, and it is ours. Observe and gate are different and both are partial. vm_tool_tap is a PostToolUse hook: it fires AFTER the tool ran and exit-0s unconditionally, so it is telemetry and structurally cannot gate. The direct-session tier has no tap at all. GatePolicy has exactly one enforcement site — the chat loop — and its approvals key on (session_id, message_id), which no mission phase can produce. WHY THE CONTAINER TIER LOOKED TOOL-FREE `claude_cli` runs `claude -p --output-format json`, which returns a single final result object, and the provider hardcodes `tool_calls: Vec::new()`. The calls happen; the transport discards them. The comment reading that emptiness as "§15 by construction: agents are provisioned tool-free" was inferring a design property from a serialization choice. Verified against the deployed Claude Code 2.1.228 rather than assumed: `--output-format stream-json --verbose` emits `tool_use` blocks with the tool name and `tool_result` blocks. The calls are fully observable; we ask for the wrong format. THE DOOR WE ALREADY BUILT AND NEVER PLUGGED IN claude_cli.rs is OURS — upstream zeroclaw-labs/zeroclaw has no such file — and so is 88eef99d4 "claude_cli --mcp-config + allow/disallow tools (act via door)". The provider already accepts mcp_config (claude's own MCP client reaches our door), tools, and disallowed_tools (lock out the natives so the gated door is the ONLY actuator). agent.config.example.toml documents the whole shape. In the live runtime: clawmates-mcp.json does not exist, there is no [providers.*] block, and every mission claw binds to claude_cli.default which sets none of it. My earlier "claude_cli cannot reach MCP, therefore the skills server is unreachable" was wrong in its reasoning — the capability is built, documented by us, and never deployed. Related: we set `agents.<alias>.mcp_bundles`, which configures ZeroClaw's OWN MCP client for its native loop. A claude_cli agent's actuator is the claude subprocess, which reads `mcp_config` on the PROVIDER. We were turning a knob wired to a loop that does not run. UPSTREAM 218 commits behind. No upstream work on claude_cli (the file is ours). ACP already exists in the fork; the three new commits are workspace-default and localization fixes, not new capability. The one item worth pulling is "feat(plugins): add shared egress policy foundation (#9137)" — a network guard with DNS pinning and metadata-address blocking, defence for the egress problem we have not solved. Stale claims corrected in place, in topology_exec.rs and the runtime config, so the codebase stops asserting the thing that is false. Recommended order, cheapest first: stream-json for observability; the PreToolUse hook for a real gate (it FIRES under claude -p per vm_stop_gate, and has zero call sites); then deploy the door. The executor swap is NOT recommended — the blockers are structural, not wiring, and the cheap fixes deliver what it was wanted for. Co-Authored-By: Claude Opus 5 <[email protected]> |