c66c3c6377a1a27af7711dc69c26f18ce60708a4
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c66c3c6377 |
feat(ui): the microVM path is reachable from the mission wizard
Everything built today — Firecracker missions, the four backends, the local GPU one — was unreachable from the dashboard. The wizard offered `zeroclaw` and `local_herdr` and nothing else, so a mission created in the UI could not be a microVM mission at all, and `local-ornith`/`glm`/`kimi` were API-only. Testing "our workflows in the UI" would have exercised none of it. Adds the runtime option and a backend picker, fed by a new `GET /api/fleet/backends` that returns `mission_roster::available_backends` verbatim — the SAME list the roster planner is handed, not a second one. Its two rules are both load-bearing and neither is visible from a node's capabilities alone: the image must be built on an online node, and the backend must have a credential contract. `agent-terminal` passes the first and fails the second — bootable, with nothing for the agent inside to authenticate with — so offering it would produce a mission that validates, launches, and dies at the agent turn. Ids are deployment vocabulary, so the picker labels them: a user choosing between `local-ornith` and `canary-claude` should not have to know which company each one bills. An empty list says why (no rootfs built) instead of showing an empty dropdown, and no node is chosen for a microVM mission because `vm_placement` picks it per phase. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
1f6108f769 |
feat(gc): reclaim the mission tree on the gateway
`cleanup_sweeper` prunes ROWS. Deleting a row has never deleted a directory, and `teardown_container` only runs while a mission still exists to tear down — so a mission removed by any path that skipped teardown left its tree behind permanently, on the smallest disk in the fleet (150 GB, shared with postgres and every checkout). 106 mission directories are sitting there now. Filesystem-first, deliberately: the DB is the PREDICATE, never the enumerator. Enumerating from the database is exactly how these became invisible — a directory whose row is gone is the one a row-driven sweep cannot see. Three reapers, one deletion path. Orphan mission dirs (no row, past a 2h grace), scratch trees (_bench/_gate/_verify/_merge past 6h — all four have leaked before), and _outputs past 90d, whose artifact rows are marked only AFTER the files are gone, because the other order claims artifacts are reaped while they are still on disk. The single removal path escalates: the server is uid 65532 and cannot delete what the per-mission daemon leaves as root, so PermissionDenied falls back to `root_copy::purge` and shouts if the tree survives even that. A GC that cannot collect is the thing being fixed, so failures are counted and reported, never swallowed. Guards worth naming: `_cargo` is a SHARED cache every mission writes to and lives under the same root, so an underscore-prefixed sibling treated as an orphan mission would delete it out from under running work and look like a slow cargo build. Only a well-formed mission id is ever a candidate — a directory whose name is not an id can have no row by construction, so without that gate every unrecognised directory looks orphaned. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
3c3d01c8d1 |
fix(llm): the chain preflight printed nothing at all
Deployed, and the report simply did not appear — from the tool built to stop
things failing silently. Two causes, both worth keeping:
There was no timeout anywhere in the probe, so one slow provider swallowed the
entire report. Each link is now bounded at 60s (generous: `complete_or` spends
up to 30s in its own backoff, so a tighter cap would report a merely throttled
link as hung) with `TimedOut` as its own state, and every line is emitted AS IT
RESOLVES rather than collected and printed at the end — a later link that hangs
must not be able to hide the ones already checked.
The first attempt at the timeout awaited the probe and then wrapped the result:
let probe = complete_or(...).await;
timeout(PROBE_TIMEOUT, async { probe }).await
That compiles, reads correctly, and bounds nothing. The timeout has to wrap the
future.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
c7c3eeab46 |
fix(test): the colon-vs-spec test did not compile
Committed and deployed while its test compile was failing: the verify step was `cargo test | grep -E "^error|test result" && git commit`, and grep exits 0 when it MATCHES, so finding the error is what let the commit proceed. The library built fine, so the deploy was sound, but the check that was supposed to gate it did the opposite of gating. The error itself was a borrow in a test closure; a plain fn fixes it. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
d9c5300859 |
fix(llm): the preflight found two broken links on its first live run, one its own
fallback chain (6 link(s), 4 usable):
claude-opus-4-8 ok
claude-sonnet-4-6 ok
claude-haiku-4-5-20251001 ok
kimi:kimi-k2.7-code BROKEN: 400 ... role 'system' must not be empty
glm:glm-4.7 ok
local:ornith-fleet:9b UNREGISTERED — resolves to the DEFAULT provider
Neither link was actually broken.
The probe sent an EMPTY system prompt, which Kimi rejects outright. A probe has
to look like the traffic it stands in for, or it measures itself.
The second is the one worth keeping. `resolve_provider` returns a spec unchanged
when it does not recognise the provider, and the part after the FIRST colon when
it does — so the obvious test, "does the model half still contain a colon",
reads correctly and is wrong the moment a model id has one. `ornith-fleet:9b`
has one. The probe reported a provider the server had just finished registering
as UNREGISTERED.
`evaluator::cross_provider_judge` had the identical check, and would therefore
have refused a local judge as "not independent" — silently falling back to a
same-family one, which is the exact claim that path exists to make honestly.
Both now compare against the whole spec.
That bug was written into the codebase before a model name with a colon existed,
was correct at the time, and became wrong when one arrived. Nothing would have
reported it; a boot-time probe of every link did, on its first run.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
c3ad5672fc |
feat(llm): six-link fallback chain, and a preflight that proves it
opus -> sonnet -> haiku -> kimi -> glm -> local. The order is capability first, then independence: three Anthropic tiers on one account (a throttle usually hits a tier, so stepping down often clears it), then two separately funded accounts (now an outage, not just a throttle, is survivable), then our own GPU (nothing left to be down). Every id was probed on this deployment and answered 200. The preflight is the more important half. Configured is not working, and this chain has a specific way of lying: `resolve_provider` falls back to the DEFAULT provider when it does not recognise a provider name, so a typo in `kimi:` does not error — it quietly runs on Anthropic, and a chain that reads as three accounts is really one. A reachability-only probe calls that link green. So `preflight` checks resolution and reachability separately, eight tokens per link through the REAL call path, and reports four states. `Throttled` is deliberately not a failure: a 429 means the spec resolved, the credential authenticated, and there was no capacity this second — the exact condition the chain exists to route around, and painting it red would train an operator to ignore red. `Unregistered` and `Broken` are failures, and they get different words because they need different fixes. It runs at boot alongside validator_preflight and runtime_preflight, spawned so it cannot delay startup. A chain is the one piece of infrastructure nobody looks at until the day it has to work, so it is now checked on the days it does not. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
5afcf63324 |
fix(harness): count what the roster run ADDED, not what the file holds
Forcing the planner onto the local link produced a green chain and a red assertion: roster: the planner sized this mission at 1 member(s) PASS roster: ROSTER.md has 3 line(s) for a 1-member roster FAIL The model was right and the check was wrong. ROSTER.md does not start empty — the auto-merge work put an earlier run's two lines onto main — so a 1-member roster that correctly appended one line delivered three, and the scenario reported a model that had ignored its own proposal. It now measures the DELTA against main. Any assertion against a scratch repo that accumulates has to, or it decays into a test of how many times it has been run before. Proven on the local model end to end: opus 429 -> local:ornith-fleet:9b answered -> `mission_roster: ... local:ornith-fleet:9b proposed 1 member(s)` -> the composed graph ran -> the branch added exactly one line. 5/5. CLAWMATES_MODEL_FALLBACK is removed from gw-04's .env again; it was set only to force the last link for this test, and the deployed default is the full chain. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
b18e62041b |
feat(llm): the fallback chain's last link runs on our own hardware
`local:ornith-fleet:9b` joins opus -> haiku -> glm as the final link. Every entry above it depends on somebody else's account staying funded and unthrottled; this one depends on a GPU in the next room. It is last because it is the weakest model, and present because a chain whose every link is external is not a fallback chain, it is one outage in a trench coat. Three small changes make it work: - `build_provider_registry` accepts a provider with an empty `api_key_env`. A model on our own hardware has nothing to authenticate to, and the old behaviour SKIPPED a keyless provider — leaving the chain quietly one link shorter than it reads, which is the failure mode this whole area keeps producing. - `provider_family` learns `ornith`/`ollama` for BARE names. A qualified `local:` spec was already answered by the split, but a bare one fell through to "unknown", and `cross_provider_judge` would then refuse a judge that is genuinely a different family from the Anthropic implementer. - A test pins that the last link survives `resolve_provider`'s split-on-FIRST- colon: `local:ornith-fleet:9b` is provider `local`, model `ornith-fleet:9b`. Splitting on the last colon would ask for a provider named `local:ornith-fleet`, and the symptom would be a silent fall back to the default provider. Infra: Ollama on tank and architect now binds 0.0.0.0 so the gateway (which has no GPU) can reach it. `tailscale serve` cannot — Ollama rejects a non-local Host header as a DNS-rebinding guard and OLLAMA_ORIGINS is CORS-only, so it 403s. 0.0.0.0 still includes loopback, so the microVM vsock pipe is unaffected; verified on both nodes. This is an explicit trade: Ollama has no auth and its API can pull and delete models, so it is now reachable from the LAN as well as the tailnet. The drop-in carries the ufw one-liner to close the LAN side. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
774f17d194 |
test(fleet): a mission served entirely by the node's own GPU
`local-ornith` scenario, green on its first real run against tank: local-ornith: a locally-served model delivered a guest kernel (6.1.128) local-ornith: no Anthropic egress from a locally-served mission local-ornith: the node bound its local-model socket for this VM local-ornith: checkout has exactly one writer (uid=65532) Three things had to be true at once and only a real run shows all three: the agent reached a model at all (a pipe to a closed port produces a turn that HANGS rather than errors, which is why this is a scenario and not a unit test), the work came back and landed on a branch, and the VM still could not reach api.anthropic.com. That last one is not theoretical. The node log for this VM is a column of `egress DENIED api.anthropic.com` — Claude Code's own telemetry, correctly refused — while the model traffic went through the vsock pipe and Ollama logged loading ornith-fleet:9b at 100% GPU with CONTEXT 131072. A local backend that quietly kept Anthropic egress would be a credential path nobody asked for. The egress check asks the NODE's proxy log rather than the agent, for the same reason the GLM measurement did: a model's account of where its tokens came from has no evidential value, and the proxy's record of what it dialled does. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
f56d41f5b7 |
feat(backend): local-ornith — a mission backend served by the node's own GPU
Claude Code pointed at the Ollama already installed on every GPU node. Ollama has served a native Anthropic-compatible /v1/messages since v0.14, so this is an env contract rather than a translation layer — the fourth variation on the same idea as agent-glm and agent-kimi. The route is NOT the egress proxy, and that is the design. `egress` speaks CONNECT, takes a destination from the guest, resolves it and decides; every one of those powers is a liability, which is why it refuses non-443 ports and IP literals after a unit test caught them being bypassed. Routing a local model through it would have meant relaxing both. `local_model` is the opposite shape: there is no destination in the protocol. fcagent listens on guest 127.0.0.1:11434 and pumps to vsock 9003; the node splices that onto its own 127.0.0.1:11434 and copies bytes. A compromised guest cannot redirect it because there is nothing to redirect — it is a pipe, not a proxy, and strictly narrower than anything an allow-list could express. The bytes never touch a network, so there is no wire for TLS to protect, and Ollama stays bound to loopback rather than being exposed on the tailnet. The socket is bound only for a backend declared to use a local model, so a `local-ornith` VM reaches the forge through egress and nothing else, while every other backend's guest port simply refuses. Both halves have negative controls. `scripts/fleet-model-setup.sh` exists because of one measurement: stock ornith:9b reported input_tokens=2050 for a 48000-word prompt and answered as though nothing had been dropped. Ollama's default window is ~2K whatever the model card says, and it truncates silently — the exact failure an agent turn would hit and never report. The script pins num_ctx=131072 into a derived tag and then PROVES both the window and tool calling before declaring success. Verified on architect: ~65536 words -> 65604 input tokens, stop_reason=tool_use. Placement needs no new capability key: building the rootfs only on GPU nodes means `nodes::online_for_backend`'s existing `rootfs @> ["local-ornith"]` predicate does the affinity, so morpheus never offers the backend. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
e96c5143bc |
test(eval): a local judge, and the 2K context window that would have hidden it
Phase 1 of the local-model plan: prove the model before writing any plumbing. `JUDGE=local` runs the existing done_when eval against Ollama on a GPU node. Requests originate on that node rather than the gateway, because the model is bound to 127.0.0.1 deliberately — it has no network exposure at all — and the gateway has no GPU. MEASURED on tank, 3 draws per case, against the incumbent on the same cases: local (ornith-fleet:9b) 14/15 — one UNPARSED, never a wrong verdict glm (glm-4.7) 13/15 — two WRONG verdicts on kernel-ok kernel-ok is the case production actually hit and the one this script's header says is expected to fail on glm-4.7. A 5.6 GB model on hardware we already own did not get it wrong once in three draws. The tag is `ornith-fleet:9b`, not `ornith:9b`, and that is the finding worth keeping. Ollama defaults to a ~2K window whatever the model claims: stock ornith:9b reported input_tokens=2050 for a 48000-word prompt and answered as though nothing had been dropped — silent truncation, confidently. The fleet tag pins num_ctx=131072, which measures 9.3 GB resident of a 16 GB card (the full 262144 also fits, at 13.6 GB, 100% GPU). These eval cases are a few hundred tokens, so this eval would have passed either way; that is exactly why the tag under test has to be the one production would use. Also measured: Anthropic /v1/messages returns well-formed tool_use with stop_reason=tool_use on both nodes; the reported count_tokens?beta=true hang is absent in 0.31.1 (clean 404, server unaffected); ~60 tok/s generate, ~2800 tok/s prefill, 120072-token prompts accepted end to end. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
f68fc019e4 |
fix(teardown): a mission dir with root-owned files is now actually removed
The server runs as uid 65532, so `remove_dir_all` on a mission directory returns PermissionDenied the moment anything root-owned is left in it — and the old code logged that at the same level as "file not found" and moved on. The directory then lived forever. After the seed-copy fix a mission holds 3281 files owned by 65532 and 26 owned by root: `.claude.json` and the session jsonl the per-mission ZeroClaw daemon writes itself, after the copy has been chowned. Twenty-six files is small enough to keep every mission directory alive without anyone noticing why. PermissionDenied now falls back to `root_copy::purge`, which deletes from inside the container as root — the same escape hatch `container_exec` keeps for exactly this, clearing debris a root process created. And if the directory survives even that, it says so, because a cleanup that silently failed is the thing being fixed. Removing the last 26 properly means running the per-mission daemon as 65532, which needs `/mission` pre-created in the image with that ownership — the daemon creates it at boot today and cannot at a lower uid. That is an image change, deliberately not bundled here. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
4967b9b8fd |
fix(runtime): the seed copy reads as root and hands the result to 65532
Running the seed copier as 65532 broke mission launch, and broke it quietly. The seed dir is root-owned with parts at mode 0600 (`.claude.json`, `clawmates-mcp.json`), so uid 65532 cannot READ them: `cp` failed on the first unreadable entry, `set -e` abandoned the rest, and the mission came up with a runtime-data holding `.zeroclaw` and nothing else — no Claude credentials, no door config. The daemon then never created its agents' workspace, and the phase failed 200 lines later on "Could not find the file /mission in container", which points nowhere near the cause. It was quiet because `seed_runtime_data` polled for the container to STOP and returned Ok without ever reading its exit code. A copier that died on a permission error and one that finished cleanly were indistinguishable. It now reads the status and says what went wrong. So: root for the read, `chown -R 65532:65532 /dst` for the result. Both halves matter and they pull opposite ways — root is needed to read the seed, and 65532 is needed because everything else in the missions tree is 65532 and a GC running as 65532 cannot delete what root left behind. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
8ddea454d1 |
fix(runtime): the seed copy ran as root too, ~3200 files per mission
The uid fix landed and the CHECKOUT came back completely clean — 0 non-65532 files under `repo/` after a benchmark run that builds and tests Rust. But the same mission still held 3247 root-owned files, all under `runtime-data/`. `seed_runtime_data` spawns a throwaway container to `cp -a` the runtime seed into the mission's directory and never set `user`, so it ran as root — the identical absent-`user` omission `container_exec` had, in a container create instead of an exec. The seed source is 65532-owned and the destination is created by the server (which itself runs as 65532), so the copy never had a reason to out-rank either. This is the tree a gateway GC has to be able to delete, and a GC running as 65532 cannot remove root-owned files — the cleanup-that-cannot-clean-up shape, found before writing the GC rather than after. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
dcd9514622 |
fix(exec): mission work runs as uid 65532, so it stops creating debris it cannot delete
`CreateExecOptions` never set `user`. Not a wrong value — an ABSENT one: the daemon defaults to root, and twelve callers inherited that without any of them choosing it. That single omission is the origin of four separate patches — root-owned `target/` directories inside a checkout owned by 65532, `root_copy` existing at all, and a cleanup that had to re-enter the container as root to undo its own mess. The rule is positional and lives in ONE place: an exec whose workdir is inside `missions_root()` runs as 65532; anything else (preflight probes, image checks) keeps the daemon default so unrelated call sites cannot break. Twelve callers each remembering to pass a uid is twelve chances to forget, and the one that forgets leaves debris the other eleven cannot remove. Non-root needs an environment the image does not provide. Measured in the deployed image: uid 65532's HOME (/zeroclaw-data) and /usr/local/cargo are both root-owned and unwritable, so this would otherwise break every cargo call — the benchmark runner, the judge's sandbox, the delivery test gate — far more quietly than the leak it fixes. The missions root IS bind-mounted and writable by 65532, so HOME/CARGO_HOME move there and the cargo cache is shared across missions rather than re-fetched per mission. Verified on gw-04: a clean `cargo build` as 65532 with those three variables produces output owned entirely by 65532. Root remains reachable only through `exec_as_root`, whose name says so, and which exists solely to clear debris earlier root execs left. `runtime_preflight` now probes the whole policy at boot, so an image that moves or tightens that mount fails loudly instead of failing every cargo call for a reason no error message would connect to a uid. evaluator_tools' inlined fourth copy of the purge is replaced by `root_copy::purge`. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
42108c840d |
docs(placement): the drain half of that fix was never the broken half
drain-midmission passed 3/3 twice, but the "re-placing this phase" line the last commit added never appeared in the log. It cannot: `online_for_backend` filters on `status = 'online'`, so a draining node is not a candidate, never reaches `unfit`, and the pin simply falls through to ranking — on the old code as well as the new. So the scenario passes either way and proves the affinity decision, not the `TargetUnfit` bug. The path that genuinely used to fail a phase is "the previous phase's node has since FILLED UP": that puts it in `unfit`, which returned a non-transient error, which never reached the queue. That is what the unit test now says, in place of a claim about draining the harness does not support. The accidental mission-to-node affinity was real and unconditional either way. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
13a35138e9 |
fix(placement): a drained previous node re-places the phase instead of failing it
`drain-midmission` found this. `choose` treated its `want` argument as a hard
requirement, and the only caller passes `missions.target_node_id` — which is
not an operator's choice, only where the PREVIOUS phase happened to run. Two
consequences, both wrong:
- A node drained or filled between phases produced `TargetUnfit`, which
`is_transient()` says false to, so `phase_runner` FAILED the phase rather
than queueing or moving it. The queue silently did not apply to the second
phase of any mission.
- While the node stayed fit, every later phase went straight back to it
regardless of ranking — accidental mission-to-node affinity, which this
module's own header says must not exist.
Mission state lives on the gateway (inject -> run -> collect -> destroy), so
re-placing costs nothing. The pin is now advisory: preferred while it fits,
and when it does not, the reason is logged and ranking proceeds. `TargetUnfit`
is deleted rather than left unconstructed, so it cannot come back as a
non-transient failure by accident.
The scenario had its own race: it waited for phase 0 to COMPLETE before
draining, but warm phases finish in ~80s against a 10s placement sweep, so
phase 1 was often already placed — and the run then blamed the platform for
running on a node that was not yet drained. It now drains while phase 0 is
still running, which does not disturb a live VM and is the more faithful test.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
d4af58be85 |
fix(harness): the capacity burst got faster than the thing watching it
The first bursts took 25 minutes because every VM paid a cold 2.4 GB rootfs copy. Warm, the same 16 missions finish in 70-140s each and the whole burst is over in about two minutes — so a sampler that waited ~90s for its launch check and then ticked every 15s caught three samples of the tail and reported "architect peaked at 1 of 6" for a run that sat at 6/6/2. Sampling now starts at the first tick, runs every 5s, and folds the launch check into the same query so verifying the launches costs no observation window. The 10-sample floor that produced the last NORUN is gone; it was measuring how long the burst took, not how well it was watched. And "nothing queued" no longer has one verdict for two causes. If the fleet never actually filled — a slot can free before the sweep reaches the 15th mission — the queue was not reached and this scenario did not test it: NORUN, naming the high-water mark. Only a burst that DID saturate can call an absent queue a failure. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
eacd3ee085 |
fix(harness): a blind sampler must not report an idle fleet
The burst re-run printed "architect peaked at 1 of 6" and "nothing ever queued" for a run I could watch sitting at architect=6 tank=6 morpheus=2 with 2 phases queued. The fleet was right; the sampler was blind. Three separate ssh+psql calls per 15s tick, each with stderr to /dev/null, and under the load of 16 concurrent missions most came back empty. Empty was then read as "nothing running" — absence encoded as a legitimate value, which is the exact seam the header of this file was written about, reproduced in a scenario added to catch it. One query per tick now, returning done/blocked/per-node in a single row, and unreadable samples are COUNTED rather than silently treated as zeroes. Fewer than ten usable samples is NORUN: a sampler that barely looked must not be able to describe itself as a fleet that was idle. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
91fbd2dc88 |
refactor: the missions root has one definition, not five
`mission_workspace::missions_root()` is now the only place that answers "where
does mission state live". It had fragmented into five: this function, private
`env::var("CLAWMATES_MISSIONS_ROOT")` copies in security_scan, benchmark_runner
and mission_outputs, and a hardcoded `MISSIONS_HOST_ROOT` const in
mission_runtime that read no env at all.
They agree on the deployed value, so nothing has broken. The risk is entirely
in what comes next: anything that sweeps or reclaims this tree has to be
sweeping the same tree the writers use, and five definitions cannot promise
that — a reaper written against one would silently leave the others' directories
behind forever, which is how the orphans got there in the first place.
A source-walk test fails any module outside `mission_workspace` that reads the
env var itself.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
e5f097c291 |
fix(harness): the capacity burst verifies its own launches
The re-run reported FAIL-NORUN "the burst did not finish in 1800s". The fleet was fine — 3 of the 16 missions were still in `draft`. Each PATCH-to-running is an ssh plus a `docker run curl`, and 16 at once does not reliably land; the response was going to /dev/null, so a launch that never happened spent the full timeout looking like a platform stall. That is precisely the swallowed-error shape this file was written to catch, committed inside the file itself. Launches are now verified against the mission rows, retried once for the stragglers, and reported as "the burst never happened" rather than as a timeout — a scenario that did not run must not be able to describe itself as a slow one. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
4fedfcec30 |
fix(placement): a young VM's unconsumed memory was handed out twice
The capacity harness scenario, on its first full run, caught what it was written to catch: capacity: architect peaked at 6 of 6 slot(s) FAIL capacity: 'morpheus' peaked at 3 concurrent VM(s) with only 2 slot(s) capacity: tank peaked at 6 of 6 slot(s) PASS capacity: the over-capacity missions QUEUED PASS capacity: all 16 queued/placed missions completed `capacity_of` inferred the host's own footprint by subtracting the VMs' FULL 8 GiB claim from observed usage — which assumes they have already consumed it. A VM booted seconds ago holds about an eighth. On morpheus (31757 MiB total, 4314 MiB idle, 2 slots) with 2 young VMs at ~6314 MiB observed, the inference 6314 - 16384 goes negative, clamps to the 2048 floor, and invents 2266 MiB — exactly enough for a third VM on a two-slot node. The footprint is only honestly MEASURABLE when nothing is committed, so remember it then: `nodes.mem_baseline_mib`, sampled by `survey` whenever it observes an idle node with fresh health. When VMs are committed, take the LARGER of the remembered reading and the old inference — a node that was once idle at 4 GiB and is now running a 20 GiB build must not be scored as idle, which would be the same over-commit arrived at from the other direction. Both directions have a test; the second is the one that would otherwise rot. Raising HOST_BASELINE_FLOOR_MIB would have made this one node's numbers pass and drifted the moment the fleet changed shape. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
2056bb1d9e |
test(fleet): prove the queue and the spread under a real burst
Phase 1 shipped placement-at-phase-launch and a queue made of `start_pending_phases` leaving a phase `pending`, both deployed unproven under load — the exact condition this project keeps getting burned by: the code is right, the system is wrong, and nothing errors. `capacity` launches `slots + 2` microVM missions simultaneously and asserts two things. That no node ever exceeds the slots `vm_placement` gave it — overcommit does not fail loudly, it swaps, and every mission on that node gets slow rather than dead. And that the excess QUEUES: a burst that drops the extras and one that wedges them both look identical to any check that only reads the end state. `capacity_blocked_since` is cleared the instant a phase is placed, so the evidence only exists mid-flight; the scenario samples while it runs. Capacity comes from `/api/fleet/capacity`, never recomputed here — a bash copy of the slot arithmetic would drift from the scheduler and then agree with itself. A burst that does not exceed capacity is reported NORUN, per rule 3. `drain-midmission` drains the node phase 0 ran on, before phase 1 is placed, and asserts phase 1 lands elsewhere AND still reads phase 0's file. That is the test of the affinity decision: mission state lives on the gateway, so re-placement is free — if it were not, this would either strand the mission or silently lose the earlier work, and "silently lose" is what a status-only check calls success. The node is restored before any assertion runs, so a failure cannot leave the fleet permanently one node smaller. Smoke-checked at CAPACITY_BURST=2: sampling, spread and completion all report, and the queue check correctly returned NORUN rather than a green tick. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
dc0443de34 |
feat(fleet): GET /api/fleet/capacity returns the scheduler's own survey
Pulled forward from the observability phase because the capacity harness scenario needs it. A test that recomputed the slot arithmetic in bash would drift from `vm_placement` and then agree with itself while the scheduler did something else — the same shape as every silent-success bug in this codebase. Returns `survey()` + `rank()` unmodified, and keeps `unfit` as its own list: "the fleet is full" and "we could not read the fleet" send an operator to different places, so they must not be summed into one number. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
d48bdbc9a7 |
fix(llm): two modules were posting to Anthropic behind the providers' back
The research scenario passed 4/4 and the log underneath it said: phase_summarizer: ... failed: anthropic 400 Bad Request: "Your credit balance is too low to access the Anthropic API" `phase_summarizer` and `mission_refiner` each built their own reqwest POST to the Messages API with `x-api-key: $ANTHROPIC_API_KEY`. No audit of `.complete(` call sites could have found them — they never touched a provider — so every phase summary and every mission-brief refinement on this deployment had been failing against an empty account while the phases themselves ran fine. The summarizer even persisted an error row per phase, which is why nothing ever retried loudly enough to notice. Both now go through `subscription::complete_with_fallback`, so they inherit the subscription-first credential choice, the 429 backoff, and the opus -> haiku -> glm chain. The summarizer records the model that ANSWERED in mission_phase_summaries.model rather than the one it asked for. The guard is a source WALK, not a file list: any .rs under cm-api/src that mentions the Messages API host or `x-api-key` fails the test. A hand-listed set of files is exactly what let these two hide. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
52500a689c |
fix(door): say so when the security governor is failing open
`Runtime::judge` returns "governor unreachable (fail-open)" whenever the provider never answers, and the door caller drops `reason` on every allow — so a judge model that is rate limited or uncredited turns the governor into a rubber stamp with nothing anywhere saying so. Fail-open stays (a governor outage must not halt agents), but it is now loud. Found while removing the metered key as a dependency: the governor reads CLAWMATES_JUDGE_MODEL, which was `claude-opus-4-8` — a model that is 429 on this deployment's subscription. gw-04's .env now points it at `glm:glm-4.7`, matching CLAWMATES_VALIDATOR_MODEL: funded separately, uncapped, and a different family from the agent it judges. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
9c9439a271 |
feat(llm): the subscription is the default provider, with a recorded fallback chain
Two changes so an empty metered account stops being a platform outage. 1. `build_provider` prefers the subscription token over ANTHROPIC_API_KEY. A bare model name resolves to whatever this returns, so making it the subscription means no server-side call can reach the metered key by construction — rather than by a source-grep test that already missed four call sites once. The metered key remains a fallback and now warns loudly when it is the one in use; boot no longer requires it at all. 2. `complete_with_fallback` walks a declared chain when a model has no capacity: opus -> haiku -> glm:glm-4.7 by default, overridable via CLAWMATES_MODEL_FALLBACK, empty to disable. Measured on gw-04 today: opus and sonnet return 429 on the subscription while haiku, GLM and Kimi all return 200, so a capped window no longer means "the planner is gone". The chain returns the model that ANSWERED, and every caller persists it — mission_plan_proposals.author_model, mission_team_proposals.author_model, and the swarm's step role. A plan drafted by the third link and filed as an opus plan is a silent quality change, which is the failure shape this project keeps paying for. Two negative controls hold the design: the chain never retries the model that just failed as its own fallback, and it steps down ONLY for a capacity failure — walking it on a malformed prompt would ask three models the same bad question and report the third one's confusion. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
ee5a939ce6 |
fix(planner): the other four server-side calls were still on the metered key
The test that was supposed to prevent this grepped for the literal `runtime.complete(` and passed while the phase planner (`mission_plan.rs`), both swarm calls, and a second enhance path in `claws.rs` still billed the pay-as-you-go account. They spell the receiver `state.runtime` or wrap the call across lines, so the receiver name was never the thing to match. The test now matches the METHOD, and covers all five files. `complete_or` gains the rule that makes it safe to apply everywhere: a `name:model` spec is an operator's explicit provider choice — the swarm worker model is configured exactly that way — and is passed straight to `Runtime::resolve_provider` untouched. Only a bare name is ambiguous, and a bare name is precisely what resolves to the default provider. Hijacking a chosen Kimi or GLM model onto Anthropic would be the same silent-substitution bug pointed the other way. `validator_preflight` and the evaluator judge keep calling the runtime directly, on purpose: both exist to exercise the CONFIGURED spec. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
deed591da6 |
fix(roster): a rate-limited subscription is a 503 with a reason, not a 500
The retry landed and still failed: all four attempts returned 429. A bare
16-token probe with the same token, straight from gw-04, also returned 429
with `x-should-retry: true` — the Claude Code subscription itself is limited
right now, and no amount of backoff inside one HTTP request will outlast it.
So stop pretending it is a server bug. New `ApiError::Unavailable` → 503,
carrying the one sentence the operator can act on ("clears on its own; try
again shortly"), instead of an opaque `internal error` that sends them into
the logs. The harness now prints the response body rather than the generic
"the planner produced no usable proposal", which is what hid both walls —
first the credit balance, now this.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
c3c4447810 |
fix(planner): wait out a rate limit instead of failing the whole proposal
Moving the roster and planner onto the subscription removed the credit wall and revealed the next one: the harness went from `400 credit balance too low` to `429 rate_limit_error`. A one-shot proposal call had no retry — there is no retry convention anywhere in cm-llm — so a limit that clears in seconds killed the "propose a team" button outright. Four attempts, 2/8/20s backoff, and only for errors that can actually clear: 429/5xx/transport. A 400, 401 or 404 returns immediately, because retrying those is a 30s hang ending in the identical message, which reads to an operator as a stall rather than a bad request. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
72046e7985 |
fix(planner): server-side model calls run on the subscription, not the metered key
The roster planner died with `400 — "Your credit balance is too low to access the Anthropic API"` while every mission on the same machine kept running. Two Anthropic credentials reach this server and they bill differently: `ANTHROPIC_API_KEY` (sk-ant-api, metered, runs out) and the Claude Code subscription token (sk-ant-oat) that every VM already uses. `Runtime::complete` with a BARE model name — "claude-opus-4-8" — resolves to the default provider, which is the metered key. Three server-side callers did that: the roster planner, the Master Planner, and the claw enhancer. Missions were never affected because `mission_runtime` deliberately sends only the subscription token into a guest; the server had no equivalent rule. `subscription::complete_or` is now that rule, and it is the ONE place a subscription token becomes a provider — `evaluator::subscription_judge` had its own copy, and two of them is how one ends up with a prefix check the other lacks. The `sk-ant-oat` prefix is checked rather than the variable name trusted: an API key pasted into the OAuth slot would authenticate, work, and bill the metered account — the same failure again, discovered weeks later. `web_search` is carried explicitly rather than defaulted. The Master Planner and the claw enhancer both pass `true`, and a helper that quietly dropped it would have taken web search away from two features while every test still passed. `validator_preflight` deliberately keeps `Runtime::complete`: it probes whatever validator spec is configured (today `glm:glm-4.7`), and forcing it onto Anthropic would make it prove the wrong thing. A test pins both halves — no other server-side caller may regress to the metered key, and preflight must keep probing the configured spec. 258 lib tests. |
||
|
|
d84d17207f |
feat(placement): place per phase, and let a full fleet queue
Phase 1b: wires the capacity model from
|
||
|
|
3a2d76aa43 |
feat(placement): capacity model for the fleet — observed memory is not capacity
Phase 1a of the fleet-intelligence plan: the arithmetic and the inputs. Nothing is wired to it yet; the launch path still picks `capable.first()`. Placement has been `ORDER BY last_seen DESC` + `.first()` — the most recently heartbeated node. Among healthy nodes all heartbeating every 5s that is arbitrary, and it consults nothing about load, so two missions launched together land on the same machine. It did not matter while tank held the only rootfs image. All three nodes serve `claude` as of today. THE correctness point, and the reason this is not a sort change: a VM that booted 30 seconds ago holds a fraction of its 8 GiB claim, so `mem_pct` reports a sold-out node as nearly idle. `capacity_of` takes the WORSE of observed usage and committed usage. The negative control pins it with the measured case — tank at 60 GiB total / 12 GiB observed / 5 VMs booted: utilisation alone says 5 more fit, the node has room for 1. Booking those five is a node in swap, which slows every VM on it together. Commitments are unioned BY IDENTITY, never added: `vm_list` reports booted VMs, `nodes::pinned_microvm_phases` reports phases chosen but not yet booted (a window of seconds in which a real 8 GiB claim exists that no node can report). The deterministic `vm_id_for` is what lets the same phase be recognised in both — counting it twice would shrink the fleet by the number of phases starting. `EvalRow::headroom()` finally gets a caller. It was written with the doc comment "for placement ranking" and has had zero callers since. It is a TIEBREAK, not a gate: ranking is slots first (spread, don't stack), then live headroom, then node id so the same fleet state yields the same answer twice — which `last_seen DESC` could never promise. Fail-closed per house convention: draining, stale health (>30s, tuned just above the 20s offline sweeper), and an unanswerable `vm_list` are all INELIGIBLE rather than low-scoring. Stale Beszel metrics are the one exception — they demote a node to zero headroom instead of excluding it, because they only ever break ties. `FleetAtCapacity` and `FleetUnreadable` are separate variants with a test asserting the second never says "at capacity": an operator sent hunting a load problem that is really a dead daemon wastes the outage. Also names the two nodes that were both called "New node" (tank, morpheus) — a capacity report naming two machines identically is one nobody can act on. 257 lib tests. |
||
|
|
5c5f1ced33 |
refactor: the judge's sandbox joins the other three copy sites on root_copy
Four places copy a mission checkout so a ROOT command can run against it without touching the live tree: the judge, the benchmark runner, the on_green_tests gate, and — until now — the judge again, with its own implementation predating the shared one. `evaluator_tools::Sandbox::for_checkout` now builds through `root_copy::RootCopy`. Same packer, same exclusion list, same reasoning in one place. It needs the copy to OUTLIVE the handle, because the judge has not run when `for_checkout` returns and a firing `Drop` would delete the tree out from under it. That is `into_workdir`, a method rather than a `mem::forget` at the call site: the transfer of cleanup responsibility is then visible in the type instead of implied by a leak. `Sandbox::purge` remains what actually clears it, since only a container running as root can remove the root-owned `target/`. What did NOT move: `pack_dir` in the microVM inject/collect path. That marshals a tree to and from a guest over vsock — a transport, not a host-side copy — and folding it in would merge two things that only look alike. 249 lib tests. |
||
|
|
8b12245e79 |
chore(image): promote Claude Code 2.1.226 after the canary passed
Canary first, production second — the point of pinning is that the upgrade is a decision, and the point of `canary-claude` is that the decision has evidence. Taken for 2.1.225's fix to a transient 401 that replaced a long-lived CLAUDE_CODE_OAUTH_TOKEN with a short-lived one and broke HEADLESS sessions until restart. Our sessions are headless and our VMs are per-mission, so "until restart" reads as a failed phase. The 2.1.225 workspace-trust prompt does NOT apply: `--help` states the dialog is skipped in non-interactive mode (`-p`, or stdout not a TTY) and we satisfy both. Read from the CLI in a booted 2.1.226 VM rather than inferred from the changelog. Verified on the production path after the rebuild: guest kernel 6.1.128, delegation to a subagent, `--settings` stop gate installed, judge independent (glm-4.7), single writer. 6/6. `rootfs-canary-claude.ext4` is left on tank as the mechanism for the next candidate, not as a leftover. |
||
|
|
099a716bfd |
fix(egress): a backend is defined in two maps, and the canary only had one
First canary run failed: phase failed, nothing delivered, and the streamed log said exactly why — "Failed to authenticate. API Error: 403 api.anthropic.com is not on the egress allow-list". Not a 2.1.226 regression. `canary-claude` was added to the server's credential map and not to the node's `provider_hosts`, so the VM booted with a valid subscription token and a door that only opened onto the forge. The fail-closed branch was working correctly: a backend nobody taught that function about reaches no model API, deliberately, so it cannot silently borrow another provider's door. Both maps now name it, each pointing at the other, with a test asserting the canary reaches the same provider as `claude` AND that unknown backends still resolve to nothing. Worth noting what made this a five-second diagnosis instead of an afternoon: the live log streaming built earlier today. The failure was a 403 inside a microVM that no longer exists, and its reason was sitting in the run's checkpoint. |
||
|
|
4193ae2cda |
feat(missions): a canary backend for testing a CLI version on the real path
Claude Code 2.1.223 -> 2.1.226 is worth taking (2.1.225 fixes a transient 401
that replaced a long-lived CLAUDE_CODE_OAUTH_TOKEN with a short-lived one and
broke HEADLESS sessions until restart — which for us means a failed phase). But
the image every mission uses is not the place to find out whether a new CLI
still delegates, still accepts `--settings`, and still finishes.
`canary-claude` is a real rootfs built from the candidate version, credentialed
identically to `claude`, so a mission can exercise it through the production
path: egress, stop gate, delegation, delivery, streaming. Testing a new CLI
against a different provider would not be testing the thing we are about to ship.
Named explicitly rather than matched on a prefix. An unrecognised backend must
still be refused at launch — that is what `backend_can_run_a_mission` and the
harness's `microvm-negctl` scenario assert — and loosening the credential map is
exactly how that guard gets softened by accident. A test pins both halves.
Already cleared by direct measurement in a booted 2.1.226 VM, before this:
- `--settings` and `--agents` still exist
- the workspace trust prompt added in 2.1.225 does NOT apply: `--help` states
the dialog is skipped in non-interactive mode (`-p`, or stdout not a TTY).
We use both.
|
||
|
|
8c93cd8569 |
fix(runs): the composed worker's checkpoint wiped the live log on every node
Composed missions streamed ZERO bytes while solo missions streamed fine. Same executor, same command, same guest — `HubVms::run` is a straight passthrough — and the node logged a tail starting for all five graph nodes against the correct outer run id, with no errors. The bytes simply were not there at the end. Two writers, one column. `fleet.rs` appends live output under `checkpoint.log`; `topology_runs::checkpoint` wrote `SET checkpoint = $2`, replacing the whole object. A composed run checkpoints after EVERY graph node, so each node's progress silently erased the log written during it. A solo run has no second writer, which is exactly why it looked like it worked. Now merged with `||`. The keys are disjoint, so the progress object still wins for everything it owns. I was wrong about the cause twice before finding this. First I blamed the guest agent's serial accept loop — real, fixed, and not this. Then I blamed pipe buffering racing the abort at turn end — plausible, and the drain fix is right on its own merits, but composed still streamed zero afterwards, which is what ruled it out. The thing that actually located it was noticing solo and composed differ by a WRITER, not by a code path. |
||
|
|
09afa7e7ff |
fix(node): aborting the tail at turn end raced the flush that matters most
Composed runs streamed NOTHING while solo runs streamed fine — same code path, `HubVms::run` is a straight passthrough, and the node logged a tail starting for all five graph nodes with the correct outer run id. The difference was timing. `claude -p ... | tee` makes stdout a PIPE, so the CLI block-buffers and flushes at EXIT. The most valuable output — the agent's summary of what it did — arrives in the instant the turn ends. The node aborted the tail the moment `handle_op` returned, so that flush was a race: a solo turn (minutes long, output already flushed by size) won it and streamed 337 bytes; each node of a composed run (~20s) lost it and streamed zero. The tail now DRAINS. A flag is set when the turn returns, and the loop exits only after a pass that read nothing new — checked AFTER a read, never before one, because exiting on the flag alone would drop exactly the bytes this exists to capture. Bounded by a 20s timeout with the abort kept as a backstop rather than the mechanism, so a VM that stopped answering cannot hold the task open. Worth naming: 5 tails started, 5 logged cleanly, 0 bytes arrived. Every individual step reported success and the feature did nothing — the same shape as the empty Live tab this whole thread began with, one layer down. |
||
|
|
5b49d5a1a8 |
feat(merge): gate publication on the merged tree's own tests
The other half of the merge button. Merging told you the branch went in; nothing
checked that what came out still worked.
Verified BEFORE publishing, not reverted after. `merge_locally` and
`push_merged` are separate functions so the caller can run the project's tests
between them, which means a merge that breaks the base is simply never pushed —
`main` is not broken for however long it takes someone to notice. A test asserts
`merge_locally` contains no push, because the moment it does, verification
becomes after-the-fact and the guarantee is gone.
Outcomes, all reported to the operator rather than swallowed:
Passed -> published
NoSuite -> published, and SAID so; a repo with no tests is a fact about the
repo, not a pass
Failed -> not published, exit code reported, branch untouched so it can be
fixed and merged again
CouldNotRun -> not published. Fail closed: a suite that could not run has not
passed, and publishing on "we could not check" is how a green
main stops meaning anything.
`verify_tests` runs `cargo test` as ROOT in a container, so the merge workdir
ends up holding a root-owned `target/` the server (uid 65532) cannot delete —
the same leak found three times today. Purged through the container before the
ordinary cleanup.
248 lib tests.
|
||
|
|
28090d1de0 |
fix(node): the log tail gave up before the turn wrote its first byte
First live test of the streaming path: mission passed 6/6, `checkpoint.log` was 0 bytes, and the node logged nothing at all. `stream_vm_log` treated "no progress" as "the turn finished writing". But the guest's `tail` reports EOF after every idle window, and the FIRST idle window is always the one before any output exists — the VM is still booting and the CLI still starting. So the tail returned `at == 0`, the node concluded the turn was done, and it stopped seconds into a run that then went on for minutes. The abort is the terminator, not idleness: the caller already aborts this task when the exec returns, so waiting cannot outlive the turn. No-progress now sleeps and retries instead of returning. Also logs when a tail STARTS. The bug was invisible in exactly the way this session keeps finding: silence on the success path, silence on the give-up path, and an empty Live tab that looked identical to a feature nobody had wired. Method note, since it cost time: I tried to confirm the deployed binary by grepping it for `vm_out` and found zero — then found zero for `pty_out` and `vm_exec` too, in a binary whose PTY streaming demonstrably works. Binary-grep is not a reliable presence test for these literals; `stream_vm_log` and `tail of` being present is what actually showed the code had shipped. |
||
|
|
0b89b8316c |
feat(observability): stream a microVM turn's stdout/stderr to the platform live
The Live tab showed nothing while a turn ran, and the agent's own account of it
went to stderr on the node and nowhere a user could reach. This is the path that
carries it.
The blocker was the guest agent. `fcagent` handled one connection at a time,
inline, so during an hour-long turn the VM accepted nothing — which is why every
existing probe (subagents, stop-gate blocks, cap) runs AFTER the turn rather than
during it. It now spawns a thread per connection, wrapped in `catch_unwind`
because this process is pid 1: a panic used to take the accept loop with it, and
an unbootable VM is a far worse outcome than a missing log. A failed spawn logs
and keeps accepting rather than dropping the listener.
PROVED against a live VM before building on it, since "sound reasoning about this
system" and "measurement" have diverged repeatedly today. Patched rootfs, booted
under Firecracker, ran an 8s exec and a concurrent tail:
exec took 8.0s ok=True
+0.0s 'line1\nline2\n' +1.2s 'line4\n' +3.2s 'line6\n' +6.0s 'DONE\n'
VERDICT: CONCURRENT — tail returned data before exec finished
The rest is the pattern the terminal already uses. New `tail` op streams a file
by OFFSET (so a dropped link resumes instead of replaying, and the tail always
terminates — one that never returns pins a thread for the life of the VM). The
node follows the log alongside the turn and pushes `Uplink::VmOut { run_id, at,
data }` over the WebSocket it already holds, mirroring `PtyOut`. The server does
what `PtyOut` deliberately does not: it APPENDS to the run's checkpoint as well
as fanning out, because a terminal has no history worth keeping and a mission log
is the record of what the agent did. `run_events_sse` emits the new bytes as
`step` events, which the live pane already renders — no frontend change.
The turn is `tee`d, not redirected: the file feeds the live stream and stdout
still becomes `VmOutcome::summary`. A redirect would have produced a live view
and an empty summary, which is the same green-and-empty shape as the bug this
fixes. Tested, along with the log living outside the collected tree so it never
lands in a user's delivered diff.
246 lib tests, 20 binaries; node and fcagent build clean.
|
||
|
|
62509a5090 |
fix(missions): a solo microVM run showed the operator an empty Live and Output tab
Found by a frontend wiring sweep, then confirmed in the database.
Everything the UI shows of a run's CONTENT reads
`topology_runs.checkpoint.records`: `/api/missions/{id}/documents` behind the
output reader, and `/api/topology-runs/{id}/events` behind the live pane. The
`team` and `microvm_graph` tiers write those records. The SOLO microVM path
never did — it updated `status` and nothing else:
tier | checkpoint_null | records
microvm_graph | f | 2-5
team | f | 5
microvm | t | 0 <-- every one
So a single-phase microVM mission ran real work, delivered a real branch, and
showed an empty Live tab and an empty Output tab. The agent's own account of the
turn went to stderr via eprintln and nowhere a user could reach.
Note what was NOT broken, since that was the initial suspicion: the SSE path
matches (`/api/topology-runs/{id}/events` on both sides), and a sweep of all 130
frontend `/api/` calls against the 164 registered routes found zero genuinely
missing endpoints. The wiring was fine; the data was absent.
The run now persists its turn as one record shaped exactly like the ones those
two readers already parse — `node_id`, `role` (the phase kind), `phase`,
`output` — so no reader changes. Written with `checkpoint || $3::jsonb` so a
future writer of other checkpoint keys is not clobbered.
246 lib tests.
|
||
|
|
3616bc4733 |
feat(missions): an operator button to merge a mission's branch into main
`MergePolicy::Never` — the default for anything touching code — has always meant
"do not merge on your own", deferring to a human. There was no way for that human
to say yes: `auto_merge` was reachable only from the paper-harvest path, no
workflow template declares `merge_policy`, and every mission ended at a branch.
`POST /api/missions/{id}/merge` is that yes, with a button on the artifacts tab.
The additive-only gate does NOT apply here, deliberately: an operator reading a
code change is exactly the judgement the policy was holding out for.
What is not waived:
- the branch comes from the artifact delivery RECORDED, not rebuilt from the
mission id, and must have `pushed: true`. A phase that never pushed shows no
button instead of one that cannot work.
- an empty branch is refused. A button reporting success for merging nothing
is worse than no button.
- a conflict refuses, aborts, and leaves the repo clean rather than forcing.
It works in a FRESH CLONE under `_merge/<mission>`, never the mission checkout:
that directory is reaped on a timer after a mission ends, so a merge using it
would succeed right after a run and fail inexplicably an hour later. The clone is
made by the server process, so nothing runs as root and ordinary cleanup works —
unlike the copies in `root_copy`.
`merge_and_push` is split out so the operator path and the automatic path run the
SAME git commands; only the gates differ. A test asserts both call it, that the
operator path does not re-apply the additive gate it exists to bypass, and that
it still refuses an empty branch.
Harness 43/43 across all five recipes before this change, with `_gate`, `_bench`
and `_verify` all at zero.
246 lib tests, 20 binaries, 89 frontend tests, clean build.
|
||
|
|
a8b8efba6a |
fix(delivery): the on_green_tests gate ran the suite in the live checkout
Fourth instance of the same defect, and the last of the three commands that run as root against a mission tree. `verify_tests` execs the project's test command with `workdir = repo` — the live checkout — inside a container running as ROOT. `cargo test` writes `target/`, so the checkout ends up owned by two uids and the next phase's cargo hits permission-denied. The harness reported `uids=0,65532` the first time this gate ever ran end to end. It survived because it had never run. Every one of the ten harness fixtures used `commit_policy: "always"`; `on_green_tests` and `on_reviewer_approval` were parsed, implemented, and never exercised — and `Gate`'s own doc already records that three recipes carried this policy while it "did precisely nothing" for want of a reader. A policy that is never exercised is indistinguishable from one that is ignored. Consolidated rather than fixed a third time. `root_copy` now owns the pattern — copy through `mission_fs::pack_dir` into a SIBLING of the mission dir, run there, and purge FROM INSIDE THE CONTAINER, because the copy's `target/` is root-owned and the server (uid 65532) cannot delete it. `benchmark_runner` moved onto it; `evaluator_tools::Sandbox` keeps its own copy logic for now (it carries an allow-list and a judge-facing API, so folding it in is a larger change than this moment warrants — noted, not done). The gate fails CLOSED if the copy cannot be made: an unverifiable suite must not license a push. Also adds the `refactor` scenario, which is what found this. I had written it off as "structurally identical to four existing scenarios" — wrong: it is the only recipe carrying `on_green_tests`, and that made it the only one testing this code path at all. 245 lib tests, 20 test binaries. |
||
|
|
a4b4d05b8d |
fix(evaluator): the verification sandbox leaked for the same reason the bench copy did
Found by checking `_verify` after fixing the identical bug in `_bench`: 16 MB
stranded across two copies, the oldest hours old.
`Sandbox::Drop` calls `std::fs::remove_dir_all` as uid 65532. The judge runs
`cargo test` in a container as ROOT — that is the entire point of the sandbox —
so the copy's `target/` is root-owned and the removal fails on it, leaving the
whole tree. The error was logged to a stream nobody reads, so the sandbox that
exists to protect the checkout quietly filled the disk instead.
Its doc comment also claimed "the copy lives under `_verify/<mission>`, which
the next pass clears anyway". That was wrong for exactly the same reason:
`for_checkout` removes a stale root before copying, with the same uid, and fails
the same way. A leaked copy was permanent, not transient.
`Sandbox::purge` removes it from inside the container, as root, where it was
written. `evaluate` now wraps its body so the purge runs on EVERY exit — that
function returns from several branches, and cleanup only some paths reach is the
same as no cleanup on the others. `Drop` stays as a fallback for the early paths
where nothing has run as root yet, and its comment no longer claims otherwise.
This is the third instance today of the same shape: cleanup that cannot clean up,
invisible because the failure was swallowed. The others were the leaked agent
containers in the runtime tests and the bench copy in
|
||
|
|
e89a32ffef |
fix(benchmark): the bench copy leaked because only root could delete it
The copy fix in
|
||
|
|
a93a4111e1 |
fix(benchmark): the baseline runner was writing root-owned files into the checkout
Caught by the harness: `benchmark: checkout has multiple writers (uids=0,65532)`.
The previous full run passed that same check, so this was introduced by wiring
`benchmark_runner` into the sweep one commit ago.
`docker_exec` enters a container running as ROOT with the missions root
bind-mounted, and `cargo bench` writes `target/`. Run in the live tree it leaves
root-owned build output in a checkout owned by uid 65532 — the single-writer
invariant broken, and the next phase's cargo hitting permission-denied on a
directory it cannot write.
This is the SAME defect `evaluator_tools::Sandbox` was written for, found by the
same probe, and fixed the same way: benchmark a COPY. `BenchCopy` packs the
checkout through `mission_fs::pack_dir` (so it excludes exactly what the
delivered diff excludes — one exclusion list, now four consumers) into
`<missions_root>/_bench/<mission>`, a SIBLING of the per-mission dirs like
`_verify` and `_outputs`, so a mission reap cannot race a running bench. Removed
on drop, including on error paths.
The operator-triggered path (POST /api/missions/{id}/benchmark) had this bug
from the start and is fixed by the same change — it shares `run`.
Worth naming the pattern: measurement must not mutate what it measures. It
applies to the judge, to the `verifier` subagent that has no Edit or Write, and
now to the benchmark runner.
243 lib tests, zero warnings.
|
||
|
|
0d8db7ff0b |
fix: close the three remaining gaps, and repair a test I silently disabled
FIRST, the self-inflicted one. My edit in |
||
|
|
2a9a62c784 |
test(harness): cover security_hardening — 4 of 5 recipes now run end to end
The third recipe whose defining phase is not `coding`, and so the third that nothing could fail before `PRODUCING_KINDS` widened: a `security_scan` phase that ran no scanner and wrote nothing reported success. One phase, not the recipe's full scan->research->code chain — what is under test is the phase KIND, and the other two kinds are already covered. Two assertions, because the first alone is weak. "Delivered a file" is satisfied by an agent that writes "I scanned it, all clear" and runs nothing — the letter-not-purpose shape this codebase keeps paying for. So the delivered patch must also carry the scanner's OWN output. Verified against the real run: the agent produced gitleaks' banner, INF/ERR lines, byte counts and exit code, not a claim about them. Only `refactor` is now uncovered, and deliberately: its single phase is `coding`, structurally identical to chain/multirole/microvm/noop. It would add runtime and no new signal. security 4/4 against the live fleet. |
||
|
|
6dd7937ece |
test(harness): cover the two recipes that had none — research_only and benchmark
The portal offers five workflow recipes. Every one of the harness's seven
fixtures was `research_and_code`, so four recipes had never run end to end —
and that is not a theoretical gap. `research_only` DESTROYED its output for as
long as it existed: `requires_repo = false`, so the capture query's
`AND m.repo_id IS NOT NULL` skipped it, the container was reaped unread, and
eight ClawHDF5 research documents were lost while the mission reported
`completed`. Nothing in 550+ tests could see it, because nothing ran the recipe.
`research-only` asserts the whole chain the loss ran through, not just the
happy end of it:
- the phase completes
- document artifacts exist AT ALL (the missing thing)
- the agent's seven identity files (SOUL.md, MEMORY.md, …) are NOT published
— the first live capture published all seven, because `.git/info/exclude`
cannot protect a mission with no `.git`
- the captured text reads back through the content endpoint, since an
artifact row pointing at nothing is a 404 with no explanation
`benchmark` covers the other half: a benchmark mission is ONE benchmark phase,
and while `empty_delivery_is_a_failure` tested `kind == "coding"` that phase was
exempt — nothing in the platform could fail it. The scenario asserts it both
completes AND delivers files.
Also: `run_scenario` takes an optional `no-checkout`. The single-writer uid probe
is a property OF A CHECKOUT, and a repo-less mission has none by design, so
probing reports a platform fault that is really a category error. It is declared
per scenario rather than inferred from a missing directory — that inference would
silently excuse a repo-BACKED mission whose checkout was reaped early, which is
the exact condition the probe exists to catch.
research-only 4/4, benchmark 3/3 against the live fleet.
|
||
|
|
a20702d55b |
fix(evaluator): a bare validator model name claimed independence it never had
`CLAWMATES_VALIDATOR_MODEL=gemini-2.5-flash` (or any bare model name) produced
an Anthropic judge grading Anthropic work, recorded `independent = true`.
The chain:
- `provider_family` reads the SPEC. A bare `gemini-2.5-flash` matches none of
the known needles, so it returns "unknown" — deliberately NOT "anthropic",
so it passes the `family == IMPLEMENTER_FAMILY` guard.
- `Runtime::resolve_provider` (runtime.rs:224) falls back to the DEFAULT
provider for any spec it cannot route. A bare name has no `provider:` to
route on, so it silently returns the house Anthropic provider.
- The existing "no provider registered" guard checks `model.contains(':')`.
That works for `glm:glm-4.7` — an unrouted colon-spec comes back carrying
its colon — and can NEVER fire for a bare name.
So the one guarantee this path exists to make (the judge is not the implementer)
was reported as satisfied while being violated. That is the same shape as the
Goodhart incident the independent judge was built after: not a wrong answer, a
wrongly-trusted one.
A validator spec must now name its provider. `names_a_provider` is a named
predicate rather than an inline `contains(':')` so the rule is testable and the
reasoning has somewhere to live.
Found while auditing my own Gemini removal — which turned out to be
behaviour-neutral here (a gemini spec went from family "gemini" to "unknown",
both non-anthropic, same verdict). The bug is pre-existing and independent of
it; removing Gemini only made the bare `gemini-*` spelling more likely to be
left behind in someone's env.
Live config is `glm:glm-4.7`, a proper registry spec, so production behaviour is
unchanged. Negative control: make `names_a_provider` return true unconditionally
and `a_validator_spec_must_name_its_provider` fails.
241 lib tests pass.
|
||
|
|
87f188ae73 |
refactor: strip Gemini from the platform, and level up the architecture_mapper
Two things.
1. The architecture_mapper proposal, applied AND made durable.
The GLM proposal (019fddd9) was accepted in full: the agent's system_prompt now
carries the Mermaid-first constraint and its brain was rewritten. Both verified
against the live row and the .h5 file.
But `apply_identity` writes `UPDATE agents SET system_prompt` and
`apply_brain_consolidation` writes that agent's brain — neither touches the team
TEMPLATE. That agent is mission-scoped, so the improvement would have died with
the mission. The model's actual insight was sharp and worth keeping: "Mermaid
diagrams beat prose" lived in the brain SEED and not in the system PROMPT, so it
only applied when the agent happened to consult its brain. That constraint is
now in templates/teams/codebase_research.toml, where every future Codebase
Research team inherits it.
(The proposal's second item mostly restated anti-patterns the seed already
lists, so the seed is unchanged. Applying an LLM's suggestion is not the same as
agreeing with all of it.)
2. Gemini is gone.
Removed: the `gemini.default` provider alias and its `is_exact_provider_match`
prefix, GEMINI_API_KEY forwarding to agent containers, the evaluator's
gemini->gemini family row, the model selectors in claws/teams/planner and in
TeamWizard + AgentComputer, and the commented provider block in the runtime
config example (whose ZEROCLAW_AGENT_MAP example still mapped a worker_gemini
that no longer existed).
`provider_alias_for("gemini")` now returns claude_cli.default via the
unrecognised-model branch, which LOGS. A stray gemini binding degrades visibly
rather than resolving to a provider row we no longer ship. A test pins that, and
another pins that GEMINI_API_KEY is forwarded in NEITHER auth mode, so adding it
back to the list is a visible change rather than an accident.
Avatar generation is DELETED, not disabled — it called Gemini's image model, and
there is no alternative: Claude and Kimi are text-only, and z.ai answers
"Unknown Model" for cogview-3-flash and cogview-4 on our plan (measured, not
assumed). AvatarModal keeps UPLOAD, which never needed a provider; only the
prompt-generation half is gone.
240 backend lib tests, 89 frontend tests, clean tsc + eslint, build succeeds.
|
||
|
|
f6c3ddbf81 |
refactor: no feature depends on Gemini any more
Depleted Gemini prepayment credits took out PDF rendering. The same key was the
only thing standing between level-up proposals and the same fate, so both are
off it.
- `pdf_renderer` is DELETED, not disabled. Nothing sets `render_pdf: true` since
markdown became the deliverable (
|
||
|
|
821cbb8622 |
feat(missions): hold every producing phase to delivering, and read markdown instead of PDFs
Two changes the portal review asked for.
1. `benchmark` and `security_hardening` had no delivery guarantee.
`empty_delivery_is_a_failure` tested `kind == "coding"`, on the reasoning that
"research phases legitimately write nothing to the tree" — which the research
directive three modules over contradicts, since it tells the agent to save
findings under /mission/repo/research/. The cost: a `benchmark` mission is ONE
benchmark phase, and with that phase exempt nothing in the platform could fail
it. Same for `security_hardening`, whose first two phases are security_scan and
research.
Now keyed on PRODUCING_KINDS = coding, research, benchmark, security_scan.
`review` stays exempt — a reviewing phase that changes nothing has done its job,
the same distinction `vm_stop_gate::per_node` makes. The test that encoded the
old rule is rewritten rather than deleted, with the reasoning that replaced it.
All 8 harness fixtures are coding phases, so harness behaviour is unchanged.
2. PDFs are dropped; markdown is the deliverable.
Rendering a PDF meant asking an LLM to convert markdown to HTML — a paid API
call per document, on the critical path of "let me read my research", which
failed on depleted Gemini credits and left every artifact unreadable. Styling at
render time is free, offline, instant and cannot 429.
- `mission_outputs` no longer requests a render.
- New `GET /api/missions/{id}/artifacts/{artifact_id}/content`. The frontend had
no way to READ an artifact at all: it listed paths and offered a PDF preview
that never rendered (and whose `rendered_pdf_path` had no route serving it).
Two containment rules, both enforced: the artifact must belong to a mission in
the caller's workspace, and the CANONICALISED path must stay under `_outputs`
— canonicalise first, because checking the string before resolving `..` is the
classic hole.
- `MarkdownBlock` now uses react-markdown + remark-gfm + rehype-slug. It was a
deliberate zero-dep renderer for "the subset the refiner emits", and that
subset stopped matching reality: agent briefs are largely GFM pipe tables,
which it showed as literal pipes. MissionOutputReader and RefineDiffModal use
the same component and gain tables for free.
- Heading ids come from rehype-slug and `outlineOf` slugs with the same
GithubSlugger, so the outline rail's anchors still resolve. A test pins that
invariant, including duplicate headings.
Styles live in globals.css under `.md-view`: the markup is generated so there
are no class hooks, and this project has no styled-jsx registry — the app-router
requirement is documented in next/dist/docs/01-app/02-guides/css-in-js.md, which
frontend/AGENTS.md exists to make me read.
The artifacts tab moved to `MissionArtifacts.tsx`. MissionCanvas was 1341 lines
against a 1250 limit BEFORE this change — already failing lint; it is now 1248.
238 backend lib tests, 20 backend test binaries, 89 frontend tests, clean tsc,
clean eslint on every file touched, production build succeeds.
|
||
|
|
da889f83ab |
fix(missions): an empty repo-less phase was re-processed on every tick forever
The guard added in
|
||
|
|
c28c7a148f |
fix(pdf): the renderer resolved every artifact path against the wrong root
`render_one` joined `missions_root()/<mission_id>/` before the artifact path, producing `<root>/<mission>/_outputs/<mission>/<phase>/...` — the mission id twice, and no such file. Artifact paths are relative to the MISSIONS ROOT. All three registration sites write `_outputs/<mission>/<phase>/...`, and `_outputs` is deliberately a sibling of the per-mission directories so it survives their reaping; joining the mission id first put the lookup inside the very directory `_outputs` exists to escape. It went unnoticed because until now the only artifacts on the system were `code_diff` rows registered with `render_pdf: false`, which this worker never reads. `produces = ["md","pdf"]` was inert, so nothing ever asked for a render. The first artifacts to ask were the first to find it — both failed with ENOENT on the doubled path. Negative control: restore the extra join and `an_artifact_path_resolves_against_the_missions_root` fails. The worker's error handling is sound and needed no change: it recorded `render_pdf_status = 'failed'` with the full path in `render_pdf_error`, which is how this was diagnosed in one read. 237 lib tests pass. |
||
|
|
89bc53b53d |
fix(missions): repo-less capture was publishing the agent's own identity files
First live run of `capture_repo_less_phases`: 9 artifacts, of which 2 were the user's research. The other 7 were AGENTS.md, HEARTBEAT.md, IDENTITY.md, MEMORY.md, SOUL.md, TOOLS.md and USER.md — the agent runtime's identity scaffolding, seeded into the workspace root because that root is pinned to the repository root. The codebase already knew about these files and already had the list. What it did not have is a defence that works without a repo: `ignore_agent_scaffolding` writes them to `.git/info/exclude`, and a mission with no repository has no `.git`. So the exact files that once got committed into a user's repo and pushed (the reason that list exists) came back through a new channel. `AGENT_SCAFFOLDING` is now `pub(crate)` and `mission_outputs` filters on it directly — one list, two consumers, so the next file the runtime starts seeding is excluded from both at once rather than from whichever was remembered. Negative control: replace the filter with `&& true` and `research_documents_are_kept_and_scaffolding_is_not` fails. Found by running it against a live mission, not by reading it. The unit tests passed the whole time — they seeded a tree that did not contain the scaffolding, because I did not know it would be there. |
||
|
|
ceab28b902 |
fix(missions): a repo-less mission threw away everything its agents wrote
`capture_finished_coding_phases` selects `AND m.repo_id IS NOT NULL`. Every
`research_only` mission is repo-less by design (`requires_repo = false`), so the
whole capture path — including the `sync_out` that copies the agent's work OUT
of the container — never ran, and the container was reaped unread.
Measured on the real mission `019fdc35` ("ClawHDF5 Research"): four agents, 9.5
minutes, EIGHT research documents — an HDF5 parser design, a Rust ecosystem
survey, a seven-crate dependency map, tracing and fuzzing strategy. Result:
`mission_artifacts` = 0, mission `completed`. Not recoverable: no container, no
volume, nothing under the missions root.
The platform did not merely fail to save the work — it INSTRUCTED it. The task
preamble tells every agent "/mission/repo ... is the mission's git checkout",
whether or not one exists, and the research directive says to save findings
there. One agent recorded the contradiction verbatim: "No git repo — file is
written." It looked, saw no repo, complied anyway.
Three changes, one per link in that chain:
1. `mission_outputs::capture_repo_less_phases` — copies `/mission/repo` out of
the container and registers each file as an artifact under `_outputs/`,
which is a SIBLING of the mission dir and survives `teardown_container`.
This is also the code that finally reads `produces`, until now an inert key:
`produces = ["md","pdf"]` now drives `render_pdf` into the existing
pdf_renderer worker.
2. The preamble is conditional. A repo-less mission is told its workspace is
scratch, that git_operations has nothing to act on, and — the part that
matters — that files left there ARE collected and published. An agent told
only "there is no repo" has no reason to write anything to disk.
3. A repo-less phase that produced no files is FAILED, unless it declares
`allow_empty`. The same rule `empty_delivery_is_a_failure` applies to coding,
for the only channel these phases have. Note this is NOT that guard widened:
it keys on `files_changed`, which is meaningless with no checkout, and would
not have saved the ClawHDF5 documents.
Negative controls, each ablated and confirmed failing: ignore `has_repo` and the
preamble test fails; empty the skip-list and the capture test keeps `.git` and
`node_modules`; write artifacts inside the mission dir and the survives-the-reap
test fails.
236 lib tests pass.
|
||
|
|
bcf4866abc |
test(harness): a gate that gives up, proven against a real VM
The unit tests prove the plumbing GIVEN `released_at_cap: Some(true)`. They cannot prove the guest writes the marker, that the probe reads it back across the vsock, or that the phase lands `failed` for the right reason — and every one of those is where this class of bug has actually lived. The check is `exit 1`: impossible by construction, so the run exercises the release path rather than hoping to catch it. `blocks` reaching the cap is deliberately NOT the assertion. A healthy agent blocked three times and succeeding on the fourth reports the same 3. The phase STATUS is the assertion; the block count and the failure reason are corroborating checks, so a phase that failed for some unrelated reason cannot pass this. Measured on gw-04 against |
||
|
|
b36ae00ea5 |
fix(missions): a gate that gave up completed the phase green
`done_when_check` is run in exactly one place: the Stop hook inside the guest. Nothing outside it has ever re-run the command — not the evaluator (which judges the PROSE `done_when`), not capture, not delivery. The hook is capped at MAX_BLOCKS so a stuck agent cannot wedge the turn. At the cap it logs `cap: <reason>` and exits 0, releasing the agent with its check still failing. That release was invisible: rc was 0 and the work collected, so both signals the run status was decided from said "fine", and the phase completed. Green phase, unmet condition, no error anywhere — the same silent-success shape this project keeps paying for. The block COUNT cannot fix it. Three blocks then a stop that finally passed and three blocks then a surrender both report `blocks: 3`, and they are opposite outcomes. So the gate now writes a `capped` marker file, probed back out of the guest alongside the block count, and `Some(true)` fails the run on BOTH paths — solo (phase_runner) and composed (microvm_turn_executor). A marker file rather than grepping the log: a block reason embeds the check's own output, so an output line starting `cap:` would read as a release that never happened. Also corrects the comment in `per_node` that sent me looking. It claimed "the phase-level check still runs post-hoc", conflating two mechanisms — that is true of `require_changes` (via `empty_delivery_is_a_failure`) and was never true of `check`. Negative controls, both ablated and confirmed failing: drop the enforcement and `a_node_whose_gate_gave_up_fails_the_run` fails; stop writing the marker and `a_gate_that_gives_up_records_that_it_gave_up` fails. And the control against over-strictness — `a_node_that_was_blocked_and_then_succeeded_passes` — is why this keys on the marker instead of the count. 231 lib tests pass. |
||
|
|
cd4d76a8c3 |
test(harness): pick the done_when wording by measuring the judge, not arguing with it
The microvm scenario's judge assertion failed four runs straight. I blamed the
wording twice and rewrote it twice; the second rewrite made it worse. That was
guessing.
With scripts/judge-eval.sh in place the question is cheap to settle. Three
candidate conditions, three draws each, same evidence and same system prompt:
"its second line is …" MET UNMET MET flaky
"records the kernel version …" MET MET UNMET flaky
"contains both … and …" MET MET MET stable
So it was never noise in general — it is a reproducible weakness with
POSITIONAL and EXCLUSIVE phrasings. "its second line is X and nothing else"
invites this judge to invent requirements about the other lines, which is
exactly the reason it kept citing ("the first line contains 'test result: ok'").
Both fixtures now state what the file CONTAINS. The composed one was checked in
both directions — 3/3 MET on good evidence, 3/3 UNMET when the versions are
missing — because a wording that always answers MET would look stable and prove
nothing.
The eval keeps `kernel-ok` failing on purpose; it is the case production hit,
and tuning it green would turn a measurement into a decoration.
Harness: 24/24, including the assertion that had failed four times.
|
||
|
|
5c066afa7b |
test: stop leaking a container per run, and add the project's first eval
TWO FINDINGS, one from cleaning up and one from refusing to keep guessing. THE LEAK. `./scripts/test.sh` left three containers running every time — 289 had accumulated. The cause was a comment that lied: `warm_pool.rs` said "Shutdown destroys assigned AND pooled sandboxes", while `SandboxManager::shutdown` drains the POOL only. Its own doc says why — assigned sandboxes persist deliberately so a redeploy can reuse them, and production reaps the strays with `reconcile_orphans` at boot. A test has no next boot, so each one that assigned a sandbox simply left it running. The three tests now call the `release_agent` that already existed, and the comment says what the code does. Verified: 0 leaked, where the same run leaked 3 before. THE EVAL. The independent judge failed the same correct phase FOUR times, each time citing a different invented requirement. I blamed the condition's wording twice and rewrote it twice — the second rewrite made it worse, by naming a command a tool-using judge then ran in its own container. Then a control showed the same model answering MET to the same question asked directly, and a third wording test showed a STRICTER phrasing scoring UNMET. Prose wording was not the variable. Continuing to iterate would have been fitting the fixture to noise. `scripts/judge-eval.sh` measures the thing instead: five cases drawn from real incidents, each with an answer a careful human would agree with. This project has 557 tests and had zero evals, which is backwards — a test pins OUR code, an eval pins the MODEL, and the model changes without us touching anything. The result is why it was worth building: glm-4.7 4/5 — wrong on kernel-ok: says UNMET when MET kimi-for-coding 4/5 — wrong on goodhart: says MET when UNMET Identical scores, opposite failure modes. GLM fails good work; KIMI passes work where 14 assertions were deleted and the failing module removed to make a suite "pass" — the exact incident the verifying judge was built after. Swapping the validator to Kimi because it passes our failing case would have installed a rubber stamp. Keep GLM: a judge that is too strict costs a re-run, a judge that is too lenient costs the guarantee. The eval also caught a bug in itself before I trusted it: Kimi answers with a `thinking` block first, and a 160-token budget was consumed entirely by it, which the harness scored as NO-ANSWER. An eval that misreads a model is worse than no eval, so it now reads thinking blocks as a fallback and has room to answer. 557 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
d24823b6f3 |
fix(missions): a failed phase stranded its mission at running forever
Found by counting containers during a cleanup, not by a test. gw-04 was holding
a per-mission runtime container for a mission whose only topology run had failed
three days earlier — phases `pending,failed`, mission still `running`.
The interaction, which lived entirely between two queries' predicates:
`start_pending_phases` launches a phase only when EVERY lower-order phase is
`completed`, so once one fails the phases after it can never run. They stayed
`pending`. `close_finished_missions` closes a mission only when NO phase is
outside ('completed','failed','skipped') — so a `pending` phase that would never
run kept the mission `running` indefinitely. And `mission_runtime`'s sweeper
fires N minutes after a TERMINAL state, so the container was never reaped.
One leaked container per failed multi-phase mission, accumulating silently, with
nothing in any log saying so. Neither query is wrong alone; the bug is that
nothing marked the phases the failure had made unreachable.
`skip_unreachable_phases` says it: a `pending` phase with a `failed` phase at a
LOWER order_idx becomes `skipped` — strictly earlier, because order is what makes
a phase unreachable, and a failure later in the list says nothing about one still
queued ahead of it. `skipped` is not a new concept: `close_finished_missions`
already treats it as terminal, and it is the honest word for a phase that was
never run, as distinct from one that failed.
RETRY HAD TO MOVE WITH IT, or this trades one bug for another. `retry_phase`
required the mission to be `running`, so closing failed missions would have made
the one outcome you would actually want to retry the one you could not. It now
accepts `failed` too, and in one transaction: resets the phase, REOPENS the
phases its failure had skipped (without that, a retry runs the phase and stops,
because everything after it is terminal-by-skip), and puts the mission back to
`running` — every launcher and closer keys off that status. `completed` and
`cancelled` stay refused; reopening those is a different decision.
557 tests pass, clippy clean. Three DB tests against real SQL, including that a
phase queued BEFORE the failure is untouched and that a draft's phases are never
swept.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
bf2055e725 |
fix(evaluator): the anti-Goodhart clause was failing work that RECORDS a value
Three consecutive production verdicts failed a phase that had done exactly what
its condition asked, each time with a different invented reason: "6.1.128 is not
a kernel release string like 'Linux 6.1.128'", then "line 2 should be 27.0.0",
then "line 1 must be empty or unrelated". I reworded the condition twice, and the
second rewording made it worse.
THE CONTROL THAT SETTLED IT: asked the same question with the same file and the
same condition — but WITHOUT our system prompt — glm-4.7 answered MET, citing the
exact line. The model judges this correctly. Our prompt does not.
The cause is a clause we wrote on purpose. `EVAL_SYSTEM_VERIFYING` is
deliberately adversarial because an earlier evidence-only judge was gamed by an
agent that emitted the string the judge had asked for, and it says to fail "a
required string or value hard-coded, stubbed, or printed rather than produced by
working code". A condition asking for a kernel version to be written into a file
IS that shape, read literally. The judge was obeying us.
Two clauses now, because each without the other is a known failure:
- the trap stays: work that satisfies the letter and not the purpose — tests
weakened, assertions fitted to wrong output, values stubbed — is not met.
- some conditions are satisfied BY a recorded value, and for those, writing the
value IS the work: a measured baseline, a scan report, a recorded environment
fact. Hard-coding is cheating only when the condition is about behaviour code
must produce.
And the other failure from those three verdicts: "judge the condition AS WRITTEN;
do not re-derive the expected value yourself" — a condition may describe a
DIFFERENT machine, an earlier run, or a remote environment, and the value the
judge would measure where it stands is not the one under judgement. That is
exactly what produced "line 2 should be 27.0.0": a tool-using judge ran `uname`
in its own container and compared.
This is not a niche fixture problem. The model-authored plans shipped today write
BASELINE.md and security-findings.md and gate on them — every one of those is a
recorded-value condition, and every one would have been rejected.
555 tests pass, clippy clean. A test pins both clauses, since removing either
reintroduces a failure this project has already paid for.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
72f8bdc87c |
test(harness): a done_when naming a COMMAND invites the judge to run it
My previous attempt at this made it worse, which is the useful part. The condition said "a Linux kernel release string" and the judge rejected `6.1.128` as "not a Linux kernel release string such as 'Linux 6.1.128'". I rewrote it as "the exact output of `uname -r`" — and the next verdict was that line 2 should be `27.0.0`. The judge has a sandbox and allow-listed commands, so naming a command told it to RUN that command, in ITS OWN container, and compare the file against the answer it got there. The file records a microVM's kernel; the judge was comparing it against the machine the judge runs on. Those are different machines by design — that is the entire point of the assertion. So a `done_when` for a tool-using judge must describe the VALUE's shape, never a command that produces it: "a bare kernel version of the form MAJOR.MINOR.PATCH (for example 6.1.128) and nothing else", plus an explicit instruction not to run uname and not to compare against the local machine, because the file records a different one. The general rule, worth carrying into how `done_when` is written anywhere: a condition phrased as "the output of X" is ambiguous about WHERE X runs, and a judge with tools resolves that ambiguity by running X where it stands. Conditions about a remote or past environment must be stated as properties of the recorded value. The scenario's real proof that the agent ran in a guest is unchanged: a separate comparison of that line against the actual gateway and node kernels, which has passed on every run including the two where the judge disagreed. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
e2f576ec02 |
fix(evaluator): the judge verifies a COPY, never the mission's own tree
The harness's uid probe caught this the moment cross-provider validation came back: `checkout has multiple writers (uids=0,65532)`. All 63 root-owned files were under `repo/target/`. The mechanism, confirmed rather than guessed: the judge's verification sandbox execs into `clawmates-runtime`, which runs as ROOT with the missions root bind-mounted, and its workdir was the mission's LIVE checkout. So when the judge ran `cargo test` to check a condition — which is the entire point of the verifying evaluator — cargo wrote `target/` into the checkout as uid 0, in a tree otherwise owned by the server. The next phase's `cargo` would then hit permission-denied on a directory it cannot write, which is the uid-split failure class copy mode exists to eliminate. IT WAS LATENT ALL DAY. While the z.ai credential was dead the judge never ran a single check, so the uid probe kept passing; restoring the credential surfaced it on the first gated mission. A guard that only holds while a dependency is broken is not a guard, and this one was only visible because the harness measures the invariant rather than the feature. Running the checks as the checkout's uid was the obvious fix and is the wrong one: `CARGO_HOME` is root-owned 0755 in that image, so a non-root uid fails, and the evaluator treats "could not run" as unverified — trading a polluted tree for phases that fail closed for a reason unrelated to their work. So the sandbox verifies a copy, made through `mission_fs::pack_dir` so it carries exactly what the delivered diff carries (no `target/`, no `node_modules/`) — one exclusion list, three consumers. The copy lives at `_verify/<mission>`, a sibling of the swept per-mission directories, and is removed on drop. This is the rule the codebase already applies to the `verifier` subagent, which has no Edit and no Write, stated for the judge: verification must not mutate what it verifies. A judge that can change the tree it is judging can make its own verdict true. NEGATIVE CONTROL, run: pointing the sandbox back at the live checkout fails `the_judge_verifies_a_copy_and_never_the_mission_tree`. The test seam (`Sandbox::at`) is never `owned` and never deletes, so a destructive constructor cannot masquerade as a plain one. 553 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
1b556c5849 |
test(harness): say what the condition means, after the judge read it strictly
The restored GLM judge failed a phase that had done the work: MICROVM.md existed with two lines and the second was `6.1.128`, and the verdict was "a kernel version number, not a Linux kernel release string such as 'Linux 6.1.128'". The judge is wrong on the fact — `6.1.128` is exactly what `uname -r` prints, and "release" is the term for it — but the CONDITION was ambiguous, and it is our fixture. "A Linux kernel release string" can be read as either `uname -r` output or `Linux x.y.z`, and a stricter reader is entitled to the second. Both scenarios now say what they mean: the exact output of `uname -r`, a bare version, no prefix. This is not weakening the assertion. The scenario's own kernel check — the one that proves the agent ran in a guest rather than on a host — is a separate, unchanged comparison against the real host kernels, and it PASSED on the same run. What changed is only that the mission-level `done_when` now describes an observable fact precisely, which is what this codebase's own plan-authoring prompt tells models to do. Worth recording rather than papering over: an over-strict independent judge is a much safer failure mode than an over-lenient one, and this is evidence the judge READS the tree instead of rubber-stamping it — the Goodhart incident that motivated cross-provider validation was the opposite failure. But it does mean a vague `done_when` can now cost a phase, which raises the value of `done_when_check` (a shell command, judged by exit status) for anything mechanical. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
521da9feb9 |
fix(deploy): the verify step is the authority, not the recreate
`scripts/deploy.sh` reported failure twice this afternoon for deploys that had succeeded. Both times the 60-second rolling timer rolled the stack onto the same `:latest` first, and the script's own `docker-compose up` then hit a container-name conflict — "already in use" once, "Renaming a container with the same name" the other — for a container the timer had already recreated correctly. A deploy signal an operator has to second-guess is precisely what this script exists to prevent. Its original reason for being was a green edge on a stale image; crying wolf trains people to ignore the alarm, which gets you the same outcome by a different route. The recreate is now best-effort and says so when it fails, and the VERIFY step decides — it compares the RUNNING image id against the resolved `:latest`, which is the only question that matters and is unaffected by which process did the roll. A genuinely failed deploy still fails there, because that check never depended on the recreate succeeding. Both false alarms were settled by hand with the binary grep (`docker exec … grep -a -c "<string only in the new code>"`), which remains the strongest check when the image id is in doubt. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
3300c9d149 |
feat(missions): say at boot whether the independent judge can be reached
The z.ai credential expired mid-session and the first symptom was a two-phase mission failing after BOTH its VMs had run — the phase completed, delivered, pushed, and then one evaluation row said "the independent validator could not be reached this pass". `cross_provider_judge` refusing to fall back to the agent's own provider is correct: a verdict from the same family is not an independent check, and producing one quietly would claim a property the verdict does not have. The cost of that refusal is that a dead validator makes EVERY `done_when` phase unmeetable — and the information needed to know that existed from the moment the server booted. Nobody was told until it was expensive. The sibling of `runtime_preflight`, and the same stance: a report, not a gate. The server must still boot with a broken validator — refusing to start turns a degraded deployment into a dead one, and a mission that opts out (`validator_model = ''`) is unaffected. Two faults, kept distinguishable because they send an operator to different places: `Unregistered` (no provider by that name — the evaluator will refuse it rather than judge with the default, so register one) versus `Unreachable` (it resolved and the call failed — fix the credential). Collapsing them into "the validator is broken" is the kind of merge that costs an hour. The probe is a real completion through `Runtime::complete` — the same resolve-then-stream path the judge itself takes. A models-list or a HEAD would pass for an expired key, a revoked key, and a key with no quota, which are exactly the cases worth catching; and a probe that dialled the provider its own way could pass while the real call fails. `NotConfigured` is reported too, and not as an error: a deployment may choose the house model. It is still worth saying out loud that the check running is not an independent one. 551 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
08dd227a45 |
feat(missions): give the planner the repository's contents, not just its names
The root listing was not enough. Given names alone the planner wrote "optimise the hot path" for a crate whose hot path is `add(a: i64, b: i64) -> i64` — a mission that was unachievable from the moment it was written, and that nothing discovered until an agent had built a benchmark harness in a VM to measure an integer addition, honestly reported no improvement was possible, and the judge correctly failed the phase. `repo_digest` fetches the whole tree (so "does this have benches/" is a fact, not an inference) and then file CONTENTS in priority order: manifests first — they say what the project is — then the README, then source ascending by size, since a planner learns more from twenty small files than from one large one. Lockfiles and build output are dropped: enormous, and they say nothing a manifest does not. THE RULE THIS ENFORCES, and the reason the rendering is its own tested module: a digest of any repository worth planning against is partial, and a model shown a partial view without being told it is partial plans as though it saw everything. So every omission is stated — how many files exist, how many were shown, what was cut from each, and "anything not shown you have NOT seen". Same distinction as `Option<u32>` for the subagent probe: "we did not look" and "there is nothing there" are different facts. Failures degrade to a stated absence rather than an empty string, and the three cases stay distinguishable: no repository, a tree that could not be read, and a tree read but no contents fetched. An unreadable tree is never rendered as an empty repository. Two more things the prompt now says, both learned from that run: plan for the repository as it IS rather than as the description implies (and if the description asks for something the code cannot support, say so in the task and plan the phase that establishes the truth, rather than a phase that must fail); and a mission agent has NO package-registry access. The agent discovered the second one mid-run and wrote a dependency-free `std::time::Instant` harness after Criterion could not be added — good adaptation, but nothing had warned it. 549 tests pass, clippy clean. The budget/priority/truncation logic is pure and tested; only the fetching touches the network. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
0aeae07db2 |
fix(missions): the planner was planning blind — show it the repository
The first real plan opened with "Identify the crate's hottest code path and run its benchmark harness". This crate has no benchmark harness. The phase ran, found nothing to baseline, delivered zero files, and the plan's second phase was left with nothing to optimise against. The planner saw the mission title, the description, and a boolean for whether a repository was bound. It never saw the repository. A plan about a codebase written without looking at the codebase is a guess that reads like a plan — and the failure surfaces two phases and one VM boot later, as an agent reporting that the thing it was told to run does not exist. The prompt now carries the repository's root listing, read from the FORGE rather than a checkout: at proposal time the mission is still a draft and `ensure_checkout` has not run, so there is nothing on disk to list. It also says outright that a phase needing something absent must CREATE it and say so in its task — the failure was not only ignorance of the tree but the assumption that missing tooling is someone else's problem. A listing that cannot be fetched degrades to "(the repository listing could not be read)" in the prompt rather than to an empty string. A model told the listing is unavailable can hedge; a model told nothing assumes — which is the same distinction as `Option<u32>` for the subagent probe, in a prompt instead of a struct. Found by running the thing end to end rather than by testing it: every unit test here passes with a planner that has never seen a repository. 543 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
a33dbdcdc3 |
feat(missions): W1/#13 — let a model author the mission's phases
The last unstarted item from the missions-as-workflows plan, and the other half
of Slice 5: that one lets a model size the TEAM, this lets it decide what the
work IS.
Every mission's phases come from one of five hand-written recipes in
`templates/workflows/*.toml`, chosen by `template_kind` before anyone saw the
mission. That is the "do it this way: 1, 2, 3" over-specification that makes a
capable model follow a worse plan than it would have chosen. The recipes stay —
they are still the default for a mission nobody proposes a plan for, and the
fallback when a proposal is refused.
Same three verbs and the same review gate as the roster, deliberately: propose
and decide are separate because only the second changes a mission, and a second
shape would be a second thing to get right. Approving REPLACES the phases (a
plan is an answer to "what is this mission", not an addition to one), draft-only.
GROUNDED IN WHAT THE PLATFORM ACTUALLY READS, which is the part that makes this
more than a copy. `phase_config::KNOWN_KEYS` already names every phase-config key
and the code that reads it — the registry built after `task` sat unread through
every mission. A plan is validated against it, so a model cannot propose a phase
whose settings nothing will act on: the failure that registry exists to EXPOSE is
one this path cannot create. Phase kinds are checked the same way, because an
unknown kind does not error — it falls through to the catch-all purpose and runs
as a generic phase that looks like it worked.
TWO THINGS THE WORK ITSELF FOUND, both the same shape:
- `done_when_check` — the stop-gate key added earlier today — was never
registered in `phase_config`, so every mission that set it has been logging
it as an unknown key. Found by a test written for a different purpose, which
is the registry doing exactly its job. Now registered with its reader.
- `done_when` and `max_iterations` are COLUMNS promoted out of config by
`missions::create`; the evaluator sweep filters on the column in SQL every
tick. My first insert wrote the config blob alone, which would have stored a
plan's completion condition where nothing judges it. NEGATIVE CONTROL run:
binding NULL instead of the promoted value fails
`an_approved_plan_replaces_the_missions_phases`.
`order_idx` comes from the array's own order rather than a field the model sets:
two sources for one fact is how a plan ends up with two phase 0s, and order_idx
is what `start_pending_phases` sequences on.
MAX_PHASES is 4 and the prompt argues for one. Each phase is a full agent run in
sequence, and splitting one change into plan → implement → test is the documented
anti-pattern — a single agent doing all three keeps the context that makes the
later steps good.
543 tests pass, clippy clean. Migration 0072.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
a48d78f8eb |
test(missions): a composed node is offered the same help as a solo one
Every composed run so far reports `subagents: 0`, and the honest question is whether that is the tasks being small or the capability being absent. It is the former, and this is what says so: a composed node's task text is built by `microvm_turn_executor` and then wrapped by the SAME `vm_prompt` inside `run_inside`, so one prompt builder serves both paths and both carry the `Agent` tool offer and the `verifier` / `explorer` roles. Asserted rather than left to code reading, because if someone gave composed nodes their own prompt without the offer, the difference would show up only as a count nobody was watching — and "the graph fanned out but no node did" is indistinguishable from "no node needed to". The roles are read from `agent_definitions()` rather than spelled out, so adding a role without mentioning it in the prompt fails here instead of shipping a role the lead is never told about. 535 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
742724e53c |
feat(fleet): Kimi as a microVM backend — the URL settled by measurement
The base URL took three measurements to find, and the first two were wrong in instructive ways. `api.moonshot.ai/anthropic/v1/messages` EXISTS and speaks the protocol — it answers with Moonshot's own structured error rather than a 404. It also rejects an `sk-kimi-` key, because it belongs to the platform.moonshot.ai account namespace. Two endpoints that both "work" for different accounts is precisely the shape that makes a guessed URL look like a broken key, and it is why this was refused rather than guessed for as long as it was. The Kimi CODE service is the one an `sk-kimi-` key belongs to: `POST https://api.kimi.com/coding/v1/messages` returns a real Anthropic Messages body — `msg_` id, `content` blocks, a `thinking` block with a signature. So `ANTHROPIC_BASE_URL=https://api.kimi.com/coding`, WITHOUT the `/v1`: Claude Code appends `/v1/messages` itself, and `/v1/v1/messages` would 404 in a way that reads as a broken image rather than a bad URL. Two more measured, each otherwise a silent failure at the first turn: `Authorization: Bearer` is accepted (so ANTHROPIC_AUTH_TOKEN is the right injection channel), and a `claude-*` model id is ACCEPTED AND ANSWERED — Kimi maps it onto `kimi-for-coding` exactly as z.ai does, so no ANTHROPIC_MODEL override is needed. Claude Code rather than Moonshot's own `kimi` CLI, deliberately. The mission harness is Claude-Code-shaped throughout: `--agents` JSON roles, the verifier's tool allowlist, the `Stop` hook behind the completion gate, the per-subagent transcripts counted as delegation evidence. `kimi` has none of those flags — its equivalents are TOML files and markdown agent dirs — so using it would mean a second executor with its own untested failure modes. TWO STALE MAPS, caught by the rootfs harness refusing to bless the image: both `fc-build-rootfs.sh` and the node's `required_cli` expected backend `kimi` to contain Moonshot's `kimi` binary. That assumption predates the measurement, and it failed a rootfs that was correct. Both now say `claude` for glm and kimi alike — the binary is the same in all three images; only the endpoint differs. Egress for `kimi` is `api.kimi.com` alone: not moonshot.ai (wrong namespace), not z.ai, not Anthropic. Asserted both ways, like the other two. The image and rootfs are built on tank and the rootfs passes all four checks (boots, git, writable /mission, `claude --version`). 534 tests, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
d3a53e7bf1 |
fix(fleet): a VM reaches its OWN provider and no other, measured not assumed
The GLM backend works — and proving it produced a better boundary than the one I shipped an hour ago. WHAT THE FIRST GLM MISSION SHOWED. It completed, and the delivered file said the model was "claude-opus-5". The node's egress log said the VM had dialled `api.anthropic.com` five times before `api.z.ai`. Either reading alone is consistent with a "GLM backend" that silently runs Anthropic — the exact silent-success shape this project keeps closing — so I did not accept either. THE ABLATION, run on tank rather than reasoned about: deny `anthropic.com` at the proxy and run the same mission again. It **completed**, dialling only `api.z.ai`. So the completions genuinely come from z.ai; Claude Code's calls to anthropic.com are its own telemetry, not its model traffic. And that same agent — served exclusively by z.ai, with Anthropic unreachable — still described itself as "Claude Opus 5 (1M context)". **A model's account of which model it is has no evidential value.** The proxy's log of which host it dialled does. This is the `uname -r` lesson again in a new place: ask the infrastructure, not the agent. So the allow-list is now PER BACKEND rather than a union: a `claude` VM reaches Anthropic and the forge, a `glm` VM reaches z.ai and the forge, and neither can reach the other's endpoint. A union was defensible when it was one host; once the measurement showed a GLM VM never needs Anthropic, keeping it would mean a credential mix-up upstream could still put one provider's secret on another provider's wire. Now it fails at a closed door instead. An unknown backend gets the forge and NO model API — it cannot run anyway, and borrowing somebody else's door is the failure this split prevents. An explicit `CLAWMATES_FC_EGRESS_ALLOW` still wins outright: an operator who set it drew a boundary on purpose. `DEFAULT_ALLOW` is deleted rather than left beside the new function, so there is one answer to "what may a mission reach" and not two. 534 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
f7f3dfe495 |
feat(fleet): GLM as a real microVM backend, and per-role models for claws
Three threads, all of which end at the same place: a mission whose verifier does
not share a model with the coder it reviews.
**GLM has a credential contract now.** `microvm_credential_for` returned one env
var name, which quietly assumed every provider reads its secret from the same
place Anthropic does. It returns a `Credential { source, target }` instead —
z.ai's key lives in the server's `ZAI_API_KEY` and Claude Code reads it as
`ANTHROPIC_AUTH_TOKEN`, and collapsing those two names is what forces a guess at
the other end. A wrong guess here sends one provider's credential to another
provider's endpoint.
`images/agent-glm` is the same CLI at the same pinned version as `agent-claude`
with `ANTHROPIC_BASE_URL` baked in. The split is deliberate: the ENDPOINT is a
property of the image, the CREDENTIAL is a property of the turn. That makes the
dangerous mix-up unrepresentable — a GLM VM cannot be handed an Anthropic
subscription token, and a claude VM cannot be pointed at z.ai. Asserted both
ways, because "the GLM VM must not carry CLAUDE_CODE_OAUTH_TOKEN" is the
property that costs a credential if it ever stops holding.
Kimi stays refused. `KIMI_API_KEY` is set and Moonshot serves an
Anthropic-compatible API, but I have not verified its base URL against the
running service, and this function is precisely where guessing a URL is
expensive. It becomes an arm the day someone measures it.
`api.z.ai` joins the node's default egress allow-list. A default that cannot
run the images we ship is a trap rather than a policy — the alternative is an
operator discovering it as a hung agent with no model access.
**Per-role models for claws** (migration 0071). `template_roles` had no model
column, so `mint_team_from_template` bound every role of every mission team to
one literal — a template whose whole point is an independent reviewer minted a
reviewer sharing a model with the coder. A role may now name its own; roles that
say nothing still take the mint's default, so every template written before this
behaves exactly as it did. The literal is now that default rather than a
hardcode.
**A harness scenario for the roster flow.** `verify-mission-delivery.sh roster`
runs the whole Slice 5 loop — planner proposes, human approves, mission runs —
and asserts the roster LANDED on the mission row rather than trusting the API's
answer. That distinction is not theoretical: the first live approval returned an
error while leaving the proposal marked approved.
Built and proven on tank ahead of the deploy: `clawmates/agent-glm:dev` reports
`2.1.223` and `BASE=https://api.z.ai/api/anthropic`, and
`fc-build-rootfs.sh … glm 8G` boots a VM from it that has git, can write
/mission, and answers `claude --version`.
533 tests pass, clippy clean. Migration 0071.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
75d09241fb |
fix(missions): the first real approval found two bugs the tests could not
Deploying Slice 5 and approving one roster in production broke it twice, in ways
528 green tests had nothing to say about.
**1. `jsonb_set` refuses a scalar.** A mission created through the API without a
`config` stores jsonb `null` — a scalar — and `jsonb_set` fails on it with
"cannot set path in scalar". The guard was `coalesce(config, '{}')`, which
protects against SQL NULL; this is a perfectly good JSON null of the wrong shape,
and coalesce passes it straight through. Every test wrote `'{}'::jsonb` because
that is what a test author types. Production types nothing at all.
**2. The approval was not atomic, and failing halfway is permanent.** The claim
and the mission write were two statements, claim first, so when the write failed
the proposal stood `approved` with nothing applied — and the partial unique index
then makes that state unrecoverable: no other proposal for that mission can ever
be approved. The mission ran solo with `team_engine` still NULL while its
proposal said otherwise.
`approve_and_apply` is now one transaction: claim, write, commit or roll back.
The type guard is `CASE WHEN jsonb_typeof(config) = 'object' THEN config ELSE
'{}'::jsonb END`, which answers the question that was actually being asked.
Both regressions are tested in the shape production had, and both NEGATIVE
CONTROLS were run rather than assumed:
- restore `coalesce` → `a_roster_applies_to_a_mission_whose_config_is_json_null`
FAILS with Postgres's own "cannot set path in scalar", the exact production
error.
- commit instead of roll back on a failed apply →
`a_failed_apply_leaves_the_proposal_undecided` FAILS with the proposal stuck
`approved`.
Worth stating plainly: the API returned 500 for that approval, so this was not
silent to the caller — but the row it left behind claimed the mission had a
roster it never received, and the mission then ran and delivered, which is the
shape that gets believed.
530 tests pass, clippy clean.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
aa470091aa |
fix(missions): a bootable rootfs is not a runnable one
Found by looking at what the fleet actually reports, not by reasoning about it: tank's `capabilities.rootfs` is `["agent-terminal", "claude", "default"]`. Slice 5 offered that list to the planner as the menu of backends and validated proposals against it — so a roster naming `agent-terminal` would have been proposed, validated, approved and launched, and then failed at the agent turn, because `microvm_credential_for` has no contract for it and refuses rather than forward an Anthropic subscription token to an unknown endpoint. Refusing at boot is correct and is exactly the wrong PLACE: it is three steps and one human approval after the point where the answer was already knowable. The menu is now the intersection of "a node can boot it" and "a mission agent can authenticate in it", which is what the question meant all along. `backend_can_run_a_mission` derives from the credential contract rather than restating it, so a backend gaining one (GLM and Kimi, when B4.6's base-URL contract is settled) becomes proposable in the same commit that makes it runnable — instead of in a second list someone has to remember. 528 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
1797669296 |
feat(missions): Slice 5 — let a model size the mission's team
`routes/planner.rs` has had Opus proposing rosters since the Master Planner
shipped, and none of it ever reached a mission: the proposal lived in React state
and died with the tab. A mission's shape came from a team template instead —
fixed roles, and every claw minted `claude-sonnet-5` from a literal in
`mint_team_from_template`. That literal is why no mission has ever run more than
one provider.
A roster is `(topology_kind, [(role, backend)])`, which is exactly what the
composed executor already consumes: `Roster::graph` builds a `TopologyGraph` with
the backend in `attrs`, and `MicroVmTurnExecutor` reads `attrs["backend"]` per
node. So a verifier on another provider's rootfs stops being a bolt-on and
becomes a graph node — the correlated-failure break the independent judge exists
for, one layer down.
Three verbs, and the split is the point. **suggest** asks the model and persists
the answer, changing nothing. **decide** approves (writes `config.roster` and
switches the mission to the composed engine) or rejects. A proposal is never
applied on arrival: a model sizing a team is a suggestion about how many VMs to
boot, and this codebase treats model output that costs money as evidence for a
decision, not the decision.
Fail-closed at every seam, because each of these otherwise surfaces much later
and much more expensively:
- a backend no ONLINE node can boot is refused when PROPOSED, naming the ones
the fleet actually has. Placement would refuse it too — at launch, after the
roster was approved and someone believed the mission would run. The model is
handed that same list in its prompt, so the usual case never arises.
- an invented `topology_kind` is refused, not defaulted. `parse_topology_kind`
defaults to hub-spoke, which is right for a template we wrote and wrong for a
string a model just produced: running a `pipeline` proposal as a hub-and-spoke
changes what every node sees and nothing would say so.
- the roster is validated BEFORE it is stored, so a stored proposal is always
one that could be approved; and again at approval, against the fleet as it is
then — a node can go offline in between.
- `MAX_MEMBERS = 6`. Each member is a whole VM, not a subagent, and a model
asked to size a team proposes twelve happily.
Two properties live in SQL rather than in the handler: at most one approved
roster per mission (partial unique index — two approved rosters are two answers
to "what shape is this mission", and the executor reads one field), and
decide-once (`WHERE status = 'proposed'`, so a double-clicked approve claims
nothing the second time). Both tested against a real database, including that the
second approval is refused by Postgres rather than merely losing a race.
NEGATIVE CONTROL, run rather than assumed: with the roster preference removed
from `composed_graph`, `an_approved_roster_outranks_the_template` FAILS — 3 nodes
from the template instead of the roster's 2. A stored roster that is silently
ignored at launch is precisely the shape this project keeps paying for.
Not closed: per-role models for CLAWS. `template_roles` has no model column, so a
ZeroClaw team still mints one model for every role. The literal is now a named
constant that says so and points at the roster path, rather than sitting inline
where nobody reads it.
527 tests pass, clippy clean. Migration 0070. Not yet exercised against the
deployed stack — the route has never been called with a live model.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
abb97e6f03 |
test(harness): a composed scenario, and the stop gate asserted in a real VM
`verify-mission-delivery.sh composed` runs a `team_engine=composed` mission and checks the one property that cannot be checked any other way: a VM is inject → run → collect → destroy, so unless the tree is carried node to node, node 2 boots from the original checkout, sees nothing of node 1's work, and still reports success. The task makes each node append ONE line to STAGES.md, so the delivered file IS the evidence — a run that lost the handoff delivers one line, and no amount of agent confidence can fabricate the missing ones. It also asserts the run's tier is `microvm_graph`. A composed mission that quietly fell back to the solo path would deliver a one-line file and look exactly like a graph that ran one node. `assert_stop_gate` reads the count `phase_runner` reports and distinguishes three outcomes that matter: a number (installed, fired that often), `0` (installed, never needed), and `-` (could NOT be installed — usually a CLI in the image with no `--settings`). Wired into the microvm scenario rather than its own, because it applies to every coding phase on that path. Both ran against the deployed stack: composed — 4/4. STAGES.md carried 5 stage lines through 5 separate VMs (planner → coder → tester → reviewer → committer), each stamped with the guest kernel 6.1.128 rather than the gateway's 6.8.0 or the node's 7.0.0. The run checkpointed 5 steps on the worker, and `updated_at` stayed ~2s old mid-turn, which is the keepalive doing its job — without it `requeue_stale` flips a live run at 180 seconds. microvm — 6/6, including the gate installed in a real VM (`blocks: 0`), one subagent, the GLM judge, and the unavailable-backend negative control. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
6991e21f94 |
feat(missions): the completion gate, moved into the agent's own loop
Every check this platform makes on a phase runs AFTER the agent has stopped: the
evaluator judges `done_when`, capture notices a coding phase delivered nothing,
and either verdict costs a whole new VM — a fresh boot, a fresh inject, and an
agent starting over with none of the context that got it that far. Meanwhile the
documented failure mode of a long-running agent is that it stops too early.
MEASURED FIRST, because the plan's chosen seam does not exist here. Probing every
hook name under `claude -p` (2.1.222, hermetic `--settings` file): `SessionStart`,
`UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `SubagentStop` and `Stop` fire;
`TaskCreated`, `TaskCompleted`, `TeammateIdle`, `SessionEnd`, `Notification` and
`PreCompact` do not. The agent-teams hooks Slice 3 deferred are inert on our path
BY CONSTRUCTION — no team forms in print mode at all — so `done_when` could never
have been wired through `TaskCompleted` exit 2. `Stop` is the seam.
`vm_stop_gate` generates a POSIX `sh` hook installed via `--settings`, under
`/root/gate` and never under `/mission/repo` (anything there is collected and
arrives in the user's delivered patch). It refuses a stop when:
- the phase must deliver and the repository is untouched — asked as TWO
questions, since an agent that committed leaves a clean tree and an agent
that did not leaves HEAD alone; only both together mean nothing happened;
- `config.done_when_check` — a command the phase author wrote — exits nonzero,
in which case its OUTPUT is the feedback, not just "the check failed".
Deliberately mechanical. NOT the `done_when` verdict: that is an LLM judgement
made host-side by a different provider on purpose, and re-running it inside the
VM would put the agent's own environment in charge of grading the agent — the
correlated failure the independent judge exists to break.
THE CAP IS LOAD-BEARING. Without a ceiling a stuck agent is blocked, retries, is
blocked again, and burns the hour-long turn budget instead of failing visibly.
After 3 blocks the gate lets it stop, records that it gave up, and leaves the
verdict to the existing post-hoc path, which is unchanged.
PROVEN AGAINST A LIVE AGENT with the REAL generated artifacts, not a paraphrase:
- a read-only task → blocked 3 times with our exact message, released at
exactly the cap, and the agent took the escape hatch the message offers
("if the task genuinely requires no code change, say so explicitly") rather
than touching a file to satisfy the gate. It did not Goodhart it.
- a task that needs an edit → `blocks: 0`, log says `pass`. No false positives.
Two things that could fail silently, both closed. `--settings` is PROBED in the
image before use (`claude --help | grep`), because an unknown option is a hard
CLI error that would turn every gated phase into a failed one; a build without
it degrades to ungated and says so, since losing a check is better than losing
the work. And `stop_blocks` is reported out of the guest — `None` for no gate,
`0` for got-it-right-first-time — so a gate that never fires is distinguishable
from one that was never installed.
`require_changes` does NOT apply per node on a composed run: a graph's verifier
node is SUPPOSED to leave the tree alone, and a per-node gate would refuse its
stop three times for doing its job. `StopGate::per_node` drops it and keeps the
declared check. The phase-level rule still runs post-hoc against what the last
node collected.
An ungated phase's command is byte-identical to before, asserted by test — most
phases are gated, so the ungated path is the one nobody would notice breaking.
517 tests pass, clippy clean.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
1d554396f4 |
fix(delivery): four failures from the #55 trace — auth, prompts, truncation, retry
All four were surfaced while tracing #55 and left open. Each one on its own is
small; together they are why a two-line git rejection took hours to read.
**1. `with_ambient_auth` failed open.** It matched one literal prefix,
`https://git.redclaw.dev/`, and returned the URL unchanged for everything else
with no log line. An `http://` remote, an explicit port, a different case in the
host, an ssh remote, a URL that already carried userinfo — all came back
unauthenticated and looked identical to success. It now returns `Authed`, which
carries the URL AND why no credential reached it, and recognises the forge in
every shape a remote can be written (host parsed with userinfo stripped BEFORE
the port, or `oauth2:token@host` reports its username as the host — the first
version of this function did exactly that and failed its own test).
**2. Nothing set `GIT_TERMINAL_PROMPT=0`.** So a credential-less URL did not
fail — git opened `/dev/tty`, and in a server container that surfaces as
`No such device or address`, several layers from the missing token. Now set on
every git invocation that can reach the network. And `push_url_for` refuses
outright when the URL is on OUR forge and unauthenticated: that push cannot
succeed, and letting it proceed only buys a symptom that looks like something
else.
**3. The truncation fix went to the wrong path.**
|
||
|
|
12147a1e01 |
feat(missions): Slice 4 — the two engines composed, with the file handoff proven
`team_engine='composed'` (the third name migration 0069 anticipated) runs a
mission as a durable ZeroClaw graph whose every node is a whole
Claude-Code-in-a-microVM session. Engine Z owns checkpoint/resume, cancellation
and per-node heterogeneity; Engine C owns shared context and cheap fan-out;
neither has the other's asset, which is why this is a composition and not a
compromise.
`MicroVmTurnExecutor` implements the existing `TurnExecutor`, so it inherits the
planners, the checkpoint, the stale-run recovery, `close_finished_phases`, the
evaluator, capture and delivery unchanged — the same trick `SubTopologyExecutor`
already plays with a heavy `run_turn`. Producer side emits ONE `queued` row
carrying the real graph and lets the worker claim it: the durability IS being
worker-driven, and the solo path's `tokio::spawn` has none of it. Still exactly
one `topology_runs` row per unit of work and one completion path — `finish()` is
now that one place, shared by every tier.
THE TRAP, solved and proven. A VM is inject → run → collect → destroy, so a
per-node VM with text-only handoff silently loses every file an earlier node
wrote: node 2 boots from the original checkout, sees nothing, and still reports
success. The mission's host checkout is the medium — every node injects from it
and collects back over it — and two properties make that safe rather than lucky:
`execute_resumable` is strictly sequential, so two VMs never write one directory;
and the vm id is deterministic per (phase, iteration, step), so a duplicate is
refused by the node ("vm already exists") instead of becoming a second writer.
NEGATIVE CONTROL, run rather than assumed: with `repo` swapped for a private
per-node workspace, `a_later_node_sees_an_earlier_nodes_files` FAILS with
`saw:[]`; restored, it passes. The `PhaseVm` seam exists for exactly this — it
models inject/collect through the real `mission_fs` tar path in milliseconds.
Two durability traps this tier walks into, both closed:
- `requeue_stale` fires at 180s on `updated_at`, and one node here can run for
an hour. `SubTopologyExecutor` keeps its parent alive from each leaf step;
there is nothing between the start and end of a VM turn, so the turn holds a
ticker that touches `updated_at` every 30s and aborts on drop. Without it a
healthy composed run is requeued mid-node and boots a second VM.
- the 15-minute stuck-run reaper asks "any step records since it was CREATED?",
which describes a healthy composed run as readily as a wedged one. Hence
`REAPABLE_TIERS` — worker-driven minus this tier. Reaping it would be #54 in
a different costume.
`on_launch` mints no team for a microVM mission, deliberately: claws in
containers are what a VM mission does not use. So `mission_orchestrator::
composed_graph` builds the shape from the team template directly — nodes, roles
and pattern, zero claws provisioned. Per-node `attrs["backend"]` and
`attrs["node_id"]` override the mission's, which is what makes a validator node
on another provider's image a first-class graph node; a malformed `node_id`
fails the node rather than quietly running it where the graph did not ask.
Refusals are recorded as a failed run, not returned as an error: `launch_phase`
is swept every ten seconds, so a returned error is a phase that retries forever
while the log repeats itself.
501 tests pass, clippy clean. NOT yet proven end to end: no composed mission has
run on the fleet, so the resume-after-a-killed-worker leg is argued from the DB
test and the step-numbering test, not from a real two-node run.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
e31688bac5 |
fix(delivery): keep the TAIL of a push error — git prints its reason last
A failed push recorded two identical auth lines, a URL, and a branch name cut off mid-word. The reject reason was on the next line and the 500-char head clamp ate it, so the artifact preserved the noise and dropped the answer. That is what left #55 unresolvable: the evidence needed to distinguish "no credentials" from "non-fast-forward" had been truncated away. `evaluator_tools::clamp_output` already existed for exactly this — head AND tail with a byte count of what it dropped, on char boundaries so multi-byte output cannot panic. Reused rather than reinvented. Investigation notes recorded on #55. Two hypotheses were disproved by measurement: push credentials are rebuilt per push from GITEA_TOKEN + repos.clone_url and never live on disk (the clone-time scrub guarantees it, and every checkout on the host — including ones that pushed — has an identical credential-free origin), and neither of the two ways `with_ambient_auth` can silently return an unauthenticated URL applies here: the clone_url matches its required prefix and the token is non-empty in a container that predates the failure. 489 tests pass, clippy clean. |
||
|
|
66f730ad16 |
test(db): a regression net for the three-minute bug, with its negative control
The #54 fix had no test that could see it. Its defining property is that it only appears past 180 seconds, and `verify-mission-delivery.sh microvm` runs a 90-second mission — so the end-to-end harness written to catch silent failure was structurally blind to this one. A unit test asserting the allowlist's membership helps, but would not notice a NEW sweeper added without the filter. `crates/cm-db/tests/self_driven_runs.rs` tests the real SQL against a migrated database, in milliseconds instead of eight minutes: - a `microvm` and a `session` run, 30 minutes idle and still `running`, must be left alone by `requeue_stale` — that is the bug, in one assertion - a `team` run in the SAME state must still be requeued, so the fix is "sweep the right rows" and not "stop sweeping" - the worker must not CLAIM a queued self-driven row, which is what turned a healthy run into "missing or invalid graph" - the allowlist names only worker-driven tiers NEGATIVE CONTROL, run rather than assumed: with the tier filter removed from `requeue_stale`, `requeue_stale_leaves_self_driven_runs_alone` FAILS; restored, it passes. A guard that cannot detect the bug it was written for is decoration, and this project has shipped one of those before. 489 tests pass, clippy clean. |
||
|
|
d49acaed5e |
fix(missions): exclude build output from COLLECT too, not just inject
The other half of the same bug. The previous commit filtered `mission_fs::pack_dir` (the inject side) and left the guest's `op_get` tarring everything, so the re-run that proved the #54 fix — it survived 480s where it used to die at 210 — still lost its work to `vm_collect ... node timed out`. Two modules written, four subagents used, nothing delivered. `op_get` now takes an `exclude` list, sent by the host from `mission_fs::transport_excludes()` — the same list `mission_delivery` uses for the diff. Policy in one place, applied at both ends of the wire. Matched on directory NAME at any depth, so a workspace's per-crate `target/` dirs are all covered, with a test that plants a nested one and asserts it does not come along. Also proven by that run: the worker no longer kills a live microVM run. It ran 480 seconds straight through the 180s requeue window and the 210s mark where mission 019fd43e died, untouched. And `subagents: 4` — the team addendum did drive real fan-out this time, which is the first evidence the Slice 3 switch does anything. 483 tests pass, clippy clean. Still to prove: a >3-minute mission that actually DELIVERS. The collect fix is tested in isolation but has not yet carried a real mission's work back, and the guest agent needs rebuilding into the rootfs before it can. |
||
|
|
4efcde9d4f |
fix(missions): #54 — the worker was killing live microVM runs at 180 seconds
My hypothesis in #54 was WRONG, and it was wrong because I built it on a bad measurement: `grep -c 'microvm phase'` returned 0, so I concluded the completion log never printed and blamed the 15-minute reaper. The line was there all along, at 14:17:45. The real cause is worse. `requeue_stale` has NO TIER FILTER. A microvm run's `updated_at` is written once at insert and never again — it is driven by a `tokio::spawn` that owns it start to finish, and nothing in `microvm_executor` writes `topology_runs`. So at 180s the sweeper declared a perfectly healthy run stale and flipped it to `queued`; `claim_next_queued` (no tier filter either) handed it to the worker; `run_job` tried to parse the microvm graph placeholder, which `TopologyGraph` cannot deserialize; and it failed the run with "missing or invalid graph". Mission 019fd43e: run created 14:11:16, mission failed ~14:14:46. 210 seconds — the 180s window plus a tick. The agent went on working and finished at 14:17:45 with three modules written, by which time the phase was already dead and the VM was orphaned. A firecracker process was still alive 1h37m later. THE UNCOMFORTABLE PART: every microVM mission that appeared to work this session did so only by finishing inside three minutes. The 90-second ones dodged this. The harness scenario dodges it. Nothing about that was visible. `WORKER_DRIVEN_TIERS` (team, company, org, swarm, compare) is now the allowlist for all three sweep paths — claim, requeue, reap. An allowlist rather than a denylist so the next self-driven tier is safe by default instead of exposed until someone remembers the file. `tier='session'` had exactly the same exposure and is covered too. A unit test asserts microvm and session are NOT in it, next to the code that inserts them. Two more fixes from the same wreckage: - `destroy` reported `killed: pgid.is_some()` — true whenever there was a pgid to signal, whether or not anything died. It now sends the signal, polls /proc for the group leader, retries, and reports what it OBSERVED; `signalled` keeps the old meaning so "nothing to kill" is distinguishable from "it would not die". - the run-status update is now guarded with `AND status <> 'cancelled'`. An operator cancelling is a decision; this task reporting an outcome minutes later is an observation, and it must not overwrite one with the other. And the root cause of the collect timeout itself: `mission_fs::pack_dir` shipped `target/` in both directions. `mission_delivery` has excluded build output from the DIFF since day one; the TRANSPORT never knew. The host checkout was 9.4 MB of which 8.9 MB was `target/`, tarred and base64'd over vsock each way. `EXCLUDED_PATHS` is now one list shared by both layers, matched on directory name at any depth so a workspace's per-crate `target/` dirs are all covered. 483 tests pass, clippy clean. |
||
|
|
0d25a94a84 |
fix(missions): agent teams do not form in print mode — say so where it is set
MEASURED, against the CLI in our own image (2.1.223): with
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 and an explicit request to "spawn two
teammates", `claude -p` did the work with two SUBAGENTS, wrote both files, and
created no ~/.claude/teams/ directory at all. The docs allow for it — "Claude may
sometimes use subagents instead of creating a team" — and headless appears to be
always: the whole feature is described around an interactive agent panel, which a
print-mode session does not have.
So Slice 3's switch, as written yesterday, set a flag with no mechanism behind it.
The first team mission caught it, because the probe was built to look for teammates
rather than to assume them.
Corrected rather than removed:
- `team_env` documents the measurement at the point the flag is set, so the next
reader does not have to rediscover it. The flag stays: harmless, and free if a
later version supports teams non-interactively.
- the prompt addendum now asks for parallel DELEGATION rather than naming
teammates, which is what print mode can actually deliver — and it keeps the two
anti-patterns worth stating (own different files; do not split one change into
stages).
- the "no teammates" warning was blaming the flag and the config path. It now
judges on the SUBAGENT count, which is the mechanism in play, and a zero
teammate count is documented as expected rather than as a fault.
What the switch buys today is real but smaller than the plan assumed: it changes
the prompt so the lead parallelises across files instead of working through them
alone. Whether that beats solo on our own missions is still unmeasured, and the
plan's prediction — that it will not be faster — stands untested.
482 tests pass, clippy clean.
UNEXPLAINED, filed as #54: that team run's `topology_runs` row is `failed` while
the log line that sits immediately before the UPDATE never printed — zero matches
for 'microvm phase' in the container's whole log. The prime suspect is
`topology_worker`'s stuck-run reaper, which fails runs that are `running` with no
step records and does not filter by tier; a microvm run has no step records by
design. If that is it, any sufficiently long VM phase is failed out from under
itself. The system failed safely here — the empty-delivery guard caught that
nothing was produced, and nothing false was reported — but the cause is not known
and it is not being written up as if it were.
|
||
|
|
cb48f7ff3b |
feat(missions): Slice 3 — agent teams behind a per-mission switch, solo by default
`missions.team_engine` (0069): NULL = solo, `'claude_code'` = Claude Code agent teams inside the mission's VM. Solo stays the default deliberately — Anthropic measure multi-agent at 3-10x the tokens with wall-clock often LONGER, since the benefit is thoroughness rather than speed — so a mission that said nothing does not get a team. In-process teammates live in the lead's process, so ONE VM hosts the whole team. That is why this is a prompt-and-env change rather than an orchestration one: no N-VM fan-out, no placement per teammate, no new completion path. The lead decides its own team size and there is no flag that limits it, so the cap (4) is stated in the prompt. The addendum also carries the two anti-patterns from Anthropic's guidance, because they are exactly the shapes our pipeline templates have: teammates must own DIFFERENT FILES (two in one file overwrite each other), and one change must not be split into stages across teammates (a handoff loses context at every step). And: wait for your teammates — a summary written before they report is the lead's own guess. A solo mission's prompt and env are byte-identical to before this change. That is enforced by test, not by intention: the comparison between solo and team is only meaningful if the solo side did not also move. Evidence, because a team mission that forms no team is silently just a solo run that looked fine and spent fewer tokens: a second probe counts members in `~/.claude/teams/*/config.json` (minus the lead), reported separately from the subagent count, and a team mission with zero teammates logs loudly with the two likely causes. The teammate path is DOCUMENTED BUT NOT YET VERIFIED in our image, unlike the subagent transcript path which was measured — so a zero there means "no evidence found", and the first real team mission is what turns it into a fact. `Option<u32>`: None means no team was asked for or the probe could not run. Hooks (`TaskCompleted` / `TeammateIdle` exit 2, which would move `done_when` from post-hoc into the agent's own loop) are the highest-value part of this slice and are deliberately NOT here — they deserve their own pass rather than a rushed tail. 482 tests pass, clippy clean. |
||
|
|
c840688adb |
feat(missions): choose the independent validator per mission (#53)
`CLAWMATES_VALIDATOR_MODEL` is deployment-wide, so proving Slice 2 put a second
provider on the critical path of EVERY phase verdict. `cross_provider_judge`
deliberately does not fall back when the independent judge fails — a verdict
quietly produced by a same-family model would claim a property it does not have —
so a z.ai outage makes phases unmeetable rather than merely unverified. That is a
per-mission trade, not a per-deployment one.
`missions.validator_model` (0068), settable at create, with three distinct states
because an empty string and NULL mean opposite things in a nullable text column:
NULL use the deployment default
'' explicitly NO independent validator — judge with the house model.
The default must not quietly reinstate independence a mission was
told to skip.
'glm:glm-4.7' this spec, subject to the same three refusals as before:
same-family rejected, unregistered provider rejected, and a failed
independent judge does not fall back.
Whitespace counts as empty: a column hand-set to " " meant to say nothing.
478 tests pass, clippy clean. Behaviour is unchanged for existing missions — they
have NULL and so keep following the deployment default.
|
||
|
|
b17e18aa67 |
fix(harness): the verdict check matched psql's display form, not the query's
`select met || ' ' || independent` casts the booleans to `true`/`false`, but the
pattern matched `t`/`f` — psql's *column display* form. So the check reported "no
verdict recorded for the phase" while the row sat in the table saying met=true,
independent=true, glm-4.7.
A check that fails for a reason unrelated to what it checks is worse than no check:
it trains you to ignore the output. The booleans are cast explicitly now so the
shape cannot drift again, and the failure message prints what it actually got.
`verify-mission-delivery.sh microvm` now passes 5/5 against production:
- the agent ran under guest kernel 6.1.128, not the gateway's 6.8.0-124 or the
node's 7.0.0-28 — the one assertion that cannot pass by accident
- the lead delegated to 1 subagent
- the condition was met and judged INDEPENDENTLY by glm-4.7
- the checkout has exactly one writer (uid 65532)
- negative control: a backend no node can run is refused at launch
|
||
|
|
9aed20b6d0 |
fix(missions): capture a failed phase's work; harness gains a microvm scenario (#51)
A REGRESSION I INTRODUCED ONE COMMIT AGO. `capture_finished_coding_phases`
selects on `mp.status = 'completed'`, so the moment an unmet phase correctly began
reporting `failed`, its diff stopped being captured, committed or pushed — the work
was silently discarded. Found by the new harness scenario, whose phase legitimately
missed its condition and then had no artifact at all.
What was produced, and whether the goal was met, are different facts. The artifact
records the first; `mp.status` records the second. Capture now covers terminal
phases (`completed`, `failed`), so a phase that did real work and missed its goal
still delivers a reviewable diff — which is exactly what the next pass needs.
`scripts/verify-mission-delivery.sh microvm` — the regression net this session was
missing. Everything the microVM track proved by hand was guarded by nothing:
- THE KERNEL LINE is the assertion that cannot pass by accident. Every other
check would also pass if the phase had quietly run in a container on the
gateway; only the kernel says WHERE it ran. Compared against the real gateway
and node kernels read at start-up rather than pinned to a version, so
upgrading vmlinux does not manufacture a failure.
- subagent count > 0, from the server's own count of Claude Code's per-subagent
transcripts. Before `Agent` was in the allowlist this was structurally
impossible and nothing said so. A probe that could not run reports "?" and
FAILS the check rather than reading as zero.
- the verdict's judge and whether it was independent.
- negative control, observed passing: a mission whose backend no node can run is
refused at launch and stays draft. Without it the positive scenario would pass
just as well against a scheduler that ignored `backend` entirely — which is
what it did until the first real microvm mission landed on a node with no such
rootfs.
Also fixed in the harness: `api` now sends the JSON body on STDIN (`curl -d @-`)
instead of interpolating it into a single-quoted argument inside a double-quoted
ssh command. A task description containing "the crate's test suite" ended the
quoting and killed the remote shell; two attempts to escape it were themselves
wrong, because the backslashes must survive bash AND sed AND sh. Removing the
interpolation removes the class, and the next author does not need to know that
apostrophes were forbidden.
475 tests pass, clippy clean.
|
||
|
|
bb807c2f3a |
fix(missions): an unmet goal condition is no longer reported as success
Found by the Goodhart test for the independent judge, which is exactly what it was built to find. The test: a phase whose `done_when` demanded a passing suite, and a task that deliberately left a failing test. glm-4.7 judged it, ran `cargo test` itself, saw `parity_is_wrong_on_purpose ... FAILED` (exit 101), and returned met=false quoting the assertion — while the agent's own summary said "All three steps are implemented exactly as specified and independently verified". The verdict and the agent's account diverged, which is the whole point of an independent judge. And then the mission closed `completed`. `if verdict.met || last_pass` marked BOTH outcomes completed, so a phase that ran out of passes without ever meeting its condition reported success — and through `close_finished_missions`, so did the mission. The verdict said met=false in a column nobody reads before believing a green status. Anything consuming mission status rather than digging into the verdict saw a goal that was never reached as a goal achieved. Exhausted-and-unmet is now `failed`, and the log names the judge and whether it was independent. This changes observable behaviour: missions that would previously have finished green with an unmet condition now finish failed. That is the correction, not a regression — but it is worth knowing before the next scheduled run. Also: `Verdict.independent` had no column. The field existed in the struct and in the logs, so the audit question the mechanism exists to answer — was this checked by something other than the model that wrote it? — could not be asked of the database. Migration 0067 adds it, defaulting to false, which is the truth about every row written before now. Verified in production before the fix: glm-4.7, 4 checks all executed, the real cargo failure quoted, met=false. 475 tests pass, clippy clean. Note for whoever rebases: `sqlx::migrate!` embeds migrations at COMPILE time, so a new migration needs cm-db rebuilt (`touch crates/cm-db/src/lib.rs`) or the integration tests fail on a column that exists in the file and not in the binary. |
||
|
|
8796fbbcbb |
feat(evaluator): Slice 2 — an independent judge, from a different provider, with the same teeth
Claude writes the code and Claude judges it. That is a correlated failure: the
model that talked itself into a shortcut is the one disposed to accept it, and it
is the structural cause of the "early victory" failure Anthropic documents and of
our own Goodhart incident.
`glm` and `kimi` are both already registered in production, so the fix needed no
new credential path.
THE UNLOCK: `judge_with_tools` took `&AnthropicProvider`, but `LlmProvider` is a
single method — `stream(ChatRequest)` — and the loop only ever used that. The
concrete type was incidental. Widening it to `&dyn LlmProvider` means a
cross-provider judge runs the SAME allow-listed command loop. Before, independence
and real verification were mutually exclusive: the tool loop existed only on the
subscription path and every other route "judged claims only", so choosing an
independent judge meant giving up the checks that make a verdict evidence. GLM is
registered in anthropic format, so tool calling reaches it unchanged.
`CLAWMATES_VALIDATOR_MODEL` (e.g. `glm:glm-4.7`) selects it. Three refusals, each
protecting the claim the field makes:
- a spec in the implementer's own family is rejected, not used — `opus` judging
`sonnet` is not independence, they share a lineage and most failure modes
- a spec naming a provider this deployment never registered is rejected.
`Runtime::resolve_provider` silently falls back to the DEFAULT provider when
the registry has no such name, which would hand back Claude while the caller
believed it had GLM. Detectable because the returned model keeps its `name:`
prefix, so it is checked rather than trusted.
- an independent judge that FAILS does not fall through to the house judge. A
verdict quietly produced by a same-family model would claim a property it does
not have. The pass stays unmet, says why, and the next sweep retries.
`Verdict.independent` records it, `#[serde(default)]` so verdicts stored before
this field read back as not independent — which is what they were. An unrecognised
model family resolves to "unknown", never to ours: guessing would report
independence nobody established.
474 tests pass, clippy clean. Not yet enabled in production — the env var is unset,
so behaviour is identical until it is set deliberately.
|
||
|
|
11b274edc6 |
chore(images): Claude Code 2.1.223, and make the verifier foreground
Reviewed the changelog rather than bumping on principle. 2.1.220 → 2.1.223 for one
reason that bears on how we use subagents:
2.1.222 — "Fixed PreToolUse auto-allow hooks bypassing tool restrictions in
background agent tasks."
Subagents run in the background by default since 2.1.198, and the `verifier`
role's entire guarantee is a TOOL restriction — no Edit, no Write. So on 2.1.220
the one property we rely on was the one that bug could undo. 2.1.221 also fixes
`--mcp-config` servers not connecting before the first turn in print mode, which
is the mode we run and will matter when the MCP door reaches a VM.
Two findings from the changelog that we already had at 2.1.220, both worth knowing:
- 2.1.219: subagents can nest to depth 3 (was 1), so our roles can delegate
further than assumed.
- 2.1.212: a subagent inherits the parent's permission mode, which confirms the
verifier's read-only property must come from `tools` and not from permissions.
That is how it was written; now the reasoning is recorded next to it.
And a correctness fix that follows from the background default: the verifier is now
`background: false`. A background verifier lets the lead carry on and write its
report before the check has finished — the finding would arrive after the
conclusion it was supposed to inform.
Verified on tank: image reports 2.1.223, rootfs rebuilt, `--vm-selftest` all green
including a real agent turn on subscription auth, egress allow and deny both firing.
|
||
|
|
2dee941080 |
feat(missions): Slice 1 — a microVM agent can delegate, and we can see that it did
`microvm_executor` passed `--allowedTools Read Edit Write Bash`, which omits the
`Agent` tool, so Claude Code could not spawn a single subagent in any of our VMs.
The tool existed, the model knew how to use it, and the allowlist quietly removed
the ability. Nothing in any output said so.
Now: `Agent` in the allowlist, two roles supplied as `--agents` JSON, and a probe
that counts what actually ran.
Roles are JSON on the command line, not files, because `/mission/repo` is
collected and diffed — a role definition written into the checkout would arrive in
the delivered patch as if the agent had authored it.
Two roles only, and the choice is the research talking:
- `verifier` — the one multi-agent pattern Anthropic endorses for coding work.
It gets Read/Grep/Glob/Bash and deliberately NOT Edit or Write: an agent that
can fix what it is checking will fix it and report success, and the report is
then about a tree nobody reviewed. Its prompt demands the COMPLETE suite,
which is the counter to the "early victory problem" — the same failure as our
own Goodhart incident.
- `explorer` — context protection, read-only.
Roles like "tester" or "committer" are absent on purpose: splitting sequential
phases of the same work is a named anti-pattern, and it is the shape our pipeline
templates already have.
THREE THINGS THE IMAGE CORRECTED, none of which review would have caught:
1. `CLAUDE_AGENT_SDK_DISABLE_BUILTIN_AGENTS=1` (in the plan) removes EVERY agent
type, including the ones `--agents` defines. Measured: the lead reported "an
empty available-agents list" after trying four role names and — to its credit —
refused to fabricate a subagent result. Worse, the unit test asserting
"builtins off is paired with our own roles" PASSED throughout, because the
pairing holds in our code and not in the CLI. Dropped, and the test rewritten
to assert only what a unit test can speak to.
2. `--forward-subagent-text` refuses to run without `--output-format=stream-json`,
which would change how this module reads output. Dropped.
3. `--append-subagent-system-prompt` does not exist in 2.1.220 despite being
documented. The anti-shortcut rule is inlined per role instead — better anyway,
since a verifier and an explorer need different wording.
Evidence instead of assumption: Claude Code writes a per-subagent transcript at
`<session>/subagents/agent-*.jsonl`, so the guest is asked to count them before
collection (they live in /root, outside the collected tree). `VmOutcome.subagents`
is `Option<u32>` and the phase log prints it: `None`/"?" means the probe could not
run, which is a different fact from "delegated to nobody" and only one of those is
about the agent.
Verified in a container against the real CLI on tank before any of this shipped:
`FANOUT-OK`, a subagent transcript on disk, and zero errored Agent calls.
469 tests pass, clippy clean.
|
||
|
|
7696009b25 |
fix(fleet): name the BACKEND in a placement refusal, not just "microvm capability"
Observed on the negative control: a mission with backend='kimi' was correctly refused, but the message read "no online node reports microvm capability" — and both nodes do report it. What one lacked was the image. That first clause would have sent an operator to reinstall firecracker on a node that already had it. The refusal now names the backend, and the remedy still names both halves. |
||
|
|
d9f53a3f96 |
fix(fleet): placement requires the backend's rootfs image, not just KVM
The first real microVM mission was placed on morpheus because it reports
{"microvm": true}, while only tank had rootfs-claude.ext4. It failed by name
rather than booting the wrong image — but whether a mission ran came down to
which capable node was listed first, which is a coin flip dressed as scheduling.
`missions.backend` was invisible to the scheduler.
The node now enumerates the images on its disk and reports them as a `rootfs`
ARRAY. `microvm::available_backends` lives beside `rootfs_for`, its inverse,
because the two must agree on what a backend name means; split apart, one drifts
and the scheduler starts promising images the booter cannot find. It only
advertises names `rootfs_for` would accept, and reports an empty array rather than
omitting the key — set_capabilities REPLACES, so a deleted image stops being
advertised instead of leaving a stale claim.
`nodes::online_for_backend` requires microvm AND that the node's list contains the
mission's backend. A node on an older daemon has no `rootfs` key and matches
nothing: unknown is not permission, the same treatment every other capability
gets. `backend_key` maps the three spellings of "the default image" to the one
name the node advertises, and is tested — a mismatch there would reject every node
for an ordinary mission with no backend set.
The launch error now names both halves of the fix, since "no capable node" was
true but unhelpful when the node was capable and merely lacked the image.
Mission gains `backend` on the domain struct; it was a column the executor read
from the phase query while the struct that placement uses could not see it.
464 tests pass, clippy clean.
|
||
|
|
1cd81a8b2a |
fix(missions): place a microvm mission before returning from on_launch
Self-inflicted, one commit old, and found by running a real mission: the early return I added for "a microvm mission materialises no team" sat ABOVE the microVM placement block in the same function, so on_launch returned before ever choosing a node. The mission then failed with the executor's own guard — "mission has no target_node_id ... a microvm mission cannot run on the gateway, which has no /dev/kvm" — which is the guard working exactly as designed, on a cause one layer further up. Placement now runs first. Worth noting the shape: adding an early return to a long function silently skipped everything below it that the same runtime_kind depends on. 461 tests pass, clippy clean. |