Files
clawmates/scripts/verify-mission-delivery.sh
T
Omar SobhandClaude Opus 5 794f2124bc
deploy / test (push) Successful in 5m18s
deploy / build (push) Successful in 5m54s
feat(microvm): Claude Code 2.1.276 rootfs pins, and a VM run that says what it ran
Every rootfs on the fleet had sat on Claude Code 2.1.223–2.1.226 since August
while the container tier moved to 2.1.276, and nothing recorded either. GLM
and Kimi exist only as microVM backends, so "have we upgraded GLM and Kimi"
is this change and the rebuild it drives.

Pins. All four agent-* images pin 2.1.276 — as separate ARGs, since Docker has
no include and each file has to stay reproducible alone — and
scripts/fc-build-rootfs.sh refuses to build if they disagree, naming the odd
one out. They had already drifted (claude 226, the rest 223) under comments
saying "same version on purpose". Between 2.1.226 and 2.1.276, 2.1.265 and
2.1.275 each broke every turn on ANTHROPIC_BASE_URL endpoints, which is how
glm and kimi reach `claude` inside a VM; the container-tier verification never
exercised that path, so the VM runs on those backends are the real test.

Provenance. `VmOutcome` carries the rootfs the node reported booting and the
guest's own `claude --version`; `launch_microvm_phase` persists both as
`checkpoint.vm` beside `records` (the two readers parse only `records`) and
names them in its log line. "Which image and CLI did this mission run on" is
a query now.

Independence. `evaluator` derived the implementer family from a constant
`"anthropic"`, true while every backend was Claude on Anthropic. With glm and
kimi rootfs it made a glm mission judged by glm:glm-5.3 read as
`independent = true` — the one claim that path exists to make honestly.
`implementer_family(missions.backend)` mirrors `microvm_credential_for`; the
subscription judge is now independent exactly when the agent did NOT run on
Anthropic.

Harness. `verify-mission-delivery.sh glm|kimi` run the microvm scenario on
each backend and add the proof the mission itself cannot give: the placed
node's journal must show the VM dialling that provider's host, never being
denied it, and dialling nothing else but the forge — a model's self-report is
measured worthless here. `assert_cli_version` reads checkpoint.vm. The stale
scratch-repo default (dead since the 09-14 wipe) is the re-synced id.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
2026-09-18 20:16:06 -05:00

1676 lines
84 KiB
Bash
Executable File

#!/usr/bin/env bash
# Verify mission delivery end-to-end against the deployed stack.
#
# Every verification of the delivery chain so far has been a throwaway bash
# script, written fresh per run and discarded. One of them printed
#
# --- host .git owner (should be one uid) ---
# UNKNOWN
#
# and that `UNKNOWN` is the whole reason this file exists: the probe could not
# read its subject, and said so in a way that looked like output rather than
# like failure. It would have printed `UNKNOWN` just as happily if the uid
# split had come back. That is seam 4 — absence encoded as a legitimate value —
# reappearing inside the tool built to detect seam 1.
#
# So the rules here are structural, not stylistic:
#
# 1. A probe returns a value or exits non-zero. There is no third outcome,
# no placeholder, no empty string that a caller might read as "fine".
# 2. The uid probe is self-tested against a mission KNOWN to have the split,
# before any result from it is believed. A probe that cannot see the
# known-bad case has not verified the good one — it has only failed to
# look. `--selftest-only` runs that check alone.
# 3. A scenario that never ran is FAIL-NORUN, never PASS. An absent branch
# is indistinguishable from a dead container, and once scored PASS.
#
# Usage:
# scripts/verify-mission-delivery.sh selftest # probe self-test alone
# scripts/verify-mission-delivery.sh uids <mission> # uid probe, one mission
# scripts/verify-mission-delivery.sh chain # phase continuity
# scripts/verify-mission-delivery.sh multirole # 3 roles + real tests
# scripts/verify-mission-delivery.sh noop # empty phase must FAIL
# scripts/verify-mission-delivery.sh microvm # runs in a guest kernel + fans out
# scripts/verify-mission-delivery.sh glm # microvm on the z.ai backend; provider proven by the node's egress log
# scripts/verify-mission-delivery.sh kimi # microvm on the Kimi backend; same proof
# scripts/verify-mission-delivery.sh capacity # a burst > the fleet must QUEUE
# scripts/verify-mission-delivery.sh drain-midmission # a drained node hands the mission on
# scripts/verify-mission-delivery.sh all # everything
#
# Environment:
# CLAWMATES_HOST ssh host running the stack (default gw-04)
# CLAWMATES_OWNER_EMAIL account to mint a session for (default om.sobh@…)
# CLAWMATES_UID_CONTROL mission id/prefix known to have a uid split;
# auto-discovered when unset
# CLAWMATES_REPO_ID scratch repo for the delivery scenarios
# CLAWMATES_TEAM_TEMPLATE team template for the delivery scenarios
# MISSION_TIMEOUT seconds to wait for a mission (default 1800)
# CAPACITY_BURST missions to launch in the capacity burst
# (default: fleet slots + 2 — smaller is reported
# NORUN, since it never exercises the queue)
set -uo pipefail
HOST="${CLAWMATES_HOST:-gw-04}"
OWNER="${CLAWMATES_OWNER_EMAIL:-om[email protected]}"
MISSIONS_ROOT="${CLAWMATES_MISSIONS_ROOT:-/var/lib/clawmates-missions}"
# Real host kernels, read at start-up. The microvm scenario asserts the agent's
# kernel differs from BOTH, which proves it ran in a guest without pinning a
# vmlinux version that an upgrade would invalidate.
GW_KERNEL=$(ssh "$HOST" 'uname -r' 2>/dev/null | tr -d '[:space:]')
NODE_KERNEL=$(ssh "${FLEET_NODE:-osobh@tank}" 'uname -r' 2>/dev/null | tr -d '[:space:]')
# The scratch repo is re-registered whenever the Gitea connection re-syncs, and
# gets a new id each time (f8bbe4d7-… died in the 2026-09-14 wipe). Override
# with CLAWMATES_REPO_ID, or find it: select id from repos where name =
# 'clawmates-delivery-scratch'.
REPO_ID="${CLAWMATES_REPO_ID:-01a0052f-dc7f-7b73-8afd-2984a91338bb}"
TEAM_TEMPLATE="${CLAWMATES_TEAM_TEMPLATE:-7e453826-41c4-4425-bab5-8f11fd0a14d7}"
MISSION_TIMEOUT="${MISSION_TIMEOUT:-1800}"
FAILURES=0
CHECKS=0
# ── Reporting ────────────────────────────────────────────────────
#
# Deliberately only three verdicts, and NORUN is one of them. "The scenario
# did not execute" must not be able to borrow PASS's vocabulary.
pass() { printf 'PASS %s\n' "$*"; CHECKS=$((CHECKS + 1)); }
fail() { printf 'FAIL %s\n' "$*"; CHECKS=$((CHECKS + 1)); FAILURES=$((FAILURES + 1)); }
norun() { printf 'FAIL-NORUN %s\n' "$*"; CHECKS=$((CHECKS + 1)); FAILURES=$((FAILURES + 1)); }
info() { printf ' %s\n' "$*"; }
die() { printf 'ABORT %s\n' "$*" >&2; exit 2; }
# ── Session + API ────────────────────────────────────────────────
mint_session() {
local secret hash
secret="verify-$(openssl rand -hex 16)"
hash=$(printf '%s' "$secret" | openssl dgst -sha256 -binary \
| openssl base64 -A | tr '+/' '-_' | tr -d '=')
local rows
rows=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"insert into auth_sessions (user_id, token_hash, expires_at) \
select id, '$hash', now() + interval '90 minutes' \
from users where email='$OWNER' limit 1 returning 1;\"" 2>/dev/null \
| head -1 | tr -d '[:space:]')
# `head -1`: psql emits the returned row AND its `INSERT 0 1` command tag,
# and collapsing both gave `1INSERT01`, which failed this check on a mint
# that had in fact worked.
#
# No row inserted means no such user — without this check the script would
# carry on and report every API call as a delivery failure.
[ "$rows" = "1" ] || { echo "could not mint a session for $OWNER (no such user?)" >&2; return 1; }
printf '%s' "$secret"
}
api() { # api <token> <METHOD> <path> [json]
local t="$1" m="$2" p="$3" b="${4:-}"
# The JSON body travels on STDIN (`curl -d @-`), not embedded in the command.
#
# It used to be interpolated into a single-quoted `-d '...'` inside a
# double-quoted ssh command, which works only for bodies containing neither
# apostrophes nor anything else the two shells rewrite. A task description
# saying "the crate's test suite" ended the quoting and the remote shell died
# with "unexpected EOF"; two attempts at escaping it were themselves wrong,
# because the backslashes have to survive bash AND sed AND sh. Removing the
# interpolation removes the whole class.
if [ -n "$b" ]; then
printf '%s' "$b" | ssh "$HOST" "docker run --rm -i --network clawmates_core \
curlimages/curl:latest -s -X $m \
-H 'Authorization: Bearer $t' -H 'Content-Type: application/json' -d @- \
http://clawmates_server_1:8080$p"
else
ssh "$HOST" "docker run --rm --network clawmates_core curlimages/curl:latest -s -X $m \
-H 'Authorization: Bearer $t' http://clawmates_server_1:8080$p"
fi
}
# ── The uid probe ────────────────────────────────────────────────
#
# The structural claim copy mode makes is that the host checkout has exactly
# one writer. This is what tests it. It prints a sorted, comma-separated uid
# list on stdout and exits 0, or prints nothing and exits 1.
probe_uids() { # probe_uids <mission-id-or-prefix>
local mission="$1" dir uids
dir=$(ssh "$HOST" "ls -d $MISSIONS_ROOT/$mission* 2>/dev/null | head -1" | tr -d '\r')
[ -n "$dir" ] || { echo "no mission dir under $MISSIONS_ROOT for $mission" >&2; return 1; }
ssh "$HOST" "sudo test -d '$dir/repo'" 2>/dev/null \
|| { echo "$dir/repo is not a directory" >&2; return 1; }
# -printf '%U' over stat: one process for the whole tree, and it reports the
# numeric uid even when the host has no passwd entry for it.
uids=$(ssh "$HOST" "sudo find '$dir/repo' -xdev -printf '%U\n' 2>/dev/null | sort -un | paste -sd, -" | tr -d '\r')
# An empty result is not "no uids", it is a failed read. A checkout always
# contains files; if find returned nothing, find did not work.
[ -n "$uids" ] || { echo "uid read produced no output for $dir/repo" >&2; return 1; }
printf '%s' "$uids"
}
# Find a mission whose checkout still shows the pre-copy-mode split. Used as
# the probe's negative control.
discover_split_control() {
ssh "$HOST" "for d in $MISSIONS_ROOT/*/repo; do
[ -d \"\$d\" ] || continue
n=\$(sudo find \"\$d\" -xdev -printf '%U\n' 2>/dev/null | sort -un | wc -l)
if [ \"\$n\" -gt 1 ]; then basename \$(dirname \"\$d\"); break; fi
done" 2>/dev/null | tr -d '\r' | head -1
}
# ── Probe self-test ──────────────────────────────────────────────
#
# Run BEFORE trusting any uid result. A green uid report from a probe that
# cannot detect the split is not evidence of anything.
selftest_uid_probe() {
local control uids
control="${CLAWMATES_UID_CONTROL:-$(discover_split_control)}"
if [ -z "$control" ]; then
# Not a pass. Every checkout on the host is single-uid, which is the
# desired end state but leaves the probe unexercised — so say exactly
# that rather than implying the probe was validated.
info "selftest: no split-uid mission remains on $HOST to use as a control"
info "selftest: uid results below are UNVALIDATED (set CLAWMATES_UID_CONTROL)"
return 0
fi
uids=$(probe_uids "$control") || { fail "selftest: probe failed on control $control"; return 1; }
case "$uids" in
*,*) pass "selftest: probe reports the split on control $control (uids=$uids)" ;;
*) fail "selftest: control $control reports a single uid ($uids) — the probe cannot detect the split it exists to find" ;;
esac
}
check_single_uid() { # check_single_uid <mission> <label>
local mission="$1" label="$2" uids
uids=$(probe_uids "$mission") || { fail "$label: uid probe could not read the checkout"; return 1; }
case "$uids" in
*,*) fail "$label: checkout has multiple writers (uids=$uids)" ;;
*) pass "$label: checkout has exactly one writer (uid=$uids)" ;;
esac
}
# ── Mission lifecycle ────────────────────────────────────────────
create_mission() { # create_mission <token> <json> -> mission id
local out id
out=$(api "$1" POST /api/missions "$(echo "$2" | tr -d '\n')")
id=$(printf '%s' "$out" | python3 -c 'import json,sys
try: print(json.load(sys.stdin).get("id",""))
except Exception: pass' 2>/dev/null)
[ -n "$id" ] || { echo "create failed: $out" >&2; return 1; }
printf '%s' "$id"
}
await_mission() { # await_mission <token> <mission> -> final status
local t="$1" m="$2" waited=0 status
while [ "$waited" -lt "$MISSION_TIMEOUT" ]; do
status=$(api "$t" GET "/api/missions/$m" | python3 -c 'import json,sys
try: print(json.load(sys.stdin).get("status",""))
except Exception: pass' 2>/dev/null)
case "$status" in
completed|failed|cancelled) printf '%s' "$status"; return 0 ;;
esac
sleep 20
waited=$((waited + 20))
done
printf 'timeout'
}
# Emit one `order_idx status files pushed commit_error push_error` line per
# phase by joining phases to their code_diff artifacts.
phase_report() { # phase_report <token> <mission>
api "$1" GET "/api/missions/$2" | python3 -c '
import json, sys
d = json.load(sys.stdin)
art = {}
for a in d.get("artifacts") or []:
if a.get("kind") == "code_diff":
art[a.get("phase_id")] = a.get("metadata") or {}
for p in sorted(d.get("phases") or [], key=lambda p: p.get("order_idx", 0)):
m = art.get(p["id"], {})
print(p.get("order_idx"), p.get("status"),
m.get("files_changed", m.get("files", "-")),
m.get("pushed", "-"), m.get("branch", "-"),
json.dumps(m.get("commit_error")), json.dumps(m.get("push_error")))
'
}
# Read a file back from the branch the mission actually pushed, so the
# assertion is against the forge rather than against the host staging dir.
fetch_delivered() { # fetch_delivered <token> <mission> <path>
local branch repo token
branch=$(api "$1" GET "/api/missions/$2" | python3 -c '
import json, sys
d = json.load(sys.stdin)
for a in d.get("artifacts") or []:
b = (a.get("metadata") or {}).get("branch")
if b: print(b); break
' 2>/dev/null)
[ -n "$branch" ] || return 1
token=$(ssh "$HOST" 'docker exec clawmates_server_1 printenv GITEA_TOKEN' | tr -d '\r')
repo=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select owner || '/' || name from repos where id='$REPO_ID';\"" \
| head -1 | tr -d '[:space:]')
[ -n "$token" ] && [ -n "$repo" ] || return 1
local enc; enc=$(printf '%s' "$branch" | sed 's|/|%2F|g')
ssh "$HOST" "curl -sf -H 'Authorization: token $token' \
'https://git.redclaw.dev/api/v1/repos/$repo/raw/$3?ref=$enc'"
}
# Read a file from the repo's DEFAULT branch, so a scenario can measure what a
# run ADDED rather than what the file happens to contain. Prints nothing and
# succeeds when the file is absent — an empty baseline is a real baseline.
fetch_main() { # fetch_main <path>
local token repo
token=$(ssh "$HOST" 'docker exec clawmates_server_1 printenv GITEA_TOKEN' | tr -d '\r')
repo=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select owner || '/' || name from repos where id='$REPO_ID';\"" \
| head -1 | tr -d '[:space:]')
[ -n "$token" ] && [ -n "$repo" ] || return 0
ssh "$HOST" "curl -sf -H 'Authorization: token $token' \
'https://git.redclaw.dev/api/v1/repos/$repo/raw/$1?ref=main'" 2>/dev/null || true
}
# ── Scenario: a phase runs inside a microVM, and fans out ─────────
#
# Guards everything the microVM track proved by hand: that the agent ran in a
# GUEST kernel rather than on a host, that Claude Code could actually delegate,
# and that the work came back and landed.
#
# The kernel line is the assertion that cannot pass by accident. Every other
# check here would also pass if the phase had quietly run in a container on the
# gateway; only the kernel says WHERE it ran. It is compared against the real
# host kernels rather than pinned to a version, so upgrading `vmlinux` does not
# turn this into a false failure.
MICROVM_BODY=$(cat <<JSON
{"title":"verify: a phase runs inside a microVM",
"template_kind":"research_and_code",
"repo_id":"$REPO_ID",
"runtime_kind":"microvm",
"backend":"claude",
"description":"Prove a coding phase executes in a Firecracker microVM and can delegate.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"done_when":"MICROVM.md exists at the repository root and contains both a test-result summary and a recorded kernel version such as 6.1.128.",
"task":"1. Use the verifier subagent to run the COMPLETE test suite of this crate and report what it found. Do not run it yourself and report that instead — the point is an independent check.\n2. Write MICROVM.md at the repository root with exactly two lines: the first is the test-result summary line the verifier reported, the second is the kernel release from running uname -r.\nCreate no other files."}}
]}
JSON
)
# How many subagents the phase actually spawned, per the server's own count of
# Claude Code's per-subagent transcripts. Read from the log because that is
# where `phase_runner` reports it; "?" means the probe could not run, which is
# not the same as zero.
subagent_count() { # subagent_count <mission>
ssh "$HOST" "docker logs --since 60m clawmates_server_1 2>&1 \
| grep -F 'microvm phase' | grep -F '$1' | tail -1" \
| sed -n 's/.*subagents: \([0-9?]*\).*/\1/p'
}
assert_microvm() { # <token> <mission> <report>
local token="$1" mission="$2" report="$3" delivered kernel subs indep
while read -r idx status files pushed _branch cerr perr; do
[ "$status" = "completed" ] || fail "microvm: phase $idx status=$status (commit_error=$cerr push_error=$perr)"
[ "$pushed" = "True" ] || fail "microvm: phase $idx not pushed (commit_error=$cerr push_error=$perr)"
case "$files" in 0|-) fail "microvm: phase $idx delivered no files" ;; esac
done <<<"$report"
# The completion gate, in a real VM. Checked here rather than in its own
# scenario because it applies to every coding phase on this path — a gate that
# stopped being installed would otherwise cost nothing visible until the next
# phase that needed it.
assert_stop_gate "$mission" microvm
delivered=$(fetch_delivered "$token" "$mission" MICROVM.md) \
|| { fail "microvm: could not read MICROVM.md from the pushed branch"; return 1; }
# The decisive check: the second line is the kernel the agent ran on.
kernel=$(printf '%s\n' "$delivered" | sed -n '2p' | tr -d '[:space:]')
if [ -z "$kernel" ]; then
fail "microvm: MICROVM.md has no kernel line: $(printf '%s' "$delivered" | tr '\n' '|')"
elif [ "$kernel" = "$GW_KERNEL" ]; then
fail "microvm: the agent ran on the GATEWAY kernel ($kernel) — not in a VM at all"
elif [ "$kernel" = "$NODE_KERNEL" ]; then
fail "microvm: the agent ran on the fleet NODE's kernel ($kernel) — not in a VM at all"
else
pass "microvm: the agent ran under a guest kernel ($kernel), not the gateway's ($GW_KERNEL) or the node's ($NODE_KERNEL)"
fi
# Fan-out. Before the `Agent` tool was added to the allowlist this was
# structurally impossible, and nothing said so.
subs=$(subagent_count "$mission")
case "$subs" in
''|'?') fail "microvm: could not count subagents (probe did not run) — delegation UNPROVEN" ;;
0) fail "microvm: the phase spawned no subagents, so the verifier never ran" ;;
*) pass "microvm: the lead delegated to $subs subagent(s)" ;;
esac
# And the verdict: judged, and by whom. `independent` is only true when the
# judge came from a different provider family than the agent.
indep=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select met::text || ' ' || independent::text || ' ' || model from mission_phase_evaluations \
where phase_id in (select id from mission_phases where mission_id='$mission') \
order by created_at desc limit 1;\"" | head -1 | tr -d '\r')
# `true`/`false`, not psql's display `t`/`f`: concatenating a boolean with text
# casts it to its full spelling. The first version matched "t t " and therefore
# never matched a real verdict, reporting "no verdict recorded" while the row was
# sitting in the table — a check that failed for a reason that had nothing to do
# with what it was checking.
case "$indep" in
"true true "*) pass "microvm: condition met, judged independently (${indep##* })" ;;
"true false "*) info "microvm: condition met but judged by the agent's own provider family (${indep##* }) — set CLAWMATES_VALIDATOR_MODEL for an independent check" ;;
"false "*) fail "microvm: the judge says the condition was NOT met: $indep" ;;
*) fail "microvm: no verdict recorded for the phase (done_when was set), got: ${indep:-<empty>}" ;;
esac
}
# ── Negative control: a backend no node can run ──────────────────
#
# Placement requires the mission's backend image to exist ON a node, not merely
# that the node has KVM. Without this check the positive scenario above would
# pass just as well against a scheduler that ignored `backend` entirely — which
# is what it did until the first real microvm mission landed on a node that had
# no such rootfs.
scenario_microvm_unavailable_backend() {
local token mission body status
token=$(mint_session) || { norun "microvm-negctl: could not mint a session"; return 1; }
body=$(printf '%s' "$MICROVM_BODY" | sed 's/"backend":"claude"/"backend":"definitely-not-built"/')
mission=$(create_mission "$token" "$(echo "$body" | tr -d '\n')") \
|| { norun "microvm-negctl: mission create failed"; return 1; }
info "microvm-negctl: mission=$mission"
api "$token" PATCH "/api/missions/$mission/status" '{"status":"running"}' >/dev/null 2>&1
sleep 5
status=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select status from missions where id='$mission';\"" | head -1 | tr -d '[:space:]')
if [ "$status" = "draft" ]; then
pass "microvm-negctl: a backend no node can run is refused at launch (mission stayed draft)"
else
fail "microvm-negctl: a mission asking for an unbuilt image reached status=$status"
fi
}
# ── Negative control: a gate that gives up must fail the phase ───
#
# `done_when_check` is run in exactly ONE place — the Stop hook inside the
# guest. The hook is capped so a stuck agent cannot wedge the turn, and at that
# cap it lets the agent stop. Until the `capped` marker existed, that release
# was invisible: the process exits 0 and the work IS collected, so both signals
# the run status was decided from said "fine" and the phase completed GREEN with
# its check still failing.
#
# The check here is `exit 1`, which no agent can satisfy. That is the point: it
# guarantees the cap is reached, so this exercises the release path itself
# rather than hoping to catch it. `blocks` reaching the cap is NOT the
# assertion — a healthy agent blocked three times and succeeding on the fourth
# reports the same number. The phase STATUS is the assertion.
GATECAP_BODY=$(cat <<JSON
{"title":"verify: a gate that gives up fails the phase",
"template_kind":"research_and_code",
"repo_id":"$REPO_ID",
"runtime_kind":"microvm",
"backend":"claude",
"description":"Negative control for the stop gate's cap release.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"done_when_check":"exit 1",
"task":"Create a file NOTES.md at the repository root containing the single word: hello. Create no other files."}}
]}
JSON
)
assert_gate_cap() { # <token> <mission> <report>
local mission="$2" report="$3" blocks note
while read -r idx status _files _pushed _branch _cerr _perr; do
if [ "$status" = "completed" ]; then
fail "gatecap: phase $idx COMPLETED with a check that cannot pass — the cap release is invisible again"
else
pass "gatecap: phase $idx did not complete (status=$status)"
fi
done <<<"$report"
# The gate must actually have run out of blocks; a phase that failed for some
# OTHER reason would satisfy the status check above while proving nothing.
blocks=$(ssh "$HOST" "docker logs --since 60m clawmates_server_1 2>&1 \
| grep -F 'microvm phase' | grep -F '$mission' | tail -1" \
| sed -n 's/.*stop-gate blocks: \([0-9-]*\).*/\1/p')
case "$blocks" in
3) pass "gatecap: the gate spent all 3 blocks before giving up" ;;
'') fail "gatecap: no stop-gate field in the log — was a gate installed at all?" ;;
*) fail "gatecap: expected 3 blocks before the cap, got $blocks" ;;
esac
note=$(ssh "$HOST" "docker logs --since 60m clawmates_server_1 2>&1 \
| grep -F 'microvm phase' | grep -F '$mission' | tail -1")
case "$note" in
*"completion gate released"*)
pass "gatecap: the failure names the cap release as the reason" ;;
*)
fail "gatecap: the phase failed but not for the gate's reason: $(printf '%s' "$note" | tail -c 200)" ;;
esac
}
# ── Scenario: research_only — a repo-less mission keeps its work ──
#
# The portal offers five recipes; until now the harness exercised ONE
# (research_and_code). `research_only` is the recipe that lost real work: it sets
# `requires_repo = false`, so `capture_finished_coding_phases` (which selects
# `AND m.repo_id IS NOT NULL`) never looked at it, the container was reaped
# unread, and eight ClawHDF5 research documents were destroyed while the mission
# reported `completed`.
#
# Nothing in the suite would have caught that, because nothing ever ran this
# recipe. That is the gap this closes.
RESEARCH_ONLY_BODY=$(cat <<JSON
{"title":"verify: a repo-less research mission keeps its output",
"template_kind":"research_only",
"team_template_id":"$TEAM_TEMPLATE",
"description":"Prove mission_outputs captures a phase with no git checkout.",
"phases":[
{"kind":"research","order_idx":0,"config":{"max_iterations":1,
"task":"Write exactly two markdown files under /mission/repo/research/: 01_findings.md and 02_notes.md. Each 5-10 lines about Rust error handling. Create no other files."}}
]}
JSON
)
# The same repo-less research, in a microVM.
#
# This path could not run at all until now: `run_phase_in_vm` packed a checkout
# directory that does not exist for a repo-less mission, and then demanded a
# `.git` inside the guest that never would. So it failed before the agent got a
# turn, and `mission_outputs` excluded microvm entirely on the grounds that a VM
# mission "always has a checkout". Both halves are gone; this is what proves it.
RESEARCH_VM_BODY=$(cat <<JSON
{"title":"verify: a repo-less research mission in a microVM",
"template_kind":"research_only",
"team_template_id":"$TEAM_TEMPLATE",
"runtime_kind":"microvm",
"backend":"claude",
"description":"Prove a mission with no repository runs in a VM and its output comes back.",
"phases":[
{"kind":"research","order_idx":0,"config":{"max_iterations":1,
"task":"Write exactly two markdown files under /mission/repo/research/: 01_findings.md and 02_notes.md. Each 5-10 lines about Rust error handling. Create no other files."}}
]}
JSON
)
assert_research_only() { # <token> <mission> <report>
local token="$1" mission="$2" report="$3" docs names scaffold
while read -r idx status _f _p _b _c _e; do
[ "$status" = "completed" ] \
&& pass "research-only: phase $idx completed" \
|| fail "research-only: phase $idx status=$status"
done <<<"$report"
# The artifacts are the whole point: without them the mission is the silent
# loss this scenario exists to detect.
docs=$(api "$token" GET "/api/missions/$mission" | python3 -c '
import json,sys
d=json.load(sys.stdin)
print(sum(1 for a in (d.get("artifacts") or []) if a.get("kind")=="document"))')
case "$docs" in
""|0) fail "research-only: NO documents captured — a repo-less phase lost its work" ;;
*) pass "research-only: $docs document artifact(s) captured from a mission with no repo" ;;
esac
# The agent runtime seeds SOUL.md/MEMORY.md/etc into the workspace root. In a
# repo-backed mission `.git/info/exclude` hides them; a repo-less mission has
# no `.git`, and the first live run published all seven as artifacts.
scaffold=$(api "$token" GET "/api/missions/$mission" | python3 -c '
import json,sys
SEED={"AGENTS.md","HEARTBEAT.md","IDENTITY.md","MEMORY.md","SOUL.md","TOOLS.md","USER.md"}
d=json.load(sys.stdin)
print(sum(1 for a in (d.get("artifacts") or []) if (a.get("title") or "") in SEED))')
[ "${scaffold:-0}" = "0" ] \
&& pass "research-only: the agent's own identity files were not published" \
|| fail "research-only: $scaffold agent scaffolding file(s) leaked into artifacts"
# And the text must actually be readable — an artifact row pointing at nothing
# is a 404 with no explanation, which is how a reader experiences lost work.
names=$(api "$token" GET "/api/missions/$mission" | python3 -c '
import json,sys
d=json.load(sys.stdin)
a=[x for x in (d.get("artifacts") or []) if x.get("kind")=="document"]
print(a[0]["id"] if a else "")')
if [ -n "$names" ]; then
local body
body=$(api "$token" GET "/api/missions/$mission/artifacts/$names/content" | python3 -c '
import json,sys
try: print(len(json.load(sys.stdin).get("content") or ""))
except Exception: print(0)')
[ "${body:-0}" -gt 0 ] \
&& pass "research-only: the captured document reads back ($body chars)" \
|| fail "research-only: the artifact exists but its content is unreadable"
fi
}
# ── Scenario: benchmark — a non-coding phase must still deliver ───
#
# A `benchmark` mission is ONE benchmark phase. `empty_delivery_is_a_failure`
# tested `kind == "coding"`, so that phase was exempt and NOTHING in the platform
# could fail it: an agent that produced no benchmark at all reported success.
# This runs the recipe the guard now covers.
BENCHMARK_BODY=$(cat <<JSON
{"title":"verify: a benchmark phase delivers files",
"template_kind":"benchmark",
"team_template_id":"$TEAM_TEMPLATE",
"repo_id":"$REPO_ID",
"description":"Prove the delivery guard covers a non-coding producing phase.",
"phases":[
{"kind":"benchmark","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"Create BENCH.md at the repository root recording a simple timing measurement you actually ran (loop a cheap operation and time it). Create no other files."}}
]}
JSON
)
assert_benchmark() { # <token> <mission> <report>
local report="$3"
while read -r idx status files _p _b cerr perr; do
[ "$status" = "completed" ] \
&& pass "benchmark: phase $idx completed" \
|| fail "benchmark: phase $idx status=$status (commit_error=$cerr push_error=$perr)"
case "$files" in
0|-) fail "benchmark: phase $idx delivered no files — the widened guard did not fire" ;;
*) pass "benchmark: phase $idx delivered $files file(s), and a non-coding phase is now held to it" ;;
esac
done <<<"$report"
}
# ── Scenario: security_hardening — the scan phase must deliver ────
#
# The third recipe whose defining phase is not `coding`, and so the third that
# nothing in the platform could fail until `PRODUCING_KINDS` widened. A
# `security_scan` phase that ran no scanner and wrote nothing reported success.
#
# One phase, not the recipe's full scan->research->code chain: what is under test
# is the phase KIND, and the later two are kinds the suite already covers. The
# scanners are real (`cargo-audit`, `gitleaks`, `trivy`, `semgrep` are all
# present in agent-claude:dev — verified in the image, not assumed from the
# directive that names them).
SECURITY_BODY=$(cat <<JSON
{"title":"verify: a security scan phase delivers findings",
"template_kind":"security_hardening",
"team_template_id":"$TEAM_TEMPLATE",
"repo_id":"$REPO_ID",
"description":"Prove the delivery guard covers a security_scan phase.",
"phases":[
{"kind":"security_scan","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"Run gitleaks against this repository (it is installed). Write SECURITY.md at the repository root recording the exact command you ran and what it reported — including 'no findings' if that is the result. Create no other files."}}
]}
JSON
)
assert_security() { # <token> <mission> <report>
local mission="$2" report="$3" evidence
while read -r idx status files _p _b cerr perr; do
[ "$status" = "completed" ] \
&& pass "security: phase $idx completed" \
|| fail "security: phase $idx status=$status (commit_error=$cerr push_error=$perr)"
case "$files" in
0|-) fail "security: phase $idx delivered no files — a scan that records nothing is not a scan" ;;
*) pass "security: phase $idx delivered $files file(s), and a security_scan phase is now held to it" ;;
esac
done <<<"$report"
# Delivering A FILE is not the same as running the scanner. An agent that
# wrote "I scanned it, all clear" satisfies the guard above while doing
# nothing — the exact letter-not-purpose shape the evaluator exists for.
# The tool's own output is the evidence, so require it by name.
evidence=$(ssh "$HOST" "docker exec clawmates_server_1 sh -c 'find \
/var/lib/clawmates-missions/_outputs/$mission -name \"*.patch\" \
-exec grep -lc gitleaks {} \; 2>/dev/null | head -1'" | tr -d '[:space:]')
[ -n "$evidence" ] \
&& pass "security: the delivered file carries the scanner's own output" \
|| fail "security: nothing delivered mentions gitleaks — the scan may not have run"
}
# ── Scenario: refactor — the on_green_tests commit gate ───────────
#
# The fifth and last recipe, and NOT just another coding phase: it declares
# `commit_policy = "on_green_tests"`, and all ten other fixtures use `"always"`.
# So the gate that decides whether work is published on the mission branch or
# diverted for review has never run end to end — and `Gate`'s own doc records
# that three recipes carried this policy while it "did precisely nothing",
# because it had no reader at all.
#
# A policy that is parsed but never exercised is indistinguishable from one that
# is ignored. This runs it: tests pass, so the work must land on the mission
# branch, not a review branch.
REFACTOR_BODY=$(cat <<JSON
{"title":"verify: the on_green_tests gate publishes when tests pass",
"template_kind":"refactor",
"team_template_id":"$TEAM_TEMPLATE",
"repo_id":"$REPO_ID",
"description":"Prove commit_policy=on_green_tests reaches delivery.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"on_green_tests","max_iterations":1,
"task":"Add a file KEEP.md at the repository root containing one sentence about this crate. Do not modify any existing file, and do not add or change any test."}}
]}
JSON
)
assert_refactor() { # <token> <mission> <report>
local report="$3" branch_seen=""
while read -r idx status files pushed branch cerr perr; do
[ "$status" = "completed" ] \
&& pass "refactor: phase $idx completed" \
|| fail "refactor: phase $idx status=$status (commit_error=$cerr push_error=$perr)"
case "$files" in
0|-) fail "refactor: phase $idx delivered no files" ;;
*) pass "refactor: phase $idx delivered $files file(s)" ;;
esac
branch_seen="$branch"
done <<<"$report"
# The gate's whole job is WHERE the work lands. A failed gate does not discard
# work — it diverts it to a review branch — so a green gate must NOT divert.
case "$branch_seen" in
"" |-) fail "refactor: no branch recorded — the gate's outcome is unobservable" ;;
*-review|*-needs-review)
fail "refactor: work landed on a REVIEW branch ($branch_seen) though the gate should have passed" ;;
*) pass "refactor: on_green_tests published to the mission branch ($branch_seen)" ;;
esac
}
# ── Scenario: the two engines composed ───────────────────────────
#
# A `team_engine=composed` mission is a durable ZeroClaw graph whose every node
# is a whole Claude-Code-in-a-microVM session. The property that cannot be
# checked any other way is the FILE HANDOFF: a VM is inject → run → collect →
# destroy, so unless the tree is carried node to node, node 2 boots from the
# original checkout, sees none of node 1's work, and still reports success.
#
# The task is written so the delivered file IS the evidence. Each node appends
# one line; a run that lost the handoff delivers a file with one line, and no
# amount of agent confidence can fake the missing ones.
COMPOSED_BODY=$(cat <<JSON
{"title":"verify: the two engines composed",
"template_kind":"research_and_code",
"team_template_id":"$TEAM_TEMPLATE",
"repo_id":"$REPO_ID",
"runtime_kind":"microvm",
"backend":"claude",
"team_engine":"composed",
"description":"Prove a ZeroClaw graph of microVM nodes carries file work between its nodes.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"done_when":"STAGES.md exists at the repository root and contains one line per stage, each recording a stage name and a kernel version such as 6.1.128.",
"task":"Append EXACTLY ONE line to STAGES.md at the repository root, creating the file if it does not exist. The line is: your stage name, a space, and the output of \`uname -r\`.\n\nDo NOT rewrite, reorder or remove any line already in the file — earlier stages wrote those, and they are the record of this run. If the file already has lines, yours goes after them. Change no other file."}}
]}
JSON
)
# The graph nodes that actually ran, from the composed executor's own log line.
composed_steps() { # composed_steps <run-id-or-mission>
ssh "$HOST" "docker logs --since 90m clawmates_server_1 2>&1 \
| grep -F 'microvm_turn_executor:' | grep -c 'ok — subagents'" | tr -d '[:space:]'
}
assert_composed() { # <token> <mission> <report>
local token="$1" mission="$2" report="$3" delivered lines tier steps
while read -r idx status files pushed _branch cerr perr; do
[ "$status" = "completed" ] || fail "composed: phase $idx status=$status (commit_error=$cerr push_error=$perr)"
[ "$pushed" = "True" ] || fail "composed: phase $idx not pushed (push_error=$perr)"
case "$files" in 0|-) fail "composed: phase $idx delivered no files" ;; esac
done <<<"$report"
# The run must have gone through the WORKER, on the composed tier. A composed
# mission that quietly fell back to the solo path would deliver a one-line
# file and look like a graph that ran one node.
tier=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select tier from topology_runs where mission_id='$mission' order by created_at desc limit 1;\"" \
| head -1 | tr -d '[:space:]')
if [ "$tier" = "microvm_graph" ]; then
pass "composed: the run was claimed by the worker on the composed tier"
else
fail "composed: run tier is '$tier', not microvm_graph — it did not compose"
fi
delivered=$(fetch_delivered "$token" "$mission" STAGES.md) \
|| { fail "composed: could not read STAGES.md from the pushed branch"; return 1; }
lines=$(printf '%s\n' "$delivered" | grep -c '[^[:space:]]')
steps=$(composed_steps "$mission")
# THE assertion. One line per node means every node saw what the last one
# left; fewer means the tree did not survive a node boundary, which is the
# silent-success shape this whole slice exists to prevent.
if [ "${lines:-0}" -ge 2 ]; then
pass "composed: STAGES.md carries $lines stage lines — the tree survived the node boundary"
else
fail "composed: STAGES.md has $lines line(s), so a later node did not see the earlier node's file: $(printf '%s' "$delivered" | tr '\n' '|')"
fi
info "composed: executor logged $steps node turn(s); delivered:$(printf '%s' "$delivered" | tr '\n' '|')"
# The kernel line again — a composed node must still be a VM, not a container
# on the gateway.
local kernel
kernel=$(printf '%s\n' "$delivered" | tail -1 | awk '{print $NF}')
case "$kernel" in
"$GW_KERNEL") fail "composed: the last node ran on the GATEWAY kernel ($kernel)" ;;
"$NODE_KERNEL") fail "composed: the last node ran on the fleet NODE's kernel ($kernel)" ;;
"") fail "composed: no kernel on the last line" ;;
*) pass "composed: the last node ran under a guest kernel ($kernel)" ;;
esac
}
# The stop gate, inside a real VM.
#
# Reported by `phase_runner` for every microVM phase. Three distinguishable
# outcomes, and the difference matters: a number means the gate was installed
# and this is how often it sent the agent back; `-` means it could not be
# installed at all (usually: the CLI in the image has no `--settings`), which is
# a degradation the log states rather than a silent absence.
assert_stop_gate() { # <mission> <label>
local blocks
blocks=$(ssh "$HOST" "docker logs --since 90m clawmates_server_1 2>&1 \
| grep -F 'microvm phase' | grep -F '$1' | tail -1" \
| sed -n 's/.*stop-gate blocks: \([0-9-]*\).*/\1/p')
case "$blocks" in
'') fail "$2-gate: the phase reported no stop-gate field at all — old server image?" ;;
'-') fail "$2-gate: the gate was NOT installed in the VM (check the server log for the --settings probe)" ;;
0) pass "$2-gate: the gate was installed and never fired — the agent finished the work first time" ;;
*) pass "$2-gate: the gate sent the agent back $blocks time(s) inside its own turn" ;;
esac
}
# ── Scenario: a model sizes the team ─────────────────────────────
#
# Slice 5. The planner proposes a roster for THIS mission, a human approves it,
# and the mission runs the graph the model chose — not the team template's.
#
# The proof is the delivered file: one line per member the model proposed. A
# roster that was accepted, stored and then silently ignored at launch — which is
# exactly what the first live approval did — delivers the template's node count
# instead, or one line from the solo path.
ROSTER_BODY=$(cat <<JSON
{"title":"verify: a model sizes this mission",
"template_kind":"research_and_code",
"repo_id":"$REPO_ID",
"runtime_kind":"microvm",
"backend":"claude",
"description":"Add a small, self-contained change and have it independently verified.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"Append EXACTLY ONE line to ROSTER.md at the repository root, creating it if absent: your stage name, a space, and the output of \`uname -r\`. Do not remove or rewrite lines already there — earlier stages wrote them. Change no other file."}}
]}
JSON
)
# The whole Slice 5 flow, which is why this is not a plain `run_scenario`: the
# mission must be shaped by an approved proposal BEFORE it launches.
scenario_roster() {
local token mission proposal members engine nodes tier delivered lines status
token=$(mint_session) || { norun "roster: could not mint a session"; return 1; }
mission=$(create_mission "$token" "$(echo "$ROSTER_BODY" | tr -d '\n')") \
|| { norun "roster: mission create failed"; return 1; }
info "roster: mission=$mission"
# 1. Ask the model. Keep the body: "no usable proposal" hid the actual
# reason (first a credit wall, then a subscription rate limit) behind a
# generic NORUN, and the reason is the only thing an operator can act on.
local ask
ask=$(api "$token" POST "/api/missions/$mission/team-proposals" '{}' 2>&1)
read -r proposal members <<<"$(api "$token" GET "/api/missions/$mission/team-proposals" | python3 -c '
import json, sys
d = json.load(sys.stdin)
if d: print(d[0]["id"], len(d[0]["roster"]["members"]))
' 2>/dev/null)"
if [ -z "${proposal:-}" ]; then
norun "roster: no usable proposal — the planner said: $(printf '%s' "$ask" | tr -d '\n' | cut -c1-200)"
return 1
fi
pass "roster: the planner sized this mission at $members member(s)"
# 2. Approve it, and check it actually LANDED on the mission. The first live
# approval returned an error while leaving the proposal marked approved,
# so "the API said ok" is not the assertion — the mission row is.
api "$token" POST "/api/missions/$mission/team-proposals/$proposal/decide" \
'{"status":"approved"}' >/dev/null 2>&1
read -r engine nodes <<<"$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select coalesce(team_engine,'-'), coalesce(jsonb_array_length(config->'roster'->'nodes'), 0) from missions where id='$mission';\"" \
| head -1 | tr '|' ' ')"
if [ "$engine" = "composed" ] && [ "${nodes:-0}" = "$members" ]; then
pass "roster: the approved roster is on the mission ($nodes nodes, engine $engine)"
else
fail "roster: approval did not reach the mission (engine=$engine nodes=${nodes:-0}, expected $members)"
return 1
fi
# 3. Run it.
api "$token" PATCH "/api/missions/$mission/status" '{"status":"running"}' >/dev/null
status=$(await_mission "$token" "$mission")
[ "$status" != "timeout" ] || { norun "roster: mission did not finish in ${MISSION_TIMEOUT}s"; return 1; }
printf '%s\n' "$(phase_report "$token" "$mission")" | sed 's/^/ phase /'
tier=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select tier from topology_runs where mission_id='$mission' order by created_at desc limit 1;\"" \
| head -1 | tr -d '[:space:]')
if [ "$tier" = "microvm_graph" ]; then
pass "roster: the mission ran the composed graph the model chose"
else
fail "roster: run tier is '$tier' — the roster was accepted and then ignored"
fi
delivered=$(fetch_delivered "$token" "$mission" ROSTER.md) \
|| { fail "roster: could not read ROSTER.md from the pushed branch"; return 1; }
lines=$(printf '%s\n' "$delivered" | grep -c '[^[:space:]]')
# Counted as a DELTA against main, not as a total.
#
# The absolute form assumed ROSTER.md starts empty, and it does not: the
# auto-merge work put an earlier run's two lines onto main, so the next
# 1-member roster delivered three lines and was reported as a model that
# ignored its own proposal. It had done exactly the right thing. Any check
# against a scratch repo that ACCUMULATES has to measure what this run added.
base=$(fetch_main "ROSTER.md" | grep -c '[^[:space:]]')
local added=$((lines - base))
if [ "$added" = "$members" ]; then
pass "roster: this run added one line per proposed member ($added, on top of $base already on main)"
else
fail "roster: this run added $added line(s) for a $members-member roster (branch=$lines main=$base): $(printf '%s' "$delivered" | tr '\n' '|')"
fi
check_single_uid "$mission" roster
}
run_scenario() { # run_scenario <label> <json> <assert-fn> [no-checkout]
local label="$1" body="$2" assert_fn="$3" checkout="${4:-checkout}" token mission status
# Every one of these MUST go through fail()/norun(). The first version of
# this function called a `die` that lived inside `$(...)` — which exits the
# command substitution's subshell, not the script — so a run where the
# session could not be minted printed two ABORT lines, incremented nothing,
# and ended with "all checks passed" and exit 0. The harness written to
# catch silent success produced silent success on its first real run.
token=$(mint_session) || { norun "$label: could not mint a session"; return 1; }
mission=$(create_mission "$token" "$body") || { norun "$label: mission create failed"; return 1; }
info "$label: mission=$mission"
api "$token" PATCH "/api/missions/$mission/status" '{"status":"running"}' >/dev/null
status=$(await_mission "$token" "$mission")
if [ "$status" = "timeout" ]; then
norun "$label: mission did not reach a terminal status in ${MISSION_TIMEOUT}s"
return 1
fi
local report; report=$(phase_report "$token" "$mission")
if [ -z "$report" ]; then
norun "$label: no phases reported — nothing executed"
return 1
fi
printf '%s\n' "$report" | sed 's/^/ phase /'
"$assert_fn" "$token" "$mission" "$report"
# The single-writer invariant is a property OF A CHECKOUT. A repo-less mission
# has none by design (`ensure_checkout` returns Ok(None)), so probing for one
# reports a platform fault that is really a category error.
#
# Declared per scenario, never inferred from "the directory is missing": that
# inference would silently excuse a repo-BACKED mission whose checkout was
# reaped early — which is exactly the condition this probe exists to catch.
if [ "$checkout" = "no-checkout" ]; then
info "$label: no checkout to probe (repo-less mission) — single-writer check n/a"
else
check_single_uid "$mission" "$label"
fi
}
# ── Scenario: phase continuity ───────────────────────────────────
#
# Phase 1 must READ what phase 0 wrote. Independent phases cannot tell a
# preserved checkout from a wiped one — which is exactly how a `reset --hard`
# survived three green-looking runs.
CHAIN_BODY=$(cat <<JSON
{"title":"verify: phase 2 builds on phase 1",
"template_kind":"research_and_code",
"team_template_id":"$TEAM_TEMPLATE",
"repo_id":"$REPO_ID",
"description":"Build up a file CHAIN.md across two phases.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"Create a file named CHAIN.md at the repository root containing exactly one line: STEP-ONE-OK. Do not create or modify any other file."}},
{"kind":"coding","order_idx":1,"config":{"commit_policy":"always","max_iterations":1,
"task":"Read the existing file CHAIN.md at the repository root. It was written by the previous phase and must already contain the line STEP-ONE-OK. Append a second line reading STEP-TWO-SAW-STEP-ONE, keeping the first line intact. If CHAIN.md does not exist, instead create a file named CHAIN_MISSING.md containing the single line PRIOR-PHASE-WORK-WAS-LOST, and do not create CHAIN.md."}}
]}
JSON
)
assert_chain() { # <token> <mission> <report>
local token="$1" mission="$2" report="$3" delivered
local phases; phases=$(printf '%s\n' "$report" | wc -l | tr -d ' ')
[ "$phases" = "2" ] || fail "chain: expected 2 phases, got $phases"
# Here-string, not a pipe: a `while` on the right of a pipe runs in a
# subshell, so every fail() inside it would increment a FAILURES that dies
# with the subshell and the script would exit 0 having reported failures.
while read -r idx status files pushed _branch cerr perr; do
[ "$status" = "completed" ] || fail "chain: phase $idx status=$status"
[ "$pushed" = "True" ] || fail "chain: phase $idx not pushed (commit_error=$cerr push_error=$perr)"
case "$files" in 0|-) fail "chain: phase $idx delivered no files" ;; esac
done <<<"$report"
delivered=$(fetch_delivered "$token" "$mission" CHAIN.md) \
|| { fail "chain: could not read CHAIN.md from the pushed branch"; return 1; }
case "$delivered" in
*STEP-ONE-OK*STEP-TWO-SAW-STEP-ONE*)
pass "chain: phase 1 read phase 0's work and appended to it" ;;
*PRIOR-PHASE-WORK-WAS-LOST*)
fail "chain: phase 1 reported the prior phase's work was lost" ;;
*)
fail "chain: CHAIN.md does not show both lines: $(printf '%s' "$delivered" | tr '\n' '|')" ;;
esac
}
# ── Scenario: multi-role with a real test suite ──────────────────
#
# The workload that failed with `COMMIT_EDITMSG: Permission denied` under the
# bind mount. REVIEW.md must carry cargo's own summary line, so the reviewer
# had to run the suite rather than assert that it passed.
MULTIROLE_BODY=$(cat <<'JSON'
{"title":"verify: implement, test, review",
"template_kind":"research_and_code",
"team_template_id":"__TEAM__",
"repo_id":"__REPO__",
"description":"Extend the scratch Rust crate with a reviewed, tested function.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"This task needs THREE distinct roles. Use a separate agent for each role; do not do all three yourself.\n1. IMPLEMENTER: in src/lib.rs add `pub fn divide(a: i64, b: i64) -> Option<i64>` returning None when b == 0, else Some(a / b).\n2. TESTER: add unit tests in the existing tests module named `divide_works` and `divide_by_zero_is_none`, covering both branches. Run `cargo test` and make it pass.\n3. REVIEWER: read the final src/lib.rs and write REVIEW.md at the repository root containing exactly three lines: `ROLES: 3`, `TESTS: <the test result summary line from cargo test>`, and `VERDICT: <one sentence>`.\nAll three files (src/lib.rs, REVIEW.md) must be left in the working tree."}}
]}
JSON
)
assert_multirole() { # <token> <mission> <report>
local token="$1" mission="$2" report="$3" review lib
while read -r idx status files pushed _branch cerr perr; do
[ "$status" = "completed" ] || fail "multirole: phase $idx status=$status"
[ "$pushed" = "True" ] || fail "multirole: phase $idx not pushed (commit_error=$cerr push_error=$perr)"
case "$files" in 0|-) fail "multirole: phase $idx delivered no files" ;; esac
done <<<"$report"
lib=$(fetch_delivered "$token" "$mission" src/lib.rs) \
|| { fail "multirole: could not read src/lib.rs from the pushed branch"; return 1; }
case "$lib" in
*"fn divide"*) pass "multirole: divide() was delivered" ;;
*) fail "multirole: src/lib.rs has no divide()" ;;
esac
case "$lib" in
*divide_by_zero_is_none*) pass "multirole: the zero-divisor test was delivered" ;;
*) fail "multirole: no divide_by_zero_is_none test" ;;
esac
review=$(fetch_delivered "$token" "$mission" REVIEW.md) \
|| { fail "multirole: no REVIEW.md on the pushed branch"; return 1; }
# cargo's own words. A reviewer who merely claimed the tests passed cannot
# produce this line, which is the point of asserting on it.
case "$review" in
*"test result: ok."*) pass "multirole: REVIEW.md carries cargo's own test summary" ;;
*) fail "multirole: REVIEW.md has no 'test result: ok.' line" ;;
esac
}
# ── Scenario: a phase that delivers nothing must FAIL ────────────
#
# The negative control for the delivery guard, and the same discipline as the
# uid self-test: a check that has never been seen to fire has not been shown to
# work. This phase is told to change nothing, so `empty_delivery_is_a_failure`
# must catch it — and the scenario PASSES when the phase comes back `failed`.
NOOP_BODY=$(cat <<JSON
{"title":"verify: a phase that delivers nothing must fail",
"template_kind":"research_and_code",
"team_template_id":"$TEAM_TEMPLATE",
"repo_id":"$REPO_ID",
"description":"Negative control for the empty-delivery guard.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"Do NOT create, modify or delete any file. Read README.md and reply with a one-sentence summary of it as your final answer. Leave the working tree exactly as you found it."}}
]}
JSON
)
assert_noop() { # <token> <mission> <report>
local report="$3" saw_failed=0
while read -r idx status files _pushed _branch _cerr _perr; do
case "$files" in
0|-) ;;
*) fail "noop: phase $idx changed $files file(s) — the scenario did not \
exercise the guard (agent ignored the instruction; re-run)"; continue ;;
esac
if [ "$status" = "failed" ]; then
saw_failed=1
else
fail "noop: phase $idx delivered nothing but reports status=$status"
fi
done <<<"$report"
[ "$saw_failed" = "1" ] && pass "noop: an empty coding phase was failed, not completed"
return 0
}
# ── Scenario: a mission served by the node's OWN GPU ──────────────
#
# `local-ornith` is Claude Code pointed at the Ollama on the node itself, over a
# vsock pipe rather than the egress proxy (`clawmates-node::local_model`). Three
# things have to be true at once and only a real run shows all three:
#
# 1. the agent reached a model AT ALL — a pipe to a closed port produces a
# turn that hangs rather than errors, which is why this is a scenario and
# not a unit test;
# 2. the work came back and landed on a branch, so the local model actually
# drove the CLI rather than merely answering;
# 3. the VM still could NOT reach api.anthropic.com. That is the measured
# property of every other backend and the reason `provider_hosts` is
# per-backend; a local backend that quietly kept Anthropic egress would be
# a credential path nobody asked for.
LOCAL_BODY=$(cat <<JSON
{"title":"verify: a mission on the node's own GPU",
"template_kind":"research_and_code",
"repo_id":"$REPO_ID",
"runtime_kind":"microvm",
"backend":"local-ornith",
"description":"Prove a coding phase runs against a locally-hosted model.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"done_when":"LOCAL.md exists at the repository root and contains a kernel version such as 6.1.128.",
"task":"Create a file named LOCAL.md at the repository root containing exactly one line: the output of running uname -r. Create no other files."}}
]}
JSON
)
assert_local() { # <token> <mission> <report>
local token="$1" mission="$2" report="$3" delivered node
while read -r idx status files pushed _branch cerr perr; do
[ "$status" = "completed" ] || fail "local-ornith: phase $idx status=$status"
[ "$pushed" = "True" ] || fail "local-ornith: phase $idx not pushed (commit_error=$cerr push_error=$perr)"
case "$files" in 0|-) fail "local-ornith: phase $idx delivered no files" ;; esac
done <<<"$report"
delivered=$(fetch_delivered "$token" "$mission" LOCAL.md) \
|| { fail "local-ornith: could not read LOCAL.md from the pushed branch"; return 1; }
# A GUEST kernel, so this also proves it ran in a VM rather than on a host.
case "$(printf '%s' "$delivered" | tr -d '[:space:]')" in
"$GW_KERNEL"|"$NODE_KERNEL")
fail "local-ornith: LOCAL.md holds a HOST kernel ($delivered) — that phase did not run in a guest" ;;
*[0-9].[0-9]*)
pass "local-ornith: a locally-served model delivered a guest kernel ($(printf '%s' "$delivered" | tr -d '[:space:]'))" ;;
*) fail "local-ornith: LOCAL.md is not a kernel version: $delivered" ;;
esac
# The negative control. Asked of the NODE's proxy log, which is the only
# record of what was actually dialled — a model's own account of where it got
# its tokens has no evidential value here.
node=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select coalesce(n.name,'') from missions m left join nodes n on n.id = m.target_node_id where m.id='$mission';\"" \
| head -1 | tr -d '[:space:]')
if [ -z "$node" ]; then
norun "local-ornith: could not tell which node ran it — egress cannot be checked"
else
local target; case "$node" in tank) target=osobh@tank ;; *) target="$node" ;; esac
if ssh "$target" "journalctl -u clawmates-node --since '-30 min' --no-pager 2>/dev/null \
| grep -q 'egress -> api.anthropic.com'"; then
fail "local-ornith: the VM reached api.anthropic.com — a local backend must not"
else
pass "local-ornith: no Anthropic egress from a locally-served mission"
fi
if ssh "$target" "journalctl -u clawmates-node --since '-30 min' --no-pager 2>/dev/null \
| grep -q 'local model socket'"; then
pass "local-ornith: the node bound its local-model socket for this VM"
else
fail "local-ornith: the node never bound a local-model socket — the guest had nothing to talk to"
fi
fi
}
# ── Provider proof for a microVM mission ─────────────────────────
#
# Which provider served a VM's turns is answered by the NODE's egress proxy
# log and by nothing the agent wrote: a z.ai-served agent called itself Claude
# Opus 5 (memory: glm-microvm-backend). The proxy allow-lists one provider host
# per backend and logs every dial and every denial, so "dialled the right host,
# never denied it, dialled no other" is the whole proof — and it is what the
# 2.1.276 rollout had to establish for glm and kimi, whose turns ride
# ANTHROPIC_BASE_URL through a CLI that broke that path twice between 2.1.226
# and 2.1.276.
placed_node() { # <mission> → ssh target of the node that ran it, or 1
local node
node=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select coalesce(n.name,'') from missions m left join nodes n on n.id = m.target_node_id where m.id='$1';\"" \
| head -1 | tr -d '[:space:]')
case "$node" in '') return 1 ;; tank) echo osobh@tank ;; *) echo "$node" ;; esac
}
vm_journal() { # <ssh-target> <mission> → this mission's VM lines only
# The VM id is m-<first 12 hex of the phase id>-<iteration>
# (microvm_executor::vm_id_for), so a grep on that prefix cannot pick up a
# neighbour's VM the way a time window can.
local p12
p12=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select substr(replace(id::text,'-',''),1,12) from mission_phases where mission_id='$2' order by order_idx limit 1;\"" \
| head -1 | tr -d '[:space:]')
[ -n "$p12" ] || return 1
ssh "$1" "journalctl -u clawmates-node --since '-3 hours' --no-pager 2>/dev/null | grep -F 'microvm m-$p12-'"
}
assert_provider_egress() { # <label> <mission> <provider-host>
local label="$1" mission="$2" want="$3" target lines hits denied other
target=$(placed_node "$mission") || { norun "$label: could not tell which node ran it — provider UNPROVEN"; return 1; }
lines=$(vm_journal "$target" "$mission") || { norun "$label: no journal lines for this mission's VM on $target"; return 1; }
hits=$(printf '%s\n' "$lines" | grep -c "egress -> $want")
denied=$(printf '%s\n' "$lines" | grep -c "egress DENIED $want")
if [ "$hits" -gt 0 ] && [ "$denied" -eq 0 ]; then
pass "$label: the VM dialled $want ${hits}x and was never denied it"
else
fail "$label: egress -> $want ${hits}x, DENIED ${denied}x — provider NOT proven"
fi
# Dials only. A DENIED api.anthropic.com on a glm/kimi VM is Claude Code's
# telemetry being refused, which is the proxy working, not a second provider.
other=$(printf '%s\n' "$lines" | grep 'egress -> ' | grep -v -e "egress -> $want" -e 'git.redclaw.dev' \
| sed 's/.*egress -> //' | sort -u | tr '\n' ' ')
if [ -z "$other" ]; then
pass "$label: nothing but the provider and the forge was reached"
else
fail "$label: other hosts reached: $other"
fi
}
assert_cli_version() { # <label> <mission> — needs checkpoint.vm (phase_runner, 2026-09-18)
local label="$1" mission="$2" got want="${CLAWMATES_EXPECT_CLI:-2.1.276}"
got=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select coalesce(checkpoint->'vm'->>'cli_version','') from topology_runs where mission_id='$mission' limit 1;\"" \
| head -1 | tr -d '\r')
case "$got" in
"$want"*) pass "$label: the guest ran Claude Code $got" ;;
'') fail "$label: no cli_version recorded on the run (server predates checkpoint.vm, or the probe failed)" ;;
*) fail "$label: guest CLI was '$got', expected $want" ;;
esac
got=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select coalesce(checkpoint->'vm'->>'rootfs','') from topology_runs where mission_id='$mission' limit 1;\"" \
| head -1 | tr -d '\r')
[ -n "$got" ] && pass "$label: booted $got" || fail "$label: no rootfs recorded on the run"
}
assert_claude_vm() { assert_microvm "$@"; assert_provider_egress microvm "$2" api.anthropic.com; assert_cli_version microvm "$2"; }
assert_glm() { assert_microvm "$@"; assert_provider_egress glm "$2" api.z.ai; assert_cli_version glm "$2"; }
assert_kimi() { assert_microvm "$@"; assert_provider_egress kimi "$2" api.kimi.com; assert_cli_version kimi "$2"; }
# ── Scenario: a burst larger than the fleet QUEUES, and spreads ───
#
# Phase 1 of the fleet-intelligence plan shipped placement-at-phase-launch and
# a queue built out of `start_pending_phases` leaving a phase `pending`. Both
# were deployed unproven under load, which is the condition this project keeps
# getting burned by: the code is right, the system is wrong, and nothing errors.
#
# Two invariants, and the second is the one that costs money to get wrong:
#
# 1. No node ever runs more VMs than `vm_placement` said it had room for.
# Overcommit does not fail loudly — it swaps, and every mission on that
# node gets slow instead of dead.
# 2. The excess QUEUES. A burst that silently drops the extras, or wedges
# them forever, both "pass" any check that only looks at the end state.
#
# Capacity comes from `/api/fleet/capacity`, which returns `vm_placement`'s own
# survey. Recomputing the slot arithmetic here would let this test drift from
# the scheduler and then agree with itself.
fleet_capacity() { # fleet_capacity <token> -> json
api "$1" GET "/api/fleet/capacity?backend=claude"
}
CAPACITY_TASK='Create a file named BURST.md at the repository root containing exactly one line: BURST-OK. Do not create or modify any other file.'
scenario_capacity() {
local token cap slots burst i mission ids=() idlist
token=$(mint_session) || { norun "capacity: could not mint a session"; return 1; }
cap=$(fleet_capacity "$token")
slots=$(printf '%s' "$cap" | python3 -c 'import json,sys
try: print(json.load(sys.stdin)["slots"])
except Exception: pass' 2>/dev/null)
# A probe returns a value or fails. An unreadable capacity survey must not
# become a burst of size 2 that passes because it never saturated anything.
case "${slots:-}" in
''|*[!0-9]*) norun "capacity: could not read /api/fleet/capacity: $(printf '%s' "$cap" | head -c 200)"; return 1 ;;
esac
local fitnodes; fitnodes=$(printf '%s' "$cap" | python3 -c 'import json,sys; print(len(json.load(sys.stdin)["nodes"]))')
info "capacity: fleet has $slots free slot(s) across $fitnodes node(s)"
printf '%s' "$cap" | python3 -c 'import json,sys
for n in json.load(sys.stdin)["nodes"]: print(" capacity: %-12s %d slot(s), %d committed" % (n["name"], n["slots"], n["committedVms"]))'
[ "$slots" -gt 0 ] || { norun "capacity: fleet reports 0 slots — nothing to saturate"; return 1; }
# Two more than the fleet can hold, so the queue is EXERCISED rather than
# merely available. Overridable for a cheaper smoke run, but a burst that
# does not exceed capacity is reported as NORUN below, never as PASS.
burst="${CAPACITY_BURST:-$((slots + 2))}"
info "capacity: launching a burst of $burst mission(s)"
for i in $(seq 1 "$burst"); do
mission=$(create_mission "$token" "$(cat <<JSON | tr -d '\n'
{"title":"verify: burst $i of $burst",
"template_kind":"research_and_code",
"repo_id":"$REPO_ID",
"runtime_kind":"microvm",
"backend":"claude",
"description":"Placement load test — one trivial file.",
"phases":[{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"$CAPACITY_TASK"}}]}
JSON
)") || { norun "capacity: mission $i of $burst failed to create"; return 1; }
ids+=("$mission")
done
idlist=$(printf "'%s'," "${ids[@]}"); idlist="${idlist%,}"
# Launch them as close to simultaneously as the API allows. Serialising the
# launches would let each placement see the previous VM already counted,
# which is the easy case; the race is the point.
#
# Each PATCH is an ssh + `docker run curl`, and 16 of those at once does not
# always land: a first run of this scenario left 3 missions in `draft` and
# then spent the full 1800s waiting for them, reporting a platform timeout
# for a launch that never happened. Discarding the launch response is the
# same swallowed-error shape this harness exists to catch, so the launches
# are now VERIFIED against the mission rows rather than assumed.
launch_all() { # launch_all <mission…>
local m
for m in "$@"; do
api "$token" PATCH "/api/missions/$m/status" '{"status":"running"}' >/dev/null 2>&1 &
done
wait
}
launch_all "${ids[@]}"
# ── Sample while it runs ───────────────────────────────────────
#
# `capacity_blocked_since` is CLEARED the moment a phase is placed, so a
# post-hoc query cannot prove a queue ever formed. The evidence exists only
# while the burst is in flight, and that window turned out to be SHORT: the
# first runs took 25 minutes because every VM paid a cold 2.4 GB rootfs copy;
# warm, the same 16 missions finish in 70-140s each and the whole burst is
# over in about two minutes. A sampler that waited 90s to start and ticked
# every 15s caught three samples of the tail and concluded the fleet was
# idle.
#
# So: one query per tick (three separate ones came back empty under load and
# empty was read as "nothing running"), a 5s tick, sampling from the first
# moment rather than after the launch check, and the launch check folded into
# the same query so it costs nothing.
local waited=0 peak_file queued=0 overcommit="" samples=0 blind=0 relaunched=0
peak_file=$(mktemp)
local sample_sql="SELECT
(SELECT count(*) FROM missions WHERE id IN ($idlist)
AND status IN ('completed','failed','cancelled')),
(SELECT count(*) FROM mission_phases p JOIN missions m ON m.id = p.mission_id
WHERE m.id IN ($idlist) AND p.capacity_blocked_since IS NOT NULL),
(SELECT count(*) FROM missions WHERE id IN ($idlist) AND status = 'draft'),
COALESCE((SELECT string_agg(nm || '=' || c, ' ') FROM (
SELECT COALESCE(n.name,'unpinned') AS nm, count(*) AS c
FROM mission_phases p JOIN missions m ON m.id = p.mission_id
LEFT JOIN nodes n ON n.id = m.target_node_id
WHERE m.id IN ($idlist) AND p.status = 'running'
GROUP BY 1) s), '')"
local sample_one; sample_one=$(printf '%s' "$sample_sql" | tr '\n' ' ')
while [ "$waited" -lt "$MISSION_TIMEOUT" ]; do
local sample done_count blocked drafts running
sample=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"$sample_one\"" 2>&1 | head -1 | tr -d '\r')
IFS='|' read -r done_count blocked drafts running <<<"$sample"
case "${done_count:-}" in
''|*[!0-9]*)
# Unreadable, not empty. Counted, never mistaken for an idle fleet.
blind=$((blind + 1)) ;;
*)
samples=$((samples + 1))
[ "${blocked:-0}" -gt 0 ] 2>/dev/null && queued=1
for pair in $running; do printf '%s|%s\n' "${pair%%=*}" "${pair##*=}" >> "$peak_file"; done
# A PATCH that silently failed leaves a mission in draft forever, and
# then the whole burst waits out the timeout looking like a stall.
# Retried once, from inside the sampling loop so the observation window
# is not spent waiting to find out.
if [ "${drafts:-0}" -gt 0 ] && [ "$relaunched" = "0" ] && [ "$waited" -ge 20 ]; then
relaunched=1
info "capacity: ${drafts} mission(s) never left draft — retrying the launch"
# shellcheck disable=SC2046
launch_all $(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select id from missions where id in ($idlist) and status = 'draft';\"" \
2>/dev/null | tr -d '\r' | grep -v '^[[:space:]]*$')
fi
[ "$done_count" = "$burst" ] && break ;;
esac
sleep 5
waited=$((waited + 5))
done
if [ "$waited" -ge "$MISSION_TIMEOUT" ]; then
norun "capacity: the burst did not finish in ${MISSION_TIMEOUT}s ($samples good samples, $blind blind)"
rm -f "$peak_file"
return 1
fi
info "capacity: took $samples sample(s) over ${waited}s ($blind unreadable)"
if [ "$samples" -lt 3 ]; then
norun "capacity: only $samples usable sample(s) ($blind unreadable) — too blind to judge the burst"
rm -f "$peak_file"
return 1
fi
# ── Invariant 1: nobody exceeded their slot count ──────────────
local peaks
peaks=$(sort -t'|' -k1,1 "$peak_file" | grep -v '^\s*$' | python3 -c '
import sys, collections
peak = collections.defaultdict(int)
for line in sys.stdin:
line = line.strip()
if not line or "|" not in line: continue
name, count = line.rsplit("|", 1)
try: peak[name.strip()] = max(peak[name.strip()], int(count))
except ValueError: pass
for k, v in sorted(peak.items()): print(k, v)
')
rm -f "$peak_file"
if [ -z "$peaks" ]; then
norun "capacity: never observed a running phase — the burst did not execute"
return 1
fi
while read -r name peak; do
[ -n "$name" ] || continue
local allowed
allowed=$(printf '%s' "$cap" | python3 -c "import json,sys
n = {x['name']: x['slots'] for x in json.load(sys.stdin)['nodes']}
print(n.get('$name', -1))")
if [ "$allowed" = "-1" ]; then
fail "capacity: work ran on '$name', which the survey did not list as fit"
elif [ "$peak" -gt "$allowed" ]; then
overcommit="yes"
fail "capacity: '$name' peaked at $peak concurrent VM(s) with only $allowed slot(s)"
else
info "capacity: $name peaked at $peak of $allowed slot(s)"
fi
done <<<"$peaks"
[ -n "$overcommit" ] || pass "capacity: no node exceeded the slots the scheduler gave it"
# ── Invariant 2: the excess queued rather than vanishing ───────
#
# "Nothing queued" has two very different causes and they must not share a
# verdict. If the fleet never actually filled — warm nodes finish a trivial
# phase in ~80s, so a slot can free before the sweep even reaches the 15th
# mission — then the queue was never reached and this scenario did not test
# it. Only a burst that DID saturate can call an absent queue a failure.
local peak_total
# `peaks` is one `name peak` line per node; the sum is the fleet's high-water
# mark. Not exact (peaks can occur at different instants) but an OVER-estimate,
# which is the safe direction: it can only make us more willing to call an
# absent queue a real failure.
peak_total=$(printf '%s\n' "$peaks" | awk '{s += $2} END {print s+0}')
if [ "$burst" -le "$slots" ]; then
norun "capacity: burst ($burst) did not exceed capacity ($slots) — the queue was never exercised"
elif [ "$queued" != "1" ] && [ "${peak_total:-0}" -lt "$slots" ]; then
norun "capacity: the fleet peaked at ${peak_total:-0} of $slots slot(s) — the burst finished before it could saturate, so the queue was never reached"
elif [ "$queued" = "1" ]; then
pass "capacity: the over-capacity missions QUEUED (capacity_blocked_since was set)"
else
fail "capacity: $burst missions on $slots slots and nothing ever queued — placement is not counting commitments"
fi
# ── Every mission still delivered ──────────────────────────────
local completed
completed=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select count(*) from missions where id in ($idlist) and status = 'completed';\"" \
| head -1 | tr -d '[:space:]')
if [ "${completed:-0}" = "$burst" ]; then
pass "capacity: all $burst queued/placed missions completed"
else
fail "capacity: only ${completed:-0} of $burst missions completed — a queued mission must still run"
fi
# ── Spread, not stack ─────────────────────────────────────────
local used
used=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select count(distinct target_node_id) from missions where id in ($idlist) and target_node_id is not null;\"" \
| head -1 | tr -d '[:space:]')
if [ "$fitnodes" -lt 2 ]; then
info "capacity: only one fit node — spread is not assertable"
elif [ "${used:-0}" -ge 2 ]; then
pass "capacity: the burst spread across ${used} node(s)"
else
fail "capacity: $burst missions all landed on ${used:-0} node(s) — ranking is stacking, not spreading"
fi
}
# ── Scenario: draining a node mid-mission re-places the next phase ─
#
# The affinity decision made this possible and this test is what proves it was
# taken for real: mission state lives on the GATEWAY, every phase is
# inject-tar → run → collect-tar → destroy, so phase 2 owes nothing to the node
# phase 1 ran on. If that were false, draining here would either strand the
# mission or silently lose phase 1's work — and "silently lose" is the outcome
# a status-only check reports as success.
DRAIN_BODY=$(cat <<JSON
{"title":"verify: a drained node hands the mission on",
"template_kind":"research_and_code",
"repo_id":"$REPO_ID",
"runtime_kind":"microvm",
"backend":"claude",
"description":"Build up DRAIN.md across two phases while the fleet changes under it.",
"phases":[
{"kind":"coding","order_idx":0,"config":{"commit_policy":"always","max_iterations":1,
"task":"Create a file named DRAIN.md at the repository root containing exactly one line: PHASE-ONE-OK. Do not create or modify any other file."}},
{"kind":"coding","order_idx":1,"config":{"commit_policy":"always","max_iterations":1,
"task":"Read the existing file DRAIN.md at the repository root. It was written by the previous phase and must already contain the line PHASE-ONE-OK. Append a second line reading PHASE-TWO-OK, keeping the first line intact. If DRAIN.md does not exist, instead create a file named DRAIN_MISSING.md containing the single line PRIOR-PHASE-WORK-WAS-LOST, and do not create DRAIN.md."}}
]}
JSON
)
undrain() { # undrain <node-uuid>
[ -n "${1:-}" ] || return 0
ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"update nodes set status='online' where id='$1' and status='draining';\"" >/dev/null 2>&1
}
scenario_drain_midmission() {
local token mission first_node second_node waited=0 status delivered
token=$(mint_session) || { norun "drain-midmission: could not mint a session"; return 1; }
local fit
fit=$(fleet_capacity "$token" | python3 -c 'import json,sys
try: print(len(json.load(sys.stdin)["nodes"]))
except Exception: print(0)')
if [ "${fit:-0}" -lt 2 ]; then
norun "drain-midmission: needs 2+ fit nodes, fleet has ${fit:-0}"
return 1
fi
mission=$(create_mission "$token" "$(echo "$DRAIN_BODY" | tr -d '\n')") \
|| { norun "drain-midmission: mission create failed"; return 1; }
info "drain-midmission: mission=$mission"
api "$token" PATCH "/api/missions/$mission/status" '{"status":"running"}' >/dev/null
# Drain the node while phase 0 is still RUNNING on it.
#
# The first version waited for phase 0 to COMPLETE and lost the race: warm
# phases finish in ~80s and `start_pending_phases` sweeps every 10s, so
# phase 1 was routinely already placed by the time the drain landed — and the
# run then reported "phase 1 ran on the DRAINED node" for a drain that had not
# yet happened. Draining does not touch a VM already running, only future
# placement, so doing it mid-phase is both safe and the more faithful test.
#
# Polling is on the PHASE, not the mission: by the time the mission is
# terminal there is nothing left to re-place.
while [ "$waited" -lt "$MISSION_TIMEOUT" ]; do
first_node=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select m.target_node_id from mission_phases p join missions m on m.id = p.mission_id
where p.mission_id = '$mission' and p.order_idx = 0
and p.status in ('running','completed') and m.target_node_id is not null;\"" \
| head -1 | tr -d '[:space:]')
[ -n "$first_node" ] && break
# A mission that died before phase 0 completed has nothing to test.
status=$(api "$token" GET "/api/missions/$mission" | python3 -c 'import json,sys
try: print(json.load(sys.stdin).get("status",""))
except Exception: pass' 2>/dev/null)
case "$status" in failed|cancelled)
norun "drain-midmission: mission ended $status before phase 0 started"; return 1 ;;
esac
sleep 3
waited=$((waited + 3))
done
if [ -z "$first_node" ]; then
norun "drain-midmission: phase 0 never started in ${MISSION_TIMEOUT}s"
return 1
fi
info "drain-midmission: phase 0 is on $first_node — draining it now"
ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"update nodes set status='draining' where id='$first_node';\"" >/dev/null
status=$(await_mission "$token" "$mission")
# Restore the node BEFORE asserting, so a failed assertion cannot leave the
# fleet one node smaller for every later run.
undrain "$first_node"
[ "$status" != "timeout" ] || { norun "drain-midmission: mission did not finish in ${MISSION_TIMEOUT}s"; return 1; }
printf '%s\n' "$(phase_report "$token" "$mission")" | sed 's/^/ phase /'
second_node=$(ssh "$HOST" "docker exec clawmates_postgres_1 psql -U postgres -d clawmates -tAc \
\"select target_node_id from missions where id = '$mission';\"" \
| head -1 | tr -d '[:space:]')
if [ -z "$second_node" ]; then
fail "drain-midmission: the mission has no target node after phase 1"
elif [ "$second_node" = "$first_node" ]; then
fail "drain-midmission: phase 1 ran on the DRAINED node $first_node"
else
pass "drain-midmission: phase 1 re-placed onto $second_node"
fi
delivered=$(fetch_delivered "$token" "$mission" DRAIN.md) \
|| { fail "drain-midmission: could not read DRAIN.md from the pushed branch"; return 1; }
case "$delivered" in
*PHASE-ONE-OK*PHASE-TWO-OK*)
pass "drain-midmission: phase 1 read phase 0's work and appended to it" ;;
*PRIOR-PHASE-WORK-WAS-LOST*)
fail "drain-midmission: re-placement lost the previous phase's work" ;;
*)
fail "drain-midmission: DRAIN.md is neither outcome: $(printf '%s' "$delivered" | tr '\n' '|')" ;;
esac
check_single_uid "$mission" drain-midmission
}
# ── Entry point ──────────────────────────────────────────────────
ssh -o BatchMode=yes -o ConnectTimeout=10 "$HOST" true 2>/dev/null \
|| die "cannot ssh to $HOST"
case "${1:-all}" in
selftest)
selftest_uid_probe
;;
uids)
[ $# -ge 2 ] || die "usage: $0 uids <mission-id>"
selftest_uid_probe
check_single_uid "$2" "uids($2)"
;;
chain)
selftest_uid_probe
run_scenario chain "$CHAIN_BODY" assert_chain
;;
multirole)
selftest_uid_probe
body=${MULTIROLE_BODY//__TEAM__/$TEAM_TEMPLATE}
run_scenario multirole "${body//__REPO__/$REPO_ID}" assert_multirole
;;
noop)
run_scenario noop "$NOOP_BODY" assert_noop
;;
canary)
# The microvm scenario, run against the CANDIDATE CLI image instead of the
# one every mission uses. Same assertions — kernel, delegation, gate, judge,
# single writer — because the question is whether the new version still
# satisfies what the current one does.
run_scenario canary \
"$(echo "$MICROVM_BODY" | sed 's/"backend":"claude"/"backend":"canary-claude"/' | tr -d '\n')" \
assert_microvm
;;
microvm)
run_scenario microvm "$(echo "$MICROVM_BODY" | tr -d '\n')" assert_claude_vm
scenario_microvm_unavailable_backend
;;
glm)
# The microvm scenario on the z.ai backend. The provider assertion is the
# point: this CLI reaches z.ai through ANTHROPIC_BASE_URL, a path 2.1.265 and
# 2.1.275 both broke, and "the mission completed" alone cannot say who
# answered.
run_scenario glm \
"$(echo "$MICROVM_BODY" | sed 's/"backend":"claude"/"backend":"glm"/' | tr -d '\n')" \
assert_glm
;;
kimi)
run_scenario kimi \
"$(echo "$MICROVM_BODY" | sed 's/"backend":"claude"/"backend":"kimi"/' | tr -d '\n')" \
assert_kimi
;;
gatecap)
run_scenario gatecap "$(echo "$GATECAP_BODY" | tr -d '\n')" assert_gate_cap
;;
research-only)
run_scenario research-only "$(echo "$RESEARCH_ONLY_BODY" | tr -d '\n')" assert_research_only no-checkout
;;
research-vm)
run_scenario research-vm "$(echo "$RESEARCH_VM_BODY" | tr -d '\n')" assert_research_only no-checkout
;;
benchmark)
run_scenario benchmark "$(echo "$BENCHMARK_BODY" | tr -d '\n')" assert_benchmark
;;
security)
run_scenario security "$(echo "$SECURITY_BODY" | tr -d '\n')" assert_security
;;
refactor)
run_scenario refactor "$(echo "$REFACTOR_BODY" | tr -d '\n')" assert_refactor
;;
composed)
run_scenario composed "$(echo "$COMPOSED_BODY" | tr -d '\n')" assert_composed
;;
roster)
scenario_roster
;;
local-ornith)
run_scenario local-ornith "$(echo "$LOCAL_BODY" | tr -d '\n')" assert_local
;;
capacity)
scenario_capacity
;;
drain-midmission)
scenario_drain_midmission
;;
all)
selftest_uid_probe
run_scenario chain "$CHAIN_BODY" assert_chain
body=${MULTIROLE_BODY//__TEAM__/$TEAM_TEMPLATE}
run_scenario multirole "${body//__REPO__/$REPO_ID}" assert_multirole
run_scenario noop "$NOOP_BODY" assert_noop
run_scenario microvm "$(echo "$MICROVM_BODY" | tr -d '\n')" assert_claude_vm
scenario_microvm_unavailable_backend
run_scenario glm "$(echo "$MICROVM_BODY" | sed 's/"backend":"claude"/"backend":"glm"/' | tr -d '\n')" assert_glm
run_scenario kimi "$(echo "$MICROVM_BODY" | sed 's/"backend":"claude"/"backend":"kimi"/' | tr -d '\n')" assert_kimi
run_scenario gatecap "$(echo "$GATECAP_BODY" | tr -d '\n')" assert_gate_cap
run_scenario research-only "$(echo "$RESEARCH_ONLY_BODY" | tr -d '\n')" assert_research_only no-checkout
run_scenario research-vm "$(echo "$RESEARCH_VM_BODY" | tr -d '\n')" assert_research_only no-checkout
run_scenario benchmark "$(echo "$BENCHMARK_BODY" | tr -d '\n')" assert_benchmark
run_scenario security "$(echo "$SECURITY_BODY" | tr -d '\n')" assert_security
run_scenario refactor "$(echo "$REFACTOR_BODY" | tr -d '\n')" assert_refactor
run_scenario composed "$(echo "$COMPOSED_BODY" | tr -d '\n')" assert_composed
run_scenario local-ornith "$(echo "$LOCAL_BODY" | tr -d '\n')" assert_local
scenario_roster
scenario_capacity
scenario_drain_midmission
;;
*)
die "unknown scenario: $1 (selftest|uids|chain|multirole|noop|microvm|canary|glm|kimi|gatecap|research-only|research-vm|benchmark|security|refactor|composed|roster|local-ornith|capacity|drain-midmission|all)"
;;
esac
if [ "$FAILURES" -gt 0 ]; then
printf '\n%d of %d check(s) failed\n' "$FAILURES" "$CHECKS"
exit 1
fi
# Zero checks is not success. A run that asserted nothing must not be able to
# print the same closing line as a run that asserted everything.
if [ "$CHECKS" -eq 0 ]; then
printf '\nno checks ran — nothing was verified\n'
exit 1
fi
printf '\nall %d check(s) passed\n' "$CHECKS"