Two stacked issues after risk_profile was fixed:
1. Claws had file_edit + 46 other tools available, but the templates
trained the agents to expect file_read/file_write (older ZeroClaw
tool names). Result: agent output kept saying "I only have file_read"
and dumped implementations into the context window as text.
2. Even with file_edit, the sandbox pointed at
/zeroclaw-data/.zeroclaw/agents/<alias>/workspace/ — NOT
/mission/repo where the checked-out mission repo actually lives.
unrestricted_filesystem=false blocked agents from reaching it.
Fixes:
- provision_claw now takes workspace_path. mission_orchestrator passes
/mission/repo — pins the per-claw workspace via
agents.<alias>.workspace.path to the bind-mount path so file_edit /
content_search / glob_search operate on the mission's git checkout.
- phase_task_text prepends an explicit tool inventory (file_edit,
content_search, glob_search, git_operations, git_forge, ...) plus a
WORKSPACE line pinned at /mission/repo. Each phase directive is
rewritten to reference file_edit / git_operations explicitly and to
call out "do NOT paste code in your reply expecting the platform to
save it."
The topology graph shipped from team.graph only carries node.role,
not node.agent. The executor then defaults to alias_for(role) which
falls to ZEROCLAW_DEFAULT_AGENT (scout) — no such agent → 400.
Look up team_members(node_id → claw_id) at enqueue time and stamp
node.agent = claw_<hex> onto every node. Executor now dials the
specific claw provisioned for THIS teams role.
Was masked pre-C3 because the shared runtime hit the same 400 —
never noticed because no one clicked through to a real run there.
Pairing codes are single-use / expiring — a mission that reuses an
existing runtime container on a retry needs a fresh code, not the
stale one from the initial launch. Drop the runtime_endpoint gate
so ensure_container always fires, and its fast path re-mints via
/admin/paircode/new for existing containers.
The seed-mount approach didnt work: even with the shared runtimes
data dir bind-mounted, a fresh gateway instance mints a new pairing
key and requires re-pairing. The topology_worker connect returned
401 forever.
New approach — per-mission gateways self-pair:
- Provisioner tails container logs after start, extracts the
X-Pairing-Code from the boot banner
- Persists it on missions.runtime_pairing_code (migration 0059)
- topology_worker constructs ZeroClawDriveExecutor with THAT code
via from_env_for_gateway_with_code, which triggers the lazy
/pair handshake on first turn and caches the returned bearer
Drops the shared-runtime data-dir mount — each per-mission gateway
now owns its own state, restoring the C3 isolation guarantee.
- mission_runtime::spawn_sweeper: force-removes runtime containers
for missions terminal for >=30 min, clears runtime_endpoint. Wired
into clawmates-server main().
- docker-compose socket-proxy: NETWORKS=1 so bollard.connect_network
can attach containers to clawmates_edge for provider egress.
- phase_runner ordering: ensure_checkout BEFORE ensure_container so
the mission dir exists before docker mounts it.
- provisioner: mkdir_p the mission dir defensively for research-only
missions that skip checkout entirely.
When a topology_run is bound to a mission whose runtime_endpoint is
set, the worker constructs ZeroClawDriveExecutor against that URL
instead of the env-derived shared gateway. Falls back to shared for
non-mission runs and pre-C3 missions.
With slices 1-3 combined, a mission launched after this deploy will:
1. get its per-mission container spawned during on_launch
2. have its checkout dropped into /var/lib/clawmates-missions/<id>
which is bind-mounted to /mission inside that container
3. run its agents against ZEROCLAW_WORKSPACE=/mission/repo — so
they can see and edit only this missions repo, no bleed-over.
- mission_orchestrator::on_launch now calls ensure_container after
the repo checkout, persists the container_name + endpoint on the
missions row. Non-fatal — logs and continues on docker errors so
dev-mode + tests keep working.
- phase_runner::launch_phase does the same as a fallback for any
mission whose runtime_endpoint is null (pre-C3 or torn down).
Nothing reads the endpoint yet; slice 3 swaps topology_worker over.
Moves ensure_checkout into launch_phase so retries + new phase
launches all trigger the clone/fetch. mission_orchestrator still
does its own checkout at initial launch time, so first-launch
timing is unchanged; this covers the retry + additional-phase
paths.
Every re-attempted phase now starts with a clean slate:
- phase_runner::launch_phase DELETEs prior status IN ('failed',
'cancelled') topology_runs for the phase before enqueuing the
new ones. Completed runs are kept for audit; only the failure
noise from earlier attempts goes.
- POST /api/missions/{id}/phases/{phase_id}/retry — resets a
failed/cancelled phase to 'pending' (auth-scoped to the calling
workspace + guarded on mission.status='running'). phase_runner
picks it up on the next 10s tick.
- MissionCanvas phase card grows a coral 'Retry' button, visible
only when phase.status='failed' and mission.status='running'.
Click → resets + refreshes; the prior failed run rows disappear
from the card as soon as phase_runner enqueues the new attempt.
Design: auto-purge in phase_runner rather than a separate 'clear
failed runs' endpoint. Users don't have to manually clean up before
retrying; the runner does it as part of the natural work of firing
a fresh attempt.
Verified: cargo check + tsc + eslint --quiet all green.
Root-cause fix for "we hit launch, waited overnight, nothing ran."
mission_orchestrator materialized teams + agents fine, but nothing
enqueued the actual work — mission_phases stayed 'pending' forever
and topology_runs count for the mission was 0.
New crates/cm-api/src/phase_runner.rs — background worker on 10s
poll that does three things:
1. start_pending_phases — for every mission_phase with
status='pending' AND parent mission.status='running' AND all
lower-order phases already 'completed', enqueue one
topology_runs row per team whose (mission_id, purpose) matches
the phase kind:
phase=research → teams with purpose='research'
phase=coding → teams with purpose='coding'
phase=benchmark → teams with purpose='coding' (fallback)
phase=security_scan → teams with purpose 'security' | 'coding'
Each run gets a phase-kind-specific task text combining the
mission title/description + a directive for that phase.
Flips phase to 'running' after enqueue.
2. close_finished_phases — SQL sweep that flips phases whose
topology_runs are all terminal to 'completed' (or 'failed' if
any run failed).
3. close_finished_missions — same shape for missions whose phases
are all terminal.
Spawned alongside task_card_worker in clawmates-server main.rs.
Ordering enforced by mission_phases.order_idx — a coding phase
doesn't fire until its research phase completes.
Idempotent: every state transition is guarded so double-firing on a
race is safe. When a mission has no matching teams for a phase (bad
wizard state), the phase stays pending and the runner logs a skip
rather than getting stuck in a fail loop.
Existing topology_worker picks up the queued runs and drives them
through the ZeroClaw executor as usual.