fix(missions): a server restart no longer kills a running mission
Mission 01a00538 ("ClawHDF5 REsearch and Refactor") failed 19 minutes and 93,762
tokens into its research phase with `pair failed: 403 Forbidden`, and its coding
phase was then correctly skipped as unreachable. The cause was not the coding
phase and not the model — it was pairing.
A per-mission runtime is authenticated with a SINGLE-USE pairing code, and the
bearer token it returns was cached in memory only. Any restart of the server
discarded that token; the next turn re-paired with a code the gateway had
already spent and got 403 — permanently, for that mission. A deploy, a crash or
an OOM would each do it. The durable-run machinery exists precisely so work
survives a restart; pairing was the one thread that did not, and it failed
closed.
`missions.runtime_token` persists the token at the moment pairing succeeds, and
the worker seeds the executor's cache from it, so a new process reuses the
credential instead of re-pairing. Persisting is best-effort: failing to save
must not fail a turn that just paired successfully.
Verified by reproducing the original failure: launched a mission, confirmed the
token was written, restarted the server MID-PHASE, and watched the mission run
to completion with no pairing failure.
Co-Authored-By: Claude Opus 5 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
43436d7181
commit
6e8785f159
@@ -185,9 +185,16 @@ async fn run_job(
|
||||
// the missions row; else fall back to the shared env-derived
|
||||
// gateway (pre-C3 missions + non-mission runs). This is what
|
||||
// isolates agents' workspace filesystem to that mission's repo.
|
||||
type MissionBinding = (Option<String>, Option<String>, Uuid, Option<Uuid>);
|
||||
type MissionBinding = (
|
||||
Option<String>,
|
||||
Option<String>,
|
||||
Uuid,
|
||||
Option<Uuid>,
|
||||
Option<String>,
|
||||
);
|
||||
let mission_binding: Option<MissionBinding> = sqlx::query_as::<_, MissionBinding>(
|
||||
"SELECT m.runtime_endpoint, m.runtime_pairing_code, m.id, r.mission_phase_id
|
||||
"SELECT m.runtime_endpoint, m.runtime_pairing_code, m.id, r.mission_phase_id,
|
||||
m.runtime_token
|
||||
FROM topology_runs r
|
||||
JOIN missions m ON m.id = r.mission_id
|
||||
WHERE r.id = $1",
|
||||
@@ -202,7 +209,7 @@ async fn run_job(
|
||||
let tap = mission_binding
|
||||
.as_ref()
|
||||
.map(
|
||||
|(_, _, mission_id, phase_id)| crate::topology_exec::MissionTap {
|
||||
|(_, _, mission_id, phase_id, _)| crate::topology_exec::MissionTap {
|
||||
pool: pool.clone(),
|
||||
workspace_id: job.workspace_id,
|
||||
mission_id: *mission_id,
|
||||
@@ -211,10 +218,17 @@ async fn run_job(
|
||||
},
|
||||
);
|
||||
let leaf_result = match mission_binding {
|
||||
Some((Some(url), Some(code), _, _)) => {
|
||||
// Seed the cached bearer from `runtime_token` when we have one: the
|
||||
// pairing code is single-use, so after a restart it is the only way in.
|
||||
Some((Some(url), Some(code), _, _, tok)) => {
|
||||
ZeroClawDriveExecutor::from_env_for_gateway_with_code(url, code)
|
||||
.map(|e| e.with_token(tok))
|
||||
}
|
||||
// No pairing code (pre-C3 missions): the persisted token is the only
|
||||
// credential, so seed it here too.
|
||||
Some((Some(url), None, _, _, tok)) => {
|
||||
ZeroClawDriveExecutor::from_env_for_gateway(url).map(|e| e.with_token(tok))
|
||||
}
|
||||
Some((Some(url), None, _, _)) => ZeroClawDriveExecutor::from_env_for_gateway(url),
|
||||
_ => ZeroClawDriveExecutor::from_env(),
|
||||
};
|
||||
let leaf = match leaf_result {
|
||||
|
||||
Reference in New Issue
Block a user