fix(missions): the security scan phase now scans, and task upserts work
Four defects, found by checking the audit's claims instead of trusting them. Two of the audit's own findings turned out to be wrong, and the registry that exists to record which config keys are read was itself inaccurate — so the corrections are part of the change. upsert_task raised 42P10 on every call, for every caller `mission_tasks_external_uniq` is a PARTIAL unique index (WHERE external_id IS NOT NULL). Postgres will not match a partial index to an ON CONFLICT target unless the statement repeats the predicate, so the upsert failed on its first row. Both callers — the task-card parser that turns INT markers into tasks, and the security scanner — map the error to a string their caller logs. Two features were broken and nothing was red. Regression test in cm-db with a negative control: reverting the WHERE reproduces 42P10 exactly. the security scan never ran `security_scan::run` was reachable only from an operator button, so security_hardening.toml — a workflow whose entire first phase is a scan — ran an agent that was never told to scan and never fired the scanner either. phase_runner now sweeps finished security_scan phases, mirroring the benchmark baseline sweep that was added for the identical defect. Guarded on a new completion marker rather than on findings: a clean scan writes no findings, so a findings-guard would rescan forever. The marker also answers the question an operator actually asks, which is not "how many findings" but "was this looked at, by what, and when". two recipes could not fail security_hardening.toml and benchmark.toml carried no `task` and no `done_when` on any phase. A phase without done_when never enters evaluating, is never judged, and reports completed whatever it did — so a security mission could scan nothing and go green, and a benchmark mission could record no baseline that the next refactor would then compare against. Both now state the work and the condition, with inert keys annotated inline rather than deleted, so the gap between what a recipe asks for and what a phase receives stays visible. the config registry was wrong in both directions `harness` was listed NOT IMPLEMENTED while benchmark_runner reads it and phase_runner runs a baseline through it. `tools` was listed NOT IMPLEMENTED while security_scan::run reads it. A registry that exists so an operator can trust what a recipe does is worse than useless when it is inaccurate. Both corrected, `bench_name` and `cmd` added, and `test_command` deleted — it had neither a reader nor a writer, so it described a situation that could not arise. Also: CLAWMATES_JUDGE_MODEL had two different defaults (opus-4-8 in routes/topology.rs vs opus-5 in cm_runtime::judge_model) and a doc comment naming a third; topology now calls the one function. GITEA_TOKEN's absence in mission_plan is stated rather than degrading to the same "could not be read" string a private repo produces. BRAINHUB_API_KEY needed no change — hub::push already rejects an unset key with a named error. That half of the finding was overstated. Co-Authored-By: Claude Opus 5 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
4358964c05
commit
18dc0b964b
@@ -109,14 +109,17 @@ pub async fn compare_topologies(
|
||||
Json(req): Json<CompareRequest>,
|
||||
) -> Result<Json<Comparison>, ApiError> {
|
||||
// Execution turns run on the exec model (default = configured model, e.g.
|
||||
// sonnet); the judge uses the judge model (default claude-opus-4-8). Either
|
||||
// sonnet); the judge uses the judge model (cm_runtime::judge_model). Either
|
||||
// can name a registry provider as "<name>:<model>" (e.g. "glm:glm-4.6",
|
||||
// "kimi:kimi-k2") to run on GLM/Kimi instead.
|
||||
let exec_spec = std::env::var("CLAWMATES_TOPOLOGY_EXEC_MODEL")
|
||||
.unwrap_or_else(|_| state.runtime.model().to_string());
|
||||
let (exec_provider, exec_model) = state.runtime.resolve_provider(&exec_spec);
|
||||
let judge_spec =
|
||||
std::env::var("CLAWMATES_JUDGE_MODEL").unwrap_or_else(|_| "claude-opus-4-8".to_string());
|
||||
// `cm_runtime::judge_model()`, not a second read of the same variable: this
|
||||
// line and that function disagreed on the default (opus-4-8 vs opus-5), so
|
||||
// an unconfigured deployment scored topology comparisons on a different
|
||||
// model than the door governor and nothing recorded which.
|
||||
let judge_spec = cm_runtime::judge_model();
|
||||
let (judge_provider, judge_model) = state.runtime.resolve_provider(&judge_spec);
|
||||
let executor = ProviderExecutor::new(exec_provider, exec_model, state.runtime.max_tokens());
|
||||
let scorer = JudgeScorer::new(judge_provider, judge_model, 16);
|
||||
|
||||
Reference in New Issue
Block a user