5c2c63f8e8546b71a3f40e002e3333727c8e2aa5
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5c2c63f8e8 |
feat(missions): a human can finally reach the plan/roster review gate
Phase 4 of the plan, plus the PLAN_COMPLETE decision and the gitea_forge
cleanup from Phase 5.
THE REVIEW UI
mission_plan and mission_roster have been complete and reachable by curl
since they shipped, with zero frontend. That matters more than a missing
screen usually would: the decide step is not a convenience, it IS the
safety mechanism. Approving a plan replaces the mission's phases; approving
a roster flips it to the composed engine. A gate nobody can reach is a gate
that is always open or always shut.
MissionProposalDrawer, modelled on LevelUpDrawer which already does
load → review → decide. Reached from a mission's SETUP tab. Verified end to
end against the live backend, not just compiled: a model proposed a roster,
approval flipped the mission to `composed`, and approval on a non-draft
mission was refused.
The plan view shows each phase's done_when, and says plainly when one is
absent — a phase without a completion condition is never judged and reports
completed whatever it did, so its absence is the thing worth seeing.
AND THE DEFECT BUILDING IT FOUND
Every refusal path computed a precise reason — "the mission is running, not
a draft", "no node can boot that backend any more" — logged it to stderr,
and returned a bare {"error":"bad request"}. The person who needed the
sentence was the one clicking Approve; they got two words, and the reason
went to a server log they cannot read.
ApiError::Refused(String) carries it now. Same argument ApiError::Unavailable
was added for ("a 500 with 'internal error' sent them looking for a bug that
was not there"), one status code down. Live: the 400 now reads "this mission
is completed — a roster can only be approved while it is a draft, because
approving one rewrites how the mission will run".
PLAN_COMPLETE, decided
The Skill-Use measurement found that int-xx-marker-protocol documents
PLAN_COMPLETE and task_card_parser never implemented it, so an agent
following the skill exactly was silently ignored. Implemented rather than
removed from the skill: the planner needs a way to say it is done
specifying, and agents already emit it.
Marker ids are now strictly INT-<digits>. `starts_with("INT-")` accepted the
range form `INT-01..02` — observed live — which parsed into an id matching
no real item, so a task card appeared for something that did not exist while
the two items it covered stayed open. Rejecting is right: an ignored marker
is visible, a plausible row is not.
GITEA_FORGE, REMOVED
Named in nine places, defined in none. Harmless while provision_claw ignored
the bundle list; once the list was honoured, an undefined name became a
capability an agent is told it has and does not. Removed from seven team
templates, a workflow recipe, the auto-provision path, and a dropdown a user
could pick it from.
A new test asserts every bundle a template names is defined in the runtime
config — and it immediately found `web_fetch` in two templates I had missed
removing by hand. Same shape as the skill-binding test, one layer up.
Agents reach the forge through git over HTTPS with the ambient GITEA_TOKEN,
which is why nothing ever broke.
Full workspace suite green (106 binaries); frontend builds clean.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
18dc0b964b |
fix(missions): the security scan phase now scans, and task upserts work
Four defects, found by checking the audit's claims instead of trusting them. Two of the audit's own findings turned out to be wrong, and the registry that exists to record which config keys are read was itself inaccurate — so the corrections are part of the change. upsert_task raised 42P10 on every call, for every caller `mission_tasks_external_uniq` is a PARTIAL unique index (WHERE external_id IS NOT NULL). Postgres will not match a partial index to an ON CONFLICT target unless the statement repeats the predicate, so the upsert failed on its first row. Both callers — the task-card parser that turns INT markers into tasks, and the security scanner — map the error to a string their caller logs. Two features were broken and nothing was red. Regression test in cm-db with a negative control: reverting the WHERE reproduces 42P10 exactly. the security scan never ran `security_scan::run` was reachable only from an operator button, so security_hardening.toml — a workflow whose entire first phase is a scan — ran an agent that was never told to scan and never fired the scanner either. phase_runner now sweeps finished security_scan phases, mirroring the benchmark baseline sweep that was added for the identical defect. Guarded on a new completion marker rather than on findings: a clean scan writes no findings, so a findings-guard would rescan forever. The marker also answers the question an operator actually asks, which is not "how many findings" but "was this looked at, by what, and when". two recipes could not fail security_hardening.toml and benchmark.toml carried no `task` and no `done_when` on any phase. A phase without done_when never enters evaluating, is never judged, and reports completed whatever it did — so a security mission could scan nothing and go green, and a benchmark mission could record no baseline that the next refactor would then compare against. Both now state the work and the condition, with inert keys annotated inline rather than deleted, so the gap between what a recipe asks for and what a phase receives stays visible. the config registry was wrong in both directions `harness` was listed NOT IMPLEMENTED while benchmark_runner reads it and phase_runner runs a baseline through it. `tools` was listed NOT IMPLEMENTED while security_scan::run reads it. A registry that exists so an operator can trust what a recipe does is worse than useless when it is inaccurate. Both corrected, `bench_name` and `cmd` added, and `test_command` deleted — it had neither a reader nor a writer, so it described a situation that could not arise. Also: CLAWMATES_JUDGE_MODEL had two different defaults (opus-4-8 in routes/topology.rs vs opus-5 in cm_runtime::judge_model) and a doc comment naming a third; topology now calls the one function. GITEA_TOKEN's absence in mission_plan is stated rather than degrading to the same "could not be read" string a private repo produces. BRAINHUB_API_KEY needed no change — hub::push already rejects an unset key with a named error. That half of the finding was overstated. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
9c9439a271 |
feat(llm): the subscription is the default provider, with a recorded fallback chain
Two changes so an empty metered account stops being a platform outage. 1. `build_provider` prefers the subscription token over ANTHROPIC_API_KEY. A bare model name resolves to whatever this returns, so making it the subscription means no server-side call can reach the metered key by construction — rather than by a source-grep test that already missed four call sites once. The metered key remains a fallback and now warns loudly when it is the one in use; boot no longer requires it at all. 2. `complete_with_fallback` walks a declared chain when a model has no capacity: opus -> haiku -> glm:glm-4.7 by default, overridable via CLAWMATES_MODEL_FALLBACK, empty to disable. Measured on gw-04 today: opus and sonnet return 429 on the subscription while haiku, GLM and Kimi all return 200, so a capped window no longer means "the planner is gone". The chain returns the model that ANSWERED, and every caller persists it — mission_plan_proposals.author_model, mission_team_proposals.author_model, and the swarm's step role. A plan drafted by the third link and filed as an opus plan is a silent quality change, which is the failure shape this project keeps paying for. Two negative controls hold the design: the chain never retries the model that just failed as its own fallback, and it steps down ONLY for a capacity failure — walking it on a malformed prompt would ask three models the same bad question and report the third one's confusion. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
ee5a939ce6 |
fix(planner): the other four server-side calls were still on the metered key
The test that was supposed to prevent this grepped for the literal `runtime.complete(` and passed while the phase planner (`mission_plan.rs`), both swarm calls, and a second enhance path in `claws.rs` still billed the pay-as-you-go account. They spell the receiver `state.runtime` or wrap the call across lines, so the receiver name was never the thing to match. The test now matches the METHOD, and covers all five files. `complete_or` gains the rule that makes it safe to apply everywhere: a `name:model` spec is an operator's explicit provider choice — the swarm worker model is configured exactly that way — and is passed straight to `Runtime::resolve_provider` untouched. Only a bare name is ambiguous, and a bare name is precisely what resolves to the default provider. Hijacking a chosen Kimi or GLM model onto Anthropic would be the same silent-substitution bug pointed the other way. `validator_preflight` and the evaluator judge keep calling the runtime directly, on purpose: both exist to exercise the CONFIGURED spec. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
08dd227a45 |
feat(missions): give the planner the repository's contents, not just its names
The root listing was not enough. Given names alone the planner wrote "optimise the hot path" for a crate whose hot path is `add(a: i64, b: i64) -> i64` — a mission that was unachievable from the moment it was written, and that nothing discovered until an agent had built a benchmark harness in a VM to measure an integer addition, honestly reported no improvement was possible, and the judge correctly failed the phase. `repo_digest` fetches the whole tree (so "does this have benches/" is a fact, not an inference) and then file CONTENTS in priority order: manifests first — they say what the project is — then the README, then source ascending by size, since a planner learns more from twenty small files than from one large one. Lockfiles and build output are dropped: enormous, and they say nothing a manifest does not. THE RULE THIS ENFORCES, and the reason the rendering is its own tested module: a digest of any repository worth planning against is partial, and a model shown a partial view without being told it is partial plans as though it saw everything. So every omission is stated — how many files exist, how many were shown, what was cut from each, and "anything not shown you have NOT seen". Same distinction as `Option<u32>` for the subagent probe: "we did not look" and "there is nothing there" are different facts. Failures degrade to a stated absence rather than an empty string, and the three cases stay distinguishable: no repository, a tree that could not be read, and a tree read but no contents fetched. An unreadable tree is never rendered as an empty repository. Two more things the prompt now says, both learned from that run: plan for the repository as it IS rather than as the description implies (and if the description asks for something the code cannot support, say so in the task and plan the phase that establishes the truth, rather than a phase that must fail); and a mission agent has NO package-registry access. The agent discovered the second one mid-run and wrote a dependency-free `std::time::Instant` harness after Criterion could not be added — good adaptation, but nothing had warned it. 549 tests pass, clippy clean. The budget/priority/truncation logic is pure and tested; only the fetching touches the network. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
0aeae07db2 |
fix(missions): the planner was planning blind — show it the repository
The first real plan opened with "Identify the crate's hottest code path and run its benchmark harness". This crate has no benchmark harness. The phase ran, found nothing to baseline, delivered zero files, and the plan's second phase was left with nothing to optimise against. The planner saw the mission title, the description, and a boolean for whether a repository was bound. It never saw the repository. A plan about a codebase written without looking at the codebase is a guess that reads like a plan — and the failure surfaces two phases and one VM boot later, as an agent reporting that the thing it was told to run does not exist. The prompt now carries the repository's root listing, read from the FORGE rather than a checkout: at proposal time the mission is still a draft and `ensure_checkout` has not run, so there is nothing on disk to list. It also says outright that a phase needing something absent must CREATE it and say so in its task — the failure was not only ignorance of the tree but the assumption that missing tooling is someone else's problem. A listing that cannot be fetched degrades to "(the repository listing could not be read)" in the prompt rather than to an empty string. A model told the listing is unavailable can hedge; a model told nothing assumes — which is the same distinction as `Option<u32>` for the subagent probe, in a prompt instead of a struct. Found by running the thing end to end rather than by testing it: every unit test here passes with a planner that has never seen a repository. 543 tests pass, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
a33dbdcdc3 |
feat(missions): W1/#13 — let a model author the mission's phases
The last unstarted item from the missions-as-workflows plan, and the other half
of Slice 5: that one lets a model size the TEAM, this lets it decide what the
work IS.
Every mission's phases come from one of five hand-written recipes in
`templates/workflows/*.toml`, chosen by `template_kind` before anyone saw the
mission. That is the "do it this way: 1, 2, 3" over-specification that makes a
capable model follow a worse plan than it would have chosen. The recipes stay —
they are still the default for a mission nobody proposes a plan for, and the
fallback when a proposal is refused.
Same three verbs and the same review gate as the roster, deliberately: propose
and decide are separate because only the second changes a mission, and a second
shape would be a second thing to get right. Approving REPLACES the phases (a
plan is an answer to "what is this mission", not an addition to one), draft-only.
GROUNDED IN WHAT THE PLATFORM ACTUALLY READS, which is the part that makes this
more than a copy. `phase_config::KNOWN_KEYS` already names every phase-config key
and the code that reads it — the registry built after `task` sat unread through
every mission. A plan is validated against it, so a model cannot propose a phase
whose settings nothing will act on: the failure that registry exists to EXPOSE is
one this path cannot create. Phase kinds are checked the same way, because an
unknown kind does not error — it falls through to the catch-all purpose and runs
as a generic phase that looks like it worked.
TWO THINGS THE WORK ITSELF FOUND, both the same shape:
- `done_when_check` — the stop-gate key added earlier today — was never
registered in `phase_config`, so every mission that set it has been logging
it as an unknown key. Found by a test written for a different purpose, which
is the registry doing exactly its job. Now registered with its reader.
- `done_when` and `max_iterations` are COLUMNS promoted out of config by
`missions::create`; the evaluator sweep filters on the column in SQL every
tick. My first insert wrote the config blob alone, which would have stored a
plan's completion condition where nothing judges it. NEGATIVE CONTROL run:
binding NULL instead of the promoted value fails
`an_approved_plan_replaces_the_missions_phases`.
`order_idx` comes from the array's own order rather than a field the model sets:
two sources for one fact is how a plan ends up with two phase 0s, and order_idx
is what `start_pending_phases` sequences on.
MAX_PHASES is 4 and the prompt argues for one. Each phase is a full agent run in
sequence, and splitting one change into plan → implement → test is the documented
anti-pattern — a single agent doing all three keeps the context that makes the
later steps good.
543 tests pass, clippy clean. Migration 0072.
Co-Authored-By: Claude Opus 5 <[email protected]>
|