Mission 01a0cbc7 on the new clawmates-frontend-scratch repo: accessible Tabs, 10/10 tests and tsc clean when re-run locally from the delivered branch, but failed both iterations because the judge's copy excludes node_modules and clawmates_core has no route to the registry. Blocks every npm team. Also fixes the stale continuous_research row and adds self_audit to the recipe table. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
294 lines
17 KiB
Markdown
294 lines
17 KiB
Markdown
# Template maturity — what our agents can actually be asked to do
|
|
|
|
A review of the 6 workflow recipes and 12 team templates, on 2026-09-22.
|
|
Graded by evidence, not by what the TOML declares: a recipe is mature when a
|
|
mission built from it has completed and something checked the result.
|
|
|
|
## Workflow recipes
|
|
|
|
| recipe | phases | standing check | staffing | grade |
|
|
|---|---|---|---|---|
|
|
| `research_and_code` | 2 (research → coding) | **13 harness scenarios** | research→`topic_research`, coding→`rust_sdlc` | **Proven** |
|
|
| `research_only` | 1 (research, repo-less) | 2 scenarios | `topic_research` | **Proven** |
|
|
| `security_hardening` | 3 (scan → research → coding) | 1 scenario | +`rust_sdlc` | Exercised |
|
|
| `benchmark` | 1 | 1 scenario | `rust_sdlc` | Exercised |
|
|
| `refactor` | 1 | 1 scenario | `rust_sdlc` | Exercised |
|
|
| `continuous_research` | 2 (analysis → script) | 1 scenario (`continuous-research`: digest on main, episode rendered, triage fields, saturation spread) | `continuous_research` | **Proven** — end to end incl. audio |
|
|
| `self_audit` | 1 (research) | none (3 live runs, 2026-09-22) | `continuous_improvement` | Exercised — diagnosed a planted cause |
|
|
|
|
`research_and_code` is the workhorse: every delivery, gate, judge, memory and
|
|
triage scenario is built on it, so it is re-proven on every harness run.
|
|
`research_only` earned its grade differently — its staffing was **measured and
|
|
corrected**: under the old `rust_sdlc` default it delivered 5 roles, 14 skill
|
|
deliveries and ~50 KB of prompt with **1 of 9** skills applicable; on
|
|
`topic_research` it is 3 roles, 4 deliveries, 24 KB, **4 of 4**.
|
|
|
|
`continuous_research` is the outlier and the one to fix first. It has the most
|
|
moving parts of any recipe — harvest → triage → manifest → analysis → ranking
|
|
→ script → `episode.json` → vault commit — it ran twice today and produced
|
|
real output both times, and **nothing in the harness would notice if it broke
|
|
tomorrow.** Every other recipe has a standing check; this one has operator
|
|
attention, which is not the same thing.
|
|
|
|
## What the recipes declare that does not fire
|
|
|
|
`phase_config::DECLARED_BUT_UNREAD` is the honest list, and recipes lean on it:
|
|
|
|
| key | used by | status |
|
|
|---|---|---|
|
|
| `loop` | `research_and_code` (`until_no_more_int_items`), `refactor` (`single_pass`) | **inert** — iteration is `max_iterations` + `done_when` |
|
|
| `produces` | almost every recipe (`["md"]`) | **inert** — artifact rendering is not driven by it |
|
|
| `input_from_phase` | `security_hardening` | **inert** — phases share a checkout |
|
|
| `mcp_bundles` (phase level) | `security_hardening` | **inert** — bundles come from the TEAM template |
|
|
|
|
None of this is hidden: `security_hardening.toml` carries a "what is real here
|
|
and what is decoration" section and annotates its own dead keys inline. That is
|
|
the right pattern and the other five should copy it. The risk is not the dead
|
|
keys themselves — it is that a reader takes `loop = "until_no_more_int_items"`
|
|
for a loop.
|
|
|
|
## Team templates
|
|
|
|
All 12 are structurally sound and **every declared skill resolves to a real
|
|
file — 0 unbound across all of them**, which is the `skill-binding-repair` work
|
|
holding (it was 55 of 85 dangling once).
|
|
|
|
| team | slots | risk profile | skills | evidenced? |
|
|
|---|---|---|---|---|
|
|
| `rust_sdlc` | planner coder tester reviewer committer | coding_readwrite | 19 | **yes** — harness default |
|
|
| `topic_research` | lead_researcher evidence_checker report_writer | research_web_readonly | 4 | **yes** — measured |
|
|
| `continuous_research` | paper_reader signal_ranker script_writer | research_readonly | 11 | **yes** — live runs |
|
|
| `backend` | api_designer db_engineer coder tester committer | coding_readwrite | 20 | **yes** — first run 2026-09-22 |
|
|
| `frontend` | designer coder tester committer | coding_readwrite | 12 | **work verified by hand; judge cannot run npm** |
|
|
| `mobile` | designer coder tester committer | coding_readwrite | 11 | no |
|
|
| `gpu` | arch_analyst kernel_author bench_engineer coder committer | coding_readwrite | 13 | no |
|
|
| `threejs` | scene_designer coder shader_author perf_engineer committer | coding_readwrite | 13 | no |
|
|
| `codebase_research` | code_archeologist architecture_mapper flow_tracer vault_scribe | research_readonly | 11 | **yes** — first run 2026-09-22 |
|
|
| `papers_research` | domain_scout paper_reader library_curator | research_web_readonly | 8 | **yes** — first run 2026-09-22 |
|
|
| `insight_research` | implementation_tracker novelty_hunter publication_drafter | research_readonly | 8 | no |
|
|
| `continuous_improvement` | brain_inspector improvement_proposer improvement_evaluator | research_readonly | 7 | **evidenced** — `self_audit` recipe; diagnosed a planted cause and found a real one |
|
|
|
|
**Six of twelve are evidenced.** `backend` was exercised on 2026-09-22,
|
|
the first time anything had been staffed from it: all five roles
|
|
provisioned with real agents (api_designer, db_engineer, coder, tester,
|
|
committer), and the mission delivered cursor pagination — `Paged<T>`,
|
|
`paginate<T: Clone>`, module declared, `cargo test` 3 passed — judged
|
|
**met on the first pass**, 2 files pushed. It also gave the task-permission
|
|
shadow its first confirmed `Edit` call.
|
|
|
|
`codebase_research` was exercised the same day, against the real
|
|
`clawmates` repository rather than the toy scratch crate, with a task that
|
|
cannot be faked from filenames: trace every hop of a mission agent's tool
|
|
call from the guest hook to the recorded event, naming file, function and
|
|
whether each hop runs in the guest or on the host. It produced a 189-line
|
|
`research/GATE-MAP.md` describing code committed **the same day** — the
|
|
four-field `NODE_EXTRACT` including `agent_type`, the shadow-mode
|
|
`would-deny.jsonl` semantics, `install_with` — and **all six of its line
|
|
citations verify exactly** (`hook_script_with` 562, `NODE_EXTRACT` 789,
|
|
`TaskPolicy` 297, `ROLE_POLICIES` 261, `settings_hook` 809,
|
|
`install_command_with` 819). Nothing hallucinated. **Five of twelve.**
|
|
|
|
`papers_research` followed: asked for a bounded library on microVM and
|
|
sandbox isolation for agents, it delivered a README index and four notes
|
|
with complete frontmatter, judged met. The judge could only check the
|
|
frontmatter was *present*; a paper team whose container has no
|
|
pdf-to-text tool is exactly where invented citations live, so every
|
|
arXiv id was checked against arXiv itself — **all four exist, with exact
|
|
title matches**. One of them, `2603.02277` (SandboxEscapeBench), is a paper
|
|
the operator's own research pass had already cited, which is independent
|
|
evidence it found on-topic work rather than plausible filler. **Six of
|
|
twelve.**
|
|
|
|
`continuous_improvement` ran too, and gets a different grade, because
|
|
running and working came apart. All three roles handed off with
|
|
attributed commits, and it was scrupulously honest: it filed **zero**
|
|
level-up proposals, its proposer committed *"no proposals warranted,
|
|
evidence base too thin"*, and its audit opens by stating the problem
|
|
exactly — *"No `.brain` files are accessible within the repo or mission
|
|
filesystem — brains are held in the platform, not checked in."*
|
|
|
|
That is the defect, and it is structural rather than a matter of data.
|
|
The template's subject is "every project agent's `.brain`"; those live in
|
|
the server's `/data/brains` volume, and **nothing delivers them into a
|
|
mission** — the same shape as `research-has-no-delivery-channel` and
|
|
`skills-had-no-delivery-channel`. So it audited the only agent-shaped
|
|
thing in reach, a `ROSTER.md` in the scratch repo, and read that file's
|
|
`6.1.128` (a guest kernel version) as an agent version, having no way to
|
|
know better. It is not counted as evidenced.
|
|
|
|
A channel alone would not rescue it. Mission crews are minted per mission
|
|
(reuse is off by decision), and missions write memory to the *repo* brain,
|
|
so every agent brain is a ~2 KB seed with no history to audit. The
|
|
template assumes long-lived agents that accumulate a record; the platform
|
|
makes disposable ones. That is a design decision to resolve, not a bug to
|
|
patch.
|
|
|
|
**Retargeted (2026-09-22).** The decision taken: audit the *repo* brain,
|
|
which is where missions actually write what they learn. The server now
|
|
exports it as `/mission/memory/PROJECT-MEMORY.md`, one line per judged
|
|
phase, and installs it into every mission that has a repository.
|
|
|
|
Two runs followed, and they separate the channel from the steering:
|
|
|
|
- **v2 (01a0cad8) — channel works, steering does not.** The 2.4 KB record
|
|
was installed. The agent made 50 tool calls, 0 of them touched it, 13
|
|
went to the roster, and the judge passed it. The brief and `done_when`
|
|
were reused from v1 ("read each agent's brain"), and the template's
|
|
retargeted role prompts are **inert on mission turns** (see
|
|
`prompt-injection-two-paths`). Rewriting `system_prompt` cannot aim a
|
|
mission; only the brief can, and nothing supplies one for this team.
|
|
- **v3 (01a0cade) — right answer, failed by the judge.** With a brief and
|
|
`done_when` pointing at the record, 3 of 7 tool calls read it. The audit
|
|
states the record exactly (4 lines, MET 4 / UNMET 0), finds no pattern,
|
|
and files **zero** proposals ("a single UNMET line would be an incident;
|
|
this record has zero"), the correct result on an all-MET record. The
|
|
judge failed it, correctly by the letter: the condition required each
|
|
finding to *quote* the record lines, and the audit summarised them in a
|
|
table instead. `max_iterations` was 1, so there was no retry.
|
|
|
|
So the subject is now reachable and the reasoning holds up. The unproven
|
|
part is the case this team exists for, a record *with* failures. That
|
|
needs a repository whose history contains UNMET verdicts, and a default
|
|
brief so no caller has to write one. Still not counted as evidenced.
|
|
|
|
**The `self_audit` recipe, and a record with real failures (2026-09-22).**
|
|
The brief now lives in `templates/workflows/self_audit.toml`, so a mission
|
|
naming only `template_kind = "self_audit"` and a repo is staffed and aimed
|
|
without the caller writing anything. To give it something to find, two
|
|
missions were planted on the scratch repo whose `done_when` required a
|
|
*Limitations* section their brief never mentioned. The judge marked both
|
|
UNMET for exactly that. The record then held 7 lines: 4 MET, 3 UNMET, the
|
|
third being v3's quoting failure, a natural single-incident control.
|
|
|
|
Mission 01a0cb3f, launched with no brief of its own, read the record (8 of
|
|
25 tool calls), stated 7 / 4 / 3, found the one pattern, quoted both lines
|
|
verbatim, and named the quoting failure as an incident with no proposal. The
|
|
judge passed it on iteration 1 of 2.
|
|
|
|
Its **diagnosis was wrong**, and the record is why. It concluded the agents
|
|
"produced the main content but omitted the final required section" and
|
|
proposed a write-then-check checklist. The recorded prompts show the agents
|
|
were never told: none of the three mentions *Limitations*, because working
|
|
agents are not shown their completion condition. The checklist would not have
|
|
helped. The record carried the condition and not the brief, so a
|
|
brief/condition mismatch was invisible to any reader of it. Three fixes
|
|
followed:
|
|
|
|
- verdict lines now carry the phase's **brief** beside its condition;
|
|
- the mission id in a line is the uuid **tail**: the two planted missions,
|
|
launched 34 s apart, both rendered as `01a0cb38` (a UUIDv7 leads with its
|
|
timestamp), and the audit read them as one mission;
|
|
- the recipe has the auditor **copy the record** into the checkout, since the
|
|
judge cannot read `/mission/memory`. Iteration 0 failed as "could not verify
|
|
the quotes" until the agent did this unprompted.
|
|
|
|
Counts, pattern-vs-incident discipline and verbatim evidence: shown. A
|
|
correct diagnosis of a cause: not yet shown, and until now not possible.
|
|
|
|
**Re-run with the fixes (mission 01a0cb54) — evidenced.** A second planted
|
|
pair, whose lines now carry their briefs, joined the record (11 lines: 5 MET,
|
|
6 UNMET). Launched again with nothing but `template_kind` and a repo, the
|
|
audit copied the record into the checkout, and the judge confirmed its counts
|
|
by grep against that copy. It passed on iteration 0. Its findings:
|
|
|
|
- **The planted pattern, diagnosed correctly.** *"The requirement appeared in
|
|
the condition but not in the brief, so the agent was never asked to write
|
|
it … a setup defect."* The proposal is the right one: the brief, old wording
|
|
quoted, new wording asking for the section.
|
|
- **A pattern nobody planted.** Its own two earlier UNMET audits (v3, and
|
|
01a0cb3f's first attempt) were judged for the same missing thing. It noted
|
|
that their briefs predate the record carrying briefs, so it could not tell
|
|
what those agents were told, and proposed "copy the record, quote verbatim".
|
|
That is the fix made by hand in `9ca71e5`, found independently from the
|
|
record alone.
|
|
|
|
**Seven of twelve.** Limits: the planted pattern is one shape (brief omits a
|
|
condition's requirement); the pre-fix lines still render two missions as
|
|
`01a0cb38`, which the audit counted correctly as two lines; and it proposes
|
|
into a markdown file. Nothing reads `research/IMPROVEMENT-AUDIT.md` or files
|
|
`level_up_proposals` from it.
|
|
|
|
**Two of the remaining six are still unevidenced**, and the other four
|
|
are blocked on a target stack. The other nine are well-formed scaffolding:
|
|
roles, prompts, brain seeds and resolving skills, and no run behind any of
|
|
them. They will probably work — they are structurally identical to the three
|
|
that do — but "probably" is the word, and this codebase has a name for the gap
|
|
between a thing being wired and a thing being proven.
|
|
|
|
Five of the nine need something we do not have: `frontend`, `mobile`,
|
|
`threejs` and `gpu` target stacks with no repository in the harness to point
|
|
them at, and `insight_research` needs a vault plus commit history. Those are
|
|
blocked on a target, not on the template. The remaining four —
|
|
`codebase_research`, `papers_research`, `continuous_improvement`, `backend` —
|
|
could be exercised against repositories we already have.
|
|
|
|
## `frontend` — the work is right; the judge cannot check it (2026-09-22)
|
|
|
|
Target: `osobh/clawmates-frontend-scratch` (private, disposable), a Vite +
|
|
React 19 + TypeScript + Tailwind v4 app with Vitest and Testing Library,
|
|
checked locally before any mission touched it. Task: an accessible Tabs
|
|
component (ARIA tablist/tab/tabpanel wiring, roving tabindex, Arrow/Home/End
|
|
with wrap-around, Tailwind focus style), with the brief and `done_when`
|
|
stating the same requirements.
|
|
|
|
Mission 01a0cbc7 delivered one commit, +191 lines, three files. Re-run
|
|
locally from the delivered branch: **10/10 tests pass (9 new), `tsc -b`
|
|
clean**. The component is correct on every stated requirement; one nit, the
|
|
tab buttons lack `type="button"`. The agents ran `npm test` and the
|
|
typecheck repeatedly in the mission container, all green.
|
|
|
|
The mission **failed**, both iterations, on the judge's side. The judge
|
|
verifies in `clawmates-runtime` against a copy that, by design, excludes
|
|
`node_modules` (the transport packer's list). It cannot reinstall:
|
|
`clawmates_core` has no gateway, so `npm ci` gets `EAI_AGAIN`. Every npm
|
|
project, which covers `frontend`, `mobile` and `threejs`, therefore fails any
|
|
`done_when` that asks for a passing suite, however good the work. Rust
|
|
escapes it only because the scratch crates have no dependencies. The judge
|
|
read the source correctly both times; it failed only on the suite it could
|
|
not run.
|
|
|
|
Not counted as evidenced until the judge can run the suite itself. How it
|
|
should get the dependencies (offline from the lockfile, registry egress, or
|
|
not at all) is an operator decision about the verifier's isolation.
|
|
|
|
## One thing both runs showed: agents reach for Bash
|
|
|
|
Cumulative tool calls across every mission since the policy shipped:
|
|
**Bash 74, Read 34, Write 15, Glob 3, Edit 1, Agent 1**. A read-only
|
|
mapping mission that could have used `Grep` and `Glob` used `Bash` 33
|
|
times and `Read` 5. That matters for task permission: the allowlist's
|
|
dedicated-tool entries (`Grep`, `ToolSearch`, `WebFetch`, `TodoWrite`)
|
|
may simply never be exercised, so "unconfirmed by a real run" will not
|
|
converge for them — and it means the surface that actually needs
|
|
governing is `Bash`, which the floor rules already cover.
|
|
|
|
## What this means we can ask for today
|
|
|
|
**With confidence:** a repo-backed research→code loop on a Rust project, with
|
|
a judge, a commit gate, a verifier subagent and delivery to a branch. A
|
|
repo-less research report with sourced claims. Both are re-proven on every
|
|
harness run.
|
|
|
|
**With supervision:** a security scan and plan, a benchmark, a refactor. Each
|
|
has completed once under a standing check; none has the depth of evidence the
|
|
first two have.
|
|
|
|
**With attention:** continuous research. It works, it produced two good
|
|
digests today, and it has no guard.
|
|
|
|
**Not yet:** anything staffed from the other nine teams. Nothing is known to
|
|
be wrong with them; nothing is known to be right either.
|
|
|
|
## The two things worth doing next
|
|
|
|
1. **A `continuous-research` harness scenario.** The recipe with the most
|
|
moving parts is the only one with no standing check, and it now also
|
|
carries the paper triage (`kind` + `evidence`). A scenario that launches
|
|
it, waits, and asserts the manifest has tags and a spread, `analysis.md`
|
|
covers every paper, and `episode.json` parses would turn operator
|
|
attention into a guard.
|
|
2. **Exercise the four unblocked teams once each**, against repositories we
|
|
already have, and record what came out. Not to prove them mature — one run
|
|
is not maturity — but to find out whether they run at all, which nobody
|
|
currently knows.
|