Run against the real clawmates repo with a task that cannot be faked from filenames: trace every hop of a mission agent's tool call from the guest hook to the recorded event, naming file, function, and guest vs host for each. It produced a 189-line research/GATE-MAP.md describing code committed the SAME DAY — the four-field NODE_EXTRACT including agent_type, the shadow-mode would-deny.jsonl semantics, install_with — and all six of its line citations verify exactly. Nothing hallucinated. That is a stronger result than the judge's verdict, which could only check that the structure was present. Also recorded: agents reach for Bash. 74 Bash calls against 34 Read, 3 Glob, 1 Edit, 1 Agent — a mapping mission that could have used Grep used Bash 33 times. So the allowlist's dedicated-tool entries may never be exercised, and the surface that actually needs governing is Bash, which the floor rules already cover. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
8.2 KiB
Template maturity — what our agents can actually be asked to do
A review of the 6 workflow recipes and 12 team templates, on 2026-09-22. Graded by evidence, not by what the TOML declares: a recipe is mature when a mission built from it has completed and something checked the result.
Workflow recipes
| recipe | phases | standing check | staffing | grade |
|---|---|---|---|---|
research_and_code |
2 (research → coding) | 13 harness scenarios | research→topic_research, coding→rust_sdlc |
Proven |
research_only |
1 (research, repo-less) | 2 scenarios | topic_research |
Proven |
security_hardening |
3 (scan → research → coding) | 1 scenario | +rust_sdlc |
Exercised |
benchmark |
1 | 1 scenario | rust_sdlc |
Exercised |
refactor |
1 | 1 scenario | rust_sdlc |
Exercised |
continuous_research |
2 (analysis → script) | none | continuous_research |
Runs live, unguarded |
research_and_code is the workhorse: every delivery, gate, judge, memory and
triage scenario is built on it, so it is re-proven on every harness run.
research_only earned its grade differently — its staffing was measured and
corrected: under the old rust_sdlc default it delivered 5 roles, 14 skill
deliveries and ~50 KB of prompt with 1 of 9 skills applicable; on
topic_research it is 3 roles, 4 deliveries, 24 KB, 4 of 4.
continuous_research is the outlier and the one to fix first. It has the most
moving parts of any recipe — harvest → triage → manifest → analysis → ranking
→ script → episode.json → vault commit — it ran twice today and produced
real output both times, and nothing in the harness would notice if it broke
tomorrow. Every other recipe has a standing check; this one has operator
attention, which is not the same thing.
What the recipes declare that does not fire
phase_config::DECLARED_BUT_UNREAD is the honest list, and recipes lean on it:
| key | used by | status |
|---|---|---|
loop |
research_and_code (until_no_more_int_items), refactor (single_pass) |
inert — iteration is max_iterations + done_when |
produces |
almost every recipe (["md"]) |
inert — artifact rendering is not driven by it |
input_from_phase |
security_hardening |
inert — phases share a checkout |
mcp_bundles (phase level) |
security_hardening |
inert — bundles come from the TEAM template |
None of this is hidden: security_hardening.toml carries a "what is real here
and what is decoration" section and annotates its own dead keys inline. That is
the right pattern and the other five should copy it. The risk is not the dead
keys themselves — it is that a reader takes loop = "until_no_more_int_items"
for a loop.
Team templates
All 12 are structurally sound and every declared skill resolves to a real
file — 0 unbound across all of them, which is the skill-binding-repair work
holding (it was 55 of 85 dangling once).
| team | slots | risk profile | skills | evidenced? |
|---|---|---|---|---|
rust_sdlc |
planner coder tester reviewer committer | coding_readwrite | 19 | yes — harness default |
topic_research |
lead_researcher evidence_checker report_writer | research_web_readonly | 4 | yes — measured |
continuous_research |
paper_reader signal_ranker script_writer | research_readonly | 11 | yes — live runs |
backend |
api_designer db_engineer coder tester committer | coding_readwrite | 20 | yes — first run 2026-09-22 |
frontend |
designer coder tester committer | coding_readwrite | 12 | no |
mobile |
designer coder tester committer | coding_readwrite | 11 | no |
gpu |
arch_analyst kernel_author bench_engineer coder committer | coding_readwrite | 13 | no |
threejs |
scene_designer coder shader_author perf_engineer committer | coding_readwrite | 13 | no |
codebase_research |
code_archeologist architecture_mapper flow_tracer vault_scribe | research_readonly | 11 | yes — first run 2026-09-22 |
papers_research |
domain_scout paper_reader library_curator | research_web_readonly | 8 | no |
insight_research |
implementation_tracker novelty_hunter publication_drafter | research_readonly | 8 | no |
continuous_improvement |
brain_inspector improvement_proposer improvement_evaluator | research_readonly | 7 | no |
Five of twelve are evidenced. backend was exercised on 2026-09-22,
the first time anything had been staffed from it: all five roles
provisioned with real agents (api_designer, db_engineer, coder, tester,
committer), and the mission delivered cursor pagination — Paged<T>,
paginate<T: Clone>, module declared, cargo test 3 passed — judged
met on the first pass, 2 files pushed. It also gave the task-permission
shadow its first confirmed Edit call.
codebase_research was exercised the same day, against the real
clawmates repository rather than the toy scratch crate, with a task that
cannot be faked from filenames: trace every hop of a mission agent's tool
call from the guest hook to the recorded event, naming file, function and
whether each hop runs in the guest or on the host. It produced a 189-line
research/GATE-MAP.md describing code committed the same day — the
four-field NODE_EXTRACT including agent_type, the shadow-mode
would-deny.jsonl semantics, install_with — and all six of its line
citations verify exactly (hook_script_with 562, NODE_EXTRACT 789,
TaskPolicy 297, ROLE_POLICIES 261, settings_hook 809,
install_command_with 819). Nothing hallucinated. Five of twelve.
Three of the remaining seven are still unevidenced. The other nine are well-formed scaffolding: roles, prompts, brain seeds and resolving skills, and no run behind any of them. They will probably work — they are structurally identical to the three that do — but "probably" is the word, and this codebase has a name for the gap between a thing being wired and a thing being proven.
Five of the nine need something we do not have: frontend, mobile,
threejs and gpu target stacks with no repository in the harness to point
them at, and insight_research needs a vault plus commit history. Those are
blocked on a target, not on the template. The remaining four —
codebase_research, papers_research, continuous_improvement, backend —
could be exercised against repositories we already have.
One thing both runs showed: agents reach for Bash
Cumulative tool calls across every mission since the policy shipped:
Bash 74, Read 34, Write 15, Glob 3, Edit 1, Agent 1. A read-only
mapping mission that could have used Grep and Glob used Bash 33
times and Read 5. That matters for task permission: the allowlist's
dedicated-tool entries (Grep, ToolSearch, WebFetch, TodoWrite)
may simply never be exercised, so "unconfirmed by a real run" will not
converge for them — and it means the surface that actually needs
governing is Bash, which the floor rules already cover.
What this means we can ask for today
With confidence: a repo-backed research→code loop on a Rust project, with a judge, a commit gate, a verifier subagent and delivery to a branch. A repo-less research report with sourced claims. Both are re-proven on every harness run.
With supervision: a security scan and plan, a benchmark, a refactor. Each has completed once under a standing check; none has the depth of evidence the first two have.
With attention: continuous research. It works, it produced two good digests today, and it has no guard.
Not yet: anything staffed from the other nine teams. Nothing is known to be wrong with them; nothing is known to be right either.
The two things worth doing next
- A
continuous-researchharness scenario. The recipe with the most moving parts is the only one with no standing check, and it now also carries the paper triage (kind+evidence). A scenario that launches it, waits, and asserts the manifest has tags and a spread,analysis.mdcovers every paper, andepisode.jsonparses would turn operator attention into a guard. - Exercise the four unblocked teams once each, against repositories we already have, and record what came out. Not to prove them mature — one run is not maturity — but to find out whether they run at all, which nobody currently knows.