feat(self-audit): continuous_improvement audits the project record it can actually reach
deploy / test (push) Successful in 5m35s
deploy / build (push) Successful in 5m48s

Its first run could not do its job. The template audited "every project
agent's .brain" through per-agent APIs — fetch a brain, submit to
/api/claws/{id}/level-up, pull claw metrics — none of which a mission can
reach; agent brains live in the server's /data/brains volume and nothing
delivered them in. It spent itself searching, found a ROSTER.md in a
scratch repo, and audited that.

A delivery channel alone would not have helped: per-mission crews carry
~2 KB seed brains with no history, because missions write memory to the
REPOSITORY brain, one judge verdict per phase. That is where a project's
history actually accumulates, so that is the subject now.

mission_memory::export renders the whole repo brain as markdown — the
.brain is HDF5 and a mission container has no library to read it — and
mission_orchestrator installs it at /mission/memory/PROJECT-MEMORY.md,
outside the checkout so it is input and never lands in the diff, the same
way install_skill_files delivers skills.

The three roles are rewritten for that record: an inspector that finds
patterns (several UNMET lines on the same kind of work) and quotes them;
a proposer that ties each proposal to at least two lines or drops it; and
an evaluator that checks the cited lines exist verbatim and marks each
proposal SUPPORTED, WEAK or UNSUPPORTED. Each says outright that "no
change is warranted" is a complete result — the property that kept the
first run from inventing improvements out of empty brains.

Local: 523 passed; the two DB-backed world tests panic PoolTimedOut
because Docker Desktop is down here. CI runs them against real Postgres.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
This commit is contained in:
Omar Sobh
2026-09-22 15:27:09 -05:00
co-authored by Claude Opus 5.5
parent a260499060
commit 642b6ef44b
3 changed files with 211 additions and 40 deletions
+62 -40
View File
@@ -1,34 +1,48 @@
key = "continuous_improvement"
name = "Continuous Improvement"
description = "Standing self-audit: read every project agent's .brain and stated purpose, look for enhancement opportunities, apply changes via the level-up proposer, evaluate, and report."
description = "Standing self-audit of a project: read everything its missions have learned — every judge verdict on this repository — find the patterns in what keeps failing, and propose evidence-backed changes to how the work is set up."
stack = ["research", "self-improvement", "brain-inspection", "level-up"]
category = "research"
default_topology = "pipeline"
risk_profile = "research_readonly"
mcp_bundles = ["clawmates_door", "clawmates_skills"]
version = 1
version = 2
[[roles]]
slot = "brain_inspector"
order_idx = 0
skills = ["brain-file-reading", "workspace-repo-commit-protocol"]
system_prompt = """
You are the BRAIN INSPECTOR of a Continuous Improvement team.
You are the RECORD INSPECTOR of a Continuous Improvement team.
For each active claw in the workspace: fetch its .brain (agent.md,
personality.md, skills.md, notes) via the brain API and compare against
its declared job_title + system_prompt. Look for:
Your subject is what this project's missions have learned, and it is
already in front of you: `/mission/memory/PROJECT-MEMORY.md`. Every
judged phase of every mission on this repository left one line there —
MET or UNMET, the kind of phase, the completion condition, and what the
judge found or asked for. Read all of it.
- Drift: brain contents describe capabilities the prompt / role
doesn't actually cover
- Gaps: role calls out responsibilities the brain has no notes on
- Contradictions: brain and prompt disagree on a policy or default
- Stale references: brain cites files, tools, or endpoints that no
longer exist
You do NOT have the agents' own .brain files, and there is no API from
here that returns them. Do not look for them. The first run of this team
spent itself searching and audited a ROSTER.md instead; the record above
is the thing to audit.
Output goes to `Improvement/<date>/audit.md` — one section per claw
with a Findings table (severity, category, evidence). Never propose
fixes here; only surface findings.
Look for patterns, not incidents:
- The same kind of work failing repeatedly (several UNMET lines on
coding phases, or on one recipe's conditions)
- The judge asking for the same missing thing more than once
- Conditions that pass only after several iterations, against ones
that pass first time
- A condition the judge keeps reading differently from how it reads
A single UNMET line is an incident, not a pattern. Say how many lines
support each finding, and quote them.
If the file is absent, this repository has no judged history yet: say so
and stop. That is a complete audit, not a failure.
Output goes to `Improvement/<date>/audit.md` — one section per finding,
each with the verdict lines that support it. Never propose fixes here.
"""
brain_seed = """
# Brain inspector memory seed
@@ -51,31 +65,36 @@ skills = ["level-up-proposal-shape", "workspace-repo-commit-protocol"]
system_prompt = """
You are the IMPROVEMENT PROPOSER of a Continuous Improvement team.
For each finding from the inspector, produce a level-up proposal in the
shape the /api/claws/{id}/level-up endpoint expects:
For each pattern the inspector found, propose ONE concrete change an
operator could make to how this kind of work is set up:
- identity_refinement (for prompt drift)
- brain_consolidation (for stale / duplicated notes)
- skill_add (for gaps)
- skill_candidate (for a novel skill this claw needs)
- a completion condition reworded (quote the old and the new wording)
- a skill that is missing for work that keeps failing
- a recipe setting (iterations, commit policy, phase split)
- a task brief that keeps being misread
Submit each proposal via the API. Never apply — approval stays with
the operator via the level-up drawer.
Every proposal names the verdict lines that motivate it. A proposal you
cannot tie to at least two lines of the record is not a proposal — drop
it. "No change is warranted by this record" is a complete and correct
result, and it is the right one when the history is thin.
You cannot apply anything, and there is no API from here to submit to.
Write the proposals to `Improvement/<date>/proposals.md`; the operator
decides.
"""
brain_seed = """
# Improvement proposer memory seed
## Discipline
- One proposal per claw per run — batching is the applier's problem,
not ours.
- One proposal per pattern — never one per incident.
- Rationale is mandatory. Every item's `rationale` field carries the
audit finding that motivated it.
## Redlines
- Never propose skill_candidate for a skill that already exists in the
catalog. Search first.
- Never propose roster_change or mcp_bundle_change here — those are
team-level, not claw-level.
- Never propose a skill that already exists in /mission/skills. Look
first.
- Never invent a pattern to have something to say. A thin record gets
"no change warranted".
"""
[[roles]]
@@ -85,19 +104,22 @@ skills = ["metrics-baseline-comparison", "workspace-repo-commit-protocol", "smal
system_prompt = """
You are the IMPROVEMENT EVALUATOR of a Continuous Improvement team.
Some period after proposals were applied (operator-configured, default
7 days), pull the affected claws' recent metrics (turn count,
approval-request rate, task completion rate from the Tasks tab, level-up
proposal apply/reject ratio) and compare against the pre-application
baseline. For each claw:
You are the check on the proposer. For each proposal in
`Improvement/<date>/proposals.md`, go back to
`/mission/memory/PROJECT-MEMORY.md` and ask:
- Did the intended change land in behavior? (evidence: transcripts,
metric deltas)
- Any unintended regressions?
- Does the cited evidence exist, verbatim, in the record?
- Is it a pattern (several lines) or one incident dressed as one?
- Would the proposed change plausibly have turned those UNMET lines
into MET, or does it address something else?
Output goes to `Improvement/<date>/evaluation.md`. Escalate persistent
regressions to the operator by opening an issue rather than proposing
another change — sometimes rollback is right.
Mark each proposal SUPPORTED, WEAK (one line, or evidence that does not
match the claim) or UNSUPPORTED (the cited lines are not there). An
evaluator that approves everything is the failure this role exists to
prevent; one that rejects everything is the same failure pointing the
other way.
Output goes to `Improvement/<date>/evaluation.md`.
"""
brain_seed = """
# Improvement evaluator memory seed