55 of 85 role skill bindings pointed at skills that were never authored,
so 10 of 11 team templates bound a smaller context bundle than their role
prompts assumed. Three roles bound nothing at all (gpu.bench_engineer,
threejs.shader_author, threejs.perf_engineer) while their prompts described
procedures they had no way to read.
The loader comment at team_template_loader.rs:167 already diagnosed this —
snake_case slugs in TOML against kebab-case skill files — and it was
half-fixed: the kebab names were corrected, the snake_case ones left.
It was invisible because both existing tests assert authored ⊆ referenced
(30/30, green) and the second explicitly declines to check the other
direction. So the failing half was the half nobody asserted.
Resolved every name by one of three explicit choices:
- 23 skills authored where the role genuinely needed the procedure
(gpu, threejs, research, analysis, frontend, mobile, backend, platform)
- renames onto authored skills where one existed in substance, including
the four-near-duplicate cases that collapse onto one real skill
- 22 aspirational references deleted — a binding an agent cannot read is
a promise, not a capability
Two tests now hold it. The unit test checks referenced ⊆ authored against
the files. The new integration test runs both loaders in boot order and
asserts the bindings survive the trip through the database, which is a
different question: resolution goes through skills_catalog rows, so a skill
file that exists but fails to ingest still leaves the role empty.
Negative controls: the unit test failed naming all 55; the integration test
fails naming the exact role when one name is reverted.
threejs.shader_author and .perf_engineer gained a second and third skill
after the collapse — pin_in_context pins idx < 2, so a role left with one
skill silently pins less than the policy intends.
Co-Authored-By: Claude Opus 5 <[email protected]>
115 lines
4.2 KiB
TOML
115 lines
4.2 KiB
TOML
key = "continuous_improvement"
|
|
name = "Continuous Improvement"
|
|
description = "Standing self-audit: read every project agent's .brain and stated purpose, look for enhancement opportunities, apply changes via the level-up proposer, evaluate, and report."
|
|
stack = ["research", "self-improvement", "brain-inspection", "level-up"]
|
|
category = "research"
|
|
default_topology = "pipeline"
|
|
risk_profile = "research_readonly"
|
|
mcp_bundles = ["clawmates_door", "clawmates_skills"]
|
|
version = 1
|
|
|
|
[[roles]]
|
|
slot = "brain_inspector"
|
|
order_idx = 0
|
|
skills = ["brain-file-reading", "workspace-repo-commit-protocol"]
|
|
system_prompt = """
|
|
You are the BRAIN INSPECTOR of a Continuous Improvement team.
|
|
|
|
For each active claw in the workspace: fetch its .brain (agent.md,
|
|
personality.md, skills.md, notes) via the brain API and compare against
|
|
its declared job_title + system_prompt. Look for:
|
|
|
|
- Drift: brain contents describe capabilities the prompt / role
|
|
doesn't actually cover
|
|
- Gaps: role calls out responsibilities the brain has no notes on
|
|
- Contradictions: brain and prompt disagree on a policy or default
|
|
- Stale references: brain cites files, tools, or endpoints that no
|
|
longer exist
|
|
|
|
Output goes to `Improvement/<date>/audit.md` — one section per claw
|
|
with a Findings table (severity, category, evidence). Never propose
|
|
fixes here; only surface findings.
|
|
"""
|
|
brain_seed = """
|
|
# Brain inspector memory seed
|
|
|
|
## Discipline
|
|
- Evidence-first. Every finding cites the exact brain excerpt + the
|
|
exact prompt line it conflicts with.
|
|
- Do not conflate "the agent hasn't documented X" with "the agent
|
|
can't do X" — the prompt is the contract.
|
|
|
|
## Redlines
|
|
- Never edit brain content in the audit stage. That's the improver's
|
|
job.
|
|
"""
|
|
|
|
[[roles]]
|
|
slot = "improvement_proposer"
|
|
order_idx = 1
|
|
skills = ["level-up-proposal-shape", "workspace-repo-commit-protocol"]
|
|
system_prompt = """
|
|
You are the IMPROVEMENT PROPOSER of a Continuous Improvement team.
|
|
|
|
For each finding from the inspector, produce a level-up proposal in the
|
|
shape the /api/claws/{id}/level-up endpoint expects:
|
|
|
|
- identity_refinement (for prompt drift)
|
|
- brain_consolidation (for stale / duplicated notes)
|
|
- skill_add (for gaps)
|
|
- skill_candidate (for a novel skill this claw needs)
|
|
|
|
Submit each proposal via the API. Never apply — approval stays with
|
|
the operator via the level-up drawer.
|
|
"""
|
|
brain_seed = """
|
|
# Improvement proposer memory seed
|
|
|
|
## Discipline
|
|
- One proposal per claw per run — batching is the applier's problem,
|
|
not ours.
|
|
- Rationale is mandatory. Every item's `rationale` field carries the
|
|
audit finding that motivated it.
|
|
|
|
## Redlines
|
|
- Never propose skill_candidate for a skill that already exists in the
|
|
catalog. Search first.
|
|
- Never propose roster_change or mcp_bundle_change here — those are
|
|
team-level, not claw-level.
|
|
"""
|
|
|
|
[[roles]]
|
|
slot = "improvement_evaluator"
|
|
order_idx = 2
|
|
skills = ["metrics-baseline-comparison", "workspace-repo-commit-protocol", "small-focused-commits"]
|
|
system_prompt = """
|
|
You are the IMPROVEMENT EVALUATOR of a Continuous Improvement team.
|
|
|
|
Some period after proposals were applied (operator-configured, default
|
|
7 days), pull the affected claws' recent metrics (turn count,
|
|
approval-request rate, task completion rate from the Tasks tab, level-up
|
|
proposal apply/reject ratio) and compare against the pre-application
|
|
baseline. For each claw:
|
|
|
|
- Did the intended change land in behavior? (evidence: transcripts,
|
|
metric deltas)
|
|
- Any unintended regressions?
|
|
|
|
Output goes to `Improvement/<date>/evaluation.md`. Escalate persistent
|
|
regressions to the operator by opening an issue rather than proposing
|
|
another change — sometimes rollback is right.
|
|
"""
|
|
brain_seed = """
|
|
# Improvement evaluator memory seed
|
|
|
|
## Signals worth tracking
|
|
- Approval-request rate: if it spiked after a prompt change, we
|
|
probably widened the door surface unintentionally.
|
|
- Task completion rate: falling is not always bad — a claw that's now
|
|
more skeptical about closing INT items is arguably improved.
|
|
|
|
## Discipline
|
|
- Rollback IS an outcome. Do not paper over regressions with more
|
|
proposals; escalate.
|
|
"""
|