248948cc847b4d229291fa65785d940b02fd36ca
45
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2a3409ec53 |
docs: the judge is back and the subagent path is no longer a claim
The z.ai quota reset on schedule. glm-5.3 answers on the same key and has since passed a real done_when — phase completed on iteration 0, with the verdict naming the arXiv ids it checked rather than waving the phase through. Mission 01a07498 was the failed validation run plus one change, and it closed the honest negative the last handoff recorded: 87 tool calls, 43 from the main turn and 44 across 4 general-purpose subagents, 4 distinct subagent_ids against 4 Agent spawns. Before this the field was correct in unit tests and had never been watched writing. The one change was the finding. The earlier task invited delegation and got none; naming the tool and forbidding the single-turn shortcut produced four spawns from the same recipe and the same delivery arm. A fan-out path that is merely invited measures nothing. Also records that postgres is clawmates-postgres-1 locally and clawmates_postgres_1 on gw-04 — the wrong one reports "No such container", which reads like a down stack rather than a typo. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz |
||
|
|
daf8d12157 |
docs: hand off — what shipped, what is verified, and what is blocked
Seven commits this pass, all deployed. The handoff leads with the thing that will otherwise waste the next session's first hour: `glm-5.3` hit a hard z.ai quota on 2026-08-29 (code 1310, resets 09-04), it is the DEFAULT validator on both stacks, and the `ZAI_API_KEY` fingerprints are identical — so every mission declaring a `done_when` fails its evaluation on local and production alike, with its artifacts fully delivered and correct. That failure is not a bug to fix. `evaluator.rs:480` refuses to fall back to the agent's own provider because a same-family verdict would claim an independence it does not have. It is also NOT the malformed-prompt 429 we hit before: this one carries a code and a reset date. Records the validation run honestly rather than as a clean sweep. Three of four things confirmed live — `always_inject` delivering a body beside an index entry in one prompt, retrieval still firing through the door, the corrected gate installed and quiet against 23 body-free Bash calls, attribution 34/34. The fourth did not happen: those agents never delegated, so the subagent field is written and null, and the path that motivated it has still never been watched populating `mission_events`. A task that invites delegation does not force it; the next attempt should instruct it outright. Also carried forward: local test state that production does not share (`workspace-repo-commit-protocol.always_inject = true`, set by hand), the three fork items sitting behind one runtime image rebuild with tank offline, and two silent-discard defects found by sweep and left unfixed — `container_tool_hooks:: install`'s outcome is recorded nowhere, which makes "did this mission run gated?" unanswerable once the container is reaped. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz |
||
|
|
563b074116 |
docs: ZeroClaw upstream, scanned against what we actually run
331 behind, 54 ahead. The previous scan said 218 and its conclusion about the egress commit was wrong, so it is marked superseded rather than edited. Merge cost is smaller than the number suggests: 660 files changed upstream, 52 by us, and **18 overlap**. `claude_cli.rs` — the provider every mission runs through — exists in our tree and in zero upstream files, so it cannot conflict. The find worth recording is not a feature. Upstream defaulted skills to compact injection on 2026-08-05 (#8313), then restored the full default for v0.8.x on 2026-08-13 (#9913). Eight days. That is our `index` arm, tried at larger scale and pulled back out of the stable line — evidence bearing directly on our own open question of whether to flip the default, and with our own data at n=1 per arm it argues for more pairs before flipping, not fewer. Their documentation also states plainly what ours should: "Compact mode reduces prompt size; it is not an isolation boundary for untrusted skill sources." Progressive disclosure is a token optimisation. It is not a security control. Also noted, as a documented limit rather than a surprise: upstream fixed case-insensitive allowlist matching (#9568) and symlink-escape path resolution (#9384) in their command gate. Ours resolves no paths, so a symlink to `curl` defeats it. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz |
||
|
|
fde1341618 |
docs: what a mission container can reach, and why item 4 could not fix it
Plan item 4 said to pull upstream's `0db7d999a` egress policy as "defence for the egress problem we have not solved". It cannot see the problem. `claude_cli` runs the claude binary as a SUBPROCESS (`Command::new(&self.binary_path).spawn()`), so every mission tool call happens inside that child. `net_guard`'s only call sites upstream are `link_enricher`, `helpers/domain_guard` and `plugins/egress` — ZeroClaw's own Rust HTTP. A mission agent's `curl` never touches the guarded stack. Pulling the commit hardens the CHAT tier; it leaves mission egress exactly as it is. Item 4 is corrected in place rather than deleted, because the reasoning is the useful part. What is actually true, measured on gw-04 with controls in both directions: positive 1.1.1.1:443 REACHABLE negative 192.0.2.1:80 TEST-NET blocked tailnet gw-02 100.84.218.70:22 REACHABLE host SSH docker gw 172.23.0.1:22 REACHABLE 169.254.169.254 REACHABLE postgres REACHABLE, password-required LAN 192.168.1.1 blocked A mission agent reaches the entire tailnet and SSH on its own host. It matters more here than it would elsewhere: these agents run model-generated shell over content fetched from the open web — 151 of 158 production Bash calls were curl/wget — so the instruction stream and the data stream are one stream. The first run of this probe attached only `clawmates_core`, reported "no internet", and was discarded: its positive control failed, so it measured nothing. A mission container is on BOTH networks and that is what must be reproduced. Recorded in full, including the half that is fine — postgres refuses unauthenticated TCP and no database credentials are forwarded into a mission container — because a report that lists only the bad half is not a measurement. Remediation is written down and deliberately NOT applied, on the operator's call. It is DOCKER-USER rules dropping the private world with the core subnet accepted first; never a public host allow-list as the opening move, because `JEPA Research` alone fetched a dozen hosts nobody would have pre-approved and a mission that cannot read cannot do research. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz |
||
|
|
7bcf7865f0 |
docs: prod ran a mission, and the whole chain held
`ClawHDF5` and `JEPA Research` were launched from the UI. The first completed,
and everything shipped over the previous two passes engaged correctly on its
first execution anywhere outside the local stack:
skills door installed — api_origin() derived the host from the server's
own container id, which had never run where it was not tested
staffing Topic Research, 3 roles / 4 deliveries, not rust_sdlc's 5 / 14
drain 92 tool calls
attribution 92 of 92, across a phase with TWO passes and six turns — the
case attribute_sessions had never met, and it attributes
nothing at all unless the counts match exactly
boundary all 8 Write/Edit paths under /mission/repo
arm inline, 0 retrievals; prod leaves the env unset
judge pass 0 met=false "zero URLs — grep -c http returns 0"
pass 1 met=true "57 http references"
The judge line is the one worth rereading: the loop converged on the exact
mechanically-checked defect it named, and pass 0 would otherwise have shipped
a report whose every claim was unattributed while reporting `completed`.
Two traps recorded rather than smoothed over:
- The drain selects phases `IN ('completed','failed')`, so a phase on its
second pass shows zero tool calls and reads as broken while being correct.
- I reused a diagnostic query with no `WHERE mission_id`. That was fine while
prod held one mission and silently wrong the moment a second launched — it
compared one mission's tap against two missions' events. The production
drain query is correctly scoped; the diagnostic was not.
Unexplained: one agent called the `Agent` tool 4 times. Mission agents are
spawning subagents and nothing in our design accounts for it.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
|
||
|
|
accae7fa94 |
docs: the delivery A/B, and the first retrieval nobody asked for
Runs 9 and 10: identical task text, one server process, and a task that never
mentions skills, MCP or retrieval. Run 8 demonstrated the instrument, but its
retrieval was instructed by the task — it showed the pipe worked, not that an
agent would judge relevance.
Under `index`, two of four skills were fetched, and attribution is the part
that matters:
Solveig (lead_researcher) -> web-search-triage
Olamide (report_writer) -> scientific-writing-conventions
Each agent reached for the skill bound to its OWN role and neither reached for
another's. An agent that fetched all four would have shown only that it could.
The regression the A/B existed to catch did not appear: 34% fewer tokens, 59
tool calls against 89, both arms passed the independent judge, and the
deliverables came out slightly larger rather than thinner.
Two readings the data does not support, recorded because the first draft of
this section made one of them:
- Every `tool.call` in a phase carries the DRAIN timestamp, not the call time.
All 59 rows of run 10 read `12:48:12`. Ordering by that column said the
report writer had fetched both skills; `agent_id` says otherwise.
- The prompt saving is 15-43%, not an order of magnitude. Skill bodies are a
minority of a turn prompt. Progressive disclosure is worth doing for Trigger,
not for context economy.
`workspace-repo-commit-protocol` scores Trigger=FAIL beside boundary=pass: it
behaved correctly without reading the rule. That verdict is left standing and
argued with in the text rather than tuned away.
n=1 per arm. A signal, not a rate.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
|
||
|
|
f52cff3e04 |
feat(skill-use): progressive disclosure, as an arm and not a switch
Trigger — did the agent reach for the skill when it applied? — cannot be measured while every body is inlined into the prompt. Nothing was reached for. `skill_use` has been reporting `NotObservable` for that reason, and it was right to. The skills door made retrieval possible; this makes it a delivery arm. `index` sends each pinned skill's name, description, `when_to_use` and the uri that returns its body, and the agent fetches what it judges relevant. `inline` is unchanged and stays the default. An A/B rather than a switch, because `index` can only cost Compliance: under `inline` the procedure sits in front of the model whether or not it noticed it applied. Trading a measured axis for an unmeasured regression in another is not an improvement, so both arms stay runnable and the arm is recorded on the mission row. Three things the mechanism refuses to do: - `index` without a door falls back to `inline`. An index names bodies and says how to fetch them; with no `clawmates_skills` server reachable that is a list of dead ends, and it fails as an agent ignoring its skills rather than as a missing config. `install_skills_door` now returns whether it installed, because the caller needs the answer and not just the log line. - The scorer reads the arm off the recorded PROMPT, not off the mission row. The row says what the mission is configured to do now; the score is being computed against a turn that ran then. - Under `index`, a skill that was offered and never read is a Fail, not the inline arm's `NotObservable` — but only where the skill had a checkable consequence in that phase. Reusing the inline text would have said "this skill was inlined into the prompt" about a skill whose body was never sent, and scoring a real miss as a structural blind spot is the failure this measurement already made once. The arm is per mission (`config.skill_delivery`), not only per deployment. Both arms run against one server process; restarting between them would put a confound in the comparison that the numbers would not show. 829 tests, 108 binaries, green. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> |
||
|
|
72eda8b3d2 |
docs: handoff reflects the pushed state
22 commits pushed, CI green, deployed. Items 1-3 of the previous list are done: staffing, attribution, and the door. Trigger is measured and red-first turned out observable from run outputs rather than from the diff — the previous list was wrong about that, and the skill says why. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
42e014976d |
docs: confirm the door from inside a mission, and correct a count I took from a transcript
Run 8, all three agents, per-agent attributed: Pedro / Ebele / Ahmad — ListMcpResourcesTool, ReadMcpResourceTool The agent's own report: 53 resources from `clawmates_skills`, and `skill:global/workspace-repo-commit-protocol` read back as `# Mission repo + commit protocol`. So the wiring works end to end, not just the mechanism. And a correction to the commit before this one. It recorded "58 MCP resources" as a measurement. That number was the model's paraphrase in a probe transcript, not an observation. `resources/list` returns 53 and `select count(*) from skills` is 53. Noted in the doc rather than quietly changed, because it is the same error this project keeps making — a model's self-report treated as evidence — and I made it in the very document arguing for measuring things. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
02d5f5a8c9 |
docs: the door is deployed, and what it does not buy
Proven against the real binary in the runtime container — connect, list (58 resources) and read (`# Mission repo + commit protocol`, the correct first heading). That probe is a two-minute loop; I reached for the ten-minute rebuild-and-run-a-mission one first, and it would have found the container-name bug sooner. No `--allowedTools` change was needed. Recorded because the guess would have been wrong in an expensive way: with no config read on the daemon, "adding" the MCP tools meant overwriting the seed's `tools` list and stripping Write and Bash from every mission agent — to solve a problem that does not exist. The §3 claim that this was "config, not code" is corrected in place: it needed a credential narrow enough to leave in a container an untrusted agent reads, and the measured proof that the credential IS narrow (same token: 58 skills from /mcp/skills, 401 from /api/missions). And what it does not buy, stated plainly: Trigger is still unmeasured, because delivery still inlines. The door makes retrieval possible; making Trigger real means switching to progressive disclosure, which could regress Compliance and so wants an A/B rather than a flip. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
3f26dfeaca |
docs: suite is 796 tests across 107 binaries after this pass
Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
19c4de36e4 |
docs: the staffing fix, measured
Run 5 is run 3's task against the new staffing: 5 roles → 3, 14 skill deliveries → 4, 50KB of prompt → 24KB, and 1 of 9 delivered skills applicable → 4 of 4. The agents produced exactly the structure the new team's task specifies — questions.md, evidence.md, REPORT.md — with zero writes outside /mission/repo. The baseline says plainly that the SCORES barely moved, because they did: run 5 is one `pass` and three `not_applicable`. What changed is what `not_applicable` means — "no machine-checkable consequence" rather than "this skill had nothing to do with this phase". Halving the prompt is real but incidental. The finding is that the denominator was wrong: seven of run 3's nine skills were never applicable, so any ratio over them measured staffing, not skill use. Handoff item 1 is closed and the orphan-container section now records what was actually in it. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
6f2b0a8f43 |
docs: record the verified suite numbers in the handoff
107 test binaries, 792 tests, zero failures across the workspace — run, not estimated from the cm-api figure. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
9560aaec41 |
test(skill-use): the coding run, and the parsing bug it found
Run 4 (`research_and_code`, real repo) is the first mission that could
have violated the TDD and commit checks. It exercised both, and found a
bug in one.
Claude Code writes a multi-line commit message as a heredoc inside a
command substitution:
git commit -m "$(cat <<'EOF'
INT-01 Add slugify function to src/lib.rs
…
EOF
)"
`commit_subjects` read the first line of the `-m` value, which is the
heredoc OPENER. Every commit check was scoring `$(cat <<'EOF'` — a string
the agent never wrote. It reported no violation only because that string
is not one of the never-merge messages, which is luck rather than a check.
Regression test built from the exact command in `mission_events`.
The TDD verdict came back `not_observable`, which is the honest answer and
also a real limit worth stating: the agents edited `src/lib.rs` once —
implementation and `#[cfg(test)] mod tests` in the same write — then ran
`cargo test` five times. In Rust the unit test lives in the file under
test, so that ordering is exactly what following the skill precisely looks
like from outside. The check detects "wrote source, never ran a test" and
cannot confirm red-first. Confirming it needs the diff, not the tool order.
Every one of run 4's 33 tool calls stayed inside /mission/repo.
Handoff and baseline updated: production has never run a mission (both
tables empty), a mission container has leaked since 2026-08-12 that no
reaper can see, and `research_only` staffs a five-role Rust SDLC crew on a
repo-less markdown mission — which is what "most skills score
not_applicable" has been measuring all along.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9
|
||
|
|
c209e654d9 |
fix(skill-use): a research phase writing markdown is not a TDD failure
The first live scoring of run 3 reported `cargo-test-driven-development` and `tdd-red-green-refactor` as compliance=FAIL: files were written and no test ever ran. Wrong, and wrong in the way this module exists to prevent. The phase wrote fifteen markdown notes and a helper script; there was no code to test-drive. Reporting it as an agent failure is a system defect wearing an agent's name — and it would have buried the actual finding, which is that a repo-less `research_only` mission is staffed with a Rust SDLC crew whose coder, tester, reviewer and committer have nothing to do. The check is now scoped to files with a source extension in the languages the skill itself names. Shell is deliberately excluded: a helper script written during a research turn is not behaviour-adding code, and the false failure costs more than the missed one. Recorded in SKILL-USE-BASELINE.md as finding 8 rather than quietly corrected. A measurement that hides its own false positives cannot be trusted about anyone else's. Also in the doc: the Trigger reason is half false now (the transport can surface a tool call; we simply still inline), and the architecture doc's observe/gate table said the container tier was ungated and unobserved, which shipped work has made wrong. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_018i9Ten1LU4jUr5d7TAWda9 |
||
|
|
0b4d91889a |
docs: hand-off refresh — container-tier work shipped, stale guidance corrected
TOOL-CALL-ARCHITECTURE.md said "switch claude_cli to stream-json" as the cheapest fix. That was wrong and is now marked so, with what actually happened: zero tool.call events with the parser working perfectly, because TurnEvent::ToolCall only fires for tools ZeroClaw itself executes. Hooks sidestep that entirely, and the doc now leads with the resolution rather than the theory. A fresh session is pointed at this file, so leaving the wrong recommendation on top would have sent it down the same path. NEXT-SESSION.md: state header, and the ordered list rewritten — items 1-3 are done or superseded. "Give the direct-session tier a tap" is dropped with its reason: that tier is dormant (CLAWMATES_MISSION_EXECUTOR unset), and checking before building saved the work. New top item is watching the first production mission, since the gate and tap are proven locally and unproven in prod. Added an operational section for the things that cost the most time: the 403 actions-log API, gw-04's legacy docker-compose, the socket proxy, disk contention between manual builds and CI, and Clerk-only prod auth. Also flagged that SKILL-USE-BASELINE.md's Trigger column is now stale in a good way — tool calls are observable on the container tier, so Trigger can be scored from behaviour instead of prose. That is the highest-value follow-up. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
5a11fae0d6 |
docs: container-tier gate and telemetry shipped; CI failures were disk
Records the verified result (10 tool.call, 4 file.touch on a real mission), how hooks succeed where stream-json could not, the production state and its rollback, and the three same-shaped bugs the live test found. Also records that CI's build failures were disk pressure from my own manual runtime builds on gw-04 — not code — and that a docs-only commit was the first casualty, which made it look like a regression. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
930c7e0b67 |
docs: the PreToolUse gate is verified end to end
Ran it against the real claude binary with the real settings document and the real hook script. Both halves. It blocks: asked to `curl -X POST`, the agent attempted the Bash call, the hook fired FROM --settings, the call was refused, and denied.jsonl recorded the payload with hook_event_name PreToolUse and the exact command. The agent relayed the reason accurately — the text from vm_tool_gate::RULES reached the model, which is the point of writing reasons rather than bare refusals. It allows: `echo` and a harmless `rm -rf ./scratch-nonexistent` both ran and denied.jsonl stayed empty. A gate that blocked everything would have passed the first test; this is the half that rules that out — and two of this gate's four bugs produced exactly that failure. So the last unproven link in the chain is closed, and the gate is real in production rather than plausibly real. One finding worth keeping: asked to `git push --force`, the model refused on its OWN before ever calling Bash, so the hook never fired and the test was inconclusive. A gate test must use a command the model will actually attempt. The model's judgement is not the gate, and testing against something it already refuses measures nothing. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
0be932fd83 |
test(gate): a fixture emitter for the live PreToolUse check, and what it proved
Tried to close the last open question — does the PreToolUse gate actually
fire in a guest — and got most of the way.
Established:
- the generated script blocks and allows correctly under DASH, not just
macOS sh: force-push and `cd /tmp && rm -rf /` return 2, while
`grep -rn 'rm -rf /' docs/` and ordinary work return 0
- without node it allows and writes the `inert` marker, so a gate that
cannot parse is distinguishable from one that matched nothing
- `claude` in the runtime image supports `--settings` (SETTINGS-OK)
- PreToolUse DOES fire under `claude -p` in this image — measured by an
earlier session and recorded in vm_stop_gate.rs:36
Unproven, and now precisely scoped: whether Claude Code honours a
PreToolUse hook supplied via `--settings <path>` specifically, with a real
agent turn. The live attempt hit the weekly subscription rate limit, and
`claude doctor` does not report hooks, so there is no non-LLM confirmation
available.
`emit_guest_assets` (ignored by default) writes the real hook script and the
real settings document to /tmp so the check can be run against the actual
binary in one docker command — no microVM, no fleet. The exact command is in
docs/NEXT-SESSION.md.
Worth stating plainly: if that link is broken, the gate is inert in
production and looks exactly like a gate that found nothing — which is the
failure mode this whole session has been about.
Full workspace suite green: 107 binaries.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
afb1e29bf3 |
docs: streamjson2 built but deliberately not deployed
The corrected runtime image exists on gw-04 and stays there. It delivers no observability until TurnEvent::ToolCall can be emitted for observed calls, so deploying it alone would be a provider output-format change carrying risk for no benefit. Production stays on the known-good :v084. The harmful v1 image was deleted from both hosts so it cannot be redeployed by accident. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
536adddd0f |
docs: stream-json tested live — it does not deliver observability, and v1 was harmful
Deployed the amd64 build to gw-04 and drove a real mission. The agent used Bash and the standard tools; no tool.call events appeared, and the gateway's unmatched-frame histogram still showed only session_start. The reason is structural: TurnEvent::ToolCall is emitted from tool_execution.rs, only for tools ZeroClaw itself runs. Claude Code runs its tools in its own subprocess, so the event never fires. A provider that knows about the calls changes nothing by itself. The first version was also harmful — it returned the observed calls as tool_calls, so the loop tried to execute Claude Code's tool names and fed "Unknown tool: Bash" back to the model. Fixed in the fork; both runtimes rolled back to the known-good image in the meantime. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
ac4fa0b8f7 |
docs: CI green and deployed — record the verified production state
Run 498 passed and deployed. Confirmed on gw-04: 53 skills, 11 templates, zero unresolved bindings, self-authoring announced ENABLED, the new gateway_preflight answering, and migration 0080 applied. Also records that run 497 was cancelled by the concurrency guard rather than failing, and that the stream-json runtime image is still NOT shipped by this pipeline. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
689a5e14a3 |
docs: CI root cause was an apostrophe, not any of the three theories
Records both real causes (run 490 stomped by an overlapping run; 491-496 killed by an apostrophe closing a single-quoted sh -c block), the guard that now catches the second class locally, and what to check when the in-flight run settles — including that a successful build is the FIRST time these commits reach production. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
d23f30e929 |
docs: record the CI investigation honestly, including what is still unknown
Establishes what is verified (the code passes on the runner host, with cargo's real exit code), what is narrowed (493/494 die inside the Rust step before cargo starts; 495 died before step 1), the three theories that were wrong, and the cheapest next experiment. Also records the two things that made this expensive: the actions-log API returns 403 for our token, and my first reproduction piped cargo into `tail` and reported tail's exit code. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
72ba4ba523 |
fix(ci): two runs stomped each other, and the logs blamed the tests
Runs 490 and 491 both failed `test`. Neither failure was in the code.
Runs 490 and 491 started 16 minutes apart and a full suite takes longer
than that, so they overlapped. The first thing a run does is
`docker rm -fv cm-ci-pg` — a name every run shared — so the newer run
deleted the older run's database mid-suite. Both failed, and the failures
read as test failures.
Verified before changing anything: the exact CI command, on gw-04, against
the same warm cargo volumes and a Postgres started exactly as CI starts it,
passes on
|
||
|
|
b653dbfe72 |
docs: hand-off note for the next session
Records the state of the tree, the one step not taken (the stream-json runtime image is built and never deployed, so no mission has confirmed tool.call rows end to end), the ordered next steps, the decisions that are the operator's, and what was deliberately left undone with reasons. Also records the two corrections made this session — "missions can't call tools" was wrong, and raw test counts are a bad coverage metric — because both were confidently stated here before being checked. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
ea0b989b3f |
docs(research): missions DO call tools — the claim was wrong, and the truth is worse
Deep research into "missions can't call tools at all", which I wrote and which is false. docs/TOOL-CALL-ARCHITECTURE.md has the full findings. WHAT IS ACTUALLY TRUE Three of the four mission paths end in `claude -p` with Claude Code's own toolset and permissions PRE-ACCEPTED: solo microVM Read Edit Write Bash Agent --permission-mode acceptEdits composed microVM same, per node same direct session Read Edit Write Bash acceptEdits So the position is not "no tools". It is: mission agents run Bash and Write with permissions pre-accepted, and nothing in this platform can gate them. That is a stronger finding than the one it replaces — "can't call tools" sounds like a missing feature; "calls tools freely, ungated, and mostly unobserved" is a security posture, and it is ours. Observe and gate are different and both are partial. vm_tool_tap is a PostToolUse hook: it fires AFTER the tool ran and exit-0s unconditionally, so it is telemetry and structurally cannot gate. The direct-session tier has no tap at all. GatePolicy has exactly one enforcement site — the chat loop — and its approvals key on (session_id, message_id), which no mission phase can produce. WHY THE CONTAINER TIER LOOKED TOOL-FREE `claude_cli` runs `claude -p --output-format json`, which returns a single final result object, and the provider hardcodes `tool_calls: Vec::new()`. The calls happen; the transport discards them. The comment reading that emptiness as "§15 by construction: agents are provisioned tool-free" was inferring a design property from a serialization choice. Verified against the deployed Claude Code 2.1.228 rather than assumed: `--output-format stream-json --verbose` emits `tool_use` blocks with the tool name and `tool_result` blocks. The calls are fully observable; we ask for the wrong format. THE DOOR WE ALREADY BUILT AND NEVER PLUGGED IN claude_cli.rs is OURS — upstream zeroclaw-labs/zeroclaw has no such file — and so is 88eef99d4 "claude_cli --mcp-config + allow/disallow tools (act via door)". The provider already accepts mcp_config (claude's own MCP client reaches our door), tools, and disallowed_tools (lock out the natives so the gated door is the ONLY actuator). agent.config.example.toml documents the whole shape. In the live runtime: clawmates-mcp.json does not exist, there is no [providers.*] block, and every mission claw binds to claude_cli.default which sets none of it. My earlier "claude_cli cannot reach MCP, therefore the skills server is unreachable" was wrong in its reasoning — the capability is built, documented by us, and never deployed. Related: we set `agents.<alias>.mcp_bundles`, which configures ZeroClaw's OWN MCP client for its native loop. A claude_cli agent's actuator is the claude subprocess, which reads `mcp_config` on the PROVIDER. We were turning a knob wired to a loop that does not run. UPSTREAM 218 commits behind. No upstream work on claude_cli (the file is ours). ACP already exists in the fork; the three new commits are workspace-default and localization fixes, not new capability. The one item worth pulling is "feat(plugins): add shared egress policy foundation (#9137)" — a network guard with DNS pinning and metadata-address blocking, defence for the egress problem we have not solved. Stale claims corrected in place, in topology_exec.rs and the runtime config, so the codebase stops asserting the thing that is false. Recommended order, cheapest first: stream-json for observability; the PreToolUse hook for a real gate (it FIRES under claude -p per vm_stop_gate, and has zero call sites); then deploy the door. The executor swap is NOT recommended — the blockers are structural, not wiring, and the cheap fixes deliver what it was wanted for. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
771092b165 |
fix(skills): a pinned skill contradicted the platform inside the same prompt
Extending the Skill-Use mechanical checks, per the baseline's own next step, found something bigger than a missing check. THE DEFECT `workspace-repo-commit-protocol` told agents that `/workspace/repo` was "the ONLY path where source-modifying edits belong". The platform mounts and advertises `/mission/repo` — 26 references in the code; `/workspace/repo` appears in none of them. The skill is bound on 29 role bindings and was delivered TWICE in the run already measured, so an agent received the real path in its tool preamble and a skill contradicting it a few hundred tokens later, in one prompt. An agent that obeyed the skill wrote source into a directory nothing collects — the phase then delivers nothing, and looks like an agent that did no work. The same skill instructed `file_read` / `file_write` / `shell`: ZeroClaw's names, the exact ones `phase_task_text` was fixed to stop advertising after five agents on a single mission spent 7.4k tokens describing the mismatch instead of working. The prompt was corrected and the skill kept saying it. Rewritten against what the code actually does, including the repo-less case (`/mission/repo` exists, is collected as artifacts, has nothing to push). THE CLASS, AND THE GUARD The skills were never checked against the platform they describe. Nothing compared them, so a skill could contradict the prompt it ships inside and stay that way indefinitely — the same shape as PLAN_COMPLETE being documented and never implemented. Two tests in `skills_loader::contradiction_tests` now hold it: no skill may name a repo path the platform does not mount, and none may instruct a tool the agent's subprocess does not expose. The second matches backticked instructions and skips corrective lines, so a skill may still WARN against the wrong names — as this one now does. Both negative-controlled by restoring the old wording. AND THE CHECK THAT STARTED IT `workspace-repo-commit-protocol` now has a Boundary check: writing outside `/mission/repo` fails, and the message names the consequence — a phase that delivers nothing — rather than just the wrong path. docs/SKILL-USE-BASELINE.md records this as the fourth defect the measurement found, and corrects the "next unit of work" note now that this one is done. Full workspace suite green: 106 binaries, zero build errors. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
113de610ec |
fix(security): a signed Slack request could be replayed forever
Phase 5. The headline is not the coverage work — it is what looking for coverage found. A CAPTURED SLACK REQUEST AUTHENTICATED INDEFINITELY `slack_signature_valid` verified the HMAC correctly, and nothing anywhere checked how old the timestamp was. The timestamp is an input to the basestring, so an old request's signature verifies exactly as well as a fresh one — meaning anyone holding a single captured signed request (a proxy log, a mirrored packet, a leaked webhook body) could replay it forever, and every replay would authenticate. Slack's documented 5-minute window is now enforced IN THE BROKER, not the caller: the broker does not trust its caller (§15), and a check the caller can forget to make is one that will eventually be forgotten. Symmetric, so a far-future timestamp cannot mint a request valid for as long as the attacker chooses. Seven unit tests over the pure function with the clock injected, and the HTTP-level test now asserts an hour-old but validly signed request is refused. Negative control: removing the window fails the stale and future cases specifically. The existing slack_inbound test used the literal timestamp "12345" — a 1970 date — which passed only because nothing checked freshness. That is the shape of the whole finding: the fixture could not have failed, so it never told us anything. COVERAGE, RE-EXAMINED The review ranked crates by raw test count. That metric was misleading and found the wrong crates: cm-safety's seven tests already cover the decide CAS, grant double-consume, expiry and the approved/rejected split, and the audit_log immutability trigger is tested over in cm-db. Reading the API surface against the tests found the real gaps — verify_slack_signature above, and `credits_for_tokens`, pure pricing arithmetic that every existing billing test went through the database to reach without ever checking directly. Now pinned: the round-up contract, the deliberate one-credit floor, and that an absurd token count cannot wrap into a negative charge (a refund granted by an overflow). Still genuinely thin: cm-brain, where 6 of 9 tests need live clawbrainhub.com. Stubbing it means reproducing an external registry protocol we have no spec for — its own piece of work, not a coverage chore. Recorded rather than faked. GATEWAY PREFLIGHT ZEROCLAW_GATEWAY_URL and ZEROCLAW_TOKEN have no defaults and are read at FIRST USE, so a deployment missing them boots clean, serves every page, and fails the first time someone presses run. Third sibling of runtime_preflight and validator_preflight, same stance: a report, not a gate. The message names the consequence — "container-tier missions cannot run" — rather than only the unset variable. One process note: `cargo test -p cm-secrets` passed while the LIBRARY build was broken, because `time` is a dev-dependency there and my reference to it only resolved under cfg(test). Switched to std. Checking `cargo build --workspace` as well as the test profile is the guard. Full workspace suite green: 106 binaries, zero build errors. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
5c2c63f8e8 |
feat(missions): a human can finally reach the plan/roster review gate
Phase 4 of the plan, plus the PLAN_COMPLETE decision and the gitea_forge
cleanup from Phase 5.
THE REVIEW UI
mission_plan and mission_roster have been complete and reachable by curl
since they shipped, with zero frontend. That matters more than a missing
screen usually would: the decide step is not a convenience, it IS the
safety mechanism. Approving a plan replaces the mission's phases; approving
a roster flips it to the composed engine. A gate nobody can reach is a gate
that is always open or always shut.
MissionProposalDrawer, modelled on LevelUpDrawer which already does
load → review → decide. Reached from a mission's SETUP tab. Verified end to
end against the live backend, not just compiled: a model proposed a roster,
approval flipped the mission to `composed`, and approval on a non-draft
mission was refused.
The plan view shows each phase's done_when, and says plainly when one is
absent — a phase without a completion condition is never judged and reports
completed whatever it did, so its absence is the thing worth seeing.
AND THE DEFECT BUILDING IT FOUND
Every refusal path computed a precise reason — "the mission is running, not
a draft", "no node can boot that backend any more" — logged it to stderr,
and returned a bare {"error":"bad request"}. The person who needed the
sentence was the one clicking Approve; they got two words, and the reason
went to a server log they cannot read.
ApiError::Refused(String) carries it now. Same argument ApiError::Unavailable
was added for ("a 500 with 'internal error' sent them looking for a bug that
was not there"), one status code down. Live: the 400 now reads "this mission
is completed — a roster can only be approved while it is a draft, because
approving one rewrites how the mission will run".
PLAN_COMPLETE, decided
The Skill-Use measurement found that int-xx-marker-protocol documents
PLAN_COMPLETE and task_card_parser never implemented it, so an agent
following the skill exactly was silently ignored. Implemented rather than
removed from the skill: the planner needs a way to say it is done
specifying, and agents already emit it.
Marker ids are now strictly INT-<digits>. `starts_with("INT-")` accepted the
range form `INT-01..02` — observed live — which parsed into an id matching
no real item, so a task card appeared for something that did not exist while
the two items it covered stayed open. Rejecting is right: an ignored marker
is visible, a plausible row is not.
GITEA_FORGE, REMOVED
Named in nine places, defined in none. Harmless while provision_claw ignored
the bundle list; once the list was honoured, an undefined name became a
capability an agent is told it has and does not. Removed from seven team
templates, a workflow recipe, the auto-provision path, and a dropdown a user
could pick it from.
A new test asserts every bundle a template names is defined in the runtime
config — and it immediately found `web_fetch` in two templates I had missed
removing by hand. Same shape as the skill-binding test, one layer up.
Agents reach the forge through git over HTTPS with the ambient GITEA_TOKEN,
which is why nothing ever broke.
Full workspace suite green (106 binaries); frontend builds clean.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
91a6b4e304 |
feat(skills): the first Skill-Use measurement, and the three defects it found
Scored on the paper's three axes against two real missions on the local
stack. docs/SKILL-USE-BASELINE.md has the numbers, the method, and the
limits.
Trigger is reported as NOT OBSERVABLE, never zero
The paper measures progressive disclosure: the agent sees a name and
description and must retrieve the body, and that retrieval is the Trigger
event. We inline full bodies, because mission claws run on claude_cli which
cannot surface a tool call — there is nothing to retrieve with. So the
agent never reaches for a skill, it simply holds one.
Scoring that zero would report a delivery-model property as an agent
failure, which is the same confusion that kept 55 empty bindings invisible
for months. The verdict type carries NotObservable(reason) as a distinct
case from Fail for exactly this.
Compliance is checked by running the REAL task_card_parser rather than a
copy of its rules — a second implementation would drift, and then the score
would pass while the mission loop still stalled. Skills without a
machine-checkable consequence score not_applicable rather than a guess.
WHAT THE MEASUREMENT FOUND
1. The prompt format made its own record unparseable. Skills were
introduced with `## <name>` and skill bodies are markdown full of `##`
headings, so run 1 scored "Sizing heuristic" and "The output shape" —
subheadings inside decompose-int-items — as skills with no catalogue
row. Now an unambiguous `--- SKILL: <name> ---` marker, with both
writers sharing one renderer so the reader cannot drift from the writer.
2. A prompt was recorded that was never sent. My own Phase 1 work recorded
the phase prompt at the dispatch fork, before the tier was chosen — and
the container tier does not send that text, it sends the bare task and
appends skills per turn. Every container mission logged a `solo` prompt
that reached no agent. A provenance record of something that did not
happen is worse than no record: it is the wrong answer, delivered
confidently. Recording now happens inside each tier, with a test that
every launcher records the prompt it actually sends.
3. int-xx-marker-protocol documents a marker the platform never
implemented. PLAN_COMPLETE is in the skill's ladder and task_card_parser
has no such kind and never has, so an agent following the skill exactly
emits a marker that is silently ignored. Observed live: run 2's planner
emitted `PLAN_COMPLETE: INT-01..02`, which is also the range form — on
the kinds that ARE parsed that yields the id `INT-01..02`, a task card
for an item that does not exist while the two real items stay open.
This is a skill/implementation mismatch, not an agent failure, and it is
exactly what the measurement exists to find: the agent did what it was
told and what it was told was wrong. Both shapes now score as failures.
The reconciliation — implement PLAN_COMPLETE or drop it from the skill —
is left as a decision rather than guessed at.
The boot log now shows what the plan asked for: 53 skills, 11 templates,
every one `N role skills bound` with NO unresolved clause. Live missions
confirm per-role delivery — the planner receives decompose-int-items, the
coder receives write-rust-current-edition.
GET /api/missions/{id}/skill-use exposes the scores, and says in its
payload whether an empty result means "nothing delivered" or "the evidence
was reaped" — those have very different causes and must not look the same.
n = 2. No spread is reported because two runs cannot establish one, and the
document says so rather than letting the number be quoted as a baseline it
is not.
Full workspace suite green: 106 binaries.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
769e002bb3 |
feat(skills): deliver on every tier, record what agents receive, let them self-author
Three phases of the approved plan, plus a correction to what the last one
claimed.
CORRECTION: skills reached ONE tier, not all of them
The previous commit said "skills can now reach a mission agent". That was
true only for the container/ZeroClaw tier — the fall-through that queues a
topology_runs row for topology_worker, which drives the executor that was
patched. compose_turn_prompt/pinned_skills_text had exactly one production
caller, and phase_runner's three other paths (composed microVM, solo
microVM, direct session) never called it. CAPABILITY-REVIEW.md said the
broad thing too; both are corrected.
Those three tiers share one task string and have no per-turn alias, so
their skills resolve per PHASE from the mission's crew and are appended
there. The container tier deliberately still injects per turn, with the
running node's own role — appending in both places would put every crew
member's skills in every turn twice.
The behavioural tests prove phase_skills_text and compose_turn_prompt work.
They cannot prove the three launch_* calls pass the composed string, and
that substitution is a one-word edit that would silently return all three
tiers to delivering nothing with every test still green. So there is also a
source-level assertion on the call sites, following the precedent in
mission_events::the_cap_is_enforced_in_one_statement. Its negative control
names the exact tier.
PROVENANCE: what an agent received, and what it said it did
Both were unanswerable. The prompt was never stored anywhere on any tier —
re-deriving it later re-runs the skill lookup against a catalogue that has
since changed, and once agents author their own skills it certainly will
have. The reasoning rows were durably write-only: pushed live once, then
never read from the database again by anything except the GC that deletes
them.
- prompt.composed records the exact bytes, on all four tiers
- the session tier writes its checkpoint record and a reasoning row,
instead of eprintln! and nothing — the same defect the solo microVM
path was fixed for, in the last tier that still had it
- narrative_for_mission reads both back
Found while doing it: the 400-event per-phase cap counted EVERY kind, so a
busy phase could push out its own phase.completed and its own provenance.
The cap now counts only the two unbounded kinds it was written for.
Negative control confirms the old behaviour dropped the prompt.
Retention is now a per-mission hold (0080) rather than a raised global —
with a test asserting unheld missions are still reaped, because an
exemption that applies to everything is not an exemption.
SELF-AUTHORING: agents apply their own skill drafts, no human click
By operator decision. level_up has generated complete drafts from a model
since it shipped; only a checkbox stood between propose and apply.
What replaces the gate is not another gate but four properties, each held
by a test:
- workspace-scoped, so a hand-authored skill can never be modified
- a draft cannot take a hand-authored skill's name. Ids are scoped and
bindings resolve by skill_id, so it could not overwrite or shadow one
anyway — but two procedures under one name means nobody reading a
transcript can tell which the agent followed, and that ambiguity is
fatal in a system where the skill is the standard being graded against
- every revision appends a skill_versions row, so it can be reverted and
a past run can be read against the text it was actually judged under
- approved_by = NULL. An agent's decision is never attributed to a person
who did not make it
Only skill_candidate applies autonomously. identity_refinement and
brain_consolidation still wait for a human: they change what an agent IS
rather than adding a procedure it can consult. State is announced at boot,
because a safety gate that changes silently is one nobody notices changed.
CLAWMATES_SKILL_SELF_AUTHORING=0 restores it.
Also: the test Postgres ran out of /dev/shm mid-suite (Docker's 64MB
default) and surfaced it during MIGRATIONS, which reads like a schema fault
and is not one. --shm-size=1g, and a pointer to the `clean` subcommand that
already existed for the 779 leaked test databases.
Full workspace suite green: 106 binaries, no failures.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
e3247fee4b |
chore(runtime): define the skills MCP bundle the templates now ask for
provision_claw honours the template's bundle list as of the previous commit, but a bundle an agent is assigned and the runtime config does not define resolves to nothing — so the assignment had to be made to mean something on the MCP side too. Carries the caveat that matters at the point of use: this channel only works for a provider that can surface tool calls, and mission claws run on claude_cli, which is text-only. Their skills arrive as prompt text instead. The entry is for tool-capable agents, and so that an assigned name resolves. gitea_forge is left UNDEFINED on purpose, with a note. Six templates name it and nothing defines it; a plausible-looking definition pointing at the wrong URL would turn a name that resolves to nothing into a server that fails at call time, which is harder to notice rather than easier. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
e4942ce985 |
fix(missions): skills can now reach a mission agent at all
Repairing the 55 broken skill bindings made the catalogue correct. This
makes it reachable, which it was not — for any skill, on any mission, since
the catalogue was built.
The skills had exactly ONE delivery channel: the `clawmates_skills` MCP
server. A mission claw could not reach it for three independent reasons:
1. `provision_claw` wrote the constant `["clawmates_door"]` and ignored
the template's mcp_bundles — which mission_orchestrator had already
resolved and stored on the team row.
2. The runtime config defines no `clawmates_skills` bundle. The live
local config defines no bundles at all, not even the door.
3. Mission claws run on `claude_cli`, which the runtime's own config
comments document as text-only: it cannot surface a tool call, so no
MCP server is reachable from a mission turn regardless of bundles.
And a mission turn's whole system context is two sentences synthesised from
the role slot in topology_exec::build_prompt. The template's role prose is
not used either — mission_orchestrator documents this, and it means the
role prompts describing which procedures to follow were never read.
Two doc comments in cm-runtime describe the mission path as already having
the summary-and-fetch contract. It never did. The belief was written down
twice and checked zero times, which is why nobody looked — and it is why
the Skill-Use measurement this review planned could only ever have returned
a trigger rate of zero. That would have read as a finding about the agents.
- provision_claw takes the bundles, with clawmates_door always added: a
template that forgets to list it must not get an ungated agent
- all 11 templates now request clawmates_skills; web_fetch removed, since
a list that is honoured must not name a bundle that does not exist
- the re-provision sweep re-asserts the team's own stored bundles rather
than a constant, which would have silently stripped a capability
mid-mission
- pinned skill BODIES are injected into the mission prompt, bounded and
with truncation stated. Bodies, not an index: there is no `skills.read`
tool on this path, so an index would advertise a capability that does
not exist — the exact failure this whole change is about
Three tests: the body reaches the prompt, an agent with no skills adds no
heading (an empty "Your skills" section announces skills the agent does not
have), and the composition is exercised separately from the lookup, because
`pinned_skills_text` working and `run_turn` calling it are different claims
and the second is the one that was false.
Also adds the three review documents: CAPABILITY-REVIEW (inventory, what
was repaired, what is deferred and why), PROVENANCE-ASSESSMENT (assess
only, per decision — what each store answers and the two candidate paths),
and RESEARCH-SWEEP (the fortnight's papers and what we did about each,
including the ones we deliberately did nothing about).
Full workspace suite green.
Co-Authored-By: Claude Opus 5 <[email protected]>
|
||
|
|
18dc0b964b |
fix(missions): the security scan phase now scans, and task upserts work
Four defects, found by checking the audit's claims instead of trusting them. Two of the audit's own findings turned out to be wrong, and the registry that exists to record which config keys are read was itself inaccurate — so the corrections are part of the change. upsert_task raised 42P10 on every call, for every caller `mission_tasks_external_uniq` is a PARTIAL unique index (WHERE external_id IS NOT NULL). Postgres will not match a partial index to an ON CONFLICT target unless the statement repeats the predicate, so the upsert failed on its first row. Both callers — the task-card parser that turns INT markers into tasks, and the security scanner — map the error to a string their caller logs. Two features were broken and nothing was red. Regression test in cm-db with a negative control: reverting the WHERE reproduces 42P10 exactly. the security scan never ran `security_scan::run` was reachable only from an operator button, so security_hardening.toml — a workflow whose entire first phase is a scan — ran an agent that was never told to scan and never fired the scanner either. phase_runner now sweeps finished security_scan phases, mirroring the benchmark baseline sweep that was added for the identical defect. Guarded on a new completion marker rather than on findings: a clean scan writes no findings, so a findings-guard would rescan forever. The marker also answers the question an operator actually asks, which is not "how many findings" but "was this looked at, by what, and when". two recipes could not fail security_hardening.toml and benchmark.toml carried no `task` and no `done_when` on any phase. A phase without done_when never enters evaluating, is never judged, and reports completed whatever it did — so a security mission could scan nothing and go green, and a benchmark mission could record no baseline that the next refactor would then compare against. Both now state the work and the condition, with inert keys annotated inline rather than deleted, so the gap between what a recipe asks for and what a phase receives stays visible. the config registry was wrong in both directions `harness` was listed NOT IMPLEMENTED while benchmark_runner reads it and phase_runner runs a baseline through it. `tools` was listed NOT IMPLEMENTED while security_scan::run reads it. A registry that exists so an operator can trust what a recipe does is worse than useless when it is inaccurate. Both corrected, `bench_name` and `cmd` added, and `test_command` deleted — it had neither a reader nor a writer, so it described a situation that could not arise. Also: CLAWMATES_JUDGE_MODEL had two different defaults (opus-4-8 in routes/topology.rs vs opus-5 in cm_runtime::judge_model) and a doc comment naming a third; topology now calls the one function. GITEA_TOKEN's absence in mission_plan is stated rather than degrading to the same "could not be read" string a private repo produces. BRAINHUB_API_KEY needed no change — hub::push already rejects an unset key with a named error. That half of the finding was overstated. Co-Authored-By: Claude Opus 5 <[email protected]> |
||
|
|
3511c3ca10 |
docs/clerk: social OAuth (Google/GitHub/Apple) dashboard setup steps
Co-Authored-By: Claude Opus 4.8 <[email protected]> |
||
|
|
bd982943b8 |
chore: commit ZeroClaw per-tenant runtime spike + architecture doc
Saves earlier-phase artifacts that were sitting untracked: - docs/agent-engine-architecture.md — per-tenant containerized ZeroClaw runtime decision doc (clawmates = §15 control plane; zeroclaw = per-tenant runtime). - deploy/clawmates-runtime/ — slim runtime Dockerfile, dev compose, example agent config, README (the proven Phase-1 drive recipe). - tools/runtime-spike/drive.mjs — Node WS drive client for the spike. No secrets (only env-var names / commented placeholders). Co-Authored-By: Claude Opus 4.8 <[email protected]> |
||
|
|
e93a3cfb53 |
feat(topology): executors for all 12 kinds (mesh + debate; mappings)
Every TopologyKind now runs, mapped to five execution patterns: - hierarchical ← hub_spoke, star_moe, market - pipeline ← ring - swarm ← flat, holacratic - mesh (new) ← blackboard (two peer-exchange rounds + aggregate) - debate (new) (propose → critique → revise → judge) execute()'s match is now exhaustive (adding a kind upstream forces an executor), so the Unsupported error is gone. Benchmark spans all five distinct patterns. 14 tests with --features provider; clippy clean. Doc updated. Co-Authored-By: Claude Opus 4.8 <[email protected]> |
||
|
|
817d8c712c |
feat(topology): cm-topology crate + architecture doc (Phases 0–1)
Foundation for the dynamic agentic-topologies platform (see docs/topology-platform.md), porting agentorg's topology modeling into pure Rust: - TopologyKind: curated 12-kind taxonomy (hierarchical, flat, pipeline, swarm, mesh, hub_spoke, ring, star_moe, market, blackboard, debate, holacratic). - TopologyGraph: role-slot nodes + typed edges, with validation. - adapter: normalize a loose JSON spec → validated graph (fills edge kinds). - classifier: structural metrics (density, hub dominance, clustering, diameter, hierarchy score) → inferred kind + confidence (tree→hierarchical, line→pipeline, cycle→ring, star→hub_spoke, complete→mesh, empty→flat). - heuristics: per-kind role distributions (ported from topology_manager.py). Pure, offline, dependency-light (serde/thiserror). 17 unit tests, clippy clean. Co-Authored-By: Claude Opus 4.8 <[email protected]> |
||
|
|
48730abd78 |
docs: design specs + reference assets; move roadmap into docs/
Add the WorkClaw design-reference docs (specs, manifests, tokens, motion, icon/asset references, gifs) under docs/, and move the platform spec/PRD/roadmap from the repo root into docs/. Co-Authored-By: Claude Opus 4.8 <[email protected]> |
||
|
|
b9fdec9173 |
Clerk deployment smoke: validated against a real instance, both halves
Last open item from the roadmap + post-1.0 list. Run against the live Clerk instance closing-seasnail-39.clerk.accounts.dev. - Backend (crates/cm-auth/tests/live_clerk.rs, CM_LIVE_CLERK=1): pulls REAL discovery + JWKS from the live instance, mints a REAL session JWT via Clerk's Backend API (create user -> open session -> session token), and runs it through AuthService::authenticate — verify + JIT provision (keyed on the real sub), duplicate-subject suppression, tamper rejection against the live JWKS. Decodes the instance domain from the publishable key; cleans up the test user after. PASSING - Frontend: built with AUTH_MODE=clerk + real keys, next start serves Clerk's <SignIn /> at /login wired to the instance (instance domain + data-clerk attributes present in the HTML). Both halves confirmed end to end against production Clerk - docs/clerk.md: documented the smoke procedure for both halves 166 Rust tests (+6 live, key-gated). Keys used via env only, never stored — rotate them (they passed through chat). Co-Authored-By: Claude Fable 5 <[email protected]> |
||
|
|
add4f79fed |
Rebrand: TeamClaw -> Clawmates (clawmates.work)
Full-depth rename per the approved plan; the 'claw' product vocabulary (claws, /claws routes, clawId, Claw Chat) stays — it is now the brand. - Display brand: Clawmates (manifest, titles, hero, login/rail logo 'clawmates'); default host app.clawmates.work; registry ghcr.io/clawmates - Crates tc-* -> cm-* (16 crates + all imports); binaries clawmates-server/broker/bundler; images clawmates/*; env prefix CLAWMATES_* (+ CM_TEST_DATABASE_URL / CM_LIVE_LLM); config clawmates.toml; helm chart deploy/helm/clawmates with clawmates-* resources; db names clawmates*; sockets /run/clawmates; cookie cm_session; kind cluster clawmates-test; seccomp node profile clawmates-agent-profile.json - All 9 Playwright brand assertions updated in lockstep; historical spec document left untouched as the only remaining 'TeamClaw' - Local env migrated: dev pg clawmates-dev-pg/clawmates_dev, shared test server clawmates-test-pg, kind cluster recreated with image + profile, compose images rebuilt under clawmates/* Verified end to end: 161 Rust + 68 frontend tests, 29 Playwright journeys, 4 live kind tests, helm/install/LOC/placeholder gates, and the clean-room install rehearsal serving the clawmates login page from a signed bundle of the rebuilt images. Co-Authored-By: Claude Fable 5 <[email protected]> |
||
|
|
ceca21ca79 |
Clerk frontend integration: one image, runtime-switched identity
- src/lib/auth/bearer.ts is the single identity dispatch for both server-side token consumers (RSC apiFetch and the /api proxy route): local -> httpOnly tc_session cookie; clerk -> Clerk getToken() session JWT. The Clerk SDK is imported lazily, so the air-gapped/local path never loads it - Runtime env (AUTH_MODE / CLERK_PUBLISHABLE_KEY / CLERK_SECRET_KEY), deliberately NOT build-time NEXT_PUBLIC_*: the same standalone image serves both deployment targets - Conditional <ClerkProvider> in the root layout (publishableKey passed at render from runtime env); /login renders Clerk's <SignIn /> in clerk mode and the local form otherwise; proxy.ts middleware delegates to clerkMiddleware() only when active - Helm: frontend deployment injects the Clerk keys from a Secret when auth.mode=clerk - mode.ts unit-tested (default local, exact-match clerk, loud failure without the publishable key); the local path stays proven by all 29 journeys; the Clerk branch is thin delegation to the SDK, exercised in deployment smoke per docs/clerk.md 157 Rust + 68 frontend tests + 29 Playwright journeys. Co-Authored-By: Claude Fable 5 <[email protected]> |
||
|
|
cbc8d35a2e |
Clerk authentication: hosted-identity session JWTs as a first-class mode
- tc-auth JwtVerifier: OIDC discovery -> JWKS, RS256 with the issuer pinned, 5s leeway (the crate's default 60s would double the life of Clerk's 60s session tokens), key cache with one refresh on unknown kid (Clerk rotates). Serves auth.mode = clerk AND generic oidc — a Clerk instance IS an OIDC issuer, so one verifier covers both - AuthService.authenticate dispatches: JWT-shaped bearers take the hosted-identity path, everything else stays a local opaque session. External users JIT-provision keyed by the stable sub claim (users.auth_subject, unique partial index in migration 0007); an existing local account with the same email is LINKED, not duplicated; role tracks the issuer claim every request (org:admin -> Owner) - Config auth.mode = "clerk" (requires issuer_url; validated), server pins the issuer at boot, Helm values/configmap accept mode=clerk - Tests with REAL crypto, no mocks: fresh RSA keypairs, a live local issuer publishing real discovery + JWKS docs, Clerk-shaped tokens — JIT + role mapping, repeat-subject no-dup, expired refused (leeway regression), wrong-key forgery refused, foreign issuer refused, and the full router round trip with Authorization: Bearer <session JWT> - docs/clerk.md: dashboard session-token customization (email + org role claims), config, @clerk/nextjs getToken() wiring, what CI proves 157 Rust + 63 frontend tests + 29 journeys. Air-gapped installs keep local auth — Clerk is a cloud-only alternative, not a replacement. Co-Authored-By: Claude Fable 5 <[email protected]> |
||
|
|
0afb359183 |
P0: workspace scaffold, CI gates, tc-domain, tc-config, tc-db vs real Postgres
- Cargo workspace with 1250-line and no-placeholder CI gates wired first - tc-domain: id newtypes, SessionKey codec (proptest round-trip), Role, GatedCategory (spec §15), AccessPolicy, core entities - tc-config: figment TOML+env config, DeployTarget/provider/auth selection with semantic validation - migrations/0001: full spec §14 schema incl. DB-enforced append-only audit_log - tc-db: compile-time-checked sqlx repos (workspaces, users, agents+policies, credits, audit) with committed .sqlx offline metadata - tc-testkit: per-test real-Postgres databases (testcontainers or TC_TEST_DATABASE_URL), embedded migrations Co-Authored-By: Claude Fable 5 <[email protected]> |