248948cc847b4d229291fa65785d940b02fd36ca
312
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d8c8793c4a |
ci fixes: cargo fmt, eslint entities, max-lines split
CI on
|
||
|
|
6ffbe978b2 |
missions: extract MissionLivePane to stay under 1500-LoC CI budget
Previous push hit the gates job's file-size gate — MissionCanvas.tsx was 1517 lines (17 over). The Live Pane xterm subcomponent is cleanly separable (no shared state with the parent, just takes nodeId + visible props), so it lifts into its own file at zero behavior cost. frontend/src/components/dashboard/MissionLivePane.tsx (new, 106 LoC) frontend/src/components/dashboard/MissionCanvas.tsx (1517 → 1424) Also drops the xterm.css + useResilientTerminal imports from MissionCanvas since only MissionLivePane needs them now. |
||
|
|
d4efb0dba2 |
missions sidebar: wrench toggle + multi-select bulk delete
Matches the AGENT tier's manage-mode pattern. Three states on the
sidebar toolbar:
- default: [Wrench] [Refresh] [+]
- select mode active: [Wrench] (highlighted) [Refresh] [+] + rows
grow a checkbox on the left and click-toggles selection instead
of opening the mission
- selection made: [Wrench] [Trash · N] [Refresh] [+] where the
trash pill shows the count and fires window.confirm → bulk
DELETE /api/missions/{id} loop
On success: exits select mode, calls onDeleted with the ids, parent
Dashboard clears missionsSel if it was in the batch. Failures stay
selected + a red error line surfaces the count.
Reuses the CRUD backend added earlier (a1c… PATCH/DELETE) — no new
server work.
|
||
|
|
bf4af48c80 |
herdr phase 3: INFRA tier Herdr sessions surface
New INFRA category "Herdr sessions" (purple sparkles icon between
Fleet and Local hardware). Shows a card per online fleet node with:
- Node name + hostname + IP
- Per-workspace agent state pills (working / blocked / done /
idle / unknown), colored dots + pane count
- "Open" button → renders that node's full Herdr TUI inline via
xterm.js (same nodeHerdrConnector + WebRTC-with-fallback the
MissionCanvas Live Pane uses)
Backend:
- node daemon: herdr_workspaces + herdr_snapshot ops
(`herdr workspace list`, `herdr api snapshot`)
- fleet_herdr::snapshot helper on top of hub.call_timeout
- GET /api/nodes/{id}/herdr/session route
Fetch flow: /api/nodes filtered to status='online' → for each,
/api/nodes/{id}/herdr/session in parallel. Snapshot errors surface
per-card without failing the whole grid.
The "Open" xterm is separate from the MissionCanvas Live Pane —
this one is scoped to the whole node's Herdr TUI (any workspace),
not a specific mission's pane. Operator toggles between nodes via
the buttons.
Verified: cargo check --workspace + tsc --noEmit both green.
|
||
|
|
a5588b0289 |
herdr phase 2: Live Pane tab (xterm.js → node's herdr TUI)
The killer UX feature: click a mission's Live Pane tab and watch the
actual Herdr TUI on the target node in the browser — cursor, colors,
tool output, all live. WebRTC DataChannel direct where the browser
can reach the node peer-to-peer, WS-relayed fallback otherwise
(same auto-negotiation the INFRA node terminal already uses).
Zero new deployment infra — reuses the existing terminal_ticket +
terminal_ws + PTY-over-control-channel machinery. The one primitive
we grew: PtyTarget::Command variant so the node can spawn an
arbitrary program (\`herdr\`) in the PTY instead of the login shell.
Node daemon (clawmates-node):
- PtyTarget grows a Command { argv } variant
- spawn_command_pty resolves bare names against user + system bin
dirs (matches how tool_update finds claude/kimi)
- PtyTarget::from_frame reads the `command` array from the pty_open
frame; precedence Command > Container > Host
cm-api:
- NodeHub::open_pty grows an optional command argv; when set, the
frame carries it and the daemon spawns the program directly.
- routes::nodes::TermCtrl gains a `command: Vec<String>`; the
fallback branch threads it through.
Frontend:
- core.ts::webrtcConnector takes an optional commandOverride
that ships inside the fallback frame
- nodeHerdrConnector(nodeId) — mints the standard ticket + WS URL
but overrides command to ["herdr"]
- MissionCanvas grows a "pane" tab, visible only when
runtime_kind='local_herdr'. LivePane subcomponent uses xterm.js
(already a workspace dep) via useResilientTerminal, shows a
connecting/relayed/direct pill in the corner.
To watch a mission live: pick "On a fleet node (Herdr)" + target
node in the wizard, launch, click Pane tab → node's Herdr TUI
appears. Navigate to the mission workspace in the Herdr sidebar
(mouse or prefix+w) to zoom into the mission's pane.
Focus-a-specific-pane-directly is a later enhancement — Herdr has
no CLI arg for it yet, so operator navigates the sidebar for now.
Verified: cargo check --workspace + tsc --noEmit both green.
|
||
|
|
2b3ec27757 |
herdr phase 1c: wizard runtime picker + on_launch auto-dispatch
Closes the operator loop for the second-runtime path. Missions
created with runtime='local_herdr' now spawn a Herdr pane on their
target_node automatically on draft→running.
Frontend (MissionWizard):
- Step 4 grows a "Runtime" section above Schedule
- Radio: "Hosted (ZeroClaw)" default | "On a fleet node (Herdr)"
- Local-Herdr shows a dropdown of ONLINE nodes only (from
/api/nodes filtered by status='online')
- canNext blocks Next when local_herdr picked without a node
- Review step shows "Runtime: Herdr on <node-name>" or "Hosted"
- Empty-online-nodes state hints "Connect one from INFRA first"
Backend:
- mission_orchestrator::on_launch grows a NodeHub param; when
mission.runtime_kind='local_herdr' + target_node_id set +
hub present → calls fleet_herdr::dispatch(). Non-fatal:
logs and continues so a research_only mission with a Herdr
runtime chosen accidentally still boots the team.
- routes::missions::set_status passes state.node_hub through.
- Test call sites updated to pass None for the new param
(integration tests don't drive real fleet nodes).
CLI stub: on_launch currently hard-codes cli="claude" for the
Herdr pane. Phase 4 will read that from the team template so a
research team → kimi, gpu team → claude, etc.
Verified: cargo check --workspace + cargo test
-p cm-api --test mission_orchestrator + tsc --noEmit all green.
|
||
|
|
d6dbd044c8 |
herdr phase 1a: missions runtime_kind + target_node schema
First slice of the second-runtime path. Missions now carry
runtime_kind ('zeroclaw' | 'local_herdr') + target_node_id (FK to
nodes) so the mission_orchestrator + phase executors can dispatch
differently depending on where the operator wants execution.
Migration:
- 0055_missions_runtime_kind.sql — adds runtime_kind (NOT NULL
DEFAULT 'zeroclaw' + CHECK), target_node_id (nullable FK ON
DELETE SET NULL). All existing missions backfill to 'zeroclaw'
so behavior is unchanged.
- topology_runs also grows herdr_workspace_id / herdr_tab_id /
herdr_pane_id text columns so a resumed run can reattach to the
same Herdr pane instead of spawning a duplicate.
Code:
- cm-db::repo::missions — Mission + NewMission carry the two new
fields; all SELECTs updated; INSERT COALESCE-defaults
runtime_kind to 'zeroclaw' when unspecified.
- routes::missions::create — validates runtime_kind and requires
target_node_id when kind='local_herdr' (400 otherwise).
- lib/api/missions.ts — RuntimeKind type; Mission carries both;
CreateMissionRequest optional fields.
Behavior is opt-in: no path exists yet to actually create a
local_herdr mission — that lands in Phase 1c (wizard picker). This
commit just makes the schema + validation in place so Phase 1b's
fleet_herdr dispatch module can key on it.
Tests: mission_orchestrator integration test still green.
|
||
|
|
1f0117e35a |
mission canvas: add / edit / delete toolbar controls
Top-right toolbar grows three CRUD controls per your request:
- Plus (always visible) — opens MissionWizard, selects the new
mission on create
- Pencil (draft-only) — opens EditMissionModal for title +
description; PATCHes /api/missions/{id}
- Trash (always visible) — window.confirm then DELETEs; sidebar
selection clears via new onDeleted callback
Backend:
- cm-db::repo::missions::update_meta(id, ws, title?, description?)
— COALESCE-based partial patch
- cm-db::repo::missions::delete(id, ws) — hard delete, cascades
via FKs on phases/tasks/artifacts/benchmark_snapshots
- PATCH /api/missions/{id} (draft-only) + DELETE /api/missions/{id}
Frontend:
- lib/api/missions — updateMission + deleteMission clients
- MissionCanvas — three toolbar buttons, EditMissionModal
(title + textarea for description), local wizard state
- Dashboard — passes onSelect + onDeleted so sidebar reacts to
create + delete without stale selection
Edit is draft-only (backend enforces + button hidden past draft) so
in-flight missions can't have their brief mutated out from under
running agents. Delete is unconditional — operator responsibility to
Cancel first if a run is live.
|
||
|
|
5b7cb55d21 |
mount LevelUpInbox + document Gemini env vars
1. Dashboard.tsx — mount LevelUpInbox as a bottom panel on the MISSIONS tier sidebar (below MissionsList, above the tier rail). Capped at 38% height so it never crowds the mission list; scrolls independently when the proposal count grows. 2. deploy/compose/.env.example — document GEMINI_API_KEY + CLAWMATES_REFINER_MODEL + CLAWMATES_LEVEL_UP_MODEL. Refine + Level- Up silently 500 without the key; the model overrides default to gemini-2.5-flash so setting the key is the only required step. Closes the two loose ends I called out last turn (inbox not mounted, env vars undocumented). The full flow — Refine → mission launch → security scan / benchmark → level-up review — is now wireable end-to-end on a fresh deploy just by copying .env.example → .env and setting POSTGRES_PASSWORD + GEMINI_API_KEY. |
||
|
|
ad1cee0b08 |
refine polish: before/after diff view + accept/cancel/restore
Refine no longer clobbers the mission description on click. Flow:
1. Click Refine → server generates the rewrite, returns
{ original, refined } WITHOUT persisting
2. RefineDiffModal shows a side-by-side pane (raw before,
Markdown-rendered after)
3. User picks:
- Accept → PATCH /api/missions/{id}/description commits refined
- Cancel → discards the proposal, description unchanged
- Restore original → forces a write of `original` (undo path
for accidentally-accepted refines, since Accept+Cancel is
still a two-step confirmation)
Backend:
- mission_refiner::refine returns a RefineResult { original, refined }
struct instead of persisting + returning the text
- routes::missions::refine now returns { original, refined }
- routes::missions::set_description added on PATCH
/api/missions/{id}/description (draft-only)
Frontend:
- lib/api/missions — refineMission return type is now RefineResult;
added setMissionDescription
- MissionCanvas — RefineDiffModal + DiffPane subcomponents;
accept / cancel / restore handlers wired to state
Closes task #20.
|
||
|
|
278cbf90b7 | artifacts tab: inline PDF preview via iframe (task #19) | ||
|
|
267b28a762 |
mission canvas: security-scan + benchmark trigger buttons
Fills the frontend gap after slices 7 + 8 shipped the backend runners
without triggers. Each phase card in the Phases tab now grows an
action row when the mission is running or completed:
- security_scan phase → "Run scan" button (POSTs /security-scan)
- benchmark phase → "Baseline" + "After" buttons, iteration auto-
derived from the current benchmark_snapshots count
Findings from security_scan already surface in the Tasks tab via
the task-card parser (external_id = cargo_audit:RUSTSEC-... etc);
the tool prefix in the badge is enough to distinguish sources
without a separate grouped view.
Closes tasks #17 + #18.
|
||
|
|
a3d5a5a96d |
level-up UI: inbox + review drawer + per-claw/per-team triggers
Fills the frontend gap left after slice 8.5 shipped the level-up
backend without any UI. Reviewers can now:
- See pending proposals across the workspace (LevelUpInbox)
- Trigger a proposal from any claw's ClawCommandCenter header
- Trigger a team-scoped proposal from TeamObserver's header
- Review one proposal item-by-item and apply the approved subset
(or reject all) via LevelUpDrawer
The drawer preselects auto-applicable kinds (identity_refinement,
skill_add, skill_candidate, brain_consolidation) and disables the
manual-only kinds (roster_change, mcp_bundle_change) with an
inline "manual — needs team wizard" hint, matching what the
backend applier does per commit
|
||
|
|
56201a6985 |
mission canvas: add Refine button + markdown-rendered description
Adds a Refine button to the left of Refresh + Launch on the mission
detail toolbar (draft-only). Clicking it POSTs to a new endpoint that
calls Gemini 2.5 Flash to rewrite the user's freeform description into
a coherent, sectioned Markdown brief (Objective / Context / Scope /
Constraints / Acceptance Criteria / Open Questions) ready for the
research + coding agents to ingest cleanly.
Backend:
- crates/cm-api/src/mission_refiner.rs — Gemini call with a
system prompt that preserves user-provided facts, avoids
invention, and emits raw markdown (not JSON).
- POST /api/missions/{id}/refine — draft-only, 400 on empty
description or non-draft state.
- cm-db::repo::missions::set_description helper.
Frontend:
- MarkdownBlock — tiny zero-dep renderer for h1/h2/h3, bullet +
numbered lists, **bold**, `code`, paragraphs. Deliberately
small; the refiner emits a bounded subset.
- MissionCanvas — Refine button (Sparkles icon, secondary style)
to the left of Refresh; description now renders through
MarkdownBlock instead of a single <p>. Disabled while
description is empty or a refine is in flight.
- lib/api/missions — refineMission client.
|
||
|
|
4663348a0e |
slice 9: big-bang cutover — delete legacy research + loops UI
Missions has been the unified surface for a while (Slices 1-8.5).
This slice makes it the ONLY surface by deleting the legacy
research + loops UI:
Removed frontend files:
- dashboard/ResearchList.tsx
- dashboard/ResearchCanvas.tsx
- dashboard/ResearchWizard.tsx
- dashboard/LoopsList.tsx
- dashboard/LoopsCanvas.tsx
- dashboard/LoopsWizard.tsx
- dashboard/LoopStaffingStep.tsx
- dashboard/NoAgentsGate.tsx
- dashboard/ResearchArtifactPicker.tsx
- dashboard/AutoProvisionCard.tsx
- dashboard/LiveRunLogs.tsx
- lib/api/research.ts
- lib/api/loops.ts
Dashboard.tsx cleanup:
- Tier enum collapses to world | missions | claw | repos | infra
- Dropped researchSel / researchRefresh / loopsSel / loopsRefresh
state slots
- Dropped the Research + Loops crumbs from the tab bar
- Dropped the two rail icons; missions rail icon retitled
- TIER_TABS lists missions in the second slot (was research)
- Combined isInfra || isMissions || isRepos guard replaces the
six-way branch
What INTENTIONALLY stayed (soft cutover on backend):
- /api/research/* + /api/loops/* routes still respond — no UI
hits them but external integrations (webhooks, prior curl
scripts) don't hard-break on this deploy
- research_topics / loops / research_outcomes / research_topic_agents
/ research_publish_approvals tables remain — the missions
backfill from 0047 references these rows via config.legacy_*
keys, and dropping the tables now would cascade FK deletes
into topology_runs.research_topic_id / .loop_id nullouts
- cm_db::repo::{research_topics, research_outcomes, loops} +
cm_api::routes::{research, research_setup, research_pipeline,
loops} + cm_runtime::loops kept — they're internal-only now,
scheduled for a follow-up cleanup PR
Follow-up PR (dedicated cleanup):
- Drop the 5 legacy tables + null out topology_runs FKs
- Delete the ~4k LoC of backend routes/repos + their tests
- Remove research/loops crumbs from URL history
Frontend TS check: clean (16 pre-existing unused warnings in
Dashboard.tsx are unrelated).
Co-Authored-By: Claude Opus 4.7 <[email protected]>
|
||
|
|
58963d5083 |
slice 8: security scan runner + trigger endpoint
Runs the security template's tool set (cargo-audit / gitleaks /
trivy fs / semgrep) inside the mission's team container and
materializes each finding as a mission_task keyed on the tool's
canonical id. The subsequent coding phase picks up the tasks and
applies remediations; the committer closes them by emitting
`COMPLETED: <external_id>` (Slice 5's task-card parser handles it).
Rust surface:
- cm_api::security_scan::run(mission_id, phase_id)
- Per-tool runners with JSON output parsing:
cargo audit --json → vulnerabilities[].advisory.id
gitleaks detect --report-format=json → [{ Fingerprint, RuleID }]
trivy fs --format=json → Results[].Vulnerabilities[].VulnerabilityID
semgrep --config=auto --json → results[] w/ rule+path+line fingerprint
- Tool errors surface as a `warning` task instead of failing the
scan — operator sees which need installing/fixing without a
silent no-op.
- Findings map to mission_tasks with external_id = "<tool>:<id>"
(e.g. cargo_audit:RUSTSEC-2024-0001, gitleaks:<sha>,
trivy_fs:CVE-2024-1234, semgrep:<rule>@<file>:<line>).
API:
- POST /api/missions/{id}/security-scan { phase_id }
→ { findings, tasks[] } — full task list after upsert so the
canvas can render immediately.
Frontend:
- triggerSecurityScan helper in lib/api/missions.ts. Findings
show up in the existing Tasks tab (Slice 5's UPSERT path).
Container requirements (opt-in):
- Runs inside the team container via `docker exec`, so the tools
must be present in that image. Missing = warning task, not fail.
- `-w /workspace/repo` so scanners see the mounted repo. Reads
teams.zeroclaw_container (populated on first phase run).
Follow-ups:
- MCP bundle wrapping the same tools as agent-callable functions
(currently agents scan by shelling out to `cargo audit` etc.
directly; a typed MCP wrap lands with clean audit trail).
- Auto-fire from the security_hardening workflow template on
phase transition (currently manual via API trigger).
- Bundle the four tools into the runtime image (or a dedicated
security-tools image) so operators don't have to install them.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
|
||
|
|
f843c9ddb1 |
slice 7: before/after benchmark runner
Executes a benchmark harness inside the mission's team container
and records the resulting metrics as a benchmark_snapshots row keyed
on (phase_id, iteration). Baseline pass (iteration=0) captures
before_metrics; each post-iteration call captures after_metrics +
computes delta vs baseline.
Rust surface:
- cm_db::repo::missions::upsert_benchmark_snapshot / benchmark_snapshots_for
- cm_api::benchmark_runner::{baseline, after_iteration, run}
- Harness enum: Auto | Criterion | CargoBench | VitestBench |
PytestBench | Shell (each with a command() vector)
- Auto detection peeks at the repo layout inside the container
(Cargo.toml → CargoBench, package.json → VitestBench, pyproject
→ PytestBench). Falls back to a Shell echo when nothing
identifiable.
- Bencher-format line parser extracts (name, ns_per_iter,
plusminus) so criterion + `cargo bench` output become structured
samples the canvas can diff.
- compute_delta pairs samples by name, emits {before_ns, after_ns,
delta_pct, direction: improved|regressed}.
API:
- POST /api/missions/{id}/benchmark { phase_id, slot, iteration? }
triggers baseline or after run and returns the mission's full
snapshot list.
- GET /api/missions/{id} now includes `benchmarks[]` in the detail
payload.
Frontend:
- New Benchmarks tab on MissionCanvas with iteration + driver
header, plus a 4-column grid (bench / before / after / Δ%) when
delta samples are present. Improved deltas render green,
regressions red.
- TS types + triggerBenchmark() helper in lib/api/missions.ts.
Wiring notes:
- team_container_for_mission reads teams.zeroclaw_container — that's
populated by topology_worker::try_team_gateway_url on first run,
so trigger baseline AFTER the mission's first phase spawns the
container.
- Not auto-fired yet by phase execution; that's the "template phase
executor" work that spans Slices 4-8. Manual API trigger works
today; automated hook is a follow-up.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
|
||
|
|
9ba5c06a1a |
slice 3: 6 team templates seeded from TOML recipes
Team templates are the canonical rosters + tool bundles that mint
concrete teams for a mission. Every builtin ships as a TOML recipe
under templates/teams/*.toml, loaded into the DB at server boot.
Migration 0048 adds:
- team_templates (id, key, name, stack, default_topology,
risk_profile, mcp_bundles, version, source,
workspace_id)
- template_roles (m2m: template_id + slot; system_prompt,
skills[], brain_seed)
- teams gets template_id + template_version for level-up lineage
Ships 6 builtins:
- rust_sdlc — planner/coder/tester/reviewer/committer for Rust
- backend — api_designer/db_engineer/coder/tester/committer
(Postgres, DuckDB, graph DBs, wire protocols)
- frontend — designer/coder/tester/committer (React + Tailwind + ShadCN)
- mobile — designer/coder/tester/committer (Expo, RN, iOS, Android)
- gpu — arch_analyst/kernel_author/bench_engineer/coder/committer
(CUDA, Metal, ROCm from Rust)
- threejs — scene_designer/coder/shader_author/perf_engineer/
committer (three.js, WebGL, WebGPU)
Each role has a versioned system_prompt + skill list + brain_seed
markdown. Skills column is a name array today; Slice 3.5a promotes it
to a typed m2m join with the real skills catalog.
Server boot:
- team_template_loader::load_builtins reads TOML from
/etc/clawmates/templates/teams (container) or templates/teams (dev),
upserts idempotently. Deterministic uuid per template key (sha256
of a fixed namespace + key) so ids are stable across boots.
- Dockerfile copies templates/ to /etc/clawmates/templates.
Read API:
- GET /api/team-templates — list all
- GET /api/team-templates/{id} — detail with roles
Wizard:
- Step 3 rewired from a raw team_id text field to a template picker
with "LLM auto-provision" as the default option + one card per
builtin, showing stack, topology, risk profile, and description.
- Mission create now passes team_template_id (not team_id) so phase
execution knows which template to mint from.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
|
||
|
|
fc67936e33 |
slice 2: MissionWizard + MissionCanvas + MissionsList frontend
Adds the unified missions tier to the dashboard. Runs in parallel
with Research + Loops tabs until Slice 9's big-bang cutover.
New files:
- lib/api/missions.ts — typed API client + 5 template presets
(research_only, research_and_code,
security_hardening, refactor, benchmark)
- dashboard/MissionsList.tsx — sidebar list with status pills +
template-kind badges + new-mission CTA
- dashboard/MissionWizard.tsx — 5-step adaptive wizard:
1) template picker (5 cards)
2) title/description + repo (when required)
3) team (placeholder — Slice 3 wires templates + auto-provision)
4) schedule (one-shot or cron)
5) review + launch
- dashboard/MissionCanvas.tsx — 4-tab detail view (overview/phases/
tasks/artifacts), draft→running launch
Dashboard.tsx gets a new "missions" tier + crumb + rail icon (flag).
Slice 2 ships a minimal implementation that creates missions with the
template's canned phase composition. Slice 4 replaces the preset table
with real TOML-recipe dispatch on the server; the client's fallback
presets keep offline preview working.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
|
||
|
|
82b5cf4385 |
research canvas: refresh claws + make Description collapsible
Two UX cleanups on the research canvas: 1. Agent cards were showing "(missing agent)" for every slot after an in-wizard auto-provision because the Dashboard's `claws` prop is SSR- rendered and doesn't include agents created during the client session. Refetch /api/team/claws on canvas mount + refreshKey change, merge with the prop (fresh wins) so newly-provisioned agents resolve correctly without a page reload. 2. The Description block was always fully expanded. Wrap it in <details open> so the user can collapse it once they've read it, matching the existing "Original prompt" pattern. Co-Authored-By: Claude Opus 4.7 <[email protected]> |
||
|
|
efc13f2904 | trim: drop stale roster-empty comment (LoopsWizard) | ||
|
|
c9dc96c44b | fix: use TopologyKind alias to keep LoopsWizard under 1250 lines | ||
|
|
c3e8f3ee01 |
fix: match AutoProvisionedAgent shape + valid OutcomeKind for loops
Co-Authored-By: Claude Opus 4.7 <[email protected]> |
||
|
|
279d785f5b |
loops wizard: extract AutoProvisionCard to stay under 1250 lines
Co-Authored-By: Claude Opus 4.7 <[email protected]> |
||
|
|
6f1515b1b0 |
team runs modal: expose runtime posture (risk_profile + mcp_bundles)
PATCH /api/teams/{id}/runtime-config existed since the team-wizard work
but had no UI — you needed curl to swap a team's risk_profile or add
gitea_forge to its bundles. Add compact pickers to TeamRunsModal above
"Run now": a Risk profile <select> and a Bundle checkbox list.
Fetches current settings from GET /api/teams/{id} on open and PATCHes
inline on every change. No save button — the debounce is the user's
next click. Optimistic state; next-open resync corrects any drift.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
|
||
|
|
9cab32c354 |
loops wizard: auto-provision team in-wizard (parity with research)
The LoopsWizard hard-blocked on an empty roster via NoAgentsGate — new users had to bounce to the Agents page first. Match ResearchWizard's flow: remove the top-level gate, add an auto-provision panel to the staffing step (6) that derives 3-5 roles from the loop's task via Claude Sonnet 5 and preselects the returned agents into selectedAgents so the existing submit path publishes them onto the loop with zero extra clicks. Existing rosters still work — the hand-pick panel now sits below the auto-provision card and only renders when the workspace has agents and no team was auto-provisioned this session. Co-Authored-By: Claude Opus 4.7 <[email protected]> |
||
|
|
e95e251c9c |
empty state: primary CTA button instead of hunt-for-plus text
The '+ new topic' button (commit
|
||
|
|
f8589471bc |
AgentComputer: clickable model picker (dropdown → PATCH /api/claws/:id/model)
The '● <model>' text on the per-agent Computer slide-out was
read-only — swapping a claw's model required curl. Now it's a
button that opens a small dropdown of the curated fleet models:
- Sonnet 5 / Opus 4.8 / Haiku 4.5 / Sonnet 4.6 (Anthropic)
- GLM 4.6 / 5.2 (Z.AI)
- Kimi K2 (Moonshot)
- Gemini 2.0 Flash (Google)
- Llama 3.3 70B (Groq)
Click a model → PATCH /api/claws/{id}/model → picker closes, the
button relabels to the new selection. The endpoint persists the DB
binding + best-effort re-provisions the runtime; the swap applies
on the claw's next turn. Long-tail models still work via curl with
any string.
Adds an optional onModelChanged callback so callers (Dashboard's
runtime-config enrich cache) can refresh their state without a
full re-render.
|
||
|
|
ac8c689f50 |
logs + team parity: pretty step/container renderers; quota + audit on team creation
Two bundled changes:
── LiveRunLogs prettification ──────────────────────────────────
The Steps + Container tabs were plain mono lines with a single
color per event. Now they get structured layout:
Steps:
- Color-hashed actor pill (stable palette so [Distiller] and
[Novelty Analyst] each get their own hue across the session).
- Phase pill (plan=cyan, work=green, synth=amber, aggregate=purple).
- Token count pill formatted 1.2k / 14.3k / etc.
- Gated-action warning pill in amber when > 0.
- Left-border color strip keyed to the actor for at-a-glance
visual grouping.
- Long outputs collapse to their first 300 chars with a '+ N more'
toggle to expand the full text.
- 'done' events get a green (or red for error) border strip +
pill instead of blending into the stream.
Container:
- Splits '[actor] action (outcome) · msg' into colored spans —
actor pill (deterministic color), action in dim, outcome pill
green/red/dim by state.
- Non-line events (info/error/done) get their own left-border
strip so bash echoes and stack traces don't drown in the daemon
chatter.
- Timestamps switch to HH:MM:SS.mmm — dense but scannable.
Small palette (LOG constants) keeps the color budget bounded — no
new UI vocabulary, just cleaner reads of what was already there.
── Team-wizard governance parity ───────────────────────────────
build_team_with_lifecycle now matches POST /api/claws' governance:
- enforce_new_agent quota check per member (previously bypassed
workspace agent quotas entirely for team/auto-provision paths).
- audit::append('agent.created', ..., {source: 'team_wizard'}) per
member so team-created claws appear in the same audit trail as
individually-created ones. Adding a 'source' key distinguishes
provenance without changing consumers.
.brain (h5) handling was already consistent between the two paths —
both use the lazy on-first-access load_brain hook seeded from
agents.system_prompt. No change there.
|
||
|
|
956be2cf4f |
wizard: in-place auto-provision team from topic (LLM-derived, sonnet-5)
Step 5 of ResearchWizard was 'assign agents from workspace roster'.
When the roster was empty, the wizard body was hard-swapped for
NoAgentsGate — you couldn't reach step 5 at all.
Now step 5 shows a 'Team' panel:
- Big cyan card: 'Auto-provision team from this topic'. One click
runs an LLM plan pass, gets 3-5 role slots + system prompts back,
materializes claws via the existing build_team pipeline, stamps
runtime posture, returns a shape that drops straight into the
submit body's agents[]. Card flips green with the derived roster.
- Below that: the classic roster picker, but only when the workspace
actually has ≥1 claw AND auto-provision hasn't landed. Otherwise
hidden — no dead empty-state affordance.
Every gate that required agents.length > 0 to render the wizard body
or the footer is gone. canNext gains a step-5 clause: allow Next when
EITHER auto-team is ready OR the user handpicked from a non-empty
roster.
Backend
- POST /api/teams/auto-provision — accepts {title, description,
outcome_kind, topology_kind?, model?, risk_profile?, mcp_bundles?}.
Derives topology from outcome_kind (integrations → pipeline; else
hub_spoke). LLM plan pass yields a JSON roster of 3-5 roles
(role_slot, name, system_prompt). Materializes team + claws via
build_team, stamps risk_profile (default research_web_readonly) +
mcp_bundles (default [clawmates_door, gitea_forge]). Response
carries team_id + agents[] in the shape /api/research already
expects.
- Every provisioned claw runs on claude-sonnet-5 by default;
overridable via the model field.
Follow-ups (not in this slice):
- Same picker in LoopsWizard (slice C — parallel change, same API).
- Post-create 'Team' section on ResearchCanvas / LoopsCanvas so
users can rebind after the fact (slice D).
- Full Teams tier UI + Agents-page deprecation (slice E).
|
||
|
|
47a42423e3 |
wizards: unlock '+' buttons when roster is empty
The 'no agents → button disabled' guard hard-locked the wizard behind an already-populated roster, forcing users to detour to the Agents page and come back. NoAgentsGate already handles the empty state gracefully inside the wizard body; the button just needed to open it. Applied to both ResearchList and LoopsList. Follow-up (design in progress): let the wizard auto-provision a team from the topic/loop itself so users never have to leave to create agents. |
||
|
|
eb06e91ac3 |
dashboard(Agents): drop counts caption, align toolbar with title
Previously the Agents header wrapped a 'N ORGS · N CO · N TEAMS · N AGENTS' small-caps caption above the title, with the toolbar (select/history/brain) hanging off the right. Dense, noisy, and the World tier already surfaces those counts at higher fidelity. Now: single-line header — 'Agents' title on the left, toolbar flush right, both center-aligned. Comment left in place noting where the caption used to live in case we ever want it back. |
||
|
|
0b7f247b0e |
wizard: 'fresh coding team' picker for paired coding loop
Second slice of the per-loop-team arc. The paired-coding-loop checkbox
in ResearchWizard step 6 now exposes a two-option picker:
⦿ Provision a dedicated coding team (default when the loop is on)
— fresh 'Coding · <topic>' team row, risk_profile =
coding_readwrite, clawmates_door in mcp_bundles. Loop's
team_id is bound at wizard-submit time.
○ Reuse the research team (legacy) — no team_id bound; coding
iterations spawn against the research topic's container.
Frontend
- New codingTeamMode state, radio picker rendered under the checkbox.
- research.ts createTopic body gains paired_coding_team_mode?: 'fresh'|
'reuse'.
Backend
- CreateTopicRequest gains paired_coding_team_mode: Option<String>.
- materialize_topic_loops takes it through and, when 'fresh', calls
the new provision_fresh_coding_team helper — inserts a teams row
via the existing insert_team_with_lifecycle (pipeline kind, same
graph as the loop), sets its runtime-config via
set_team_runtime_config, then binds loop.team_id.
- All operations best-effort with stderr logging — a team-provision
failure leaves the loop functional under the legacy fallback.
Not shipped in this slice (deferred to runtime hookup slice):
- research_container::spawn keyed on team_id → per-team container
- Config template rewrite injecting the team's risk_profile
- Migration of existing paired loops onto their own teams
The plumbing lands now so the wizard's intent is recorded; the
runtime honors it in the next PR.
|
||
|
|
b2a38da3f9 |
loops: 'View full output' modal for a completed iteration
The expanded iteration panel now exposes the full topology_run.result in a proper viewer instead of leaving it stranded in Postgres. Reads /api/topology-runs/:id (existing endpoint, no backend change), prefers result.final_output (the produced markdown/code), falls back to a per-step transcript, last-resort a raw JSON dump. Modal renders as a fixed overlay — Escape or backdrop-click to close. Header carries kind/status/steps/tokens metadata, Copy button hits navigator.clipboard, Download button emits a .md file named run-<id-slice>.md. The button on the row is disabled until the iteration terminates (completed | failed) — running iterations already have LiveRunLogs surfacing the tail. |
||
|
|
2ba40b03b7 |
loops: per-iteration live logs (Steps + Container tabs), collapse graph JSON
The Loops canvas' IterationsTimeline was a flat status-pill list —
no way to see WHAT an iteration was actually doing. Meanwhile the
Research canvas had full LiveRunLogs with Steps + Container tabs
against the same underlying topology_run SSE endpoints. This
factors LiveRunLogs so both canvases share it.
Frontend
- LiveRunLogs gains a directRunId?: string prop. In this mode it
skips the topic-scoped active-runs + pipeline-state polls and
pins activeRun to the given id. topicId stays optional (topic
mode unchanged from ResearchCanvas' perspective).
- LoopsCanvas IterationsTimeline: each row is now a click-to-
expand card. On expand, renders <LiveRunLogs directRunId={...}/>
right below the header — Steps + Container tabs, full SSE tail,
same 320px terminal.
- Auto-opens the newest running/queued iteration so a click on
'Run now' immediately exposes the live pane.
- New CollapsibleSection helper wraps the graph JSON block so it
starts closed. Reference material stays one click away without
cluttering the canvas.
Backend
- run_container_log_sse now resolves the tail target via
loop.source_research_topic_id when the run itself has no
research_topic_id. Paired coding loops (kind=exec with a
source_research_topic_id) reuse the paired topic's team
container, so we tail its docker logs. Pure loop runs still
error with a clearer message.
|
||
|
|
ae113dd319 |
live-run-logs: add Container tab that tails filtered daemon logs
Steps summaries only fire AFTER each topology step completes — so a
stalled first turn was completely dark. Add a second tab that
streams the team runtime container's daemon log live via a new SSE
endpoint.
Backend
- GET /api/topology-runs/:id/container-log — workspace-scoped SSE
around bollard's docker.logs(follow=true, tail=200). Buffers on
newline so partial mux chunks don't truncate a log line.
- compact_container_log: parse a zeroclaw daemon line
('[actor] ... zc_action=X zc_outcome=Y ... msg') into
'[actor] action (outcome) · msg'. Framing-only continuations are
dropped; non-zc lines (bash echoes, backtraces) pass through as-is
so nothing interesting is lost. ANSI escapes stripped.
- Everything funnels through one async_stream! so early exits
(workspace check / docker connect / no bound topic) yield an
'error' event and return without breaking Sse::new's single stream
type.
Frontend
- LiveRunLogs gets a sub-tabs strip: Steps · Container.
- New useContainerLog(runId, active) hook — gated by tab so we don't
hold two open SSE streams when the operator isn't looking.
- Same terminal widget renders each container line with a level
color (info/done grey, line default, error red). Sub-tab pill
shows count + status live.
|
||
|
|
f70f6c679e |
research: errored-state card + one-click rerun (no wizard re-entry)
When a topic ends up parked in 'processing' with all runs failed and nothing in flight, the sidebar card was still spinning as if progress were happening. Now: Backend - topology_runs::run_counts_by_research_topic — batch query that returns (in_flight, failed-since-last-success) per topic. Used by the list endpoint; dynamic sqlx::query() so no prepare needed. - TopicListItem DTO gains runs_in_flight + runs_failed. - start_topic status guard relaxed: allow (standby) OR (processing AND runs_in_flight == 0). Blocks accidental double-fires on a live pipeline; permits rerun on a failed one. Same request body, same behavior once accepted, so the frontend just POSTs /research/:id/start on the RotateCw click. Frontend - ResearchList detects errored: status===processing && !in_flight && failed>0. Swaps the MiniSpinner for a red AlertTriangle and changes the status text to 'error · N failed'. - New RotateCw icon button next to the delete Trash — same button cluster, one click, no wizard re-entry required. Disables while a request is in flight; error surfaces in the sidebar's shared error banner. |
||
|
|
43f7880327 |
agent cards: disk-LED glow on live topology step
Wire SSE step events from LiveRunLogs up to ResearchCanvas via an
optional onStep callback (StepPulse: {role, phase, node_id, ts}).
ResearchCanvas keeps a role -> {phase, expiresAt} map; a 250ms tick
clears expired entries so the glow fades naturally.
Each agent card:
- 8px LED dot next to the name (green idle-off / colored when live)
- outer glow via two-layer box-shadow when the agent's role_slot
matches the last step's role, within a 2.5s window
- inline 'reading' / 'writing' hint below the role
Color mapping (disk-LED metaphor):
- Plan phase -> cyan #5ec8d8 (reading / decomposing)
- Work/Synth/Aggregate -> green #5fd08a (writing / producing)
Overlapping steps reset the timer so back-to-back activity on the
same role holds the glow. When no run is active nothing pulses —
LiveRunLogs is silent, no callback fires.
|
||
|
|
45292bf5fb |
sidebar polish: tighter chevron + '+' drops to Research line
- Dashboard chevron: 24x24 → 22x22, top:10 right:8 → top:4 right:4 so it sits in the true corner of the panel. - ResearchList header: alignItems flex-start → flex-end so the '+' button lands on the 'Research' title baseline instead of hugging the 'N TOPICS' caption line. Reserved 34px right padding on the header so the '+' clears the corner chevron. |
||
|
|
1876641415 |
dashboard: collapsible context sidebar (research/loops/repos/infra)
The earlier PR added a chevron to RosterColumn but the research /
loops / repos / infra tiers render a completely different sidebar
(ResearchList/LoopsList/…) inside a fixed 252px div in Dashboard.tsx
— the RosterColumn toggle never appeared on those pages.
Add a proper collapse on the Dashboard's own context-list wrapper:
- Overlaid PanelLeftClose chevron top-right (top: 10, right: 8) so
it doesn't fight the per-tier header content.
- When collapsed, wrapper shrinks to 32px with a PanelLeftOpen
button; canvas gets the reclaimed ~220px.
- State via useSyncExternalStore on localStorage
('cm.dashboard.sidebarCollapsed'), same pattern I used on
RosterColumn to satisfy react-hooks/set-state-in-effect.
|
||
|
|
573956e840 |
live-run-logs: fix step-record parser + surface pre-first-step phases
Two fixes based on live smoke test:
- StepRecord parser: the SSE step payload is {node_id, role, phase,
output, gated, tokens} (orchestrator::StepRecord), not the made-up
shape my first pass looked for. Render as '[role] phase · <first
line of output> · Nt · N gated'.
- Pre-first-step visibility: the checkpoint only journals AFTER each
turn completes, so the container-startup + agent-boot window (often
30-90s for the first step) was completely dark ('waiting for first
step…'). Poll pipeline-state every 2s and render its stages as
pseudo-log lines (setup 🟢 staffing · 3 agents assigned / setup 🔵
container · starting…) until real step events arrive.
|
||
|
|
2ef50bb30b |
lint(research): satisfy react-hooks/set-state-in-effect
- RosterColumn: useSyncExternalStore on localStorage instead of hydrate-in-effect setState. Same-tab writes nudge subscribers via a synthetic storage event. - LiveRunLogs: reset-on-runId-change via the React-sanctioned set-state-during-render pattern (prevRunId ref state), so the effect body no longer calls setState synchronously. |
||
|
|
1d69866bcf | research canvas: collapsible sidebar + live topology-run logs (#6) | ||
|
|
149ad992bb | wizard: materialize picked repo across clawstor fleet at step-2-next (#5) | ||
|
|
1cb643142c | research/publish: gate approve+reject on Owner role (#4) | ||
|
|
ee05037095 |
research-pipeline-diag: state-aware statuses (waiting != failing)
The diagnostic was calling any 0-outcome + 0-completed-runs state a
FAIL — red dot + 'No outcome produced yet — check the runs stage for
the failure reason'. That fires the moment a wizard-materialized
topic lands, before its very first turn even completes, and stays
red for the whole 2-3 min a legitimate coordinator turn runs. Result:
users see 'FAIL' on every fresh topic and can't tell a real failure
from a normal in-flight state.
Backend fix — introduce a `waiting` status (blue/pulsing in UI):
Runs stage:
0 runs -> skip ('No runs yet — pipeline hasn't fired')
any running/queued -> waiting ('N in flight, M completed')
all failed -> fail (with error text)
some failed -> warn
all completed no fails -> ok
Outcomes stage:
outcome_count > 0 -> ok
0 outcomes + 0 runs -> skip ('No outcome yet (pipeline hasn't fired)')
0 outcomes + any running -> waiting ('Waiting for the current run to finish…')
0 outcomes + any failed -> fail (the actual silent-bug case)
0 outcomes + all done ok -> warn (weird — completed but wrote nothing)
Also suppresses the run stage's `latest_error` detail when the run
status is `waiting` or `skip` — reporting a stale error next to an
actively-running job is what made users think the current run had
failed.
Frontend:
- PipelineStage['status'] union grows a 'waiting' arm.
- Pill color: cyan (#5ec8d8) with a pulsing scale/opacity animation
(new cm-pulse keyframe in motion.css).
- Strip summary line: 'Pipeline in flight — waiting for run to finish…'
when there's any waiting stage and no failures.
- Border tint: cyan border when waiting, coral when failing, neutral
otherwise.
Zero backend semantic changes to the outcome-write path — this is
purely UI truth-telling.
|
||
|
|
bfcdca0583 |
research-canvas: managed-by-loop UI + start_topic guard (fold cleanup)
Closes the UX gap the fold introduced: the topic canvas was still showing "Start research" for standby-state topics even when a scheduled loop already owned the runs. Clicking it would 409 (or worse: race the loop into a duplicate run). Topic status stayed at standby forever because the loop path bypassed start_topic's set_status transition. Four changes: 1. **Backend status transition** — compose_and_enqueue_iteration for kind='research' now calls set_status_if(standby, processing) on the topic before the run is enqueued. New DB helper set_status_if only advances when the current status matches the "from" arg — safe against races and re-invocations. Later iterations no-op since the topic is already past standby. 2. **has_managed_loop on TopicDetail** — get_topic hydrates a new ManagedLoop struct (loop_id, title, enabled, next_fire_at, last_run_id, schedule_summary) when a kind='research' loop is bound to the topic. summarize_schedule() derives a human string from the loop's triggers jsonb (e.g. "cron: 0 3 * * * · on new artifact", "one-shot", "manual"). New DB helper loops::research_loop_for_topic returns the row. 3. **Canvas branch** — nextAction takes a managedByLoop flag; when set + status=standby, returns null (no button). The canvas renders a "MANAGED BY LOOP" strip below the topic title showing loop name, schedule summary, next fire time, and enabled dot. Reviewer buttons (Request publish / Approve / Reject) still show normally in later states — reviewers should still promote outcomes even when a loop is producing them. 4. **start_topic guard** — refuses with 409 when a research loop already owns the topic. Closes the direct-POST hole for anyone bypassing the frontend. TS type + summarize_schedule live in the same commit so an old client hitting a new backend just ignores the extra field (no breakage), and a new client hitting an old backend renders the classic buttons (managed_by_loop is optional). |
||
|
|
3e42b1ea39 |
research-wizard: schedule step (once/nightly/manual) + paired coding loop
Completes the research/loop fold. The wizard now closes with a
"How should this research run?" step; picking any mode materializes a
kind='research' loop bound to the topic, and an optional checkbox
adds a paired kind='exec' loop that consumes each new artifact.
Frontend:
- ResearchWizard grows from 5 to 6 steps. Step 6 is the schedule
picker:
· Just once — initial_burst=1, no other triggers
· Nightly — initial_burst=1 + cron "0 3 * * *"
· Manual — webhook_enabled=true
Below the radios, an optional card offers "Also create a coding
loop that consumes each new artifact" — creates a paired
kind='exec' loop with on_artifact_update=true + initial_burst=1
bound to the same topic.
- createTopic API type extended with `schedule` + `create_paired_coding_loop`.
- Both fields ride the existing POST /api/research call; back-compat
is preserved when the wizard omits them.
Backend:
- CreateTopicRequest gains TopicSchedule + create_paired_coding_loop.
- After topic + agent attach, create_topic calls
materialize_topic_loops which:
1. Creates a research-kind loop titled "Research · <topic>" bound
to the topic. Triggers vary by schedule mode; next_fire_at
computed from cron for nightly. Falls back silently if
loop-create errors so the topic still lands.
2. Flips kind to 'research' via loops::set_kind (NewLoop doesn't
take kind directly — default is 'exec' for backward compat).
3. Optionally creates a coding loop titled "Coding · <topic>"
with on_artifact_update=true + initial_burst=1.
- cm_db::repo::loops::set_kind — trivial UPDATE helper used by the
materialize path.
D-answer callouts:
- D1 (fold): every runnable thing is now a loop. "Just once" is a
research loop with initial_burst=1 and no other triggers.
- D2 (inherit): both paired loops carry the same source_research_topic_id
— repo binding lives on the topic, not duplicated.
- D3 (coordinator resolves): the research iteration prompt (from the
earlier commit) instructs the team to preserve stable INT ids and
mark deprecations; coding loops' consumed lists stay valid across
versions.
Follow-ups queued:
- Kind pill on LoopsList cards (research=purple, exec=cyan) so users
can tell them apart at a glance.
- Extract start_topic's task-build so kind='research' iterations
reuse the same coordinator prompt shape as one-shot runs (they
currently use a simpler refresh-oriented prompt; that's fine for
MVP but a rich shared build would give better parity).
- Research topic sidebar shows "linked to N loops" badge.
|
||
|
|
5aa2de7f30 |
loops-wizard: initial_burst + on_artifact_update trigger UI
Surfaces the two new trigger fields shipped in the previous commit so
users can compose the new loop patterns without editing raw jsonb.
Step 4 (Triggers) now has five toggles instead of three:
1. **Run at create-time (initial burst)** — number input, defaults to 1
so every new loop runs at least once. Toggle off = classic behavior
(no auto-fire).
2. Cron schedule (existing)
3. On previous completion (existing) — pairs with initial_burst > 1
for a back-to-back burst chain, and with cron for continuous
iteration between schedule ticks.
4. **On new research artifact** — visible always; auto-disabled with
an explanatory label when no source topic is bound. Coalesced
server-side against in-flight runs so rapid revisions don't spawn
duplicates.
5. Webhook (existing)
Body wiring:
- Sends initial_burst > 0 → { initial_burst: N } in triggers jsonb.
- Sends on_artifact_update only when a source topic is bound (safety
belt matching the backend's fan-out filter).
Step gate:
- Advancing from step 4 now requires ANY one of the five triggers
(was just the three legacy toggles). The initial_burst default of 1
means users can happy-path through this step without deliberate
action, which is the intended nudge — every loop should have at
least one path to run.
Read/write parity:
- readInitialTriggers grows two fields so edit-mode reflects existing
loops' triggers correctly.
Next commit lands ResearchWizard's "When should this run?" step +
optional paired coding loop, and after that the kind='research'
dispatch path.
|
||
|
|
ce111273bb |
loops: reorder history collapsed under the progress pill
Completes the reorder rationale loop — the previous commit captured
REORDER: markers server-side but nothing surfaced them. Now the loops
sidebar renders a small "N reorders ▸" button under the progress pill
whenever the loop has any reorder events. Click expands to a
newest-first list of "iter <n> · <rationale>" lines so a reviewer can
see, at a glance, when the plan was adjusted and why.
Backend:
- LoopProgress DTO gains recent_reorders: Vec<Value> — newest-first,
capped at 5 so the card stays compact. Full history remains on the
loop row's reorder_events column.
- cm_db::repo::loops::recent_reorders — reads the jsonb array, returns
the last N in newest-first order.
- list_progress populates it per loop.
Frontend:
- LoopReorderEvent + recent_reorders on LoopProgress type.
- LoopsList tracks openHistoryId per-loop (one open at a time).
- Card renders history button + expanded panel styled to match the
progress pill above.
Notes:
- The event object schema is {run_id, iteration, text, ts}. Fields are
optional in the TS type so future schema tweaks don't break the
render.
- 5-item cap chosen so the sidebar card doesn't grow unbounded. If a
loop accumulates a lot of reorders, follow-up UI can render the full
history on the loop detail page.
|