774f17d1949b772c7c10187af6f84f7ca4fae240
447
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f9e8d8d779 |
research: publishing → published + artifact download (R1)
Before this commit, approve_publish left topics stuck in 'publishing' forever — the sidebar 'published' bucket was always empty and nothing surfaced the artifact. Two-part fix: State machine — approve_publish now transitions reviewing → publishing → published in one API call. Real async packaging isn't a thing yet because the artifact IS the markdown already written to research_outcomes when the last run completed (topology_worker). The intermediate 'publishing' state is preserved (schema-level trigger stamps published_at on landing there) so we keep the option to detour through it later for pdf render / mirror-to-store / etc. Download endpoint — GET /api/research/:id/artifact returns the latest outcome as text/markdown with Content-Disposition: attachment. Filename sanitizes the topic title to ascii-alnum + dash and appends the outcome version so accumulated revision drafts (post R2) don't clobber. Workspace ownership check via research_topics::get; 404 if no outcome yet (reviewers browsing before a run completes). Frontend — ResearchList shows a green Download icon-button next to Delete when status === 'published'. Clicks trigger a plain anchor download of the .md — no JS blob dance needed since the response is already an attachment. Follow-up queued: pdf render (server-side or client-side of the md), mirror-to-store (S3-ish or Obsidian vault) as an async step during the publishing→published detour. |
||
|
|
f60df36717 |
research: teardown per-topic container on publish + delete (P1)
Path B spawned a per-topic team container on start_topic (research runtime) but nothing ever stopped it. Containers accumulated on gw-04 across the lifetime of every topic — the earlier cleanup pass reclaimed 6.72 GB of these. Now the container is stopped + removed automatically at the two terminal transitions: - delete_topic — the topic row is gone, the container is meaningless. - approve_publish (reviewing → publishing) — the research work is done. Artifact writing (queued as R1) reads from the durable run_events log, so it doesn't need a live runtime. Wiring is a fire-and-forget teardown() helper in research_container.rs: - Connects Docker via the same socket-proxy shim as spawn(). - Calls the existing stop() which is idempotent (404 = already gone, 304 = already stopped, both treated as ok). - Errors log with eprintln! and don't block the API response — a topic is done whether Docker is reachable or not. Dev machines without Docker running just log a debug line and return. Path A (repo-only mount + shared clawmates_core network) is untouched so there's no Traefik dynamic-file cleanup needed. If we ever add per-topic Traefik routing, extend teardown() with the corresponding file remove. Follow-up: gw-04 will still have any legacy containers from before this commit. One-shot `docker ps -a --filter "name=research-" -q | xargs -r docker rm -f` on gw-04 is the manual sweep. |
||
|
|
763b95a253 |
api: gate create_loop + create_topic on empty workspace roster (W2)
Frontend already disables the create button when the roster is empty, but nothing stopped a direct POST from materializing an orphan loop / topic with nothing to staff it. Both handlers now check the workspace agent count up front and return 409 Conflict when it's zero. - crates/cm-api/src/routes/loops.rs: gate at top of create_loop - crates/cm-api/src/routes/research.rs: gate at top of create_topic Uses the existing cm_db::repo::agents::count_active(pool, ws) helper (count of workspace agents where deleted_at IS NULL). 409 is the right mapping: state-of-the-workspace-prevents-this, not client-input-bad. |
||
|
|
5a2340587e |
world: loop:<id> landmark orbs (V1)
Mirror the repo: landmark pattern for scheduled loops. Every enabled loop in the workspace gets a labeled amber orb in the World, whether it's currently running or between fires. Assigned agents converge on it with a soft 0.15 touch — the loop is a persistent landmark, not a transient run. Backend: - active_loops(pool, ws) query joins loops + loop_agents where enabled, returning (loop_id, title, agent_id) — one row per (loop, agent). - SSE loop emits node.activity + world.touch symmetric to the research block. Seen-once set dedupes the label emission across agents. Frontend: - New "loop" tier in the Tier alias, LEVEL_COLOR (#f0b866 warm amber), and ensureNode radius (11 — same landmark size as repo). - engine.onTouch / onNodeActivity preserve the tier from the loop: prefix (previously would have collapsed to service). - Pawn fireColor tinted amber for loop: touches. - WorldCanvas: loop tier joins the struct group for solid-at-rest glow + always-on labels. Focus mode recognizes loop: prefix (click → focused subtree, Esc to exit). Focus pill switches to "LOOP FOCUS" in amber when the selected id is a loop. Contrast with repo: (transient — only appears when a topic is in processing/reviewing/publishing). Loops are persistent because their whole point is recurrence. |
||
|
|
708f45d09b |
world: solid team + project orbs, hide ROOT, per-topic repo landmarks
Three fixes to the Live viz per user feedback:
1) Hide the ROOT sentinel — it was rendering at the origin as a red disc
with no label. It's a physics anchor, not a real orb. Skipping it in
nodes/edges/labels loops removes the mystery red circle.
2) Emit `repo:<topic_id>` project orbs for active research topics from
the world SSE loop. Labeled with the topic title, tier "repo" gets its
own soft sky-blue palette and a landmark radius (r=11 between company
and team). Assigned agents get a low-weight (0.15) convergence touch
so pawns cluster around their project's orb even at rest — no wait
for a file op to see the affiliation.
3) Solid at rest, bloom on interaction. Two engine changes:
- Structural + repo orbs now have a soft 0.08 glow at heat=0 (down
from 0.22 ambient) so the disc reads as solid until agents heat it.
- `world.touch` heat is now weight-scaled (`+w*0.6`) instead of a
flat `+0.5` regardless of intent. Soft convergence stays soft;
file ops still explode.
Click a `repo:` orb → the existing Commit F focus mode already treats
that prefix as a subtree root, so users drop straight into the
Gource-style repo detail view with only the files their agents are
touching.
Follow-ups queued: teardown of repo orbs when a topic reaches 'published'
(currently they persist until the SSE loop's status filter drops them,
which is correct); loops equivalent (loop:<id> landmark orbs).
|
||
|
|
f88e8642d9 |
world viz: repo focus mode — click a file/dir/run to enter the tree
MVP of the repo-detail sub-view. Click any file/dir/run/repo node
in Live and the viz culls to just that node's subtree — you're
now watching agents crawl the repo instead of the whole workspace.
Esc (or the pill button that appears) exits.
Backend
- routes/world.rs: file-op tool events now emit a `file:<path>`
world.touch IN ADDITION to the existing `tool:<name>` touch.
The path is pulled from the tool input's path / target / file
/ filename / url keys (same lookup summarize_input uses, but we
keep the full string so the client can build a real hierarchy).
Non-file tools are unchanged — they still hit tool:<name> nodes
as before.
Engine
- onTouch synthesizes a directory hierarchy when the id starts
with `file:`. Each intermediate path segment gets a `dir:<acc>`
node (label = the segment), parented at the previous dir; the
file itself parents at the innermost dir. Ensures the layout
spring-simulates as a tree naturally, no separate render mode
needed.
- New pawn fireColor for file: touches: coral #ff8a7a. Reads as
"file work" vs #5ec8d8 (tool convergence) vs #5fd08a (run
activity).
WorldCanvas
- Render pass now takes a `visibleNodes: Set<string> | null`.
When the selected id starts with file: / dir: / run: / repo:,
we BFS descendants via parentId and hide every non-descendant
node. Node meshes, glow sprites, hierarchy edges, and labels
all gate on the set. Pawns stay visible (agents still dart to
the focused files).
- ESC handler on window: clears the selection by calling
onSelect("") when a repo-focus id is set.
- Small "REPO FOCUS · <path>" pill lands at top-center with an
Esc button so the exit is discoverable at a glance without
learning the shortcut.
- Dashboard.onWorldSelect now treats empty string as "clear
focus" (setWorldSel(null)) so the same callback handles ESC.
Not yet: the always-on repo:<topic_id> node emitted at run
start when a research topic has a repo bound. Today the focus
works off run:<id> nodes; a repo:<id> anchor would let users
click without waiting for a first file touch. Also skipped:
per-file heat map / call-count visualization tied to touch
weight over time. Both are natural follow-ups on this bones.
|
||
|
|
aa941b72a1 |
world viz: agents grow with their brain (log-curve dot size)
Fresh agents start at "size of their letters" — a small dot in
the live viz — and visibly bloom out as their .brain file fills.
Turns "which of these agents is heavily loaded" into a glance
instead of a menu dive.
Taxonomy
- New agent.memory event: { agentId, bytes?, count? }. STATEFUL, so
a late subscriber sees the last value replayed and pawns arrive
pre-sized. count is included in the schema for a future combo
metric but not emitted yet — bytes carries the visual today.
Backend (routes/world.rs SSE loop)
- Per-agent, per-tick std::fs::metadata() on
brain_dir()/claw_<uuid>.h5. Just the inode stat — no HDF5 open,
no memory count, sub-ms per agent. Emit agent.memory { bytes }
only when the value has changed (or on first sight).
- Tracks last_bytes: HashMap<String, u64> in the SSE-stream scope
alongside the existing status HashMap.
- Missing file (agent never provisioned a brain) reads as 0 bytes
and yields scale = 1.0 downstream — pawn stays small.
Engine
- GPawn gains memoryScale (visible) + memoryScaleTarget (chased).
Base is 1.0; ensurePawn initializes both.
- memoryScaleFromBytes(bytes): 1 + log10(1 + bytes/1MB) * 0.6, cap
MAX_MEMORY_SCALE = 3.5. So 10MB ~ 1.6x, 100MB ~ 2.2x, 1GB ~ 2.8x.
Log curve keeps a heavy brain readable without a lite one being
invisible.
- onMemory(e) sets the target. stepPawns eases the visible scale
toward it at ~3/sec — a big incoming snapshot doesn't pop the
sphere; it swells in like it's inhaling.
Renderer (WorldCanvas)
- Live subscription registers agent.memory alongside the existing
status/touch/reasoning listeners.
- Pawn sphere scale = 5 * p.memoryScale (was hardcoded 8). Halo
scales in proportion (max(24, 4.25 * s)) so a memory-heavy agent
reads as a bigger presence, not a small dot with a huge halo.
- AABB bounds for the frame-camera math updated to use s instead
of 8 so the camera actually frames a big agent when it's the
outlier.
Not yet wired: comm lines between pawns when agents talk to each
other (Commit E next), and the topology-edge overlay that renders
the graph shape dimly at rest. Both build on top of this — bigger
dots make comm beams more visible.
|
||
|
|
acd2a0f287 |
structure polish: post-reify nav + ensure-chain + TeamWizard auto-parent
Two small quality-of-life fixes on top of the reify commit:
Post-reify navigation
OrphanMigrationDialog already returned team_id in its result;
Dashboard now pushes /?team=<team_id> before router.refresh() so
the user lands on the freshly-materialized team and sees exactly
where their agents just moved. Previously they had to hunt for it
in the newly-rebuilt sidebar.
Wizard auto-materialize (POST /api/structure/ensure-chain)
cm-db: ensure_chain(pool, ws, fallback_org, fallback_company) —
fast path returns coordinates of the first org+company already
bound in this workspace (workspace's oldest org, oldest company
under it). Slow path inserts a new org+company with the
fallback names ("My Workspace" / "General") + binds them via
org_companies. Returns { org_id, company_id, created }. Small
txn — leaves the workspace consistent whether it was already
wired or not.
cm-api: POST /api/structure/ensure-chain accepts optional
fallback_org_name and fallback_company_name in the body (trimmed,
else default). Returns the ids.
CreateTeamRequest gains an optional attach_to_company_id. When
set, after build_team() completes, we look up the company
(workspace ownership check enforced by companies::get), count
its existing teams for a stable n_i node id, and insert a
company_teams binding — so the team lands under the parent
atomically instead of a follow-up round-trip.
TeamWizard now calls ensure-chain before POST /api/teams and
passes the returned company_id in attach_to_company_id. Both
calls are best-effort — if ensure-chain fails (network etc.)
we still try to create the team, and the migration dialog stays
available as the fallback UX. Wizard flow now: fresh workspace's
first team is fully wired from the moment it appears in the
tree — no synthetic "My Workspace" scaffolding ever gets
rendered around it.
The Team/Company create paths not touched here (create_team_from_claws,
company create, org create, MasterPlannerModal scaffold) still
work as before — they just won't auto-parent yet. Later commits
can wire them the same way.
|
||
|
|
8b789beec0 |
structure: reify-orphans endpoint + "give these a home" dialog
Turns the four synthetic tree containers into a real migration path.
Clicking any of them ("My Workspace", "Teams", "Direct",
"Ungrouped") opens a dialog that creates a real
org → company → team chain and re-parents every orphan into it, all
in one DB transaction.
cm-db (new module structure_reify)
- orphan_agents / orphan_teams / orphan_companies: workspace-scoped
SELECTs of entities without a parent binding in team_members /
company_teams / org_companies. Used both by the dialog's counter
and internally by the migration.
- count_orphans: cheap combined-count via three subqueries in a
single SELECT so the dialog only round-trips once for the header.
- reify_orphans(pool, ws, org_name, company_name, team_name):
1. begins a tx
2. inserts a new org + company + team (all `flat`, empty graphs
— user can shape them later via the existing PATCH endpoints)
3. binds company under org (org_companies "n0")
4. binds team under company (company_teams "n0")
5. inserts team_members rows for every orphan agent (n1, n2, …)
6. inserts company_teams rows for every orphan team
7. inserts org_companies rows for every orphan company
8. commits, returns the created ids + moved counts
cm-api (routes/structure)
- GET /api/structure/orphan-counts → { agents, teams, companies }
- POST /api/structure/reify-orphans → { org_id, company_id, team_id,
moved_* }. Trims + rejects any empty name; validates before
starting the transaction so a 400 never rolls anything back.
Frontend
- New OrphanMigrationDialog: fetches counts on open, three name
fields (defaults: Organization "My Workspace", Company "General",
Team "Everyone"), POSTs on save. "Nothing to migrate" state
disables the save button when the workspace is already fully
wired. Copy explicitly notes that everything is renameable in the
sidebar afterward.
- Dashboard: onTreeSelect now branches on SYNTHETIC_TREE_IDS —
clicking a synthetic node opens the dialog instead of falling
through to the (nonexistent) selection. On successful reify,
router.refresh() so the sidebar + world viz reflect the new real
chain.
What this doesn't do yet (next commit)
- Wizard auto-materialize: when creating a team/company via wizard,
auto-create parent placeholders if they don't exist. Deferred so
this commit stays focused.
|
||
|
|
99e5207e69 |
sidebar: click-to-rename org/company/team + strip synthetics from world viz
Two related pieces of the "kill My Workspace" cleanup, landed
together because they share the same file:
Backend
- Three tiny inline-rename endpoints:
PATCH /api/orgs/{id}/name
PATCH /api/companies/{id}/name
PATCH /api/teams/{id}/name
Each takes { name: string }, trims + rejects empty, returns 204.
Backed by rename_org / rename_company / rename_team in cm-db —
single-row UPDATEs scoped to the caller's workspace, NotFound if
the id isn't visible.
- Registered next to the existing PATCH /:id (topology) routes so
they don't collide.
Frontend
- StructureTree accepts an optional onRename and canRename.
TreeRow: click on the label text of a renamable node → the span
becomes an <input>, focus + select-all, save on Enter or blur,
cancel on Escape. The rest of the row (row chevron / row body)
still navigates + selects as before, so single-click behaviour
is preserved for everything except the name text itself.
react-hooks/set-state-in-effect avoided by resetting the draft
in the enterEdit() click handler instead of inside a useEffect.
- Dashboard passes canRename={item.level !== "claw" && !synthetic}
(claws don't have a rename endpoint yet; synthetic scaffolding
gets reified into real rows in the next commit — the wizard
auto-materialize + orphan-migration dialog).
onRename fires the corresponding PATCH and calls router.refresh()
so the label lands in every consumer of the tree.
- World viz seed: new stripSynthetics(roots) helper walks the tree
and lifts children of any synthetic container up to their
grandparent's level. worldCanvasRoots feeds through this before
narrowRoots(). Result: the Live viz no longer shows "My Workspace"
or "Teams" nodes — real agents orbit the world root directly
(which is what you were asking for). Sidebar tree still shows
them so orphaned agents remain visible until the migration lands.
|
||
|
|
88c78bd16e |
research: topology_worker points executor at per-topic gateway (commit 2/3)
Commit 2 of the path-B plan. The container that commit 1 spawns
now actually receives the run's turns — up until now it was
started but unused. This is the payoff commit: research runs are
truly isolated per topic.
Backend
- topology_exec.rs: from_env() refactored to a thin wrapper over a
new from_env_for_gateway(url) helper. Same shape (env-derived
aliases + default + token) but the caller supplies the URL. The
auth token/pairing code still comes from ZEROCLAW_TOKEN /
ZEROCLAW_PAIRING_CODE on the server; research_container's
inherited_env propagates those into the team container so the
same credentials work at both endpoints.
- research_container.rs: new wait_ready(url, deadline) that polls
<url>/health with a 1.5s per-request timeout every 500ms until
it 200s or the deadline passes. reqwest-based so it doesn't need
bollard. Called by the worker after claim, before pair, to bridge
the "container is starting, gateway not yet listening" gap.
- topology_worker.rs run_job:
1. Look up research_topic_id for the claimed run.
2. If Some, load the topic and read zeroclaw_gateway_url.
3. If a URL is present:
- best-effort wait_ready(url, 30s); a timeout logs but
doesn't abort — the pair call below will just fail
faster than pinging forever
- build the leaf via from_env_for_gateway(url)
Else fall back to from_env() (workspace-wide gateway).
4. The rest of run_job is unchanged — the leaf drops into
either SubTopologyExecutor (org/company) or direct drive
(team tier) as before.
What now works end-to-end
Starting a research topic with a bound repo:
1. clone-shallow into per-topic workspace
2. docker create + start the clawmates-runtime container, name
= research-<topic>-team, joined to clawmates_core so the
server reaches it by name
3. persist container name + gateway URL on the topic row
4. enqueue the topology run tagged with research_topic_id
5. worker claims → looks up the topic → waits for the team
gateway's /health → constructs a from_env_for_gateway
executor pointed at http://research-<topic>-team:42617
6. every turn's `/ws/chat?agent=…` hits the isolated container;
agents inside see the repo at /workspace/repo (rw); each
topic's memory/state lives under its own /zeroclaw-data mount
Deploy prereqs (unchanged from commit 1)
- clawmates_server compose service needs a bind-mount of
CLAWMATES_RESEARCH_WORKSPACE_ROOT so the paths spawn() writes to
are visible on the host and the spawned team container mounts
the same underlying data.
- socket-proxy ACL needs POST + DELETE on /containers (prod ✓).
|
||
|
|
21ac35c8d4 |
research: spawn per-topic ZeroClaw team container on start (commit 1/3)
Commit 1 of the path-B (real per-topic isolation) plan. The
container spawns and its coordinates persist — nothing talks to
it yet; commit 2 wires ZeroClawDriveExecutor to prefer the topic's
URL when populated. This split keeps each landing verifiable.
Backend
- Migration 0038: research_topics gets zeroclaw_container_name +
zeroclaw_gateway_url columns. Both nullable so a topic can exist
before a spawn and teardown just NULLs them out.
- cm-db: ResearchTopic struct extended; get/list SELECTs updated;
new set_zeroclaw_container(id, workspace_id, name, url) helper
used both for spawn (Some/Some) and teardown (None/None).
- cm-api: bollard added as a workspace dep (matches cm-sandbox's
version). New research_container module:
· connect() → uses DOCKER_HOST when set (prod's socket-proxy
at tcp://socket-proxy:2375) else the local socket. Same
pattern cm-sandbox already uses.
· container_name_for(topic_id) → "research-<uuid>-team"
(deterministic so a re-start reattaches to the same
container instead of orphaning it).
· inherited_env() → propagates ZEROCLAW_*, OPENAI_*,
ANTHROPIC_*, GEMINI_*, GROQ_* from the parent server env
(provider config + tokens), stripping the server's own
ZEROCLAW_GATEWAY_URL/WORKSPACE so the team runtime doesn't
loop back on itself. Appends ZEROCLAW_GATEWAY_PORT=42617
and ZEROCLAW_WORKSPACE=/zeroclaw-data/workspace for the
team's own listener.
· spawn(docker, topic_id, repo_host_path, state_host_path):
- inspect: if the container already exists, start it if
stopped and return its coordinates (idempotent restart).
- else create with:
image = CLAWMATES_RESEARCH_TEAM_IMAGE or
clawmates-runtime:latest
cmd = [daemon, --host, 0.0.0.0]
env = inherited_env()
mounts = repo_host_path → /workspace/repo (rw)
state_host_path → /zeroclaw-data (rw)
network = CLAWMATES_RESEARCH_TEAM_NETWORK or
clawmates_core
labels = clawmates.role=research-team,
clawmates.research.topic_id=<uuid>
- creates state_host_path first so bind doesn't ENOENT.
· stop(docker, name) → stop + remove. Idempotent on 404/304.
- start_topic wires spawn after the clone completes:
· state root = CLAWMATES_RESEARCH_WORKSPACE_ROOT / <topic> /
state
· on success, persists (name, url) on the topic row so commit
2 can look them up when constructing the executor
· every failure (docker connect, docker create/start, DB
persist) is best-effort: logs and continues. A missing team
container leaves the topic pointing at the workspace-wide
gateway URL (env), preserving prior behavior.
Deploy prerequisites (not in this commit)
- The compose stack's clawmates_server service needs bind-mounts
of CLAWMATES_RESEARCH_WORKSPACE_ROOT (e.g.
/var/lib/clawmates-research:/var/lib/clawmates-research) so
paths the server writes to are visible on the host and the
spawned team container mounts the same underlying data.
- socket-proxy's ACL must allow POST + DELETE on /containers
(already the case in prod per the audited compose file).
|
||
|
|
7984e65174 |
research_topics::create: bag 8 args into a NewTopic struct (fixes clippy)
The prior signature took (pool, workspace_id, title, description, outcome_kind, topology_kind, repo_id, created_by) — 8 args, one over the clippy::too_many_arguments ceiling and blocking CI. Refactor to a NewTopic<'a> input struct mirroring the NewLoop / NewSubTopology pattern the codebase already uses. Wizard-side additions like a repo commit branch land as struct fields instead of cascading into every call site. |
||
|
|
3465bb7a6d |
research: persist bound repo + shallow-clone on start_topic
This is the minimum viable version of the "agents actually work on
a repo" architecture. Full vision (isolated ZeroClaw container per
topic, dynamic agent provisioning inside, pause/resume, commit
gate) is real weeks of work — this closes the first, most-visible
gap so the ClawHDF5 topic can actually run against its codebase.
Backend
- Migration 0037: research_topics gets repo_id UUID (nullable, FK
to repos ON DELETE SET NULL) and repo_workspace_path TEXT for
the on-disk checkout location. Index on repo_id when set.
- research_topics::create takes repo_id: Option<Uuid>. get + list
select it and repo_workspace_path. set_repo_workspace_path
persists the path once the first clone lands.
- CreateTopicRequest accepts `repo: Option<TopicRepoRef>` — the
same denormalized shape the wizard already sends. Only repo_id
is authoritative; other fields are ignored (dead_code-allowed
so serde still deserializes the full body).
- start_topic branches on topic.repo_id. When set, it calls
ensure_repo_workspace:
· resolves repo.clone_url + repo.default_branch
· target path = CLAWMATES_RESEARCH_WORKSPACE_ROOT
// <topic_id> // repo (defaults under $TMPDIR)
· runs `git clone --depth 1 --single-branch --branch <b>` via
tokio::process. Reuses the checkout if .git already exists.
· persists the path so re-starts skip the clone
· runs `git ls-files` to sample the tree (first 60 entries,
total count reported honestly so the prompt doesn't lie
about coverage)
All best-effort — a clone failure logs but still starts the run
without repo context rather than aborting.
- build_coordinator_task takes Option<&RepoContext>. When present,
the framing gets a REPO block (slug / path / branch / file
sample) and a USING THE REPO section instructing the coordinator
to ground every recommendation in a concrete file reference and
never fabricate paths. The per-topology bodies are unchanged —
the repo guidance sits above them so it applies to every shape.
What this unblocks / doesn't unblock
Unblocks: The coordinator prompt now knows the repo exists, where
it lives on disk, and what's in it. Even without file-editing
tools wired to the checkout, the coordinator can point spokes at
concrete modules and the final artifact can reference real files.
For a spec-shaped outcome like ClawHDF5's, that's the difference
between abstract advice and a spec grounded in the actual crates.
Does NOT unblock: The agents themselves editing files, running
tests, or committing. That requires either mounting the checkout
into the ZeroClaw sandbox or exposing a new MCP tool for
repo-scoped file ops — separate follow-up.
|
||
|
|
7c1af2e070 |
research: pipeline-running signal + spinners so users aren't guessing
The prior flow was ambiguous: after hitting Start research, status flipped to "processing" and a "Submit for review" button appeared immediately with no indication that anything was actually running. Users had to guess whether the pipeline was working or stalled. Backend surfaces the truth as a signal: - new topology_runs::active_runs_for_research_topic counts queued+running runs whose research_topic_id matches - TopicDetail includes runs_in_flight: i64 alongside the existing status field, so the canvas can distinguish "pipeline still working" from "runner stalled". ResearchCanvas is now honest about state: - while runs_in_flight > 0, the header status pill grows a cyan "N runs in flight" badge with an inline SVG spinner - the stage-explainer card turns cyan-bordered and shows a "pipeline is running" hint, plus copy pointing the user at the Agents tier where each teammate's activity streams live - the "Submit for review (manual)" button is HIDDEN while any run is in flight — it's an escape hatch for stalled runs only, not the happy-path action. It reappears if runs_in_flight drops to zero but the topic is still marked processing, so a stalled runner can still be nudged along. - the canvas polls getTopic every 4s while status is processing/ publishing or runs_in_flight > 0, so the spinner + outcome swap in automatically when the pipeline completes. ResearchList sidebar: - each row's status dot becomes a spinner when the topic's status is processing or publishing, matching the canvas at a glance - the list also polls every 6s while ANY topic is active, so transitions land in the sidebar without waiting on a parent bump. The poll is gated on a derived boolean to avoid effect thrash. Follow-up: same pattern belongs on LoopsList / LoopsCanvas for loop iterations in flight — same signal (queued+running runs per loop) but not wired here. |
||
|
|
316cdbf929 |
research sidebar: delete-with-confirm per topic
Mirror the row-level delete affordance the loops sidebar already
has. Loops was wired earlier; research had a bare title-only card
with no way to remove a stale topic.
- cm-db: research_topics::delete cascades via existing FK rules
(research_topic_agents, research_publish_approvals, and the new
research_outcomes all CASCADE on topic_id; topology_runs's
research_topic_id back-ref is SET NULL so historical runs stay).
- cm-api: DELETE /api/research/{id} → 204. Idempotent.
- Frontend: deleteTopic helper. ResearchList row is now a card
with the existing title/status/outcome header plus a trash icon
that flips the card into an inline "Delete topic + all outcomes?"
confirm strip. Confirm → red Delete / gray Cancel. If the
deleted topic was selected, selection clears; local counter
bumps the list refetch without waiting on a parent.
|
||
|
|
a2d3d85ebe |
research pipeline v2: topology-aware start + persisted draft
Three connected changes that turn "Start research" from a status flip into a real pipeline that produces a reviewable artifact: - Migration 0036: adds research_topics.topology_kind (default 'hub_spoke') and a new research_outcomes table (id, topic_id, version DESC, body_md, produced_by_run_id, created_at) so each run's final synthesis is versioned and persistent. - Wizard now has a topology picker in the Outcome step — hub_spoke / pipeline / hierarchical / star_moe — with copy that steers users to the right shape (Pipeline for research → distill → analyze → implement rosters, hub_spoke for the coordinator- and-specialists default). - start_topic reads the chosen topology_kind, parses it into a cm_topology::TopologyKind, and dispatches a per-shape coordinator prompt via build_coordinator_task. Pipeline explicitly tells stage 1 not to write the final artifact and propagates a "final stage MUST emit a complete markdown document with measurable acceptance criteria" instruction downstream. The graph builder is called with the topology the user actually picked instead of hard-coded HubSpoke. - topology_worker::freeze_research_outcome fires after every successful complete(). It looks up research_topic_id on the run; if set and final_output is non-empty, it inserts a new research_outcomes row (version auto-derived server-side via coalesce(max(version), 0) + 1). Best-effort — a DB hiccup logs but doesn't fail the run. - TopicDetail now includes topology_kind and latest_outcome. ResearchCanvas swaps in the outcome's body_md (rendered as pre-wrap markdown, versioned header, produced-at timestamp) whenever an outcome exists; the original prompt collapses into an "Original prompt" <details> below so it's still one click away. Pre-run topics still show the description as before. Follow-ups still open: reject-with-revision loop feeding the coordinator, publishing → published transition + real artifact export (md / pdf), and an approvals inbox surface for reviewers. |
||
|
|
7e2b02d8bb |
research: wire start_topic to actually run the pipeline
The prior start_topic only flipped the status column — no work was enqueued. Now clicking "Start research" actually launches the assigned agents through the orchestrator. - start_topic loads the topic + its research_topic_agents, picks a coordinator (first slot with role_slot containing "coordinator"; else the first slot), and swaps it to index 0. - Builds a hub_spoke topology graph via cm_topology::build with roles = [coordinator, spoke1, spoke2, …]. hub_spoke wires edges from the hub to every spoke and back, so the coordinator can address any specialist per turn. - Assembles a coordinator prompt from the topic's title, description, outcome_kind, and a roster line for each teammate — so the coordinator knows who's on the team and what each does. - Enqueues via a new topology_runs helper enqueue_run_for_research_topic that stores research_topic_id on the run row. `topology_worker::maybe_transition_research_topic` → `notify_run_completed` already picks up on that back-ref and flips the topic processing → reviewing when the last run terminates — that path was dead code until now. - Per-agent activity streams into each claw's card for free: the orchestrator journals turn events into run_events; the existing /api/world/live SSE normalizer emits agent.reasoning.delta / agent.tool.call / agent.task.update keyed by agent id, which ClawCommandCenter is already subscribed to. The `published` terminal state is still unreached (that's the "publishing → published + artifact" step from the earlier walkthrough — separate follow-up). |
||
|
|
81a436d221 |
research canvas: wider rail + stage explainer + 409 fix
Three connected fixes from a single session's feedback:
- Structure rail widened from 60→76 px, tabs 42→58 px wide with a
bit of left padding so labels ("Research", "Visualizations") no
longer bump against the active-tab indicator strip.
- Research canvas: each stage now shows a small "STAGE · <status>"
card explaining what state the topic is in and a "Next → …" hint
describing what the primary button will do. No more guessing
which of standby/processing/reviewing/publishing means what.
- Request-publish 409 fix:
- Backend TopicDetail now includes has_pending_publish_request
(SELECTs pending_for_topic when the topic loads). Frontend
TopicDetail interface + ResearchCanvas honor the flag: when
an approval is already pending the "Request publish" button
is replaced with an amber "Awaiting reviewer approval" pill,
so double-clicks can't 409 in the first place.
- runAction() also catches 409 as a signal-of-success (the
user's intent — "queue for approval" — is satisfied by the
first attempt), refetches the topic, and lets the new
awaiting-approval card render instead of surfacing a scary
error to the user.
|
||
|
|
92e923e585 |
planner: specialists mode → single agent (backend + UI polish)
The frontend already renamed the "Specialists" tab to "Agent" and
rewrote the intro to ask for one specialist, but the backend
prompt was still telling Opus to propose "2 or 3 domain
specialists". Two-facing message got confusing outputs.
- SPECIALISTS_NOTE now says "SINGLE agent … `members` MUST contain
EXACTLY ONE entry … topology_kind='flat' … team_name reads like
a personal handle." Removes the ambiguity.
- TEAM_NOTE range aligned to "4 to 6 agents" (matches the UI hint
and intro).
- Frontend right-panel now reads "PROPOSED AGENT" for specialists
mode (was "PROPOSED TEAM") and pluralization matches the count
("1 agent" vs "3 agents"). Build button switches to "Build agent"
in that mode.
|
||
|
|
984d9a1274 |
reap: cascade orgs → companies → teams → agents
Selecting a team/company/org in the Agents sidebar used to only
delete the grouping row; the agents inside survived, ungrouped.
Not what the user wanted, and it left a trail of orphaned runtime
state (containers, .brain files, DB rows) behind.
Backend — POST /api/claws/batch-delete is now a universal cascade
reaper. Body accepts { ids?, teams?, companies?, orgs? } in any
combination. The server walks org → companies_of_org →
teams_of_company → agents_of_team, dedupes against explicit ids,
and hard-purges every unique agent (deprovision ZeroClaw runtime,
tear down sandbox container, unlink .brain/.onion files,
transactional agents::hard_purge). Group rows are deleted last; FK
cascades on team_members, company_teams, org_companies, and
loop_agents/teams/orgs clean up the join tables. Every stage
streams SSE.
Three new cm-db helpers wire the walk: agents_of_team,
teams_of_company, companies_of_org — all DISTINCT selects on the
existing join tables.
Frontend — Dashboard's Agents-tier StructureTree now sets
selectLevel="*" (was "claw"), so the Wrench → checkbox affordance
appears on org/company/team/agent nodes alike; the same-level
invariant in onToggleSelect still prevents mixed batches.
ReapProgressModal collapses to a single POST regardless of kind —
body key derived from kind — and its subtitle is honest:
"Cascading through every agent inside — permanent."
|
||
|
|
c063d4bbdd | loops: cargo fmt --all (unblock CI) | ||
|
|
a6da19430f |
loops: repo picker + agent/team/org staffing + sidebar edit/delete
Adds the missing pieces the wizard needed and the sidebar controls around it: - LoopsWizard is now a 6-step flow (identity → repo → task/topology → triggers → repeat → assign agents) plus the existing secrets card. ResearchWizard picks up the same repo step and a hard gate when the workspace has zero agents. - New LoopStaffingStep with three tabs — Individual / Team / Organization — that mix freely per loop; selections persist via new loop_agents / loop_teams / loop_orgs join tables (0035 migration), each cascading on loop_id so hard-delete stays a single-row DELETE. - Backend CreateLoopRequest / UpdateLoopRequest accept the three lists and apply_staffing does a transactional replace-all; list_loops / get_loop hydrate the lists via a flattened LoopWithStaffing response. - LoopsList sidebar gains per-row enable/disable, edit (reopens the wizard prefilled with the current loop, PATCHes on submit), and delete with an inline confirm. - NoAgentsGate blocks launching a loop or research topic from a workspace with no roster; the sidebar `+` buttons also disable with a tooltip pointing at the TEAM tier. Not yet wired: the run driver still fills role slots from the workspace-wide pool; teaching enqueue_iteration to prefer loop_agents/loop_teams/loop_orgs is a follow-up. |
||
|
|
637e1bdd69 |
repos: sidebar actions (sync/edit/remove) + edit modal
Sidebar: - Each connection header now has three inline icon buttons: Sync now (spins while in flight), Edit (opens the modal), Remove (opens an inline confirm strip). Removes cascade repos via ON DELETE CASCADE. - The connection's last_sync_error surfaces as a red inline banner under the header — no more 'error status with nowhere to see why'. - Sync is POST /api/repos/connections/:id/sync (already existed); after either sync or delete the sidebar re-fetches so state stays consistent. Edit modal (RepoConnectionEditModal): - Loads GET /api/repos/connections/:id, pre-fills owner/base_url/label - PATCHes only the fields that actually changed; empty string on a Some(&str) field sends explicit null so the backend clears it - Sync-now + Remove reachable from inside the modal too - Rotating the token is out of scope: the modal says as much and points the user at delete + re-create through the wizard (the broker doesn't expose an update path, and rotating in place would require duplicating the whole broker->store_secret flow here) Backend: - GET /api/repos/connections/:id — same ConnectionSummary shape - PATCH /api/repos/connections/:id — owner/base_url use Option<Option<T>> double-nesting so 'omit = leave alone' and 'null = clear' round-trip distinctly through serde - repo_connections::update with COALESCE-per-field so the SQL matches the double-Option semantics without an OR-chain per field |
||
|
|
977a1233af |
cm-secrets: chmod 0666 broker socket after bind
Fixes 500 on any broker-touching route (POST /api/repos/connections, POST /api/apps, the OAuth callback) when broker + server run under different UIDs — which is exactly the prod topology on gw-04 (broker uid 10001, clawmates-server distroless nonroot uid 65532). Linux Unix socket connect(2) requires read+write on the socket file, and the default bind mode 0755 gives 'others' r-x only. Widen to 0666 after bind. The broker socket only lives inside the shared broker_run volume — two containers mount it, nothing else on the host can see it — so widening is safe. If set_permissions is a no-op on the target filesystem (abstract sockets on some kernels), we log and continue instead of failing serve(). |
||
|
|
e858a7f92f |
repos: Gitea provider (first-class) — sync + wizard default
Fleet's Gitea (git.redclaw.dev) hosts most of this workspace's repos,
so Gitea gets the same inline sync treatment GitHub already had.
Backend sync_gitea:
- base_url is required — Gitea has no shared 'gitea.com'; we accept
either the instance root (auto-appends /api/v1) or the fully-formed
API base if the user already included the suffix
- /orgs/{owner}/repos when owner set, /repos/search when not (with the
{data: [...], ok: bool} envelope Gitea wraps that endpoint in)
- 404 with an owner surfaces as 'org not found or PAT lacks access',
same UX as GitHub
- 50/page, capped at 20 pages (~1000 repos); short page terminates
- upsert_gitea_repo tolerates the small field-name differences
(stars_count vs stargazers_count, owner.login vs owner.username on
older versions)
Frontend wizard:
- Gitea listed first — matches the workspace's actual usage
- Default provider selection is now gitea
- Token-input placeholder tailored per provider (Gitea's is
'Settings → Applications → Generate New Token (repo)')
GitLab still returns 'not yet supported' — that's the next follow-up.
|
||
|
|
6d087bf537 |
repos: backend — schema, /api/repos routes + GitHub sync provider
Migration 0034: two tables. repo_connections carries the workspace's per-provider config (owner, base_url, label, last_synced_at, last_sync_error) and points at an app_connections row for the PAT. repos is the per-connection cache with (connection_id, external_id) unique so upsert is idempotent across re-syncs. Cascading deletes clean up cleanly on connection removal. cm-secrets grows a FetchAuthorized op — GET with the stored PAT injected as bearer, returns status + JSON body without ever exposing the credential to cm-api. This is the least-privilege door for read-only provider APIs (list repos), distinct from the InvokeHttp path that still requires a single-use approval grant for outbound writes. cm-api::routes::repos wires: - POST /api/repos/connections (broker store_secret + insert both rows + initial sync + mark_synced) - GET /api/repos/connections - DELETE /api/repos/connections/:id - POST /api/repos/connections/:id/sync - GET /api/repos (500 cap, newest provider_updated first) - GET /api/repos/:id (full detail incl. clone_url + html_url) GitHub provider inline for v1 — paginated pull of /orgs/:owner/repos (when owner set) or /user/repos (when absent), 100/page, capped at 20 pages (~2k repos) to keep first-sync latency bounded. Non-2xx surface back to the caller as sync_error; parse failures are best-effort per repo (skipped, logged, don't abort the batch). Gitea + GitLab providers land in a follow-up — mostly URL swap + response-shape adapter. |
||
|
|
806ba869e5 |
teams: ephemeral lifecycle for Scheduled + Triggered planner modes
Migration 0033: adds teams.lifecycle ('permanent' | 'ephemeral') and a
topology_runs.team_id back-ref with a partial index for the sibling-in-
flight check.
cm-db repo:
- teams::insert_team_with_lifecycle (insert_team keeps the permanent default)
- topology_runs::enqueue_run_for_team (populates team_id)
- topology_runs::check_ephemeral_teardown — atomic SELECT that only
returns Some when the team is ephemeral AND no siblings are still
queued/running; carries the workspace + bound claw ids for cleanup.
cm-api:
- topology_worker post-terminal hook maybe_teardown_ephemeral_team
runs deprovision_claw on each bound claw (best-effort; failures log
but don't block Postgres deletion), then hard_purge each agent row,
then delete_team.
- routes::teams::build_team_with_lifecycle (build_team keeps default);
run_team enqueues with team_id.
- planner ScaffoldRequest gains mode; lifecycle_for(mode) sets the team
to ephemeral for scheduled + triggered, permanent otherwise.
Frontend MasterPlannerModal passes mode in the scaffold payload so the
backend can derive lifecycle without duplicating the mode taxonomy.
Tests: 3 new (returns claws when no siblings, holds when siblings queued,
ignores permanent teams). 10/10 topology_jobs green; workspace clippy
--tests clean.
|
||
|
|
b0acdfd987 |
master planner: wire user-locked topology_kind through chat + scaffold
Frontend: send topologyKind to /api/planner/chat so the planner's user prompt gets a USER-LOCKED TOPOLOGY block telling Opus to use it verbatim. On buildTeam, override proposal.topology_kind with the user's pick (belt-and-braces — if the planner ignored the lock, we still ship the right shape). Proposal chip renders the effective kind in a lavender tint when it was overridden, with a hover title showing what was replaced. Backend PlannerChatRequest gains an optional topology_kind. Empty / absent = planner picks. Not honored for 'swarm' mode (swarm planner doesn't take a topology kind). |
||
|
|
987f4f0e84 |
master planner: add 'Team' mode + size bands per mode
Modes now: specialists (2–3 domain experts, deep prompts) · team (4–8 balanced roles, coordinator + complements) · swarm (10+ workers, self- verifying loop) · scheduled (ephemeral, cron/one-shot) · triggered (ephemeral, webhook). Backend planner_system_for() gains a TEAM_NOTE using PLANNER_SYSTEM; specialists / scheduled / triggered notes are rewritten to bake in the size + ephemeral guidance. Swarm's system prompt now targets task_count>=10 explicitly. Frontend MODES / INTRO copy match. Chat-preserving switchMode from the prior commit handles the new mode transparently — no state-plumbing changes needed. |
||
|
|
8a4e222aec |
research: auto-transition processing → reviewing on last run
Hooks the topology_worker's post-terminal path into a new notify_run_completed repo helper that atomically transitions the topic processing → reviewing when the completed run has research_topic_id set AND no siblings for that topic are still queued or running. Guarded on status='processing' so a retry, a re-fire, or a topic already past processing are all no-ops. Best-effort at the worker; DB hiccups are logged and never fail the run. The manual /submit-review endpoint stays as an escape hatch for topics that end up parked in processing with nothing to complete (updated the doc comment). |
||
|
|
6d1dda6197 |
loops: iteration timeline — GET /api/topology-runs?loop_id=X
Extend the topology-runs list route with an optional loop_id filter that returns iterations for a single loop, newest-iteration-first. Adds the iteration and finished_at columns to the summary (skip-null on the JSON so compares stay compact). Backed by list_by_loop in the repo, which uses the existing topology_runs_loop_idx partial index. LoopsCanvas fetches the runs in parallel with the loop detail and renders an iteration timeline card (iteration #, status pill, start time, duration, run id prefix) between the graph section and the actions row. |
||
|
|
973eeb272e |
research: publish approval gate + explicit state transitions
Fourth commit of the Research + Loops arc. Completes the state machine
for research topics with the publish approval gate the spec asked for.
Migration 0032 — research_publish_approvals
Dedicated small table (id, workspace_id, topic_id, requested_by,
status, decided_by/at, created_at). Keeping it separate from the
existing `approvals` table (0001) because that one is tightly coupled
to gated tool calls inside an agent run — session_key + run_id +
action_type + category + payload + preview + requested_by_agent, all
NOT NULL. Forcing those nullable would ripple through cm_safety;
cleaner to give publish approvals their own two-transition state
machine.
New endpoints
POST /api/research/:id/submit-review processing → reviewing
(v1 caller-driven; the
orchestrator hook comes
when we wire actual runs)
POST /api/research/:id/request-publish creates a pending
approval. Rejects with
409 if the topic already
has one open.
GET /api/research/publish-approvals list workspace's pending
POST /api/research/publish-approvals/:id/approve flips approval to
approved + transitions
the topic
reviewing → publishing
(which stamps
published_at)
POST /api/research/publish-approvals/:id/reject stays in reviewing; new
requests allowed
The approve/reject write is an atomic UPDATE ... WHERE status = 'pending';
the decide() repo function returns whether the caller won the race so
concurrent double-approves collapse to a single topic transition.
State machine after this commit:
standby ─POST /start─▶ processing ─POST /submit-review─▶ reviewing
─POST /request-publish + approve─▶ publishing ─(future: artifact
assembly)─▶ published
|
||
|
|
4b48c521eb |
loops: backend routes + repo + cron scheduler + HMAC webhook
Third commit of the Research + Loops arc. Lights up loops as durable
recurring topology executions:
GET /api/loops list workspace's loops
POST /api/loops create — returns webhook_token +
signing_key ONCE when webhook trigger
is enabled; never exposed again
GET /api/loops/:id detail
PATCH /api/loops/:id update definition
DELETE /api/loops/:id delete
POST /api/loops/:id/run trigger one iteration NOW
POST /api/loops/:id/enable set enabled=true
POST /api/loops/:id/disable set enabled=false
POST /webhooks/loops/:token public; HMAC-SHA256-verified
Scheduler (cm_runtime::spawn_loop_scheduler) wakes every 10s, queries the
partial index on (next_fire_at) for due loops, enqueues one topology_runs
row per fire with loop_id + iteration + parent_run_id chained back to the
previous iteration. Uses croner via the existing scheduling::next_occurrence
helper. Missed windows fire ONCE and skip the backlog — next_fire_at is
always computed strictly AFTER now(), so a late scheduler doesn't drain a
buildup.
Webhook signatures follow the same pattern as the Stripe billing webhook
(HMAC-SHA256 with constant-time hex compare). Token + signing key are
24-byte OS-RNG values; the URL uses base64-url for the token, and the
signing key is base64-std. Both surface exactly once at create time.
All three fire paths (scheduler, immediate-run, webhook) funnel through
`cm_db::repo::loops::enqueue_iteration` so the invariants stay in one
place. `iters` repeat policy is enforced by the scheduler tick; `until`
and `on_completion` land with the orchestrator hook in commit 4.
Adds cm-llm as a direct cm-api dep, getrandom for the webhook material
generator, and wires the scheduler spawn into the server binary alongside
the resume sweeper and outbox drainer.
|
||
|
|
4ffcb3d652 | chore: cargo fmt --all — clean up research.rs formatting | ||
|
|
fb047879c2 |
research: backend routes + repo + wizard refine (behind /api/research)
Second commit of the Research + Loops arc. Builds on 0030 by lighting up
the container CRUD, the agent-slot attach/detach, the standby→processing
transition, and the wizard's one-shot LLM refine.
GET /api/research list workspace's topics
POST /api/research create (accepts wizard output)
GET /api/research/:id detail (topic + attached agents)
PATCH /api/research/:id update non-status fields
POST /api/research/:id/agents attach agent (idempotent)
DELETE /api/research/:id/agents/:agent detach
POST /api/research/:id/start standby → processing
POST /api/research/wizard/refine one-shot LLM refine
The refine endpoint accumulates the workspace's default LLM provider's
stream into a JSON object (`{title, description}`) with the tight system
prompt at the top of the module. Same provider routing as agent runs
(via runtime.provider()), so a workspace already using GLM/Kimi gets it
for free.
State machine's remaining transitions (processing → reviewing on last
run_completed; reviewing → publishing via approvals gate) land with the
orchestrator hookup + approvals extension. Publish approval and loops
are separate commits still to come.
Adds cm-llm as a direct cm-api dep (previously only pulled transitively
via cm-runtime) so the refine endpoint can build a ChatRequest. Uses
sqlx::query! for compile-time verification; .sqlx cache generated on
morpheus against a fresh migrated DB.
|
||
|
|
696d8237fe |
tests(warm_pool): seed a real agent so upsert actually writes the row
Root cause of the "reuse must not drain the pool" flake: the test called
`manager.exec(AgentId::new(), ...)` with a random UUID that had no agents
row. `agent_containers::upsert` uses `INSERT ... FROM agents WHERE a.id = $1`,
which silently inserts zero rows when no agent matches — so the "assigned"
sandbox never persisted to the DB. The next exec's reuse lookup returned
None, fell through to the provision branch, and popped from the warm pool
instead of reusing the assigned sandbox. When the warmer hadn't refilled by
the time we asserted, pool_size == 1 instead of 2.
Seed a workspace + owner + agent up front (mirrors soak.rs's setup). Now
upsert commits a real row, the second exec hits the reuse branch, and the
pool stays whole — the test asserts what its name claims.
The tolerant-health-check change in cm-runtime (
|
||
|
|
04f8302871 |
cm-db: threads — drop unnecessary .iter().copied() on participants slice
Inherited from the a2a merge. Iterating `&[AgentId]` directly yields `&AgentId`; `AgentId::as_uuid(&self)` works fine on that, so the extra `.iter().copied()` was pure noise and tripped clippy's `unnecessary_to_owned` under `-D warnings`. |
||
|
|
15cffba2d7 |
cm-runtime: don't drain the warm pool on a flaky sandbox health check
`driver.health()` can return `Err(_)` on transient docker daemon hiccups (connection reset, timeout mid-inspect, daemon busy). exec() used `.unwrap_or(false)` which treated Err as "dead" and: 1. Destroyed the agent's assigned sandbox. 2. Fell through to the warm pool and popped one, draining it by 1. 3. Provisioned a fresh assigned sandbox on top. Under CI load — where several test processes hit the docker daemon concurrently — this fired as the warm_pool.rs:72 "reuse must not drain the pool" flake. The test asserts pool_size == 2 after a reuse; when health flaked, the reuse turned into a drain-and-refill and the assert raced the warmer. Match only a CONFIRMED `Ok(false)` (container dead or 404). On Err, assume alive; if it really is dead, the subsequent exec surfaces the error with a clear message instead of silent sandbox destruction and warm-pool drainage. |
||
|
|
ebb5b5780b |
chore: cargo fmt --all — clean up a2a merge's fmt violations
The a2a merge (
|
||
|
|
3b382d659a |
broker: make Postgres pool size configurable, bump default 5 → 8
Under concurrent door executions the broker was queueing on its 5-connection pool, adding latency to tool calls that fanned out from the same run. Bump the default to 8 (one connection per concurrent door before queueing) and expose CLAWMATES_BROKER_POOL_SIZE so the fleet can tune it up as the load grows without a rebuild. |
||
|
|
e2c82d9a93 |
cm-api: fleet — bound PTY sink channels so a stalled browser can't OOM the hub
pty_sinks was an unbounded mpsc, which means a runaway PTY (say `cat /var/log/huge`) with a browser that stopped consuming would accumulate megabytes in the hub's memory indefinitely. Switching to a bounded mpsc::Sender means the send fails when the buffer is full; the NodeHub then closes that session cleanly instead of holding output forever. Signal sinks stay unbounded — they carry small control frames. |
||
|
|
d2b1f0569e |
cm-api: quota — enforce max_active_runs on run enqueue
Adds a Quota.max_active_runs ceiling (queued + running topology runs at once, per workspace) to stop a single workspace flooding the shared queue. Free tier: 10, Pro: 25, Team: 100. Enforced at every /run enqueue site: run_org, run_company, run_team, and the webhook trigger. Webhooks return 429 rather than 402 so external callers can back off — the guard is what stops a leaked webhook token from being weaponized into a queue flood. A single team run also spawns a tier-tree of children, so the practical cap grows with the topology — this counts the outer runs, not every step. |
||
|
|
51d501e36e |
clawmates-node: bound write.send with 10s timeout so half-open TCP can't wedge daemon
Under a half-open TCP (server side closed, client OS still buffering writes), the daemon's write.send() inside the tokio::select! branch blocks forever. tokio::select does not preempt a running future, so the whole loop freezes — idle_tick never gets to check last_rx, no 'channel ended' log ever fires, and the daemon silently spins on a dead socket for hours. Observed on architect Jul 5 2026: daemon connected at 12:59:13, heartbeats worked for ~2.5 min, then went silent. TCP session showed ESTABLISHED on the node side, no ESTABLISHED on the gateway side. Restarting the daemon 'fixed' it — but it re-hung within minutes. Fix: wrap both write.send() call sites in a tokio::time::timeout of WRITE_DEADLINE=10s. If a send stalls past that, we log and return Ok() to trigger the main-loop reconnect. Short enough that it fires long before the 40s read-idle would (which was our only escape hatch and never triggered because the loop was frozen). |
||
|
|
d0d8e7fd40 |
Merge feat/a2a-rooms-delegation-ingress into main
Brings a2a rooms, delegation, and A2A ingress work back into main. Prod's
DB has migrations 0026 (group_rooms) and 0027 (a2a) applied from an
earlier hand-tagged fleet21 build cut from this branch, but main never got
them — so the CI-built server image from main refused to start against
prod's DB with "migration 26 was previously applied but is missing".
Landing the branch closes that gap: main + prod DB now share the same
migration state, so images built from main can safely roll onto gw-04.
Included:
- 0026_group_rooms.sql / 0027_a2a.sql — align main with prod's schema
- cm-api routes/a2a.rs + mcp_door.rs updates — A2A ingress and MCP door
- cm-runtime tools/delegate.rs + tools/chat.rs — delegation + N-way rooms
- frontend TeamObserver + FleetPanels + AgentObserver updates
- taxonomy.ts — delegation + A2A signals in the live world feed
Not included (still WIP on the local checkout):
- 0028_backfill_on_delete.sql / 0029_hot_query_indexes.sql
- broker pool + team-run quota changes
- Dashboard.tsx UI rename (Large World → Visualizations)
- Node write.send timeout (aaab663) — separate concern
The SSE resume fix (
|
||
|
|
202e8535f0 |
fix(approvals): await resume_run so channel is ready before response returns
The `decide()` helper (used by both approve and reject) wrapped `runtime.resume_run(ready).await` inside `tokio::spawn`, so the handler returned 200 to the client before the resume had even started. When a client (or test) immediately reconnected to /api/gateway with resumeFrom, the run's broadcast channel didn't exist yet — `runtime.subscribe(run_id)` returned None, the journal was still empty (no new events yet), and the SSE stream closed with zero events. Symptom: approvals_api tests flaky in CI (`unwrap on None` at line 282 of reject_over_http_executes_nothing, sometimes `missing step_finished` on the accept path). resume_run's *own* awaited portion is only the setup — claim_resume, checkpoint load, and open_channel. It internally spawns the long-running multi-step work. Awaiting it inline means we wait milliseconds for setup, then return. By the time the client reconnects, the channel exists, so subscribe() finds it and live-tail works. The sweeper still handles the crash-between-decision-and-resume case for durability. |
||
|
|
85b0e1ac33 |
feat(observe): surface delegation + A2A in the live world feed via audit poll
Edge-initiated inter-agent events (gated delegation, A2A ingress) bypass the
run loop, so the world SSE now polls the append-only audit log (cursor on the
BIGINT id, seeded to max on first pass) and emits:
- delegation.invoked -> agent.delegate {fromAgentId,toAgentId,toName,task}
- a2a.invoked -> a2a.invoked {agentId}
TeamObserver renders agent.delegate as an A->B handoff in the team timeline;
adds the agent.delegate taxonomy type. Reliable live (the synthesized-run path
never streamed — active_runs + first-sight cursor jump skip it).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
||
|
|
7589f62aca |
feat(a2a): External access panel + GET /api/a2a/settings
Adds GET /api/a2a/settings (enabled + publicBaseUrl) and an A2ASection in the Fleet overview: enable/disable A2A for the workspace, mint/list/revoke external bearer tokens (token shown once), and the discovery URL. Claw/skill publishing is configured per-claw (follow-up). Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
cbfa0ff24f |
feat: agent-to-agent platform on ZeroClaw 0.8.2 — rooms, delegation, A2A ingress
Builds on the v0.8.2 runtime. Four workstreams, all behind the §15 MCP door:
- Group rooms (Phase 1): migration 0026; N-way threads repo with a DM/room
count-guard; chat.send {room} + room.create/invite/leave tools; RoomMessage
-> room.message SSE; /api/claw-chat/rooms* APIs; Observer room badge.
- Per-claw door identity: door caller_agent resolves the X-ZeroClaw-Agent
header (set by the fork) to the specific claw, falling back to roster[0].
- Gated delegation bridge (Phase 3): clawmates__delegate door tool drives a
sibling via the existing /ws/chat ZeroClawDriveExecutor (not A2A); self-deny,
per-workspace hourly budget, audit trail, untrusted-banner result. Native
in-daemon delegation stays off (it would bypass the door).
- A2A tenant ingress (Phase 2): migration 0027 (workspace_a2a + a2a_tokens);
runtime_provision enable_a2a_server/publish_claw; routes/a2a.rs tenant-aware
proxy (per-workspace tokens, injected internal bearer, daemon stays internal,
cards URL-rewritten to the cm-api edge); a2a.invoked taxonomy.
Tests: cm-db room repos, cm-runtime chat tools, door units. sqlx cache updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
||
|
|
1a531e91cd |
Fleet tools: add Rust version card (probe rustc + nightly latest + rustup update)
Adds Rust to the per-node dev-tool cards: daemon probes rustc → reports 'rust'; nightly checker fetches latest stable from GitHub rust-lang/rust; GET endpoint maps rust→Rust (after docker); one-click update runs 'rustup update stable'. Frontend is data-driven (no change). $HOME/.cargo/bin added to probe candidate dirs. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |