94828ed8877f505d4c7542fe169e7d2dfd9298ff
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
94828ed887 |
Fleet: robust real-time connectivity + mosh-inspired reconnecting terminal
Nodes flapped online/offline and the terminal died on the first blip. WebSockets are the right transport (outbound, NAT-friendly); the fixes harden around it. Server (cm-api): - Anti-clobber connection epoch: a reconnecting daemon gets a fresh epoch; a stale run_channel's teardown only clears the hub + sets offline if it still owns the slot — so a lingering old channel can't flip a live reconnection offline (the main false-offline cause). - WS keepalive: run_channel now pings every 15s and tears down if no inbound frame (incl. pong) for 35s — dead links detected in seconds, not minutes. - Staleness sweeper backstop: spawn_node_sweeper (8s tick / 20s window) wired in clawmates-server, so a vanished node goes offline within ~28s even if its channel hangs (mark_stale_offline was defined but never called). Daemon (clawmates-node v0.3.0): - Heartbeats off the select thread (dedicated thread owns System + blocking docker/tailscale/disk CLIs) so a slow op never starves heartbeats/pongs. - Each handle_frame runs on its own task; added a 40s inbound idle deadline so a half-open socket triggers a reconnect. Frontend: - useNodes streams /api/nodes/live (SSE push) instead of a 3s poll; isLive() derives online from lastSeen freshness (<15s) so a transient column flip never shows a healthy node down. - Node terminal: clean auto-reconnect loop (re-mint ticket -> reconnect -> tmux re-attaches and redraws the live screen = mosh-style snap-to-state over TCP), replacing the [disconnected] dead-end. Mosh evaluated: harvest principles (session/transport decoupling, snap-to-state, already given by tmux), don't adopt — UDP is incompatible with our browser+CF+NAT topology and it's GPLv3. Removed temporary terminal debug traces + /api/debug route. Verified: node holds steadily online (heartbeat 1-3s, no flap) and goes cleanly offline when the daemon stops. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
b840de7c3b |
Infra: operator-console redesign (per the Infrastructure design package)
Refactor the Infrastructure tier into the dark operator console from the design package, wired to the real fleet backend (no new plumbing): - InfraNav: page-scoped left nav grouped RUN HERE FIRST (Local & Tailscale, with online count + sub-items) / NEXT (Cloud providers, SOON) + coral "Connect a host" button; coral spine on the active view. - FleetConsole: center reads top-to-bottom — kicker → "Your fleet" → 5 stat tiles (HOSTS/ONLINE/MEMORY/STORAGE/CONTAINERS) → Tailscale device card → LOCAL HOSTS host cards + add-host dashed tile. Cloud view = how-it-works + SOON. - Host cards (design system): status dot (blink when online) + hostname, IP/version line, CPU/RAM meter bars (cyan→amber→coral by load), disk/load/ctrs mini-row, and a footer that's a copyable `ssh <node>` target (online) OR a "waiting for daemon" spinner (pairing) — PLUS a Terminal button that opens the in-app shell. - FleetPill (top bar, N/M hosts online) + FleetStatusBar (ambient: daemon · tailnet · WSS · sandbox). Per the chosen reconciliation: kept the cloud-apps computer pull-out as the right column, the connect-host wizard as a modal, and the in-app terminal. Dashboard infra branch rewired (InfraNav + FleetConsole + status bar; dropped the old InfraStage/InfraConsole split). cm-blink/spin/cm-fade keyframes already existed. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
cf6c331b02 |
Fleet: node hostname/IP on register + node terminal in the infra computer
Hostname/IP: - Daemon reports the machine's hostname (sysinfo) + primary outbound IPv4 on each heartbeat. migrations/0021 adds nodes.hostname/local_ip; cm-db heartbeat stores them; node JSON exposes them. Cards now title on the real hostname (falling back to name) + show the IP, instead of the "New node" placeholder. `name` stays user-overridable (rename). Terminal moved into the pull-out computer (no more per-card modal): - New infra computer app NodeTerminalApp (computer/apps/infra) — xterm bridged to a node's host shell over the node control channel, filling the app window (mirrors the agent Terminal's layout + ResizeObserver). Added "terminal" to the INFRA_CATALOG grid; a ?node= panel param targets a specific node (picker when unset). Clicking Terminal on a node card now opens the infra computer to that node's shell instead of a separate full-screen window. Deleted NodeTerminal.tsx. Rebuilt + re-hosted both daemon binaries (hostname change). Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
fb59378aa2 |
Fleet P2b: run agent sandboxes on connected nodes (RemoteDriver + placement)
Agents can now provision their sandbox on a connected fleet node instead of the gateway host. Local stays the strict default, so existing agents are byte-for- byte unaffected until explicitly placed elsewhere. Security parity: the daemon links the REAL cm-sandbox DockerDriver and runs the typed container ops (sb_provision/sb_exec/sb_destroy/sb_health/sb_list) through it — identical hardening (cap-drop ALL, seccomp, no-net, read-only, non-root) to local sandboxes. cm-sandbox spec types are now Serialize/Deserialize so the spec crosses the channel. - cm-api: RemoteDriver (impl SandboxDriver over the node channel) + HubDriverProvider (impl cm_runtime::NodeDriverProvider, hands out a driver only for connected nodes via a sync online set) + NodeHub.call/is_connected. AppState.with_node_hub so the hub is shared with the placement provider. - cm-runtime SandboxManager: driver_for(node_id) routes by the recorded agent_containers.node_id (local default = existing driver, identical path); placement_node() reads the workspace setting and falls back to local if the node is offline; exec/release route accordingly. NodeDriverProvider trait. - DB: 0020_workspace_placement + repo (for_agent/get/set/clear). - main.rs: build the NodeHub first; inject HubDriverProvider into the agent manager + share the hub with AppState. - API+UI: GET/PUT /api/fleet/placement + a "Run agents on: Local / <node>" selector in the Fleet overview. Note: a node must be able to pull the agent image (the daemon docker-pulls it); interactive PTY for agent containers on remote nodes is not wired (Terminal app stays local) — the in-dashboard node shell already covers host access. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
33aa9c0693 |
Fleet P2b: node sandbox-readiness check (hardened workload on a node)
Proves a connected node can host hardened agent workloads end-to-end, without
touching the agent run loop (zero blast radius on existing agents).
- Daemon: typed `sb_check` op — pulls a tiny image and runs it fully locked down
(cap-drop ALL, no-new-privileges, no network, read-only rootfs, non-root,
memory/pids caps), then tears it down. Fixed command; nothing caller-supplied
runs (preserves the exec-hardening invariant).
- cm-api: NodeHub.sandbox_check + POST /api/nodes/{id}/sandbox-check.
- UI: a shield "sandbox check" button on each online node card streams the
result (✓ SANDBOX READY + container id/uname).
This validates the full provision→run→destroy mechanism on nodes. The remaining
P2 work — wiring real agent deploys to auto-place onto nodes — is its own
subsystem (a RemoteDriver reusing the local DockerDriver for security parity,
agent-image distribution to nodes, and node-routing in SandboxManager) and is
best done as a focused pass; it is intentionally NOT bundled here to keep the
core agent path untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
||
|
|
f5f96508eb |
Fleet P2a: in-dashboard remote terminal (PTY over the WSS channel)
You can now open a real shell on any connected node from the dashboard — the
daemon spawns a host PTY and streams it over the existing outbound control
channel (no inbound port, no Tailscale brokering needed).
Daemon:
- portable-pty host shell sessions: pty_open/pty_in/pty_resize/pty_close ops; a
reader thread streams base64 pty_out frames. Outbound frames now funnel through
one mpsc channel so PTY output and heartbeats interleave.
cm-api NodeHub:
- per-connection pty_sinks + sid multiplexing; open_terminal/terminal_input/
terminal_resize/terminal_close; in-memory single-use terminal tickets (the
browser WS can't carry a bearer, and the session is instance-local anyway).
- routes/nodes.rs: POST /api/nodes/{id}/terminal/ticket + GET .../terminal/ws
(bridges browser xterm <-> node PTY: binary = keystrokes, text = resize).
Frontend:
- NodeTerminal xterm modal (reuses the agent Terminal's xterm setup); a Terminal
button on each online node card opens a shell.
This proves the bidirectional streaming-over-channel mechanism the RemoteDriver
will reuse. Remaining P2: RemoteDriver + placement (run agent workloads on nodes).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|
||
|
|
7332d69f8a |
Fleet P1: BYO Tailscale + network metrics, Tailscale SSH, exec hardening
Security hardening: - The gateway no longer sends arbitrary shell to nodes. The WSS exec op is replaced by a typed `verify` op the daemon runs itself (fixed host+docker check); future container ops are typed too. cm-api NodeHub.verify() + the daemon's handle_command only dispatches vetted ops. BYO Tailscale: - migrations/0019_workspace_tailscale.sql + cm-db fleet_tailscale repo (store the user's Tailscale API key + tailnet, server-side only). - cm-api routes/tailscale.rs: POST/GET/DELETE /api/fleet/tailscale + GET /api/fleet/tailscale/devices (proxies api.tailscale.com device list). - Daemon: --tailscale-authkey → `tailscale up --authkey … --ssh` (enables Tailscale SSH for keyless user access); else `tailscale set --ssh=true`. Reports its tailscale IP (already). UI: - Fleet overview gains a Tailscale section: connect (key+tailnet) + live tailnet device status (online/last-seen/IP/os). Node cards show a copyable Tailscale SSH target (ssh <ip>). Remaining: P2 — RemoteDriver + placement (run agents on nodes) and the in-UI remote terminal (PTY proxied over the WSS channel). Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> |
||
|
|
2bdd0a23e8 |
Fleet P0: node registry + daemon + health + connect-host wizard
Users can connect their own local-hardware nodes into a fleet. Each node runs a
new Rust daemon that dials home over an outbound WebSocket, reports host health,
and runs commands we send.
Backend:
- migrations/0018_fleet_nodes.sql: nodes + node_health tables + agent_containers
(node_id, workspace_id) index. cm-domain NodeId.
- cm-db repo/nodes.rs: create/auth/list+health/get/heartbeat/set_status/delete
(unchecked sqlx, no .sqlx regen).
- cm-api fleet.rs NodeHub: live daemon channels (node_id→sender) + the WS channel
runner (heartbeat→DB upsert, exec request/response framing). routes/nodes.rs:
POST /pair, GET /nodes, SSE /nodes/live, POST /{id}/exec-test, DELETE /{id},
WS /nodes/agent (token-auth). Wired into AppState + router.
Daemon (new crate crates/bins/clawmates-node):
- sysinfo host metrics (cpu/mem/pressure/swap/disk/load/containers), outbound WSS
dial + reconnect, heartbeat loop, exec command handling, tailscale-ip probe.
install.sh convenience installer.
Frontend:
- Fleet sidebar item + FleetOverview + LocalHardware node-health cards (live via
/api/nodes, 3s poll) + ConnectHostWizard (install → verify connection →
exec-test). InfraStage dispatches fleet/local; default selection = fleet.
Deferred: P1 (BYO Tailscale + network metrics), P2 (RemoteDriver + placement so
agents actually run on connected nodes).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
|