Fleet: actionable executions — rules engine + metrics-aware placement (Phase 2)
ci / gates (push) Failing after 16s
ci / rust (push) Has been skipped
ci / sandbox-k8s (push) Has been skipped
ci / frontend (push) Has been skipped
ci / e2e (push) Has been skipped

Turn the Beszel-tapped metrics into a self-managing loop.

- migration node_rules (workspace/node-scoped: metric op threshold, for_seconds,
  action JSONB, last_fired).
- cm-db: repo/node_rules.rs (CRUD + list_enabled); node_metrics::eval_all merges
  Beszel + heartbeat scalars per node + a headroom() heuristic; nodes::status_of;
  heartbeat now PRESERVES a `draining` status across heartbeats (so a cordon sticks).
- cm-api: node_rules.rs evaluator (spawn_evaluator, 20s) — when a metric condition
  holds for the rule's window it fires drain / undrain / alert (in-memory sustained
  + cooldown tracking, modeled on the node sweeper); routes/beszel.rs rules CRUD
  (GET/POST/PATCH/DELETE /api/fleet/rules); spawned in clawmates-server.
- cm-runtime: placement_node() is metrics-aware — a `draining` node stops receiving
  new agent sandboxes (falls back to local), so the drain rule is actionable.
- frontend: FleetRules section in the Local view — build rules (node · metric · op ·
  threshold · duration → action), toggle/delete, with fired-history.

The loop: hot/overloaded node → rule drains it → placement avoids it → recovers →
undrain rule brings it back. Deployed; node_rules migration applied.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
This commit is contained in:
Omar Sobh
2026-06-25 23:43:49 -07:00
co-authored by Claude Opus 4.8
parent 36a227566b
commit c94784bab2
11 changed files with 546 additions and 3 deletions
+68 -1
View File
@@ -2,9 +2,10 @@
//! per node: scalar columns for the fleet cards + rules engine, plus a JSONB blob
//! for the per-node monitor page.
use cm_domain::NodeId;
use cm_domain::{NodeId, WorkspaceId};
use serde_json::Value;
use sqlx::{PgPool, Row};
use uuid::Uuid;
use crate::DbError;
@@ -57,6 +58,72 @@ pub async fn upsert(pool: &PgPool, node_id: NodeId, m: &NodeMetrics) -> Result<(
Ok(())
}
/// A node's current evaluatable scalars (Beszel-tapped, falling back to the
/// heartbeat health), for the rules engine + metrics-aware placement.
#[derive(Debug, Clone)]
pub struct EvalRow {
pub node_id: NodeId,
pub workspace_id: WorkspaceId,
pub status: String,
pub cpu_pct: Option<f64>,
pub mem_pct: Option<f64>,
pub disk_pct: Option<f64>,
pub gpu_pct: Option<f64>,
pub temp_max: Option<f64>,
pub load1: Option<f64>,
}
impl EvalRow {
/// Look up a metric by rule name.
pub fn metric(&self, name: &str) -> Option<f64> {
match name {
"cpu_pct" => self.cpu_pct,
"mem_pct" => self.mem_pct,
"disk_pct" => self.disk_pct,
"gpu_pct" => self.gpu_pct,
"temp_max" => self.temp_max,
"load1" => self.load1,
_ => None,
}
}
/// Free headroom heuristic (higher = more capacity) for placement ranking.
pub fn headroom(&self) -> f64 {
let used = self.cpu_pct.unwrap_or(0.0).max(self.mem_pct.unwrap_or(0.0));
100.0 - used
}
}
/// Every node's current metric scalars (merged Beszel + heartbeat health).
pub async fn eval_all(pool: &PgPool) -> Result<Vec<EvalRow>, DbError> {
let rows = sqlx::query(
"SELECT n.id, n.workspace_id, n.status,
COALESCE(m.cpu_pct, h.cpu_pct) AS cpu_pct,
COALESCE(m.mem_pct, CASE WHEN h.mem_total > 0 THEN h.mem_used::float8 / h.mem_total * 100 END) AS mem_pct,
COALESCE(m.disk_pct, CASE WHEN h.disk_total > 0 THEN (h.disk_total - h.disk_free)::float8 / h.disk_total * 100 END) AS disk_pct,
m.gpu_pct, m.temp_max,
COALESCE(m.load1, h.load1) AS load1
FROM nodes n
LEFT JOIN node_health h ON h.node_id = n.id
LEFT JOIN node_metrics m ON m.node_id = n.id",
)
.fetch_all(pool)
.await?;
Ok(rows
.into_iter()
.map(|r| EvalRow {
node_id: NodeId::from(r.get::<Uuid, _>("id")),
workspace_id: WorkspaceId::from(r.get::<Uuid, _>("workspace_id")),
status: r.get("status"),
cpu_pct: r.get("cpu_pct"),
mem_pct: r.get("mem_pct"),
disk_pct: r.get("disk_pct"),
gpu_pct: r.get("gpu_pct"),
temp_max: r.get("temp_max"),
load1: r.get("load1"),
})
.collect())
}
/// The latest metrics blob for a node (the full snapshot for the monitor page).
pub async fn latest(pool: &PgPool, node_id: NodeId) -> Result<Option<Value>, DbError> {
let row = sqlx::query(