Fleet: Beszel hub integration — rich per-node metrics + per-node monitor (Phase 1)
Tap each node's Beszel metrics (GPU/temps/disk-IO/network/per-container — beyond
our basic heartbeat) by reading the workspace's Beszel hub. The agents run in
WS-only mode with no locally-readable socket, so (per the de-risk) the server taps
the hub's PocketBase API instead of the daemon reading agents — no daemon changes.
- migrations: workspace_beszel (BYO hub URL + login, server-side only, mirrors the
Tailscale BYO pattern) + node_metrics (latest scalar columns + JSONB blob).
- cm-db: repo/fleet_beszel.rs, repo/node_metrics.rs; nodes SELECT joins node_metrics
(gpu_pct/temp_max surfaced on node_json for the live cards).
- cm-api: beszel.rs client (auth-with-password, poll `systems`, map to nodes by
hostname, upsert metrics) + a 15s spawn_poller; routes/beszel.rs (connect/status/
disconnect + GET /api/nodes/{id}/metrics with history proxied live from the hub).
- frontend: HostCard gains a GPU/temp readout + a Monitor button; NodeMonitor is a
full-width per-node page (current panel + CPU/mem/GPU/temp/net/disk charts from the
hub's 1m history); a "Beszel monitoring" connect form in the Local view.
Reachability confirmed: gw-04 → the hub over the tailnet (100.123.224.84:8090). Needs
the user to connect their hub login to activate the poller.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
4de2f31b50
commit
36a227566b
@@ -37,6 +37,9 @@ pub struct NodeRow {
|
||||
pub last_seen: Option<OffsetDateTime>,
|
||||
pub created_at: OffsetDateTime,
|
||||
pub health: Option<NodeHealth>,
|
||||
/// Latest Beszel metrics scalars (the rich bits beyond basic health).
|
||||
pub gpu_pct: Option<f64>,
|
||||
pub temp_max: Option<f64>,
|
||||
}
|
||||
|
||||
/// Register a new (pending) node with its control-channel token.
|
||||
@@ -73,8 +76,10 @@ pub async fn auth(pool: &PgPool, token: &str) -> Result<Option<(NodeId, Workspac
|
||||
|
||||
const SELECT_WITH_HEALTH: &str = "SELECT n.id, n.name, n.hostname, n.local_ip, n.status, n.agent_version, n.tailscale_ip, n.last_seen, n.created_at,
|
||||
h.node_id AS health_node, h.cpu_pct, h.mem_total, h.mem_used, h.mem_pressure, h.swap_used,
|
||||
h.disk_total, h.disk_free, h.load1, h.load5, h.load15, h.container_count
|
||||
FROM nodes n LEFT JOIN node_health h ON h.node_id = n.id";
|
||||
h.disk_total, h.disk_free, h.load1, h.load5, h.load15, h.container_count,
|
||||
m.gpu_pct AS m_gpu_pct, m.temp_max AS m_temp_max
|
||||
FROM nodes n LEFT JOIN node_health h ON h.node_id = n.id
|
||||
LEFT JOIN node_metrics m ON m.node_id = n.id";
|
||||
|
||||
/// List a workspace's nodes (oldest first) with their latest health.
|
||||
pub async fn list(pool: &PgPool, workspace_id: WorkspaceId) -> Result<Vec<NodeRow>, DbError> {
|
||||
@@ -220,5 +225,7 @@ fn map_node(r: sqlx::postgres::PgRow) -> NodeRow {
|
||||
last_seen: r.get("last_seen"),
|
||||
created_at: r.get("created_at"),
|
||||
health,
|
||||
gpu_pct: r.get("m_gpu_pct"),
|
||||
temp_max: r.get("m_temp_max"),
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user