teams: ephemeral lifecycle for Scheduled + Triggered planner modes
ci / gates (push) Successful in 6s
ci / frontend (push) Successful in 38s
ci / rust (push) Successful in 3m6s
ci / e2e (push) Has been skipped
ci / publish (push) Successful in 2m23s

Migration 0033: adds teams.lifecycle ('permanent' | 'ephemeral') and a
topology_runs.team_id back-ref with a partial index for the sibling-in-
flight check.

cm-db repo:
- teams::insert_team_with_lifecycle (insert_team keeps the permanent default)
- topology_runs::enqueue_run_for_team (populates team_id)
- topology_runs::check_ephemeral_teardown — atomic SELECT that only
  returns Some when the team is ephemeral AND no siblings are still
  queued/running; carries the workspace + bound claw ids for cleanup.

cm-api:
- topology_worker post-terminal hook maybe_teardown_ephemeral_team
  runs deprovision_claw on each bound claw (best-effort; failures log
  but don't block Postgres deletion), then hard_purge each agent row,
  then delete_team.
- routes::teams::build_team_with_lifecycle (build_team keeps default);
  run_team enqueues with team_id.
- planner ScaffoldRequest gains mode; lifecycle_for(mode) sets the team
  to ephemeral for scheduled + triggered, permanent otherwise.

Frontend MasterPlannerModal passes mode in the scaffold payload so the
backend can derive lifecycle without duplicating the mode taxonomy.

Tests: 3 new (returns claws when no siblings, holds when siblings queued,
ignores permanent teams). 10/10 topology_jobs green; workspace clippy
--tests clean.
This commit is contained in:
Omar Sobh
2026-07-07 04:30:07 -07:00
parent b0acdfd987
commit 806ba869e5
12 changed files with 412 additions and 14 deletions
+72
View File
@@ -122,6 +122,78 @@ pub async fn enqueue_run_tier(
Ok(())
}
/// Enqueue a durable run bound to a specific team. `team_id` is stored so the
/// post-terminal ephemeral-teardown hook can find the team from a completed run.
pub async fn enqueue_run_for_team(
pool: &PgPool,
id: Uuid,
workspace_id: WorkspaceId,
task: &str,
graph: &Value,
team_id: Uuid,
) -> Result<(), DbError> {
sqlx::query!(
"INSERT INTO topology_runs
(id, workspace_id, task, kind, status, graph, tier, team_id)
VALUES ($1, $2, $3, 'run', 'queued', $4, 'team', $5)",
id,
workspace_id.as_uuid(),
task,
graph,
team_id,
)
.execute(pool)
.await?;
Ok(())
}
/// Result of `check_ephemeral_teardown` when this run's terminal completion
/// should tear down its team.
pub struct EphemeralTeardown {
pub team_id: Uuid,
pub workspace_id: Uuid,
pub claw_ids: Vec<Uuid>,
}
/// If this run belongs to an ephemeral team AND no siblings are still queued /
/// running, return the team's teardown context (team id + workspace + bound
/// claws). Callers then deprovision each claw and delete the team. Returns
/// `None` if there's nothing to do (permanent team, still-running siblings, or
/// no team back-ref at all).
pub async fn check_ephemeral_teardown(
pool: &PgPool,
run_id: Uuid,
) -> Result<Option<EphemeralTeardown>, DbError> {
let row = sqlx::query!(
"SELECT t.id AS team_id, t.workspace_id
FROM topology_runs r
JOIN teams t ON t.id = r.team_id
WHERE r.id = $1
AND t.lifecycle = 'ephemeral'
AND NOT EXISTS (
SELECT 1 FROM topology_runs sib
WHERE sib.team_id = t.id
AND sib.id <> r.id
AND sib.status IN ('queued', 'running')
)",
run_id,
)
.fetch_optional(pool)
.await?;
let Some(row) = row else { return Ok(None) };
let claws = sqlx::query!(
"SELECT claw_id FROM team_members WHERE team_id = $1",
row.team_id,
)
.fetch_all(pool)
.await?;
Ok(Some(EphemeralTeardown {
team_id: row.team_id,
workspace_id: row.workspace_id,
claw_ids: claws.into_iter().map(|r| r.claw_id).collect(),
}))
}
/// Atomically claim the oldest queued job, flipping it to `running`. Uses
/// `FOR UPDATE SKIP LOCKED` so multiple workers never claim the same job.
/// Returns `None` when the queue is empty.