feat(fleet): GLM as a real microVM backend, and per-role models for claws
Three threads, all of which end at the same place: a mission whose verifier does
not share a model with the coder it reviews.
**GLM has a credential contract now.** `microvm_credential_for` returned one env
var name, which quietly assumed every provider reads its secret from the same
place Anthropic does. It returns a `Credential { source, target }` instead —
z.ai's key lives in the server's `ZAI_API_KEY` and Claude Code reads it as
`ANTHROPIC_AUTH_TOKEN`, and collapsing those two names is what forces a guess at
the other end. A wrong guess here sends one provider's credential to another
provider's endpoint.
`images/agent-glm` is the same CLI at the same pinned version as `agent-claude`
with `ANTHROPIC_BASE_URL` baked in. The split is deliberate: the ENDPOINT is a
property of the image, the CREDENTIAL is a property of the turn. That makes the
dangerous mix-up unrepresentable — a GLM VM cannot be handed an Anthropic
subscription token, and a claude VM cannot be pointed at z.ai. Asserted both
ways, because "the GLM VM must not carry CLAUDE_CODE_OAUTH_TOKEN" is the
property that costs a credential if it ever stops holding.
Kimi stays refused. `KIMI_API_KEY` is set and Moonshot serves an
Anthropic-compatible API, but I have not verified its base URL against the
running service, and this function is precisely where guessing a URL is
expensive. It becomes an arm the day someone measures it.
`api.z.ai` joins the node's default egress allow-list. A default that cannot
run the images we ship is a trap rather than a policy — the alternative is an
operator discovering it as a hung agent with no model access.
**Per-role models for claws** (migration 0071). `template_roles` had no model
column, so `mint_team_from_template` bound every role of every mission team to
one literal — a template whose whole point is an independent reviewer minted a
reviewer sharing a model with the coder. A role may now name its own; roles that
say nothing still take the mint's default, so every template written before this
behaves exactly as it did. The literal is now that default rather than a
hardcode.
**A harness scenario for the roster flow.** `verify-mission-delivery.sh roster`
runs the whole Slice 5 loop — planner proposes, human approves, mission runs —
and asserts the roster LANDED on the mission row rather than trusting the API's
answer. That distinction is not theoretical: the first live approval returned an
error while leaving the proposal marked approved.
Built and proven on tank ahead of the deploy: `clawmates/agent-glm:dev` reports
`2.1.223` and `BASE=https://api.z.ai/api/anthropic`, and
`fc-build-rootfs.sh … glm 8G` boots a VM from it that has git, can write
/mission, and answers `claude --version`.
533 tests pass, clippy clean. Migration 0071.
Co-Authored-By: Claude Opus 5 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
75d09241fb
commit
f7f3dfe495
@@ -491,7 +491,17 @@ async fn mint_team_from_template(
|
||||
.map_err(|e| format!("insert agent {}: {e}", role.slot))?;
|
||||
let claw_id = agent.id.as_uuid();
|
||||
|
||||
cm_db::repo::agents::set_model_binding(pool, agent.id, default_model)
|
||||
// The ROLE's model when the template names one, else the mint's default.
|
||||
// Before migration 0071 there was no role model at all, so every claw of
|
||||
// every mission team ran the same one — including a reviewer reviewing
|
||||
// the coder it shares a model with.
|
||||
let role_model = role
|
||||
.model
|
||||
.as_deref()
|
||||
.map(str::trim)
|
||||
.filter(|m| !m.is_empty())
|
||||
.unwrap_or(default_model);
|
||||
cm_db::repo::agents::set_model_binding(pool, agent.id, role_model)
|
||||
.await
|
||||
.map_err(|e| format!("set_model_binding {claw_id}: {e}"))?;
|
||||
|
||||
@@ -508,7 +518,7 @@ async fn mint_team_from_template(
|
||||
// out-of-band via MissionRuntimeProvisioner::pin_agent_workspaces.
|
||||
if let Some(p) = provisioner {
|
||||
match p
|
||||
.provision_claw(claw_id, default_model, &template.template.risk_profile)
|
||||
.provision_claw(claw_id, role_model, &template.template.risk_profile)
|
||||
.await
|
||||
{
|
||||
Ok(_) => provisioned_claws.push(agent.id),
|
||||
@@ -574,18 +584,13 @@ async fn mint_team_from_template(
|
||||
Ok(team_id)
|
||||
}
|
||||
|
||||
/// The model every claw a mission mints runs on.
|
||||
/// The model a minted claw runs on when its template role does not name one.
|
||||
///
|
||||
/// One model for every role, which is a real limitation and not a preference:
|
||||
/// `template_roles` has no `model` column, so a template cannot express "the
|
||||
/// verifier runs elsewhere" — and a same-model verifier is the correlated
|
||||
/// failure the independent judge exists to break.
|
||||
///
|
||||
/// The composed path closes this: an approved roster
|
||||
/// (`routes::mission_roster`) carries a `backend` per node, and each backend is
|
||||
/// a different provider's CLI in its own VM. Closing it for CLAWS as well needs
|
||||
/// a per-role model on the template or on the mint, and neither exists yet —
|
||||
/// stated here rather than left as a literal nobody notices.
|
||||
/// A DEFAULT now, not a hardcode: `template_roles.model` (migration 0071) lets a
|
||||
/// template put its reviewer on a different model from the coder it reviews,
|
||||
/// which is the correlated failure the cross-provider judge exists to break,
|
||||
/// one layer down. Roles that say nothing still land here, so every template
|
||||
/// that existed before 0071 behaves exactly as it did.
|
||||
const MINTED_CLAW_MODEL: &str = "claude-sonnet-5";
|
||||
|
||||
/// The graph a COMPOSED microVM mission runs, built from its team template
|
||||
|
||||
Reference in New Issue
Block a user