Plan item 4 said to pull upstream's `0db7d999a` egress policy as "defence for the egress problem we have not solved". It cannot see the problem. `claude_cli` runs the claude binary as a SUBPROCESS (`Command::new(&self.binary_path).spawn()`), so every mission tool call happens inside that child. `net_guard`'s only call sites upstream are `link_enricher`, `helpers/domain_guard` and `plugins/egress` — ZeroClaw's own Rust HTTP. A mission agent's `curl` never touches the guarded stack. Pulling the commit hardens the CHAT tier; it leaves mission egress exactly as it is. Item 4 is corrected in place rather than deleted, because the reasoning is the useful part. What is actually true, measured on gw-04 with controls in both directions: positive 1.1.1.1:443 REACHABLE negative 192.0.2.1:80 TEST-NET blocked tailnet gw-02 100.84.218.70:22 REACHABLE host SSH docker gw 172.23.0.1:22 REACHABLE 169.254.169.254 REACHABLE postgres REACHABLE, password-required LAN 192.168.1.1 blocked A mission agent reaches the entire tailnet and SSH on its own host. It matters more here than it would elsewhere: these agents run model-generated shell over content fetched from the open web — 151 of 158 production Bash calls were curl/wget — so the instruction stream and the data stream are one stream. The first run of this probe attached only `clawmates_core`, reported "no internet", and was discarded: its positive control failed, so it measured nothing. A mission container is on BOTH networks and that is what must be reproduced. Recorded in full, including the half that is fine — postgres refuses unauthenticated TCP and no database credentials are forwarded into a mission container — because a report that lists only the bad half is not a measurement. Remediation is written down and deliberately NOT applied, on the operator's call. It is DOCKER-USER rules dropping the private world with the core subnet accepted first; never a public host allow-list as the opening move, because `JEPA Research` alone fetched a dozen hosts nobody would have pre-approved and a mission that cannot read cannot do research. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
6.9 KiB
What a mission container can reach — 2026-08-27
Measured, not modelled. Nothing in this document changes production; it exists because plan item 4 named a fix that cannot reach the problem, and the problem turned out to be larger than the one it was written for.
Item 4 was mis-scoped
The plan said to pull upstream's 0db7d999a feat(plugins): add shared egress policy foundation (#9137) — "defence for the egress problem we have not
solved". It is good work (DNS pinning, IPv4-mapped metadata blocking,
proxy-conflict surfacing, ~2,400 lines across three crates) and it cannot
observe a single mission tool call.
Two facts settle it:
claude_cliruns the claude binary as a subprocess —Command::new(&self.binary_path)….spawn()incrates/zeroclaw-providers/src/claude_cli.rs:345. Every tool the mission agent runs happens inside that child process.net_guard's only call sites upstream arelink_enricher.rs,helpers/domain_guard.rsandplugins/egress.rs— ZeroClaw's own Rust HTTP paths.
A mission agent's curl is spawned by claude, not by ZeroClaw, so it never
touches the guarded stack. Pulling the commit would harden the chat tier's
link previews and native web tools. It would leave mission egress exactly as it
is. Worth doing on its own merits; not worth doing under the belief that it
closes this.
The topology
Mission containers are attached to two docker networks
(mission_runtime.rs:327):
| network | internal |
subnet | what it is for |
|---|---|---|---|
clawmates_core |
true | 172.20.0.0/16 | server, database, skills door |
clawmates_edge |
false | 172.23.0.0/16 | outbound provider egress |
core has no default route at all — confirmed from a container on it, where
every external address is unreachable and even the positive control fails.
edge supplies the default route, and with it everything below.
The measurement
A throwaway alpine:3.20 container attached to both networks, exactly as a
mission container is. Controls in both directions, because a probe whose
positive control fails proves nothing — the first run of this probe was
core-only, reported "no internet", and was discarded for that reason.
--- routes ---
default via 172.23.0.1 dev eth1
172.20.0.0/16 dev eth0 scope link src 172.20.0.6
172.23.0.0/16 dev eth1 scope link src 172.23.0.5
positive control 1.1.1.1:443 raw IP REACHABLE (expected)
positive control arxiv.org:443 DNS+connect REACHABLE (expected)
negative control 192.0.2.1:80 TEST-NET blocked (expected)
tailnet gw-02 100.84.218.70:22 REACHABLE <--
host SSH docker gateway 172.23.0.1:22 REACHABLE <--
link-local 169.254.169.254:80 REACHABLE <--
database clawmates_postgres_1:5432 REACHABLE (authenticated)
internet unrestricted REACHABLE
LAN 192.168.1.1:80 blocked
architect:8090 read blocked only because that host has been offline 15 days;
it is not evidence of a control.
What this means
A mission agent reaches the entire tailnet and SSH on its own host. The container's default route is the docker gateway, the host runs tailscale, and NAT forwards the rest. Nothing between the agent and 100.64.0.0/10.
That matters more here than it would for ordinary software, because a mission
agent runs model-generated shell commands over content it fetched from the
open web. Two production missions made 158 Bash calls, 151 of them
curl/wget against 12+ hosts, and delegated 12 more fetches to subagents. The
instruction stream and the data stream are the same stream.
Related: gw-01/02/04 are key-only with fail2ban, but the hosts have no firewall, and gw-02's databases are bound to localhost only. Reachability is not compromise. It is the precondition for it.
What is not exposed
Stated because a report that lists only the bad half is not a measurement:
- Postgres requires a password over TCP. Attempted from the mission network
as both
postgresandclawmates:fe_sendauth: no password supplied. - No database credentials are forwarded into mission containers. The env is
ZEROCLAW_GATEWAY_PORT,CM_MISSION_ID,GIT_CONFIG_*and provider keys — nothing else (mission_runtime.rs:735). - The LAN is not reachable, only the tailnet.
169.254.169.254is link-local on bare metal, not a cloud metadata service. Reachable, but there is no credential endpoint behind it on gw-04.
Remediation — described, deliberately NOT applied
The operator's call was to measure and report. The smallest change that closes the lateral path, for whenever that decision is made:
# exempt first: missions legitimately need the core network (skills door, API)
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 172.20.0.0/16 -j ACCEPT
# then deny the private world
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 100.64.0.0/10 -j DROP # tailnet
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 169.254.0.0/16 -j DROP # link-local
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 10.0.0.0/8 -j DROP
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 192.168.0.0/16 -j DROP
Public egress is untouched, so research missions keep working — which is the
constraint that rules out a host allow-list as the first move. JEPA Research
alone fetched arxiv, api.github.com, raw.githubusercontent.com, arrow.apache.org,
pytables, clawpack, h5py, lancedb, netcdf4, paperswithcode and more. No
pre-approved list would have contained them, and a mission that cannot read
cannot do research.
Order matters (-I prepends, so the ACCEPT must be inserted last to sit first),
the rules are not persistent across reboot as written, and they should be
verified with the same positive/negative control pair used above rather than
assumed.
One fix that was applied
mission_runtime.rs discarded the result of the edge-network attach:
let _ = self.docker.connect_network(EDGE_NETWORK, …).await;
core is internal, so a failed attach leaves a mission with no egress at
all — no provider call, no fetch — while the launch reports success and the
phase can still complete. The green-with-nothing shape this codebase keeps
meeting.
Now: on error, the container's own network list decides. Already attached is benign and logged; genuinely not attached fails the launch with a message that says what it means. The container's networks are the fact, not the return code.
Limits of this measurement
- One host (gw-04), one moment. Other deployments may differ.
- Reachability of a port, not exploitability of a service.
architectand the fleet nodes were offline, so they were not probed.- The probe container ran
alpine, notclawmates-runtime. Same two networks and the same route table; a different image cannot have more access.