Plan item 4 said to pull upstream's `0db7d999a` egress policy as "defence for the egress problem we have not solved". It cannot see the problem. `claude_cli` runs the claude binary as a SUBPROCESS (`Command::new(&self.binary_path).spawn()`), so every mission tool call happens inside that child. `net_guard`'s only call sites upstream are `link_enricher`, `helpers/domain_guard` and `plugins/egress` — ZeroClaw's own Rust HTTP. A mission agent's `curl` never touches the guarded stack. Pulling the commit hardens the CHAT tier; it leaves mission egress exactly as it is. Item 4 is corrected in place rather than deleted, because the reasoning is the useful part. What is actually true, measured on gw-04 with controls in both directions: positive 1.1.1.1:443 REACHABLE negative 192.0.2.1:80 TEST-NET blocked tailnet gw-02 100.84.218.70:22 REACHABLE host SSH docker gw 172.23.0.1:22 REACHABLE 169.254.169.254 REACHABLE postgres REACHABLE, password-required LAN 192.168.1.1 blocked A mission agent reaches the entire tailnet and SSH on its own host. It matters more here than it would elsewhere: these agents run model-generated shell over content fetched from the open web — 151 of 158 production Bash calls were curl/wget — so the instruction stream and the data stream are one stream. The first run of this probe attached only `clawmates_core`, reported "no internet", and was discarded: its positive control failed, so it measured nothing. A mission container is on BOTH networks and that is what must be reproduced. Recorded in full, including the half that is fine — postgres refuses unauthenticated TCP and no database credentials are forwarded into a mission container — because a report that lists only the bad half is not a measurement. Remediation is written down and deliberately NOT applied, on the operator's call. It is DOCKER-USER rules dropping the private world with the core subnet accepted first; never a public host allow-list as the opening move, because `JEPA Research` alone fetched a dozen hosts nobody would have pre-approved and a mission that cannot read cannot do research. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
153 lines
6.9 KiB
Markdown
153 lines
6.9 KiB
Markdown
# What a mission container can reach — 2026-08-27
|
|
|
|
Measured, not modelled. Nothing in this document changes production; it exists
|
|
because plan item 4 named a fix that cannot reach the problem, and the problem
|
|
turned out to be larger than the one it was written for.
|
|
|
|
## Item 4 was mis-scoped
|
|
|
|
The plan said to pull upstream's `0db7d999a feat(plugins): add shared egress
|
|
policy foundation (#9137)` — "defence for the egress problem we have not
|
|
solved". It is good work (DNS pinning, IPv4-mapped metadata blocking,
|
|
proxy-conflict surfacing, ~2,400 lines across three crates) and it **cannot
|
|
observe a single mission tool call.**
|
|
|
|
Two facts settle it:
|
|
|
|
- `claude_cli` runs the claude binary as a **subprocess** —
|
|
`Command::new(&self.binary_path)` … `.spawn()` in
|
|
`crates/zeroclaw-providers/src/claude_cli.rs:345`. Every tool the mission
|
|
agent runs happens inside that child process.
|
|
- `net_guard`'s only call sites upstream are `link_enricher.rs`,
|
|
`helpers/domain_guard.rs` and `plugins/egress.rs` — ZeroClaw's own Rust HTTP
|
|
paths.
|
|
|
|
A mission agent's `curl` is spawned by claude, not by ZeroClaw, so it never
|
|
touches the guarded stack. Pulling the commit would harden the **chat** tier's
|
|
link previews and native web tools. It would leave mission egress exactly as it
|
|
is. Worth doing on its own merits; not worth doing under the belief that it
|
|
closes this.
|
|
|
|
## The topology
|
|
|
|
Mission containers are attached to two docker networks
|
|
(`mission_runtime.rs:327`):
|
|
|
|
| network | `internal` | subnet | what it is for |
|
|
|-------------------|-----------|----------------|-------------------------------|
|
|
| `clawmates_core` | **true** | 172.20.0.0/16 | server, database, skills door |
|
|
| `clawmates_edge` | false | 172.23.0.0/16 | outbound provider egress |
|
|
|
|
`core` has no default route at all — confirmed from a container on it, where
|
|
every external address is unreachable and even the positive control fails.
|
|
`edge` supplies the default route, and with it everything below.
|
|
|
|
## The measurement
|
|
|
|
A throwaway `alpine:3.20` container attached to **both** networks, exactly as a
|
|
mission container is. Controls in both directions, because a probe whose
|
|
positive control fails proves nothing — the first run of this probe was
|
|
`core`-only, reported "no internet", and was discarded for that reason.
|
|
|
|
```
|
|
--- routes ---
|
|
default via 172.23.0.1 dev eth1
|
|
172.20.0.0/16 dev eth0 scope link src 172.20.0.6
|
|
172.23.0.0/16 dev eth1 scope link src 172.23.0.5
|
|
|
|
positive control 1.1.1.1:443 raw IP REACHABLE (expected)
|
|
positive control arxiv.org:443 DNS+connect REACHABLE (expected)
|
|
negative control 192.0.2.1:80 TEST-NET blocked (expected)
|
|
|
|
tailnet gw-02 100.84.218.70:22 REACHABLE <--
|
|
host SSH docker gateway 172.23.0.1:22 REACHABLE <--
|
|
link-local 169.254.169.254:80 REACHABLE <--
|
|
database clawmates_postgres_1:5432 REACHABLE (authenticated)
|
|
internet unrestricted REACHABLE
|
|
LAN 192.168.1.1:80 blocked
|
|
```
|
|
|
|
`architect:8090` read blocked only because that host has been offline 15 days;
|
|
it is not evidence of a control.
|
|
|
|
## What this means
|
|
|
|
A mission agent reaches the **entire tailnet** and **SSH on its own host**. The
|
|
container's default route is the docker gateway, the host runs tailscale, and
|
|
NAT forwards the rest. Nothing between the agent and 100.64.0.0/10.
|
|
|
|
That matters more here than it would for ordinary software, because a mission
|
|
agent runs **model-generated shell commands over content it fetched from the
|
|
open web**. Two production missions made 158 `Bash` calls, 151 of them
|
|
`curl`/`wget` against 12+ hosts, and delegated 12 more fetches to subagents. The
|
|
instruction stream and the data stream are the same stream.
|
|
|
|
Related: gw-01/02/04 are key-only with fail2ban,
|
|
but the hosts have no firewall, and gw-02's databases are bound to localhost
|
|
only. Reachability is not compromise. It is the precondition for it.
|
|
|
|
### What is *not* exposed
|
|
|
|
Stated because a report that lists only the bad half is not a measurement:
|
|
|
|
- **Postgres requires a password over TCP.** Attempted from the mission network
|
|
as both `postgres` and `clawmates`: `fe_sendauth: no password supplied`.
|
|
- **No database credentials are forwarded into mission containers.** The env is
|
|
`ZEROCLAW_GATEWAY_PORT`, `CM_MISSION_ID`, `GIT_CONFIG_*` and provider keys —
|
|
nothing else (`mission_runtime.rs:735`).
|
|
- **The LAN is not reachable**, only the tailnet.
|
|
- `169.254.169.254` is link-local on bare metal, not a cloud metadata service.
|
|
Reachable, but there is no credential endpoint behind it on gw-04.
|
|
|
|
## Remediation — described, deliberately NOT applied
|
|
|
|
The operator's call was to measure and report. The smallest change that closes
|
|
the lateral path, for whenever that decision is made:
|
|
|
|
```sh
|
|
# exempt first: missions legitimately need the core network (skills door, API)
|
|
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 172.20.0.0/16 -j ACCEPT
|
|
# then deny the private world
|
|
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 100.64.0.0/10 -j DROP # tailnet
|
|
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 169.254.0.0/16 -j DROP # link-local
|
|
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 10.0.0.0/8 -j DROP
|
|
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 192.168.0.0/16 -j DROP
|
|
```
|
|
|
|
Public egress is untouched, so research missions keep working — which is the
|
|
constraint that rules out a host allow-list as the first move. `JEPA Research`
|
|
alone fetched arxiv, api.github.com, raw.githubusercontent.com, arrow.apache.org,
|
|
pytables, clawpack, h5py, lancedb, netcdf4, paperswithcode and more. No
|
|
pre-approved list would have contained them, and a mission that cannot read
|
|
cannot do research.
|
|
|
|
Order matters (`-I` prepends, so the ACCEPT must be inserted last to sit first),
|
|
the rules are not persistent across reboot as written, and they should be
|
|
verified with the same positive/negative control pair used above rather than
|
|
assumed.
|
|
|
|
## One fix that was applied
|
|
|
|
`mission_runtime.rs` discarded the result of the edge-network attach:
|
|
|
|
```rust
|
|
let _ = self.docker.connect_network(EDGE_NETWORK, …).await;
|
|
```
|
|
|
|
`core` is internal, so a failed attach leaves a mission with **no egress at
|
|
all** — no provider call, no fetch — while the launch reports success and the
|
|
phase can still complete. The green-with-nothing shape this codebase keeps
|
|
meeting.
|
|
|
|
Now: on error, the container's own network list decides. Already attached is
|
|
benign and logged; genuinely not attached fails the launch with a message that
|
|
says what it means. The container's networks are the fact, not the return code.
|
|
|
|
## Limits of this measurement
|
|
|
|
- One host (gw-04), one moment. Other deployments may differ.
|
|
- Reachability of a **port**, not exploitability of a service.
|
|
- `architect` and the fleet nodes were offline, so they were not probed.
|
|
- The probe container ran `alpine`, not `clawmates-runtime`. Same two networks
|
|
and the same route table; a different image cannot have more access.
|