Files
clawmates/docs/MISSION-EGRESS.md
T
Omar SobhandClaude Opus 5 fde1341618
deploy / test (push) Successful in 5m15s
deploy / build (push) Successful in 5m43s
docs: what a mission container can reach, and why item 4 could not fix it
Plan item 4 said to pull upstream's `0db7d999a` egress policy as "defence for
the egress problem we have not solved". It cannot see the problem.

`claude_cli` runs the claude binary as a SUBPROCESS
(`Command::new(&self.binary_path).spawn()`), so every mission tool call happens
inside that child. `net_guard`'s only call sites upstream are `link_enricher`,
`helpers/domain_guard` and `plugins/egress` — ZeroClaw's own Rust HTTP. A
mission agent's `curl` never touches the guarded stack. Pulling the commit
hardens the CHAT tier; it leaves mission egress exactly as it is. Item 4 is
corrected in place rather than deleted, because the reasoning is the useful
part.

What is actually true, measured on gw-04 with controls in both directions:

  positive  1.1.1.1:443             REACHABLE
  negative  192.0.2.1:80 TEST-NET   blocked
  tailnet   gw-02 100.84.218.70:22  REACHABLE
  host SSH  docker gw 172.23.0.1:22 REACHABLE
  169.254.169.254                   REACHABLE
  postgres                          REACHABLE, password-required
  LAN 192.168.1.1                   blocked

A mission agent reaches the entire tailnet and SSH on its own host. It matters
more here than it would elsewhere: these agents run model-generated shell over
content fetched from the open web — 151 of 158 production Bash calls were
curl/wget — so the instruction stream and the data stream are one stream.

The first run of this probe attached only `clawmates_core`, reported "no
internet", and was discarded: its positive control failed, so it measured
nothing. A mission container is on BOTH networks and that is what must be
reproduced.

Recorded in full, including the half that is fine — postgres refuses
unauthenticated TCP and no database credentials are forwarded into a mission
container — because a report that lists only the bad half is not a measurement.

Remediation is written down and deliberately NOT applied, on the operator's
call. It is DOCKER-USER rules dropping the private world with the core subnet
accepted first; never a public host allow-list as the opening move, because
`JEPA Research` alone fetched a dozen hosts nobody would have pre-approved and
a mission that cannot read cannot do research.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
2026-08-27 09:11:26 -05:00

153 lines
6.9 KiB
Markdown

# What a mission container can reach — 2026-08-27
Measured, not modelled. Nothing in this document changes production; it exists
because plan item 4 named a fix that cannot reach the problem, and the problem
turned out to be larger than the one it was written for.
## Item 4 was mis-scoped
The plan said to pull upstream's `0db7d999a feat(plugins): add shared egress
policy foundation (#9137)` — "defence for the egress problem we have not
solved". It is good work (DNS pinning, IPv4-mapped metadata blocking,
proxy-conflict surfacing, ~2,400 lines across three crates) and it **cannot
observe a single mission tool call.**
Two facts settle it:
- `claude_cli` runs the claude binary as a **subprocess**
`Command::new(&self.binary_path)``.spawn()` in
`crates/zeroclaw-providers/src/claude_cli.rs:345`. Every tool the mission
agent runs happens inside that child process.
- `net_guard`'s only call sites upstream are `link_enricher.rs`,
`helpers/domain_guard.rs` and `plugins/egress.rs` — ZeroClaw's own Rust HTTP
paths.
A mission agent's `curl` is spawned by claude, not by ZeroClaw, so it never
touches the guarded stack. Pulling the commit would harden the **chat** tier's
link previews and native web tools. It would leave mission egress exactly as it
is. Worth doing on its own merits; not worth doing under the belief that it
closes this.
## The topology
Mission containers are attached to two docker networks
(`mission_runtime.rs:327`):
| network | `internal` | subnet | what it is for |
|-------------------|-----------|----------------|-------------------------------|
| `clawmates_core` | **true** | 172.20.0.0/16 | server, database, skills door |
| `clawmates_edge` | false | 172.23.0.0/16 | outbound provider egress |
`core` has no default route at all — confirmed from a container on it, where
every external address is unreachable and even the positive control fails.
`edge` supplies the default route, and with it everything below.
## The measurement
A throwaway `alpine:3.20` container attached to **both** networks, exactly as a
mission container is. Controls in both directions, because a probe whose
positive control fails proves nothing — the first run of this probe was
`core`-only, reported "no internet", and was discarded for that reason.
```
--- routes ---
default via 172.23.0.1 dev eth1
172.20.0.0/16 dev eth0 scope link src 172.20.0.6
172.23.0.0/16 dev eth1 scope link src 172.23.0.5
positive control 1.1.1.1:443 raw IP REACHABLE (expected)
positive control arxiv.org:443 DNS+connect REACHABLE (expected)
negative control 192.0.2.1:80 TEST-NET blocked (expected)
tailnet gw-02 100.84.218.70:22 REACHABLE <--
host SSH docker gateway 172.23.0.1:22 REACHABLE <--
link-local 169.254.169.254:80 REACHABLE <--
database clawmates_postgres_1:5432 REACHABLE (authenticated)
internet unrestricted REACHABLE
LAN 192.168.1.1:80 blocked
```
`architect:8090` read blocked only because that host has been offline 15 days;
it is not evidence of a control.
## What this means
A mission agent reaches the **entire tailnet** and **SSH on its own host**. The
container's default route is the docker gateway, the host runs tailscale, and
NAT forwards the rest. Nothing between the agent and 100.64.0.0/10.
That matters more here than it would for ordinary software, because a mission
agent runs **model-generated shell commands over content it fetched from the
open web**. Two production missions made 158 `Bash` calls, 151 of them
`curl`/`wget` against 12+ hosts, and delegated 12 more fetches to subagents. The
instruction stream and the data stream are the same stream.
Related: gw-01/02/04 are key-only with fail2ban,
but the hosts have no firewall, and gw-02's databases are bound to localhost
only. Reachability is not compromise. It is the precondition for it.
### What is *not* exposed
Stated because a report that lists only the bad half is not a measurement:
- **Postgres requires a password over TCP.** Attempted from the mission network
as both `postgres` and `clawmates`: `fe_sendauth: no password supplied`.
- **No database credentials are forwarded into mission containers.** The env is
`ZEROCLAW_GATEWAY_PORT`, `CM_MISSION_ID`, `GIT_CONFIG_*` and provider keys —
nothing else (`mission_runtime.rs:735`).
- **The LAN is not reachable**, only the tailnet.
- `169.254.169.254` is link-local on bare metal, not a cloud metadata service.
Reachable, but there is no credential endpoint behind it on gw-04.
## Remediation — described, deliberately NOT applied
The operator's call was to measure and report. The smallest change that closes
the lateral path, for whenever that decision is made:
```sh
# exempt first: missions legitimately need the core network (skills door, API)
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 172.20.0.0/16 -j ACCEPT
# then deny the private world
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 100.64.0.0/10 -j DROP # tailnet
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 169.254.0.0/16 -j DROP # link-local
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 10.0.0.0/8 -j DROP
iptables -I DOCKER-USER -s 172.23.0.0/16 -d 192.168.0.0/16 -j DROP
```
Public egress is untouched, so research missions keep working — which is the
constraint that rules out a host allow-list as the first move. `JEPA Research`
alone fetched arxiv, api.github.com, raw.githubusercontent.com, arrow.apache.org,
pytables, clawpack, h5py, lancedb, netcdf4, paperswithcode and more. No
pre-approved list would have contained them, and a mission that cannot read
cannot do research.
Order matters (`-I` prepends, so the ACCEPT must be inserted last to sit first),
the rules are not persistent across reboot as written, and they should be
verified with the same positive/negative control pair used above rather than
assumed.
## One fix that was applied
`mission_runtime.rs` discarded the result of the edge-network attach:
```rust
let _ = self.docker.connect_network(EDGE_NETWORK, ).await;
```
`core` is internal, so a failed attach leaves a mission with **no egress at
all** — no provider call, no fetch — while the launch reports success and the
phase can still complete. The green-with-nothing shape this codebase keeps
meeting.
Now: on error, the container's own network list decides. Already attached is
benign and logged; genuinely not attached fails the launch with a message that
says what it means. The container's networks are the fact, not the return code.
## Limits of this measurement
- One host (gw-04), one moment. Other deployments may differ.
- Reachability of a **port**, not exploitability of a service.
- `architect` and the fleet nodes were offline, so they were not probed.
- The probe container ran `alpine`, not `clawmates-runtime`. Same two networks
and the same route table; a different image cannot have more access.