A board power-cycle exposed two bugs that left the node dead after boot:
- recover-uno-q.sh step 3 ran `pgrep -f zeroclaw-supervisor`, which matches
the pgrep's OWN shell (its args contain the supervisor path) — so it always
concluded "already running" and never launched the supervisor. Bracket the
pattern (`[z]eroclaw-supervisor.sh`) so only the real process matches.
- The supervisor's single-instance lock only checked that the locked pid was
alive. After a reboot the stale pid can be recycled by an unrelated process,
falsely blocking startup. Now require the live pid's /proc/<pid>/cmdline to
actually be a supervisor before deferring — otherwise treat the lock as stale.
- Also export a full PATH in the supervisor for cron's minimal @reboot env.
Validated on hardware: discriminator matches a real supervisor and rejects
init's pid; fixed recover detects the running supervisor without spawning a
duplicate; board comes back healthy (gateway + llama up).
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
A sudden USB/adb drop breaks the Uno Q node in ways that don't self-heal:
adb tunnels vanish, held-shell services die, llama wedges (alive but not
serving), and flashes silently stop landing. systemd is the clean fix with
root — but the dev board's account is expired (no sudo) and has no user
session bus, so neither system nor user units run. These are the no-root
equivalents, both validated on hardware:
- zeroclaw-supervisor.sh — on-board watchdog. Polls the /health ENDPOINTS
(a wedged process passes pgrep but fails here) and restarts llama / the
daemon on death or wedge. Launches children with `setsid … exec` so they
survive the launching shell — the property `nohup … &` in adb shell lacks.
Startup grace avoids reaping llama mid-cold-load; single-instance lock;
no per-restart shell leak. Proven: kill -9 the daemon → auto-restarted.
Boot-persisted via `@reboot` crontab (no root).
- recover-uno-q.sh — host-side. After the board is back, re-does adb + both
tunnels (out :8080, back :8090 shim), ensures the supervisor is running,
and health-checks every hop by endpoint. Proven end-to-end after a
simulated tunnel drop.
README: new Resilience section documenting the no-root reality + both tools.
Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>