Files
apress/deploy/uno-q/README.md
T
Omar SobhandClaude Opus 4.8 3f6c6ab399 feat(uno-q): board self-register on boot + QR claim prefill (onboarding slice 3)
Board-side half of "preloaded + self-register + claim", plus the QR flow.

- apess-selfregister.sh (on-board): pairs locally for a token, discovers
  the LAN IP, and announces {kitId, url, token, claimCode} to APESS
  /nodes/self-register (x-fleet-secret gated). Retries until APESS is up;
  idempotent, safe on boot and on a timer.
- systemd/apess-selfregister.{service,timer}: self-register After the
  daemon, re-announce every 5 min so a DHCP lease change can't strand a
  board.
- apess-node.env.example: per-board identity (KIT_ID, CLAIM_CODE,
  FLEET_SECRET, APESS_URL).
- gen-kit-codes.sh (host): mint per-kit 6-digit codes, write the
  per-board env files, and emit the QR sticker CSV
  (…/workshop?kit=KIT-NN&code=NNNNNN).
- web: BoardClaim accepts initialCode; TeamRegistration pre-fills it from
  ?code= so scanning the kit QR fills kit + code — one tap to claim.
- deploy/uno-q/README: documents the whole self-serve onboarding path.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
2026-07-07 09:36:37 -07:00

174 lines
7.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Uno Q workshop node — provisioning
Turns an Arduino Uno Q into an APESS workshop node: a ZeroClaw daemon that
exposes an HTTP/SSE gateway, drives the on-board MCU (generate → compile →
flash), and answers with cloud-first / on-board-Qwen-fallback inference. APESS
proxies each board and streams its activity to participants and instructors.
## What a provisioned board runs
- **`zeroclaw daemon`** on `:8080` — the gateway APESS talks to (`/pair`,
`/webhook`, `/ws/chat`, `/api/events`, `/health`). Use `daemon`, **not**
`gateway start` — only the daemon/agent/channel paths register the hardware
tools (`uno_q_flash`, …).
- **`llama-server`** on `:8083` — on-board Qwen for offline/fallback inference.
- **Three agents**, selected per request via `?agent=` (APESS maps the harness
provider choice → alias; see `src/lib/harness.ts` `harnessToAgent`):
| alias | provider | behaviour |
|-------|----------|-----------|
| `default` | `custom.claude` | cloud, **fallback to on-board Qwen** |
| `cloud` | `custom.cloud` | cloud only |
| `local` | `llamacpp.local`| on-board Qwen only (fully offline) |
## Prerequisites on the board
- `/home/arduino/zeroclaw` — the ZeroClaw binary (aarch64), **with the
gateway-peripheral-registration fix** (merged to `osobh/zeroclaw` main).
- `/home/arduino/llama/llama-server` + `/home/arduino/models/qwen.gguf`.
- Arduino Zephyr core `arduino:zephyr:unoq` (0.51.0) + OpenOCD for flashing,
and `ArduinoGraphics` if using scroll text.
## Provision (dev — over USB/adb)
```sh
export APESS_ADMIN_CODE=adm-xxxxxxxx # instructor code
./provision-uno-q.sh <adb-serial> team-07 https://apess-api.redclaw.dev
```
Steps it runs: install `config.template.toml` (cloud endpoint substituted) →
start `llama-server` → start `zeroclaw daemon` → pair for a bearer token →
`POST /nodes` to APESS. Idempotent; re-run to re-provision.
Override the cloud endpoint (default is the local `claude_shim` on `:8090`):
```sh
CLOUD_URI=https://api.anthropic.com/v1 CLOUD_MODEL=claude-haiku-4-5 \
APESS_ADMIN_CODE=adm-xxxx ./provision-uno-q.sh <serial> team-07
```
> The shim on `:8090` is a bring-up convenience (one Mac). A real fleet points
> `CLOUD_URI` at a shared cloud endpoint (Anthropic / OpenRouter / LiteLLM) with
> a key, so boards don't each need a tunnel.
## Onboarding — attendee self-serve (self-register + claim)
The workshop path: boards are **preloaded** and, on boot, **announce themselves**
into APESS's *unclaimed pool*; an attendee then **claims** their board to their
team from the web app (kit + 6-digit code) — no operator, no admin code, no adb.
```
board boot ──self-register {kitId,url,token}──▶ APESS unclaimed pool
attendee: scan kit QR → team name → 6-digit code ──POST /claim──┘
(bearer token moves pool → node bridge, never touches the browser)
```
**Fleet prep (host, once).** Mint per-kit codes + per-board env files + the QR
sticker CSV:
```sh
FLEET_SECRET=<shared-secret> APESS_URL=https://apess-api.redclaw.dev \
APESS_WEB=https://apess.redclaw.dev ./gen-kit-codes.sh 15
# → kit-codes/env/KIT-NN.env (one per board) kit-codes/kit-codes.csv (stickers)
```
The same `FLEET_SECRET` must be set on the APESS API (`FLEET_SECRET` env) — it
gates `/nodes/self-register` so only your boards can seed the pool. Each sticker
QR encodes `…/workshop?kit=KIT-NN&code=NNNNNN`, so scanning it pre-fills both.
**Per board.** Drop its env file + the self-register unit:
```sh
adb -s <serial> push kit-codes/env/KIT-07.env /home/arduino/.zeroclaw/apess-node.env
adb -s <serial> push apess-selfregister.sh /home/arduino/
adb -s <serial> shell 'chmod +x /home/arduino/apess-selfregister.sh'
# with root: enable the boot + refresh timer
sudo cp systemd/apess-selfregister.service systemd/apess-selfregister.timer /etc/systemd/system/
sudo systemctl enable --now apess-selfregister.timer
# no root: run it once now (and let zeroclaw-supervisor / cron re-run it)
/home/arduino/apess-selfregister.sh
```
`provision-uno-q.sh` (admin-driven, binds a board straight to a known team) is
still there for pre-provisioning / demo boards — the two paths coexist.
## Provision (production — systemd, boots on power-up)
With root on the board:
```sh
sudo cp systemd/zeroclaw-llama.service systemd/zeroclaw-daemon.service /etc/systemd/system/
sudo systemctl enable --now zeroclaw-llama zeroclaw-daemon
```
For a LAN fleet (participants reach the board's WiFi IP directly), set in
`config.toml`:
```toml
[gateway]
host = "0.0.0.0"
allow_public_bind = true
```
Then pair + register once (steps 45 of the script) so APESS has the board's
`{ url, token }`.
## Resilience — surviving a disconnect (no-root boards)
A sudden USB/adb drop breaks things that don't self-heal: the adb tunnels
vanish (board loses the cloud shim), held-shell services die, llama can wedge
(process alive but `:8083` dead), and flashes silently stop landing while the
tool still reports success. The MCU keeps its last sketch; the paired token
survives.
The systemd units above are the clean answer **when you have root**. Some dev
boards don't — an expired account blocks `sudo` and there's no user session bus,
so neither system nor user units can run. For those, use the no-root pieces:
- **`zeroclaw-supervisor.sh`** (runs on the board) — a watchdog that polls the
`/health` **endpoints** (a wedged process passes `pgrep` but fails here) and
restarts llama / the daemon when they die or wedge. Children are launched with
`setsid … exec` so they survive the shell that started them — the property
plain `nohup … &` inside `adb shell` does **not** give you. Install + persist:
```sh
adb -s <serial> push zeroclaw-supervisor.sh /home/arduino/ && \
adb -s <serial> shell 'chmod +x /home/arduino/zeroclaw-supervisor.sh; \
setsid nohup /home/arduino/zeroclaw-supervisor.sh >/dev/null 2>&1 </dev/null &'
# boot persistence (no root; cron must be running):
adb -s <serial> shell '(crontab -l 2>/dev/null | grep -v zeroclaw-supervisor.sh; \
echo "@reboot /home/arduino/zeroclaw-supervisor.sh") | crontab -'
```
- **`recover-uno-q.sh`** (runs on the host) — after the board is physically back,
re-does adb + both tunnels, ensures the supervisor is up, and health-checks
every hop by endpoint:
```sh
./recover-uno-q.sh <serial> [cloud-shim-port] # default shim port 8090
```
See the `unoq-disconnect-recovery` note for the full failure-mode list.
## Verify
```sh
adb -s <serial> forward tcp:8080 tcp:8080
curl -s localhost:8080/health # {"status":"ok"}
# each agent alias resolves (needs a bearer token from /pair):
curl -s -X POST 'localhost:8080/webhook?agent=local' -H "authorization: Bearer $TOK" \
-H 'content-type: application/json' -d '{"message":"one word: local"}'
```
`?agent=local` routes to Qwen (offline), `?agent=cloud` to the cloud, `?agent=default`
cloud-with-fallback. In APESS, the team's provider/fallback toggle picks the alias.
## Files
- `config.template.toml` — the node config (secrets stripped; `__CLOUD_URI__` /
`__CLOUD_MODEL__` substituted at provision time).
- `provision-uno-q.sh` — one-shot provisioner (adb-driven).
- `systemd/*.service` — production units (need root).
- `zeroclaw-supervisor.sh` — on-board no-root watchdog (endpoint health + restart).
- `recover-uno-q.sh` — host-side post-disconnect recovery (re-tunnel + health-check).