Historically every bin (claw-store, claw-cargo, claw-fuse)
re-declared the same 13-line 'mod ...' block at its own root.
Working but fragile — a new bin (or a new module) needed edits
in N+1 places and stayed one edit-slip away from silently
dropping something.
Now: src/lib.rs owns the shared module tree. Cargo.toml gets a
[lib] entry so cargo picks it up. claw_fuse.rs migrated as the
proof-of-concept — 13 mod lines → 4 use lines.
claw-store + claw-cargo bins still use their internal 'mod ...'
blocks and 'crate::...' paths — migrating them is mechanical
but noisy; deferred to a follow-on so this PR stays reviewable.
Both bins + tests continue to build unchanged.
387 tests pass.
deploy/scripts/rotate-snapshots.sh creates daily-YYYY-MM-DD
snapshot + prunes daily-* older than RETAIN_DAYS. Only touches
its own daily-* namespace so hand-created snapshots (release
anchors etc) never get reaped.
02:00 timer runs ahead of the 03:15 ref-sweep + 03:30 gc so
tonight's fresh snapshot pins protect its blobs from eviction.
DRY_RUN=1 for preview.
Default remains dry-run. --apply iterates the stale set and calls
RefTracking::forget per fp. Blob eviction stays a separate step
(next cluster-gc). Errors are surfaced per-fp, batch continues.
03:15 local, Persistent=true. Runs ahead of the 03:30 GC so
operators see the stale-fingerprint report before eviction lands.
Environment= drives the Gitea URL; token expected in a drop-in
(clawstor-ref-sweep.service.d/token.conf) so it doesn't sit in
the unit file. README + install recipe updated.
03:30 local, Persistent=true. Ordered before the Sun 04:00
scrub so scrub reads a fresh post-GC layout.
Default ExecStart is orphan-chunk sweep only (safe on any
node). Fleets that want LRU size-cap eviction add a drop-in:
systemctl --user edit clawstor-gc.service
[Service]
ExecStart=
ExecStart=%h/clawstor-deploy/claw-store --config ... cluster-gc --evict-to-gb 200
README updated.
Deploy artifact. Sunday 04:00 local, Persistent=true so a
box that missed the wall-clock moment fires on next boot.
Timer + Service pair — both go under
~/.config/systemd/user/. The service is Type=oneshot; timer
drives it. Non-zero exit from cluster-scrub (integrity issue)
surfaces via systemd failed state; journalctl has the details.
README updated with install recipe.
<mount>/refs/<fp-hex> file, content = blob that fp points at
Adds RefStore::list() unioning legacy refs/ and stamped refs-v2/.
FUSE wires the new dir. On lookup, tries get_stamped first then
legacy — matches the resolution order used by claw-cargo build.
Ref files piggy-back the blob inode (same content), so a hex
readable via /blobs/<hex> and /refs/<fp> shares the inode. cheap.
status=203/EXEC means the exec path didn't resolve. systemd
expands %h (home) at parse time but does NOT expand ${VAR}
in the ExecStart executable position — only in ExecStart args.
Rewrote to use %h/clawstor-deploy/claw-fuse directly.
Deploy artifact — not built into the binary tree. `deploy/systemd/
clawstor-fuse.service` + README with the install recipe.
Unit shape:
* Type=simple, blocks on the FUSE binary (unmount = SIGTERM).
* ExecStartPre = mkdir -p mount, best-effort lazy unmount of any
stale prior mount (guards against a hard SIGKILL leaving the
kernel with a dangling mount).
* Restart=on-failure, RestartSec=5 — recover from transient IO
errors without operator involvement.
* Environment= for CLAWSTOR_FUSE_BIN / DATA / MOUNT so a
`systemctl edit` drop-in retargets without editing the unit file.
Dependency chain: clawstor-fuse.service After= + Wants=
clawstor-cluster.service. FUSE reads only the on-disk state so
this is technically not needed for correctness — but keeps the
mount from spinning up on a node where the daemon is broken.
Companion to the previous list/get fix. Same layer-split problem:
* delete() only unlinked tags/, leaving stamped tags-v2/ behind.
Result: `unpin` prints \"no such tag\" for pins created via
Phase 3c+ RPC even though the tag is right there on disk.
* contains() only checked tags/. Same false-negative.
Fix: both APIs now inspect BOTH layers. delete() unlinks
whichever files exist (either or both) AND removes the TTL
expiry sidecar if present. Returns true when anything was
actually removed.
Legacy behavior preserved: tests unchanged, 381 tests pass.
Live smoke exposed the gap: `claw-cargo pin` writes to the
stamped store (tags-v2/, Phase 3c+) but `TagStore::list()` +
`get()` only walked the legacy `tags/` layer. Modern pins were
invisible to everything using those APIs — including the new
FUSE tag layer, but also `claw-cargo list-tags`.
Fix:
* list() now unions legacy + stamped entries (deduped by key).
New private list_stamped_only() walks tags-v2/.
* get() falls through to get_stamped() when the legacy file is
absent — modern pins resolve without callers knowing which
layer stored them.
No API breakage: legacy tests still pass unchanged.
381 tests pass.
<mount>/tags/<sanitized-name> ← file, content = current tag's blob
Layers atop the Phase 6a/6b mount. Operator can now `cat` a named
build cache without translating a tag → blob-id first:
cat <mount>/tags/clawverse:main:latest-cache | tar -tvzf -
Slash → underscore for filenames (tag keys like `a/b:c` land as
`a_b:c`), tags with control chars or NUL are dropped from the
listing. Non-existent value blobs also drop.
Tag file inodes 1_000..9_999. Each getattr / read re-resolves
the tag key (they're mutable — a `pin --replace` under a tag
should show the new blob without unmount).
Reused alloc_blob_ino so tag reads share inodes with /blobs/<hex>
where possible. TagStore opened at `<data_dir>/tags-db` matching
the daemon's convention.
Extends the read-only FUSE mount with a `snapshots/` tree:
<mount>/snapshots/<name>/ ← dir per snapshot
<mount>/snapshots/<name>/<blob-id-hex> ← file, content = assembled blob
Lets an operator browse a point-in-time capture by name — `ls`,
`find`, `sha256sum` all just work.
Implementation:
* Inode partitioning: 1 = /, 2 = /blobs, 3 = /snapshots,
10_000..99_999 = snapshot dirs (lazy allocation), 100_000+
= blob files (shared with the /blobs tree — same blob has
the same inode whether reached via /blobs or /snapshots/<n>).
* lookup on /snapshots/<name> validates the snapshot exists via
SnapshotStore::get.
* lookup on /snapshots/<name>/<blob-hex> validates both that
the snapshot references that blob AND that the blob is on
disk — no stale symlinks.
* Reused the shared alloc_blob_ino helper so the /blobs and
/snapshots trees hand out identical inodes for the same blob.
No new tests: FUSE is integration-heavy and the underlying
snapshot + blob primitives are already covered.
First slice of Phase 6. Ships a minimal, feature-gated `claw-fuse`
binary that mounts the local blob store read-only as a POSIX
filesystem:
<mount>/blobs/<blob-id-hex> ← file, content = assembled blob
<mount>/blobs/ ← dir, ls shows all blob-ids
<mount>/ ← dir, contains `blobs`
Lets an operator `tar -tvzf`, `md5sum`, or grep at a cached
tarball without wiring a client. Debug + audit tool for now;
warm-tier git-worktrees + write path come in later slices.
Feature-gated so my macOS dev box doesn't need macFUSE headers
to build the rest of the tree:
* Cargo.toml declares `[[bin]] name = "claw-fuse"` with
`required-features = ["fuse"]`.
* Feature `fuse` pulls in `fuser = "0.15"`, target-restricted
to `cfg(target_os = "linux")` — dep resolution never
considers fuser on other platforms.
* `cargo build` (default) leaves claw-fuse out entirely.
`cargo build --features fuse --bin claw-fuse` on Linux builds it.
Design notes baked into the impl:
* Inode allocation is lazy — first `lookup` for a hex assigns an
inode. Avoids pre-indexing the full blob store at mount time
which would be O(blobs) fs walk before FUSE is even ready.
* getattr / read validate that the manifest exists on every
call — no stale-inode reads if a blob is GC'd out from under
us mid-mount. Extra read cost is negligible against the
per-request FUSE overhead.
* size = manifest.total_size (bytes reported without touching
chunk files) so `ls -l` is cheap.
* runtime = current-thread tokio, block_on per callback. fuser
is sync; a full tokio worker pool would just add scheduling
overhead when callbacks are already serialized by the kernel.
No new tests here — Filesystem impls are integration-heavy and
the underlying BlobStore methods are already covered. The
`fuse` feature build itself will be smoke-tested on tank.
381 tests pass unchanged (feature-gated bin doesn't affect the
existing test surface).
Same shape as Phase 8c did for cluster-peer-status + cluster-repair.
New flags: --tailscale-addr (optional) + --lan-probe-ms (default 200).
Route (LAN vs tailnet) printed on the output. Zero flag = identical
to pre-8 single-addr behavior.
Fourth of four operator-facing CLIs now routing-aware
(cluster-peer-status, cluster-repair, cluster-ping done; cluster-ping
was the last outstanding one).
No new tests: pure glue over connect_lan_first, which has its own
unit coverage.
Reclaims local target-dir disk in increasing bluntness. All
modes are LOCAL only — the fleet blob cache is untouched, so
`claw-cargo build` after smart-clean restores from peer.
Modes:
* incremental-only — remove target/*/incremental/ across all
profiles. Safest; keeps final artifacts + deps.
* soft (default) — remove target/ entirely. Blob still on peer.
* hard — soft, but requires --force. Reserved for operators who
know their build is transient. Rejected without --force even
in --dry-run so the safety belt can't be trained away.
--dry-run reports paths + byte count without touching disk.
New helpers (unit-tested in isolation):
* find_incremental_dirs(target) — walks target/*/incremental,
returns only existing entries.
* dir_size_bytes(root) — recursive byte count, silent on read
errors (used only for reporting, not correctness).
+5 tests: incremental discovery (existing only), missing target
empty, byte sum recursive, missing dir returns 0, hard-without-
force rejects.
381 tests pass (+5). Pre-existing macOS
hot::tests::test_project_target_size_bytes failure unchanged.
Live smoke on tank↔architect exposed the gap: bind_rpc_tailscale
was being *advertised* via gossip so peers learned to dial it,
but the daemon never actually LISTENED there. Tailnet dials hit
a closed port.
Fix: when both bind_rpc_lan and bind_rpc_tailscale are set (and
differ), spawn a second QuicServer on the tailnet address. Shares
the same fleet-CA identity + RpcRouter as the LAN listener —
requests from either side hit the same handlers.
If the second bind fails (e.g. tailnet interface not up), we log
a warning and keep the LAN listener alive rather than aborting
daemon startup. Standard graceful-degrade shape.
No new tests here — a live integration test would need two
network interfaces + a running tailscale, which the CI runners
don't have. Coverage happens on the tank+architect deployment:
`ss -lunp` on architect must show TWO clawstor UDP listeners
after this change (10.0.0.13:7702 + 100.104.171.32:7702).
Follow-on: cert SAN for the tailnet address. The current
fleet-CA-signed leaf only has the node name as SAN, so rustls
verification on the client side still checks against
--peer <name> which passes because CN == node name. But a
belt-and-suspenders leaf using fleet-ca-tailscale-sign (Phase
8a) would be more correct.
Live smoke on tank↔architect (both LAN) failed with 200ms probe:
LAN handshake takes longer than that in the wild (TLS 1.3 with
full cert chain + rustls startup on fresh endpoint). The old
single-addr .connect() had no deadline, so pre-8c callers never
noticed.
Fix: when `tailscale` is `None`, treat LAN as unlimited — the
probe deadline only matters as a fall-through trigger, and
there's nothing to fall through to. Callers with a real fallback
addr still get the fast-path routing behavior unchanged.
+1 test (connect_lan_first_lan_only_ignores_probe_deadline)
using a 1-nanosecond probe budget that a real handshake could
never meet — must succeed anyway because no fallback exists.
376 tests pass (+1).
Wires the operator CLIs to the Phase 8b connect_lan_first primitive.
Roaming ops (laptop on LTE, coffee-shop wifi) can now pass a
tailnet address alongside the usual --rpc-addr and get the
LAN-first-with-fallback behavior automatically.
New flags on both cluster-peer-status and cluster-repair:
* --tailscale-addr <addr> — optional tailnet RPC socket. When
set, --rpc-addr is tried first with
a short deadline, then this on
failure/timeout.
* --lan-probe-ms <ms> — LAN probe deadline. Default 200
matches the arch doc.
Zero flag → byte-identical to pre-8c behavior (single-addr dial).
Both flags → chosen route printed in the output header so
operators can see whether LAN or tailnet won.
No new tests: this is thin glue over connect_lan_first, which
already has its own unit coverage. Smoke test live on tank
against architect (LAN), and against fake unroutable + real
tailnet exercises both branches.
Second Phase 8 slice. Prior transport.connect() took a single
address; the LAN-first-then-Tailscale routing the arch doc calls
out was implicit ("pick lan_addr OR tailscale_addr from gossip
state") and never actually raced or fell through.
New: QuicClient::connect_lan_first(name, lan, tailscale, lan_probe)
* Try LAN first with `lan_probe` deadline (fleet default ~200ms).
* If LAN handshake fails OR the deadline fires → fall back to
the Tailscale address.
* Both slots None → error immediately (no hang).
Returns (connection, ConnectRoute) so callers + telemetry see
which side won. New enum ConnectRoute::{Lan(addr), Tailscale(addr)}.
+3 tests exercising the three shapes:
- lan-first when LAN reachable (never dials fake tailscale addr)
- fallback when LAN black-holes (240.0.0.1 SYN gets no response;
probe deadline fires, tailscale server wins)
- errors cleanly when both addrs absent
375 tests pass (+3). Pre-existing macOS
hot::tests::test_project_target_size_bytes failure unchanged.
Follow-ons for Phase 8 completion:
- Wire the peer-connect call sites (RPC forwarding, PeerStatus,
build-cache) through connect_lan_first with per-peer
lan/tailscale addrs from gossip state.
- Document the roaming-client config template.