3d504ee6b34837e63802610550943fc968957ab7
22
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b431475af7 |
dashboard-v2 (PR 1): design doc + backend read-only endpoints
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Successful in 17s
Kicks off the single-pane-of-glass command-center rewrite. Legacy /api/* handlers untouched — v2 is additive so cutover is safe. Doc: docs/dashboard-v2.md — design goals, endpoint spec, cutover plan. Backend: new serve_v2 module wired into serve.rs. Endpoints: GET /api/v2/node/local/status GET /api/v2/node/:name/status (fan-out: not-yet-impl) GET /api/v2/storage/blobs?limit&offset GET /api/v2/storage/tags?prefix GET /api/v2/storage/refs?limit&offset GET /api/v2/storage/snapshots GET /api/v2/storage/ref-tracking?repo V2State opens BlobStore / TagStore / RefStore / SnapshotStore / RefTracking under the daemon's blob_store_root + the conventional subdirs (tags-db, refs-db). All handlers are single-node reads; cross-node fan-out lands in PR 3. +6 tests: status w/ seeded blob+snapshot, cross-node returns 501, blobs pagination, tags prefix filter, snapshots list, ref-tracking repo filter. 395 tests pass. serve.rs merges v2 routes onto the axum router with shared CORS. Single-shot deploy on tank + architect will expose /api/v2/* alongside the existing /api/*. PR 2 = frontend rewrite consuming these endpoints. PR 3 = fleet fan-out (cross-node aggregation via QUIC RPC). PR 4 = action endpoints (POST scrub/gc/snapshot/pin). PR 5 = cutover (deprecate legacy /api/*). |
||
|
|
8a502be65c |
Polish: cargo fix pass (auto-remove unused imports)
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Failing after 3s
|
||
|
|
fa863fdc7d |
Phase 7e: claw-cargo smart-clean — 3 local cleanup modes
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Failing after 3s
Reclaims local target-dir disk in increasing bluntness. All modes are LOCAL only — the fleet blob cache is untouched, so `claw-cargo build` after smart-clean restores from peer. Modes: * incremental-only — remove target/*/incremental/ across all profiles. Safest; keeps final artifacts + deps. * soft (default) — remove target/ entirely. Blob still on peer. * hard — soft, but requires --force. Reserved for operators who know their build is transient. Rejected without --force even in --dry-run so the safety belt can't be trained away. --dry-run reports paths + byte count without touching disk. New helpers (unit-tested in isolation): * find_incremental_dirs(target) — walks target/*/incremental, returns only existing entries. * dir_size_bytes(root) — recursive byte count, silent on read errors (used only for reporting, not correctness). +5 tests: incremental discovery (existing only), missing target empty, byte sum recursive, missing dir returns 0, hard-without- force rejects. 381 tests pass (+5). Pre-existing macOS hot::tests::test_project_target_size_bytes failure unchanged. |
||
|
|
3af6390316 |
Phase 7f: claw-cargo auto-records fingerprint → (repo, git_ref)
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Failing after 3s
Closes the ref-tracking loop. claw-cargo build now records the producing (repo, git_ref) alongside every cache-put fingerprint, so cluster-ref-sweep can identify stale entries later without operator bookkeeping. New BuildArgs flags: * --repo <owner/name> (env CLAWSTOR_REPO, or GITEA_REPOSITORY / GITHUB_REPOSITORY when the CI runner sets them via that name in workflow env) * --git-ref <branch-or-tag> (env CLAWSTOR_GIT_REF) * --ref-tracking-dir <path> (env CLAWSTOR_DATA_DIR, typically /var/lib/claw-store/data — same root as cluster.blob_store_root) Semantics: * All three unset → silently skipped. Existing cache flows are unchanged. * dir doesn't exist or open() fails → logs warn, cache still valid. * record() call fails → logs warn, cache still valid. The tracking store is co-located with the daemon's data dir so cluster-ref-sweep on that host sees the annotations. Runners mount /var/lib/claw-store/data via bind-mount today. No new tests here — the primitive (RefTracking::record) already has full coverage. This is thin glue. |
||
|
|
39f9a9652a |
Phase 4e: cmd_pin --offline + drain + wal-status CLI
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Successful in 10s
Wires cmd_pin through the WalQueue built in the preceding four PRs (#48 → #52). First real caller of the client-mode WAL stack. New surfaces: claw-cargo pin --offline --blob <BlobId> [--ttl <duration>] * no peer connection is opened * enqueues the same three mutations cmd_pin would emit online: primary tag (PutTagVersioned), .fingerprint companion (PutTagVersioned), and — if --ttl — the two SetTagExpiry sidecars * requires --blob because offline mode can't do the GetRefVersioned lookup that resolves fingerprint → BlobId * prints the assigned WAL seqs + "next step: drain" claw-cargo drain --peer ... * opens a peer, drains the queue, truncates up to the last applied seq * partial-failure safe: whatever applied is truncated; anything after a hard error stays on disk for retry * exits non-zero when drain stopped mid-stream claw-cargo wal-status * read-only, no network * pending count, oldest/newest seq, storage path, decoded entries (or UNDECODABLE marker on frame errors) WAL location follows the XDG state-home pattern already used by manifest.rs: 1. $XDG_STATE_HOME/claw-cargo/wal/ 2. $HOME/.local/state/claw-cargo/wal/ 3. ./.claw-cargo-wal/ (worst-case container fallback) Tests (2): default_wal_path_honours_xdg_state_home (mirroring manifest.rs's env-var pattern) + parse_blob_id_rejects_bad_ hex_and_wrong_len. claw_cargo.rs grew from 1668 to ~1870 lines. Still under the 1300-per-*module* interpretation but this bin file has been above 1300 since Phase 5. Split-out is Phase 6 territory. Co-Authored-By: Claude Opus 4.7 <[email protected]> |
||
|
|
4630925040 |
Phase 4b follow-on: pin --ttl RPC + CLI
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Successful in 10s
Closes Phase 4b by exposing the TTL sidecar written by Phase 4b primitives on the wire and via `claw-cargo pin`. Wire additions: * Method::SetTagExpiry (0x1a) — payload `key_len:u16 || key || expires_at:u64 (LE)`. Reply single-byte OK. `expires_at == 0` clears the sidecar. * Method::GetTagExpiry (0x1b) — payload raw key bytes. Reply 8 bytes (u64 LE) on hit; NotFound when no sidecar is present. Both accept writes even when the stamped tag itself is absent, matching `TagStore::set_stamped_expiry` semantics — the sidecar takes effect the moment the tag lands. CLI: * `claw-cargo pin --ttl <duration>` — humantime-style duration (`30d`, `1h30m`, `2w`, ...). Applied to both the primary tag and its `.fingerprint` companion so eviction treats them as one lifetime. `--ttl 0` / `clear` / `none` clears an existing sidecar without touching the value. Tests: encode/decode roundtrip + malformed-input rejection for `encode_expiry_record`, method-byte stability, NotConfigured without a tag store, end-to-end set/get/overwrite/clear over QUIC, and a real-pin flow that publishes a stamped tag then attaches TTL. Duration parser is unit-tested for single/compound forms, case-insensitive units, bad input, and clock alignment. No new deps — the humantime-style parser is 60 lines in-tree. Co-Authored-By: Claude Opus 4.7 <[email protected]> |
||
|
|
5c55dd7044 |
Phase 4a hotfix: cmd_pin resolves stamped refs
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Successful in 23s
pin lookup was legacy call_get_ref only, missing refs written via call_put_ref_versioned (all Phase 3b+ builds). Try versioned first, fall back to legacy — same pattern as claw-cargo build path. |
||
|
|
2c3cd2ab38 |
Phase 3c + 3e: stamped tags + namespaced ref keys — closes Phase 3
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Successful in 20s
Ships the last two pieces from the arch doc's Phase 3 scope for the
cargo-cache use case:
## 3c: Stamped tags (CRDT-merge on PutTag)
Mirror of Phase 3a/3b for TagStore. Two concurrent `claw-cargo pin`
calls on the same tag now race deterministically instead of silently
clobbering.
* `StampedTagValue` — same 48-byte (value, clock, node) tuple as
StampedRef.
* `TagStore::put_stamped(key, StampedTagValue) -> TagPutOutcome`
and `TagStore::get_stamped(key)` — data lives under `tags-v2/`
(separate from `tags/` for cutover safety).
* New wire methods `PutTagVersioned = 0x18` +
`GetTagVersioned = 0x19`.
* `call_put_tag_versioned` / `call_get_tag_versioned` client
helpers.
* `claw-cargo pin` now writes stamped tags. Concurrent pin gets
AlreadyExists and moves on (blob content is content-addressed so
both winners agree on the payload).
## 3e: Namespaced ref keys
Opt-in `--namespace <slug>` on peer-facing subcommands. When set,
the ref key becomes `blake3("clawstor.ns.v1" || namespace || fp)`
so two runners on different namespaces (`clawverse/main` vs
`clawverse/pr-42`) don't collide on the same fingerprint. Empty
namespace = pre-3e behavior, so this is 100% backward compat.
* `refs::namespaced_ref_key(namespace, fingerprint) -> RefKey`
primitive.
* `PeerArgs::namespace: Option<String>` CLI flag flows through to
`cmd_status`, `cmd_prefetch`, `cmd_build`.
* `peer_lookup` now takes a `RefKey` directly (was `&Fingerprint`)
so the namespace resolution stays in the caller — the daemon
never sees "namespace" as a concept.
## 3d: Deferred
Full vector clocks per namespace are noted in the arch doc as a
Phase-3 goal; scalar wall-clock (clock + node stamp) is sufficient
for the cargo-cache use case (single-key LWW merge). NTP-synced
runners see monotonic ordering; skewed runners lose an ordering
but the CRDT semantics still guarantee no data corruption. Full VC
is deferred to a future phase.
+9 tests, 280 total (baseline +8: 7 unit + 1 e2e over real QUIC).
|
||
|
|
ac51e3e9b5 |
Phase 3b: thread stamped refs through claw-cargo + forwarding
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Successful in 25s
Runner + prewarm use PutRefVersioned/GetRefVersioned; daemon GetRefVersioned forwards on miss + pulls blob transparently. New GetRefVersionedLocal (0x17) prevents recursion. Backward-compat: existing GetRef/PutRef path unchanged; two on-disk namespaces coexist (refs/ and refs-v2/). +1 test, 272 total (unchanged from 3a because we reused existing scaffolding). |
||
|
|
1053930451 |
claw-cargo: default --parallel-restore back to 1 (sequential)
Pi 5 loopback measurement 2026-07-13: --parallel-restore 1 : wall 2m52s, restore 20s --parallel-restore 8 : wall 3m06s, restore 34s Sequential is 70% faster on loopback. N-way stream contention costs more than a single stream's congestion-control amortization. Same shape as Phase 5k prewarm — fanout only wins when per-stream throughput has a ceiling (WAN, tunneled links). --parallel-restore N remains as opt-in. |
||
|
|
846ecffe10 |
Pi deploy follow-ups: XDG default_path + parallel restore
Two fixes surfaced by the vision-02 Pi 5 measurement: ## XDG default_path `Manifest::default_path` was hardcoded to `/var/lib/claw-store/projects.toml`. That path is read-only under the user-mode systemd unit's `ProtectSystem=strict`, and creating it needs root — awful for a runner install. Precedence, matching XDG Base Directory: 1. `$XDG_STATE_HOME/claw-store/projects.toml` 2. `$HOME/.local/state/claw-store/projects.toml` 3. `/var/lib/claw-store/projects.toml` (system fallback) User-mode installs now write in $HOME by default; system installs (root, no HOME set) still land in /var/lib. +1 test: `default_path_honours_xdg_state_home` — covers all three precedence branches. Env mutation is process-global so the test saves + restores. ## Parallel restore on cache HIT Pi restore of 947 MiB via `BlobGetStream` took ~18s (~53 MiB/s) single-stream. Per-stream throughput ceilings on the connection type cap sequential fetches; parallel chunk fetches stack their contributions. - New `call_blob_get_parallel(conn, blob_id, concurrency) -> Option<Vec<u8>>` in `rpc/client.rs`. `JoinSet` + `Semaphore`, reassembles by chunk index at manifest-known offsets so out-of-order arrival is fine. - `claw-cargo build --parallel-restore N` (default 8). `N <= 1` falls through to `BlobGetStream` for parity. - Memory: `total_size + 4 MiB × in-flight` — dominated by the reassembly buffer, not the fanout. +1 test: `parallel_blob_get_reassembles_multi_chunk_blob_byte_equal` covers roundtrip byte-equality vs BlobGetStream, tail-chunk offset, concurrency=1 correctness, and NotFound → None. 259 tests pass (+2). Pre-existing macOS failure unchanged. |
||
|
|
84aa758fd4 |
Two pilot follow-ons: rustc drift warning + blob GC
Both surfaced by the 2026-07-12 pilot as real operator concerns: ## rustc drift warning at build time Runners silently silo their cache when rustc versions differ across peers (fingerprint depends on rustc verbose output). The pilot's first flow burned a full cold+upload before we realized the silo. - `PeerStatusReply.local_rustc_release` — new field, populated from the peer's own gossip `RUSTC_RELEASE` key via a new `ClusterGossip::self_kv(key)` accessor. - `claw-cargo build`: on cache MISS, calls `PeerStatus`; if the peer's rustc release ≠ our local `rustc --version`, emits a WARN with both versions + hint to add `rust-toolchain.toml`. - Best-effort: absence of either release string is a shrug, not an error. ## blob GC Blob store grows unbounded on a runner; disk-full is a real incident. `gc_orphan_chunks` already existed but wasn't exposed. - New CLI: `claw-store cluster-gc` — runs `gc_orphan_chunks`, prints report. Safe to run any time, safe to interrupt. - New config: `cluster.gc_interval_hours: Option<u64>`. When set to a positive integer, the daemon spawns a periodic ticker that invokes GC in-process. Skips the first tick (nothing to reclaim on boot). Errors are logged and retried next tick. - Shutdown aborts the ticker cleanly. 254 tests pass (baseline unchanged). Pre-existing macOS failure untouched. |
||
|
|
54e9da4d62 |
Phase 5k: parallel-fanout chunk transfer for prewarm
Pilot 2026-07-12 measured 109 MiB/s on the sequential prewarm path — ~11% of a 10G fabric. `quinn::Connection` is cheap-Clone (internal Arc), so we can run the has→get→put pipeline per chunk in concurrent tasks under a bounded semaphore. - `prewarm_missing_chunks_between_parallel(up, down, id, concurrency)` in `rpc/client.rs`. `concurrency <= 1` degrades to the sequential path (kept for diagnostic parity). - `claw-cargo prewarm --parallel N` (default 8). Ignored with `--buffered`. Memory ceiling: 4 MiB × in-flight = 32 MiB @ 8, 128 MiB @ 32. - Uses `tokio::task::JoinSet` + `Arc<Semaphore>`; permit held for the whole per-chunk pipeline so we never over-commit. - Retry pass on `put_manifest` mismatch stays sequential — small, correctness-critical. - Errors: JoinSet drains completely + returns first task error so a mid-fanout failure doesn't leave zombie tasks. +1 test: `end_to_end_parallel_prewarm_copies_chunks_and_matches_sequential` runs 5-chunk payload with concurrency=3, verifies byte-equal restore, then reruns with concurrency=8 → 0 uploads (has_chunk dedup), then concurrency=0 → 0 uploads (sequential fallback path). 254 tests pass (+1 from previous). Pre-existing macOS failure unchanged. |
||
|
|
cb07bfc574 |
capture: stream to a Writer instead of buffering the whole tar in RAM
Field finding 2026-07-12 (clawverse measurement): the buffered `capture_target -> Vec<u8>` path peaked at 2.8 GB RAM to capture a 6.1 GB target/debug into a 995 MiB compressed tar. Every byte crossed RAM before touching the network. * `capture_target_to_writer(target_dir, writer) -> u64` — new streaming variant. Walks the tree + writes tar+zstd straight into the caller's Writer via a small ByteCounter wrapper. Peak memory stays at ~zstd sliding window size (few MB). * `capture_target -> Vec<u8>` kept as a thin wrapper for the tests + smaller callers that don't care. * `cmd_build`: capture into a tempfile under `target/`, then open it with `tokio::fs::File` (AsyncRead + Unpin) and hand that to `call_blob_put_stream`. Same-filesystem tempfile means no cross- mount concerns; auto-unlinks on drop. +1 test: `capture_streaming_matches_buffered_and_restores_correctly` proves the streamed bytes match the buffered variant, the reported byte count agrees with the written length, and roundtrip restore from the streamed file works. Combined with PR #22 (QUIC idle timeout), this closes the two RAM/ timeout blockers surfaced by the clawverse pilot. Expected memory ceiling on a runner drops from GBs to MBs, unlocking small-runner deployments (the actual pitch use case). |
||
|
|
e70f5d74e0 |
Pilot findings: 5 real-world fixes from the 2026-07-12 deploy
Bundles the profile→dir bug (PR #20 supersede) with four new fixes discovered by running clawstor against itself + across the fabric: * target_subdir_for: `dev`/`test` → `debug/`, `release`/`bench` → `release/`, custom passes through. Was silently skipping upload. * rustc release via gossip: daemon probes `rustc --version` at start, publishes the release string as `clawstor.rustc.release`. PeerView carries it; `cluster-peer-status` prints it in a new column and emits a warning line when the fleet has mixed versions. Would have surfaced the tank/architect 1.96.1 vs 1.95.0 drift instantly. * prewarm publishes fingerprint→blob ref downstream: `pin` now writes a companion tag `<name>.fingerprint` holding the fingerprint bytes. `prewarm` reads the companion, PutTag's it downstream, then PutRef(fp→blob) so a subsequent fingerprint-based `build` HITS. Without this, prewarm was almost useless for the runner path (build always missed even with matching source + rustc). * streaming byte counters: BlobPutStream + BlobGetStream now record the transferred bytes via `record_blob_{put,get}_bytes`. Metric used to stay at 0 no matter how much you moved. * capture determinism: replaced `tar::Builder::append_dir_all` (uses `read_dir`'s native order) with `append_dir_sorted` that walks the tree recursively and sorts by filename bytes at every level. Two byte-identical trees now produce byte-identical tars regardless of filesystem ordering. +3 tests: - target_subdir_matches_cargo_layout (from #20) - fingerprint_companion_tag_uses_dotted_suffix - capture_is_order_independent_of_filesystem_readdir (guard against the exact bug we saw in the field) 252 tests pass (+1 from Phase 5h's 251). Pre-existing macOS `du -sb` failure unchanged. Supersedes #20 (also included here). Ready for re-deploy to tank + architect for the retest run. |
||
|
|
6fb16286dd |
Phase 5h: streaming chunk-level prewarm
Bounded-memory cross-peer prewarm: instead of buffering the whole blob in RAM (previous 5f path), iterate the upstream manifest chunk by chunk, ask downstream `HasChunk`, stream missing chunks one at a time. Memory ceiling is 1 chunk (4 MiB) regardless of blob size — a 5 GiB target dir no longer needs 5 GiB of mediator RAM. * rpc/client.rs: new `prewarm_missing_chunks_between(upstream, downstream, blob_id)` helper. Returns `(uploaded, total)` — the difference is the dedup save. Retries once if the downstream `PutManifest` reports missing chunks after our push (guards a narrow eviction race); a second failure surfaces as `Err`. * claw_cargo.rs: `prewarm` now streams by default; new `--buffered` flag for the old whole-blob path (kept for diagnostic comparability during rollout). Human-readable output shows the mode + dedup count. +2 tests: - cold downstream: 3-chunk payload (with a partial tail chunk) is copied exactly, reassembly is byte-equal to source - partial dedup: pre-seed 1 of 3 chunks on downstream → uploaded=2; rerun is a full no-op (uploaded=0), proving idempotence 251 tests pass (+2 from Phase 5j). Pre-existing macOS failure unchanged. Follow-ons: (1) parallel chunk transfer (uploaded chunks in fan-out) would speed multi-GB prewarms further; (2) exposing an accurate transferred-bytes counter needs router-side accounting instead of the current chunk-count × CHUNK_SIZE approximation. |
||
|
|
d36cec11a6 |
Phase 5g: cache metrics + GetMetrics RPC + peer-metrics CLI
Every RPC handler that answers a hit-or-miss question now increments
lock-free atomic counters. The GetMetrics RPC (0x13) returns a JSON
snapshot of every counter; the new claw-cargo peer-metrics CLI
prints hit rates, byte volumes, and counter uptime.
Placement engines can now poll these across the fleet to bias runner
scheduling toward whichever node has the warmest cache for a given
repo/tag combination.
## New module: cluster/metrics.rs (268 lines)
Types:
- CacheMetrics — atomic counters, all AtomicU64, Relaxed ordering
(metrics are advisory, not consistency-critical)
- MetricsReply — JSON snapshot returned by GetMetrics
Public API:
- CacheMetrics::new() — timestamped start, all counters at 0
- record_get_ref_hit / _miss
- record_get_tag_hit / _miss
- record_blob_get_bytes / record_blob_put_bytes
- record_get_chunk_hit / _miss
- record_has_chunk_hit / _miss
- snapshot() — atomic-load every field into a MetricsReply
MetricsReply derived helpers:
- get_ref_hit_rate() / get_tag_hit_rate() / has_chunk_hit_rate() —
Option<f64> so 0/0 returns None instead of NaN
## RPC method
- GetMetrics (0x13): payload = empty; reply = JSON MetricsReply
Wire-level instrumentation added to RpcRouter dispatch:
- GetRef → record_get_ref_hit / _miss
- GetTag → record_get_tag_hit / _miss
- BlobGet → record_blob_get_bytes (on hit)
- BlobPut → record_blob_put_bytes
- HasChunk → record_has_chunk_hit / _miss
- GetChunk → record_get_chunk_hit + record_blob_get_bytes on hit
/ record_get_chunk_miss
RpcRouter grows Arc<CacheMetrics> unconditionally — every router has
metrics, so a fresh node with no traffic still returns a valid
snapshot with all zeros + started_unix.
Streaming variants (BlobPutStream / BlobGetStream) don't yet track
byte counts — they'd require plumbing the count out of put_stream /
stream_to. Follow-on if it turns out to matter for placement.
## Client helper + CLI
- call_get_metrics(&conn) → Result<MetricsReply>
- claw-cargo peer-metrics [--peer ...] [--peer-addr ...] [--tls-dir ...]
Fetches + pretty-prints:
counter uptime: 42s
GetRef hits/misses: 123 / 45
hit rate: 73.21%
GetTag hits/misses: 8 / 2
hit rate: 80.00%
HasChunk hits/miss: 512 / 88
hit rate: 85.33%
GetChunk hits/miss: 47 / 12
Blob GET bytes: 1.23 GiB
Blob PUT bytes: 3.45 GiB
human_bytes() helper picks GiB / MiB / KiB / B based on magnitude.
Subcommand count now 9: build / prefetch / status / fingerprint /
pin / unpin / list-tags / prewarm / peer-metrics.
## Tests (13 new, all real filesystem / real QUIC — no mocks)
CacheMetrics (6):
- new_starts_all_counters_at_zero_except_timestamp
- recorders_increment_the_right_field (every recorder × 1-2 counts)
- hit_rates_none_when_zero_events (avoids 0/0 NaN)
- hit_rates_compute_correctly (3 hits / 1 miss → 75%)
- snapshot_round_trips_through_json
- snapshots_across_threads_are_consistent_up_to_relaxed_ordering
(8 threads × 1000 increments → exactly 8000)
RPC integration (7):
- phase_5g_method_byte_encoding
- get_metrics_returns_empty_snapshot_before_any_activity
- get_ref_records_hit_and_miss_counters (2 hits + 1 miss)
- get_tag_records_hit_and_miss_counters
- blob_get_and_blob_put_record_byte_counts
- has_chunk_and_get_chunk_record_hit_miss_counters
- **end_to_end_get_metrics_over_real_quic** — seed activity locally,
fire GetRef/GetTag/BlobGet dispatches to move the counters, then
fetch metrics through real QUIC + mTLS and verify each field
including the 0.5 hit rate calculation
238 tests pass. Pre-existing macOS-only failure unchanged.
File sizes (all under 1300-line ceiling):
- cluster/metrics.rs: 268
- cluster/rpc.rs: 818
- cluster/rpc/tests_phase5.rs: 905
- claw_cargo.rs: 894
## What this enables
Fleet-wide visibility into which peer is actually serving traffic:
# From anywhere with connectivity + fleet mTLS
claw-cargo peer-metrics --peer tank
claw-cargo peer-metrics --peer architect
claw-cargo peer-metrics --peer morpheus
Compare hit rates side by side to see which node's cache is warmest.
A placement engine can automate this — poll every 30s, feed the
scheduler.
## Follow-on
- 5h: streaming variant of prewarm (fixed-memory ceiling for many-GB
blobs)
- 5i: metrics also published via gossip so PeerView carries hit rate
without a per-peer GetMetrics roundtrip
- 5j: prometheus /metrics endpoint on the daemon for existing dash
integrations
- 6: FUSE mount for warm-tier git worktrees
|
||
|
|
db55903311 |
Phase 5f: claw-cargo prewarm — cross-peer cache copy
The last piece before "Gitea webhook triggers a cache-warm for the
CI runner before its build starts." Adds a `prewarm` subcommand that
copies a tagged cache from one peer (upstream) to another (downstream)
in one shot — same tag, same BlobId, both sides serve it after.
## New subcommand
```
claw-cargo prewarm \
--from-peer tank --from-addr 10.0.0.14:7702 \
--to-peer morpheus --to-addr 10.0.0.15:7702 \
--tls-dir /etc/claw-store/tls \
--pin clawverse:main:latest
```
Flow:
1. Connect to upstream with local mTLS identity
2. `GetTag(tag)` → BlobId; `BlobStat(BlobId)` → size + chunk count
3. `BlobGetStream(BlobId)` → download bytes
4. Connect to downstream (second QUIC endpoint, same identity)
5. `BlobPutStream(bytes)` → returns BlobId; verified equal to upstream's
6. `PutTag(tag → BlobId)` on downstream
Summary output shows tag, blob id, both endpoints, byte count,
download/upload timings, total wall clock.
Assumes upstream + downstream share the same fleet CA (the common
case). Mixed-fleet variant with distinct identities is a follow-on.
## Integrity check
`assigned_id != blob_id` after the downstream upload triggers a
bail — the two BlobIds must match because content is BLAKE3-hashed
end-to-end. If they don't, the wire path corrupted bytes and the
whole prewarm fails loud rather than silently pinning a bad blob.
## Buffered vs streamed
Current implementation buffers the whole blob in memory between
download and upload. Fine for cargo target dirs (~1-5 GB compressed);
would break for a 20 GB blob. A follow-on will pipe upstream → tokio
duplex → downstream to run at fixed memory.
## Tests (1 new, real 2-peer QUIC)
- **`end_to_end_prewarm_copies_tagged_blob_between_two_peers`**
Two full RpcRouters serving in-process (A upstream + C downstream),
each on a distinct port. Seeds A with a blob + tag, then runs the
exact sequence prewarm runs internally: `GetTag → BlobGetStream`
against A, then `BlobPutStream → PutTag` against C. Verifies that
C's blob store returns byte-equal payload and C's tag store now
points at the same BlobId. Proves the composition works.
## Housekeeping
`rpc/tests.rs` hit 1408 lines with the new prewarm test. Phase 5
tests (5b refs + 5d tags + 5e restore + 5f prewarm) split to
`rpc/tests_phase5.rs` via a second `#[path]` module in rpc.rs.
Result:
- rpc/tests.rs: 727 (phase 1-2d tests)
- rpc/tests_phase5.rs: 722 (phase 5 tests)
- rpc.rs: 777
- rpc/client.rs: 589
- claw_cargo.rs: 818
- All under ceiling.
225 tests pass. Pre-existing macOS-only failure unchanged.
## What this enables
The complete CI runner flow now works end-to-end:
```
Primary (e.g. tank):
claw-cargo build # first ever build — MISS, uploads
claw-cargo pin --name clawverse:main:latest
Fleet control plane on PR open:
gitea webhook → shell hook → claw-cargo prewarm \
--from-peer tank --to-peer $RUNNER_LOCAL \
--pin clawverse:main:latest
# runner's local daemon now serves the tag + blob
Runner picks up job:
claw-cargo build --peer 127.0.0.1:7702
# local daemon is warm → prefetch returns HIT
# cargo build runs against restored deps → workspace-crates only
# 50 min → 3 min
```
Every subcommand claw-cargo needs for this pipeline now exists:
build / prefetch / prefetch --pin / status / fingerprint /
pin / unpin / list-tags / prewarm.
## What's next
- 5g: cache hit/miss metrics into gossip so placement engines can
bias runner scheduling toward warm nodes
- 5h: streaming variant of prewarm (tokio duplex) for many-GB blobs
- 6: FUSE mount for warm-tier git worktrees
- 3: full CRDT metadata if the plain-tag model shows conflict problems
|
||
|
|
05ba800d01 |
Phase 5e: prefetch --pin <tag>
Small, focused extension to Phase 5c's prefetch: an optional
`--pin <tag-name>` flag that skips fingerprint compute entirely
and resolves the tag → BlobId via GetTag, then downloads that.
## Use case
Restore an old cache into a fresh checkout for regression testing:
$ claw-cargo prefetch --pin clawverse:main:2026-07-12
cache HIT — downloading 3221225472 bytes (768 chunks) to /path/target/dev
...
── claw-cargo prefetch ─────────────────────────────
source: --pin clawverse:main:2026-07-12
blob: 8c2f1a…
downloaded: 3221225472 bytes in 12.3s
restored to: /path/target/dev
────────────────────────────────────────────────────
Or diagnose a "why does this build fail against the pinned cache"
question by prefetching the tagged cache and then running cargo
against your current source. Cargo will detect the mismatched
.fingerprint state and rebuild affected crates — that's the point,
you're diffing behaviour between two known-good cache snapshots.
## Changes
- New PrefetchArgs struct (was reusing PeerArgs) with an optional
`pin: Option<String>` field
- resolve_pin(conn, tag) — internal helper that does
GetTag → BlobStat, returning None on either NotFound
- cmd_prefetch branches at the top: --pin → resolve_pin(); default
→ fingerprint-based peer_lookup()
- Rest of the flow is unchanged: BlobStat → BlobGetStream →
restore_target
- Summary output shows `source: --pin <tag>` instead of
`fingerprint: <hex>` when the pinned path was taken
`peer_lookup` (fingerprint path) and `resolve_pin` (tag path) return
the same `Option<(BlobId, BlobStat)>` shape so the downstream code
is identical.
## Live smoke test
`prefetch --help` now advertises --pin with full description.
Missing-tag path prints "no such tag: <name>" and exits 0
(consistent with the fingerprint-miss path).
## Tests (1 new, real QUIC)
- **`end_to_end_tag_resolve_and_stream_restore_over_real_quic`** —
seeds blob store with a 2 MiB "captured target" payload, publishes
a tag pointing at its BlobId, then runs the exact client
sequence `prefetch --pin <tag>` runs internally:
GetTag → BlobStat → BlobGetStream
Verifies bytes reassemble byte-equal to source. Also covers the
missing-tag path.
The pin flow uses the same underlying calls tested separately in
Phase 5b/5c/5d, so the new test proves the composition works rather
than re-verifying primitives.
224 tests pass. Pre-existing macOS-only failure unchanged.
## What's next
- 5f: Gitea webhook pre-fetch — daemon receives PR-open hints and
warms cache for the predicted fingerprint before CI runner starts
- 6: FUSE mount for warm-tier git worktrees so `~/projects/clawverse`
is transparently fleet-shared
- 3: full CRDT metadata layer (only if real conflicts emerge in the
simple tag model)
|
||
|
|
b7904b59a5 |
Phase 5d: named tags + pin/unpin/list-tags CLI
Human-readable pins on top of the raw 32-byte ref layer. Operators
publish `clawverse:main:latest-cache` → BlobId once, then everything
downstream (CI runners, dev laptops) references the tag instead of
passing 64-char hex hashes around.
## Module: cluster/tags.rs (433 lines)
TagStore for string-key → 32-byte-value:
- open(root) — creates layout, safe on existing stores
- put(key, value) / get(key) / delete(key) / contains(key)
- list() — sorted by key
- Atomic writes via tempfile + rename
- Key length capped at MAX_TAG_KEY_BYTES (4 KiB); empty keys rejected
On-disk record: `key_len:u16 (LE) || key_bytes || value:32bytes`.
Filename is `blake3(key)` hex so arbitrary UTF-8 keys land at
deterministic paths without shell escaping.
TagEntry type (public, serde) for list results:
`{ key, value_hex }`. Includes `decode_value() → Result<[u8;32]>`.
## RPC methods
- PutTag (0x0f): payload = encoded record → STREAM_STATUS_OK / err
- GetTag (0x10): payload = key bytes → 32-byte value / NotFound
- DeleteTag (0x11): payload = key bytes → STREAM_STATUS_OK / NotFound
- ListTags (0x12): payload = empty → JSON Vec<TagEntry>
RpcRouter grows optional Arc<TagStore> via `.with_tag_store(store)`.
## Services + config
ClusterServices auto-opens a TagStore at `<blob_store_root>/tags-db`
alongside the ref store. `tag_store` field on ClusterServices, same
enable-with-blob-store semantics.
## claw-cargo new subcommands
- `claw-cargo pin --name clawverse:main:latest`
Compute current fingerprint → look up its BlobId via GetRef →
publish TagStore mapping. Errors cleanly if the fingerprint
hasn't been built yet (nothing to point at).
- `claw-cargo unpin --name clawverse:main:latest`
Delete the tag. Prints "no such tag" if it wasn't set.
- `claw-cargo list-tags`
Print every tag with its 32-byte hex value.
Total subcommand count now 7: build / prefetch / status / fingerprint
/ pin / unpin / list-tags. All share the layered config from Phase 5c.
## Client helpers
- call_put_tag / call_get_tag / call_delete_tag / call_list_tags
- All follow the same error-mapping conventions as prior client helpers
(NotFound → Ok(None) or Ok(false), everything else → Err)
## Housekeeping
rpc.rs was pushing past the 1300-line ceiling with the tag methods
added. Client helpers moved to `cluster/rpc/client.rs` with a
re-export (`pub use client::*;`) so external callers still write
`cluster::rpc::call_*`. Result:
- rpc.rs: 773 (was 1343)
- rpc/client.rs: 593 (new)
- rpc/tests.rs: 1235
- All under ceiling.
## Tests (33 new, all real filesystem / real QUIC — no mocks)
TagStore (16 in cluster/tags.rs):
- open_creates_layout
- get_returns_none_for_missing (+ contains false)
- put_and_get_round_trip
- put_overwrites_prior_value
- delete_returns_true_for_existing_and_false_for_missing
- put_rejects_empty_key
- put_rejects_oversize_key
- list_returns_all_tags_sorted
- list_is_empty_on_fresh_store
- keys_with_slashes_and_colons_round_trip (real-world tag shape)
- encode_and_decode_round_trip (raw wire format)
- decode_rejects_short_record
- decode_rejects_length_mismatch
- decode_rejects_non_utf8_key
- tag_entry_decode_value_round_trip
- tag_entry_decode_value_rejects_bad_hex
RPC dispatch (7 new):
- phase_5d_method_byte_encoding
- tag_rpcs_return_not_configured_without_store
- put_tag_stores_and_get_tag_reads_back
- get_tag_returns_not_found_for_missing
- get_tag_rejects_empty_key
- delete_tag_removes_and_returns_not_found_after
- list_tags_returns_json_sorted
End-to-end over real QUIC (1):
- **end_to_end_pin_lookup_delete_over_real_quic** — publish tag →
look up → list → delete → confirm gone. Full round trip through
the wire layer including JSON deserialization of the list.
Also 8 downstream tests continued passing after the client.rs split
(no test moved, they were untouched).
223 tests pass. Pre-existing macOS-only failure unchanged.
## What this enables
Operator flow:
# Build once on the primary
$ claw-cargo build
→ cache MISS → cargo build (50 min) → capture + upload
→ summary: fingerprint 4a3b…, blob 8c2f…, uploaded 3.2 GiB
# Publish a friendly name
$ claw-cargo pin --name clawverse:main:2026-07-12
pinned: clawverse:main:2026-07-12
fingerprint: 4a3b2c…
blob: 8c2f1a…
# Anyone else can now find it via list-tags
$ claw-cargo list-tags
clawverse:main:2026-07-12 8c2f1a…
clawverse:main:latest 8c2f1a…
# CI runner sees the same fingerprint in its workspace state, hits
# the ref directly via GetRef — the tag is for operator visibility
## Follow-on
- 5e: prefetch --pin <tag> — bypass fingerprint compute, download
the tagged BlobId directly (useful when you want an old cache to
test regression scenarios)
- 5f: gitea webhook pre-fetch — daemon pre-warms cache for known
fingerprints before CI runner starts
- 3: full CRDT metadata layer (namespaces, versioned pointers,
vector clocks) if the plain-tag model turns out to have
real-world conflict scenarios
|
||
|
|
57d8358255 |
Phase 5c: claw-cargo UX — config files + status + prefetch
Ships the last-mile ergonomics that make claw-cargo actually usable
day-to-day: layered config files so you don't retype --peer-addr on
every invocation, plus two lightweight subcommands (status +
prefetch) for the "what's in the cache" and "warm my target dir"
workflows respectively.
## Config precedence
Later wins:
1. Built-in defaults (profile=dev, features=[])
2. ~/.claw-cargo/config.toml (per-user defaults)
3. <workspace>/.claw-cargo.toml (per-repo overrides)
4. CLI flags (per-invocation overrides)
Shape:
[peer]
name = "tank"
addr = "10.0.0.14:7702"
tls_dir = "/etc/claw-store/tls"
[build]
profile = "release"
features = ["a", "b"]
## New module: cluster/client_config.rs (496 lines)
- ClientConfig / PeerSection / BuildSection — TOML-serialisable, all
Option<> fields at every layer so partial configs are legal
- ClientConfig::from_toml_str / from_file_or_default (missing file →
default, not error)
- ClientConfig::merge — Option::Some in `other` wins over `self`
- ClientConfig::load_layered(workspace) — user → workspace
- ResolvedClientConfig — final flattened shape after CLI overrides,
with require_peer_name / require_peer_addr / require_tls_dir /
require_peer_bundle helpers that produce a specific error message
instead of "some Option was None"
- write_config_file — for tests + future `claw-cargo config init`
Ships with 11 unit tests including a load_layered test that fakes
HOME + workspace via a tempdir, writes both configs, verifies the
workspace override takes precedence.
## claw-cargo (rewritten to 469 lines)
Four subcommands with layered config:
claw-cargo fingerprint [--profile ...] [--features ...] [--workspace ...]
→ local-only, no network
claw-cargo status <peer args>
→ connect + GetRef + BlobStat, print hit/miss + size, no download
claw-cargo prefetch <peer args>
→ hit → BlobGetStream + restore_target, no cargo
claw-cargo build <peer args> [-- extra cargo args]
→ same as Phase 5b flow, now with layered config for peer args
Refactored internals:
- setup_local / setup_peer — figure out workspace, load config,
resolve CLI overrides, compute fingerprint
- connect_peer — load NodeIdentity, open QUIC connection
- peer_lookup — GetRef → BlobStat, handle the "ref points at a
garbage-collected blob" case as a miss
## Live smoke test
Verified end-to-end on this workspace:
# No config file → built-in defaults
$ claw-cargo fingerprint
profile: dev, features: (none), fingerprint: 8ee4cf…
# Add .claw-cargo.toml with profile=release + features=some-feature
$ claw-cargo fingerprint
profile: release, features: some-feature, fingerprint: 2dfcb1…
# CLI overrides just the profile; features fall through from config
$ claw-cargo fingerprint --profile dev
profile: dev, features: some-feature, fingerprint: 84a756…
# `status` without peer args → clean validation error
$ claw-cargo status
Error: peer.name not set (config file or --peer)
## Tests (11 new, all real filesystem — no mocks)
- from_toml_str_parses_full_config
- from_toml_str_handles_partial_sections (peer.name only)
- from_file_or_default_returns_default_when_missing
- merge_prefers_later_over_earlier (unset fields fall through)
- resolve_applies_cli_overrides_over_layered
- resolve_falls_through_to_default_profile_when_unset_everywhere
- require_peer_bundle_errors_when_incomplete (specific error text)
- validate_peer_errors_on_missing_field
- load_layered_reads_both_files — fake HOME + workspace, verifies
workspace override takes precedence
- user_config_path_uses_home
- write_and_read_round_trip_via_disk (nested dir creation)
199 tests pass. Pre-existing macOS-only failure unchanged.
File sizes (well under 1300-line ceiling):
- cluster/client_config.rs: 496
- claw_cargo.rs: 469
## What's next
The CLI is now usable day-to-day. Realistic next steps:
- 5d: publish cache hit/miss metrics into gossip so the placement
engine can bias runner scheduling toward warm nodes
- 5e: pre-fetch on Gitea webhook — daemon receives a "PR opened for
fingerprint X" hint and warms the local cache before the runner
even starts pulling
- 3: CRDT metadata for human-readable pins (`clawverse:main:latest`
→ fingerprint hex) so operators can pin cache versions without
passing raw hashes around
- 6: FUSE mount so `~/projects/clawverse` on any node is transparently
the tank-hosted canonical warm-tier copy
|
||
|
|
2d09b4687c |
Phase 5b: KV refs + claw-cargo CLI (the killer feature, live)
Ships the actual user-facing cargo build cache. Combined with Phase 5a
(fingerprint + capture + restore) + the whole Phase 2 blob substrate,
`claw-cargo build` now runs `cargo build` with a peer-cache lookup:
hit → download+restore, miss → build+capture+upload.
## What ships
### cluster/refs.rs (243 lines)
A dumb 32-byte-key → 32-byte-value directory-backed store. Used to map
fingerprints → BlobIds. Layout mirrors BlobStore:
<root>/
refs/<kk>/<key_hex>.ref — 32 raw bytes
.tmp/ — atomic-rename staging
Public API: RefStore::open / get / put / delete / contains. All writes
atomic via tempfile + rename. Deliberately no versioning or CRDT
semantics — that's Phase 3. Every real cargo-cache lookup is a
single-key-single-value shape.
### New RPC methods
- GetRef (0x0d): payload = 32-byte RefKey; reply = 32 bytes / NotFound
- PutRef (0x0e): payload = 32-byte RefKey || 32-byte RefValue;
reply = STREAM_STATUS_OK / error
### RpcRouter + services
- RpcRouter grows optional Arc<RefStore> via `with_ref_store`
- ClusterServices opens a RefStore alongside the BlobStore when
`blob_store_root` is configured (co-located at `<blob_root>/refs-db`)
- `blob_store_enabled()` / `ref_store_enabled()` introspection
### claw-cargo binary (319 lines)
New bin target `claw-cargo` — thin CLI wrapping the whole stack:
claw-cargo fingerprint --profile release --features "a,b"
→ prints the workspace fingerprint (no network)
claw-cargo build \
--peer <name> --peer-addr <ip:port> --tls-dir <dir> \
--profile release --features "a,b" \
-- --workspace=x --frozen ...
→ 1. compute fingerprint
2. QUIC + mTLS connect to peer
3. GetRef(fingerprint) → BlobId?
HIT: BlobStat → BlobGetStream → restore_target → cargo build
MISS: cargo build → capture_target → BlobPutStream → PutRef
4. Print summary: fingerprint, hit/miss, bytes, cargo elapsed
## Live smoke test
Ran claw-cargo fingerprint on this workspace with three profile/feature
combos — got three distinct 32-byte fingerprints. Same profile+features
on the same workspace state → same fingerprint (Phase 5a's guarantee
carried through the CLI).
## Tests (14 new, all real — no mocks)
Refs store (7):
- open creates layout
- get returns None for missing
- put + get round-trips
- put overwrites prior value
- delete removes ref + reports (false on second delete)
- distinct keys produce distinct on-disk files (bucket fan-out proof)
- rejects_wrong_length_on_disk (corruption detection)
RPC (7):
- phase_5b_method_byte_encoding
- get_ref_returns_not_found_for_missing
- put_ref_stores_and_get_ref_reads_back
- put_ref_rejects_wrong_length_payload
- get_ref_rejects_wrong_length_payload
- ref_rpcs_return_not_configured_without_store
- end_to_end_put_ref_get_ref_over_real_quic — full 2-node QUIC + mTLS
round trip proving PutRef/GetRef work at the wire level
188 tests pass. Pre-existing macOS-only failure unchanged.
File sizes (all under 1300-line ceiling):
- cluster/refs.rs: 243
- cluster/rpc.rs: 1169
- cluster/rpc/tests.rs: 1073
- cluster/services.rs: 565
- claw_cargo.rs: 319
## Where this leaves us
The distributed FS + cargo cache is functionally complete for the
happy path:
Node A builds clawverse for the first time
→ cargo build (50 min cold)
→ capture_target (a few seconds)
→ push to node B via BlobPutStream (network-bound)
→ PutRef(fingerprint → BlobId)
Node B on the same workspace state runs `claw-cargo build …`
→ compute_fingerprint (ms)
→ GetRef → hit
→ BlobGetStream (network-bound)
→ restore_target (a few seconds)
→ cargo build → sees valid deps/.fingerprint, builds only
workspace crates (~3 min instead of 50)
Same workspace state on a third machine? Same fingerprint → same
cache hit. That's the whole design.
## Follow-on
- Phase 5c: pre-fetch on Gitea webhook so CI runners never wait
- Phase 5d: metric ticker publishes cache hit rate into gossip so
the placement engine can bias runner scheduling toward warm nodes
- Phase 3: CRDT metadata for human-readable pins on top of raw
32-byte refs (`clawverse:main:latest-cache` → fingerprint hex)
- Phase 6+: FUSE mount for the warm-tier git worktrees
|