ec24f37d90a3f9b18b31e309882f6f33fdf864eb
11
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f38efc7096 |
Honor zfs_dataset = "none" instead of erroring; surface shutdown-prep panel from fleet view
Build with clawstor cache / Cargo build (clawstor-cached) (pull_request) Failing after 4s
Two problems surfaced while checking on the fleet after the shutdown- prep button PR: 1. morpheus is configured with zfs_dataset = "none" (it has no ZFS pool -- warm tier is a plain directory on the LVM root volume), but nothing in the code actually implemented that as a sentinel. cmd_snapshot/cmd_replicate and the daemon's periodic snap/repl ticks always tried real zfs/zpool calls regardless, producing "zfs: command not found" errors on every hourly tick and in the shutdown-prep report. WarmConfig::zfs_enabled() now gates all four call sites; a non-ZFS node gets a clean "nothing to snapshot/replicate" instead of a raw shell error. 2. safe-shutdown-prep.sh's zpool-health step now checks `command -v zpool` first instead of leaking "zpool: command not found" into the report. 3. "we don't see the button" turned out to be page confusion: the shutdown-prep panel lives on the per-node detail page (/v2/nodes/<name>), not the root Fleet Health landing page. Added a small "view detail · maintenance & shutdown prep →" hint to the bottom of every NodeCard so it's discoverable without already knowing to click through. Verified against tank, architect, and morpheus -- morpheus's shutdown-prep --dry-run report is now clean (no "command not found" lines) both when run locally and via cross-node RPC from tank. Co-Authored-By: Claude Sonnet 5 <[email protected]> |
||
|
|
75d822d0f2 |
Fail fast when cluster gossip bootstrap fails
Bind failures at startup are almost always a boot-time race against DHCP/network-online (bind address not yet assigned to the interface). Previously the daemon caught the error and kept running in a degraded state with no gossip, RPC, or Prometheus endpoint and no visible failure signal. Now it propagates the error so the process exits and systemd's Restart=on-failure retries once the network is actually up. Co-Authored-By: Claude Sonnet 5 <[email protected]> |
||
|
|
f43b34ad15 |
Phase 2b: Blob RPC (Stat/Get/Put/LoadManifest)
Wires the Phase 2 content-addressed blob store onto the network via
four new methods on the existing RpcRouter. Combined with the mTLS +
gossip stack from Phase 1c-1e, peers can now exchange content-addressed
blobs over real QUIC. This is the substrate the fingerprint-keyed cargo
cache (Phase 5) sits on directly.
New methods:
- BlobStat (0x03): payload = 32-byte BlobId, reply = JSON BlobStat
- BlobGet (0x04): payload = 32-byte BlobId, reply = raw bytes
- BlobPut (0x05): payload = raw bytes, reply = 32-byte BlobId
- BlobLoadManifest (0x06): payload = 32-byte BlobId, reply = JSON manifest
New error codes:
- NotFound (0xf3) — the requested BlobId isn't in the local store
- InvalidRequest (0xf4) — e.g. non-32-byte payload for a hash-keyed method
- NotConfigured (0xf5) — Blob* called on a router without an attached store
Wire-format bump: MAX_MESSAGE_BYTES 16 KiB → 16 MiB so a single 4 MiB
chunk (plus JSON overhead) fits comfortably. Anything above 16 MiB
needs the streaming variants coming in Phase 2c.
Client helpers:
- call_blob_stat / call_blob_get / call_blob_put / call_blob_load_manifest
- All map ErrorCode::NotFound to Ok(None), other codes to Err.
- call_blob_put verifies the peer-assigned BlobId matches local blake3
hash of the payload — corruption or protocol drift surfaces
immediately instead of silently accepting a mismatched receipt.
Router changes:
- RpcRouter grows an Option<Arc<BlobStore>> via with_blob_store(store).
- handle() split into handle_outcome() → HandlerOutcome enum
{Reply(bytes) | Error(ErrorCode)} for cleaner control flow across
the growing method set.
Services / daemon:
- ClusterServices::start gains blob_store_root: Option<PathBuf>.
- ClusterConfig gains blob_store_root: Option<PathBuf>.
- daemon.rs reads it from cluster_cfg + passes through.
- New ClusterServices::blob_store_enabled() introspection.
Tests (11 new, all real — no mocks):
Router (7 new):
- method_round_trips_byte_encoding — updated for 6 methods
- error_code_describe_covers_all_variants — updated for 6 codes
- decode_error_covers_all_known_codes
- blob_rpcs_return_not_configured_without_store — all four Blob*
methods return NotConfigured when the router lacks a store
- blob_stat_returns_not_found_for_missing
- blob_stat_returns_json_for_existing
- blob_stat_returns_invalid_request_for_bad_length
- blob_get_returns_content_bytes
- blob_put_stores_bytes_and_returns_hash — verifies BlobId matches
independent local hash
- blob_load_manifest_returns_json_for_existing (2-chunk case)
- blob_load_manifest_returns_not_found_for_missing
End-to-end over real QUIC (2 new):
- end_to_end_blob_put_stat_get_over_real_quic — full 4-method loop
(Put → Stat → Get → LoadManifest) + NotFound path
- end_to_end_multi_chunk_blob_over_real_quic — 6 MiB blob → 2 chunks,
proves MAX_MESSAGE_BYTES bump took effect
Services (1 new):
- services_with_blob_store_serves_blob_rpc_end_to_end — cut CA, sign
leaves, config includes blob_store_root, start ClusterServices,
dial from B over persisted mTLS, put + get through the router,
then independently verify bytes landed on A's on-disk store
Also fixed a parallel-test port collision: services `next_port()`
now increments by 2 so `port + 1` (the RPC bind) is reserved
alongside `port` (the gossip bind).
128 tests pass. Pre-existing macOS-only failure unchanged.
File sizes (all under 1300-line ceiling):
- cluster/rpc.rs: 1006
- cluster/services.rs: 535
- cluster/blob.rs: 802
- config.rs: 511
- daemon.rs: 265
Follow-on:
- 2c: streaming variants (AsyncRead/AsyncWrite) so a many-GB blob
transfers without holding it in memory
- 2d: chunk-level RPC (BlobPutChunk / BlobGetChunk) so a receiver
can request only chunks it's missing after LoadManifest
- Phase 5 (the killer feature) can now build on Phase 2b directly —
fingerprint the target dir, PutBlob the compressed tarball, and
next node calls GetBlob keyed by the same fingerprint hash.
|
||
|
|
2ae079fbd4 |
Phase 1e: daemon-integrated cluster services + PeerStatus RPC
Ties Phase 1a-1d together into a live subsystem the daemon actually
runs. When `[cluster]` is present in config, `claw-store daemon` now:
1. Starts chitchat gossip via ClusterServices::start.
2. Publishes hot.max_gb immediately and re-measures the hot dir
every 30s, updating `clawstor.hot.used`.
3. If `[cluster.tls]` is also configured, loads NodeIdentity from
PEM files and binds a QUIC RPC server that accepts + dispatches
incoming Ping / PeerStatus requests via RpcRouter.
New modules:
- cluster/rpc.rs (510 lines):
- Method enum (Ping=0x01, PeerStatus=0x02)
- ErrorCode enum (EmptyRequest / UnknownMethod / HandlerFailure)
- PeerStatusReply { local_name, local_zone, peers: Vec<PeerView> }
- RpcRouter — dispatch state (holds Arc<ClusterGossip> + local
name/zone)
- rpc_call / call_ping / call_peer_status — client helpers
- serve_connection — server accept-bidi loop
- Wire format: `method:u8 || payload:bytes` → `reply:bytes` or
single-byte ErrorCode
- cluster/services.rs (418 lines):
- ClusterServices { gossip, accept_task, metric_task }
- start(cluster_cfg, name, hot_dir, hot_max_bytes) → bootstraps
everything above
- shutdown() aborts background tasks cleanly
- Recursive dir-walker (spawn_blocking) for hot-tier metric
PeerView now derives Serialize/Deserialize so it round-trips through
JSON over the RPC.
daemon.rs integration (~35 lines added):
- Bootstraps ClusterServices before entering the select loop
- Held for daemon lifetime
- Shutdown on SIGTERM
- Gossip-less config still runs standalone (backwards compat)
CLI: `cluster-peer-status --peer <name> --rpc-addr <addr>
--tls-dir <dir>`
Loads a persistent identity, dials the peer, calls PeerStatus,
prints the peer's local view as a table.
Tests (15 new, all real — no mocks, real UDP + TLS + JSON round trip):
RPC (9):
- method_round_trips_byte_encoding
- dispatch_returns_pong_for_ping
- dispatch_returns_json_for_peer_status
- dispatch_returns_empty_request_error
- dispatch_returns_unknown_method_error
- rpc_call_rejects_oversize_payload (with real quinn connection)
- end_to_end_ping_and_peer_status_over_real_quic — full 2-node quinn
with mTLS + both RPCs
- peer_status_reflects_peer_gossip_state — 2 gossip instances converge,
RPC caller from a third identity sees the converged view
- error_code_describe_covers_all_variants
Services (6):
- dir_bytes_sync_returns_zero_for_missing_path
- dir_bytes_sync_sums_recursive_file_sizes (3-level nesting)
- services_start_without_tls_leaves_rpc_disabled
- services_start_with_tls_serves_rpc_end_to_end — full stack: cut CA on
disk, sign leaves, start ClusterServices for A, dial from B via
persisted mTLS, run ping + PeerStatus over the wire
- services_publish_hot_used_metric_periodically
- services_gossip_sees_peer_after_convergence
95 tests pass. Pre-existing macOS-only failure unchanged.
File sizes (all under 1300-line ceiling):
- cluster/rpc.rs: 510
- cluster/services.rs: 418
- cluster/gossip.rs: 579
- cluster/transport.rs: 916
- daemon.rs: 262
- main.rs: 749
Phase 1 done end-to-end. `claw-store daemon` now boots a real distributed
cluster stack when config asks for one; peers can call each other's
PeerStatus RPC and see live membership. Next: Phase 2 (content-addressed
blob store) can hook new RPC methods into the same RpcRouter with a
one-line dispatch arm.
|
||
|
|
aefa1cce58 |
URGENT FIX: stale-gc must not treat last_active=None as stale
Previous commit
|
||
|
|
643ba170b6 |
daemon: proactively deactivate stale (>48h idle) projects each tick
Answers a real operational question — 'why does /hot/targets stay near
full even when I haven't touched most of these projects for weeks?'
Previously the stale sweep was gated behind 'hot tier > 90% full', so
an idle project held NVMe until the operator manually deactivated it
or space ran out and LRU came for it.
The change:
* new claw-store/src/actions.rs — shared deactivate_project(cfg,
manifest_path, project) with the FULL flow: sync-to-peer +
hot rm + .cargo/config.toml removal + manifest update, plus a
DeactivateOutcome { synced, sync_error, freed_bytes } return.
Extracted from main.rs::cmd_deactivate so both CLI and daemon
run the same code — no drift between manual and auto semantics.
* daemon.rs poll_tick now runs the stale sweep on EVERY tick,
independent of space pressure. Each project idle > stale_hours
(48 by config) goes through the full deactivate, logs
'stale-gc: X freed N MB (synced=Y)'.
* gc_by_space (LRU) still gates on >90% full — it's the emergency
'even after the stale sweep we're still tight' path.
* main.rs cmd_deactivate is now a thin CLI wrapper that adds
println! feedback + reports freed MB in the terminal output.
Space-pressure LRU keeps its original .cargo/config.toml-preserving
semantics (rm hot only, not the shim) so a project marked 'active'
by manual activation but hard-evicted for space can still be
re-activated cheaply. Stale sweep is the aggressive one because the
project genuinely hasn't been touched.
Tests: two new unit tests in actions.rs cover the full flow +
idempotency; the freed_bytes assertion is Linux-only-safe (BSD du
returns 0 for -sb, same limitation as the existing hot test).
|
||
|
|
92b2dd751e |
daemon: auto-enqueue sync when a project's git HEAD moves
Removes the biggest workflow footgun in the architect↔tank flow: the
operator no longer has to remember to run 'claw-store sync <project>'.
An ordinary 'git commit' anywhere under /slab/projects/*/* is now
picked up automatically on the next 5-min poll tick and the peer is
notified — same code path deactivate/sync used to trigger by hand.
New module claw-store/src/head_watch.rs owns the state:
* HeadCache — {project → last-seen HEAD sha}, persisted at
/var/lib/claw-store/head-cache.toml (same directory as the
existing sync queue and manifest).
* scan_and_enqueue walks warm_root/<org>/<repo>, calls
'git rev-parse HEAD' on each, compares to cache, enqueues on
change. Non-git dirs, symlinks, and unreadable entries are
skipped (matches the project-list walker's behavior).
First-scan policy: if the cache is empty on startup we populate it
silently instead of enqueueing every project — otherwise a fresh
install with 55 warm projects would flood the peer with 55 syncs on
the first tick. Existing state gets baselined; only movement from
that point on triggers work.
daemon.rs poll_tick was one 'retry the queue' block; it's now
'scan for HEAD changes → save cache → drain queue' in that order,
so a commit from between ticks gets enqueued AND drained the
same tick.
Four unit tests cover the invariants: silent first scan, enqueue
on real HEAD move (only for the moved repo), cache roundtrip,
non-git directory skip.
|
||
|
|
1b1d66eac9 |
feat(v0.3.0): weekly snapshots, incremental replication, auth, uptime, hardening
- snapshot: implement weekly ZFS snapshots (Sunday midnight, retain_weekly config) - snapshot: incremental cold replication via zfs send -i; tracks last replicated snapshot in /var/lib/claw-store/last-replicated-snapshot - daemon: write start-time file for uptime reporting; add SIGTERM graceful shutdown - serve: daemon_uptime_secs now reads the start-time file (was hardcoded 0) - serve: validate project names in POST handlers (org/repo slug, no traversal) - serve: Bearer token auth middleware on all POST endpoints via cfg.api_token - config: add optional api_token field (backward compatible, defaults to None) - sync: add SSH timeouts (ConnectTimeout=10, ServerAliveInterval=5) to peer notify - hot/sync/main: replace unwrap() on path-to-str with proper anyhow errors - config/tank.toml: document 10G fabric IP and nightly_at field intent - .gitignore: exclude compiled dashboard assets (claw-store/static/) - Makefile: add build, install, install-systemd, install-dashboard, deploy targets - tests: 29 passing (up from 25); 4 new weekly snapshot tests Co-Authored-By: Claude Sonnet 4.6 <[email protected]> |
||
|
|
af50adec19 |
feat(v0.2.0): atomic+flocked manifest, pinned projects, serve systemd unit
The two biggest pain points coming out of the architecture review:
1. The manifest at /var/lib/claw-store/projects.toml was the only
piece of writable state but had no locking, no atomic writes,
and three concurrent writers (daemon poll tick, every CLI verb,
and the dashboard shelling out via /api/activate). Two writers
interleaving silently dropped one of them; a crash mid-write
left a corrupt half-written TOML that the next reader parsed
as an empty manifest.
2. Reboot survival: the dashboard had no systemd unit and was a
stray hand-launched process. Architect lost its dashboard on
todays reboot.
This commit lands:
- Manifest::update(path, FnOnce(&mut Manifest)) — locked-atomic
load-mutate-save in one transaction. Uses libc::flock(LOCK_EX) on
a sidecar .lock file (so the data file can be replaced by rename
without invalidating the lock) and tempfile + persist for the
rename. Concurrent writers serialise; readers see the previous
state or the new state, never a torn write. Manifest::load uses
LOCK_SH so it never races a mid-rename.
- Project.pinned: bool with #[serde(default)] so legacy manifests
parse cleanly. hot::gc_stale_targets and hot::gc_by_space both
skip pinned projects, with a WARN log when every remaining
project is pinned but were still over budget — operator intent
beats space pressure.
- claw-store pin <project> / unpin <project> CLI verbs.
status command surfaces pin marker (📌).
activate preserves an existing rows pinned flag so re-activating
doesnt silently unpin.
- claw-store-serve.service systemd unit. Type=simple, Restart=
on-failure, RestartSec=15, ProtectSystem=strict + ReadWritePaths
=/var/lib/claw-store, ProtectHome, NoNewPrivileges, PrivateTmp.
- daemon poll tick reloads the manifest from disk at the start of
each cycle (so CLI activations between ticks are visible) and
routes its GC write through Manifest::update (so it cant race a
concurrent CLI pin).
- libc + tempfile move from dev-deps into runtime deps.
- empty-manifest fallthrough on load (treat "" as default) so a
half-written tempfile crashed pre-rename doesnt hard-fail the
daemon next boot.
- 25 tests passing incl. new ones: legacy-toml-parses, update-
serializes-two-sequential-writers, pinned-survives-stale-gc,
pinned-survives-space-gc-even-when-lru.
Version bumped 0.1.0 → 0.2.0.
Co-Authored-By: Claude Opus 4.7 <[email protected]>
|
||
|
|
a9c12173fe |
feat: activate/deactivate/list/sync/pull commands + peer sync queue
- activate <org/repo>: wire hot tier, write .cargo/config.toml, register - deactivate <org/repo>: auto-sync to peer, evict hot tier, unregister - list: show all warm repos with activation status and last-active time - sync <org/repo>: git push origin then SSH peer to pull - pull <org/repo>: git pull from origin, stamp last_sync in manifest - SyncQueue: file-based retry queue (/var/lib/claw-store/sync-queue.toml) - daemon: drains sync queue on every 5-min poll tick - config: [peer] host/user for cross-node notification - manifest: Project.name is now org/repo, adds last_sync field - hot tier paths are org/repo-namespaced to avoid collisions Co-Authored-By: Claude Sonnet 4.6 <[email protected]> |
||
|
|
7ae1c4f950 | feat: daemon event loop and full CLI (init, status, gc, snapshot, restore, replicate) |