8-crate pure-Rust workspace for revision-aware HDF5 sync. ## Crates - clawhdf5-onion: ClawOnion VFD — page-level versioned HDF5 storage, binary format, writer/reader, branch DAG, GC, snapshots, provenance - clawsync-core: BLAKE3, xxHash3, FastCDC (+ SIMD NEON), zstd/lz4 - clawsync-onion: IBLT sketch, Merkle tree differ, packet differ/merger, ClawSyncManifest, SyncSelector - clawsync-hdf5: dataset-level manifest, differ, patcher, wire payload reconstruction (apply_received_payloads) - clawsync-transport: TCP, QUIC (quinn 0.11/TLS 1.3), SyncPeer abstraction, length-prefixed rkyv wire protocol (21 SyncMessage variants) - clawsync-agent: OnionMemory, SyncScheduler, TcpSyncBackend, PeerCapabilities negotiation - clawsync-fs: CDC-based delta sync for any file type; FsSyncClient/Server, W=16 pipelining, atomic writes - clawsync-cli: push/pull/serve/hdf5-sync/serve-hdf5/sync/serve-fs + all local management commands; --quic on all network commands ## Key features - IBLT pre-flight: O(revision count) vs rsync's O(file size) - W=16 sliding-window push: 13–15x speedup over stop-and-wait at WAN RTT - Dataset-granular HDF5 sync: only modified datasets transferred - CDC delta for any file type: insertion-stable chunk boundaries - Full revision DAG: branch, merge, rollback, export, snapshot, GC - QUIC transport: TLS 1.3, per-message streams via quinn 0.11 ## Tests ~573 passing (default features); ~589 with --features simd-cdc ## Performance (Apple Silicon) - Reconstruct rev=100: 68 µs (target ≤ 1 ms) - BLAKE3 Rayon 1 MB: 10.3 GiB/s (target ≥ 5 GB/s) - GC 500 revisions: 20.6 µs (target ≤ 2 s) - W=16 vs W=1 at 5 ms RTT: 14.8x speedup - No-op pre-flight at 16 MB: 4 ms vs rsync 35 ms (7.8x) Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
1.4 KiB
1.4 KiB
clawsync-core
Low-level primitives for ClawSync: checksums, content-defined chunking, delta computation, and compression.
Overview
clawsync-core provides the performance-critical building blocks used by the
higher-level ClawSync crates. All hot paths are Rayon-parallelized and target
≥10 GB/s BLAKE3 throughput on AVX2 hardware.
Modules
checksum
blake3_hash(data)— single-call BLAKE3blake3_hash_large(data)— Rayon tree hashing for inputs ≥128 KB (≥10 GiB/s)blake3_hash_batch(chunks)— parallel batch hashing viapar_iterblake3_hex(data)/hash_to_hex(hash)— lookup-table hex encodingxxh3_hash(data)— xxHash3 for fast block checksums
cdc
Content-defined chunking via fastcdc. Produces variable-length chunks with
stable boundaries for efficient delta transfer.
delta
Parallel stripe delta computation (Rayon). Produces Vec<DeltaBlock> from two
byte slices.
compress
Streaming zstd and lz4 frame wrappers for block-level compression.
Performance
| Operation | Throughput |
|---|---|
| BLAKE3 single-threaded | ~3 GiB/s |
| BLAKE3 Rayon (1 MB input) | ≥10 GiB/s |
| BLAKE3 batch (256 × 4 KB) | ≥9 GiB/s |
| xxHash3 | ≥20 GiB/s |
Requires RUSTFLAGS="-C target-cpu=native" for full SIMD throughput (set in
.cargo/config.toml of the workspace).
License
MIT — see repository root.