Files
clawsync/crates/clawsync-core/README.md
osobhandClaude Sonnet 4.6 260e15f5b6 Initial commit: ClawSync v0.1.0
8-crate pure-Rust workspace for revision-aware HDF5 sync.

## Crates
- clawhdf5-onion: ClawOnion VFD — page-level versioned HDF5 storage,
  binary format, writer/reader, branch DAG, GC, snapshots, provenance
- clawsync-core: BLAKE3, xxHash3, FastCDC (+ SIMD NEON), zstd/lz4
- clawsync-onion: IBLT sketch, Merkle tree differ, packet differ/merger,
  ClawSyncManifest, SyncSelector
- clawsync-hdf5: dataset-level manifest, differ, patcher, wire payload
  reconstruction (apply_received_payloads)
- clawsync-transport: TCP, QUIC (quinn 0.11/TLS 1.3), SyncPeer abstraction,
  length-prefixed rkyv wire protocol (21 SyncMessage variants)
- clawsync-agent: OnionMemory, SyncScheduler, TcpSyncBackend,
  PeerCapabilities negotiation
- clawsync-fs: CDC-based delta sync for any file type; FsSyncClient/Server,
  W=16 pipelining, atomic writes
- clawsync-cli: push/pull/serve/hdf5-sync/serve-hdf5/sync/serve-fs +
  all local management commands; --quic on all network commands

## Key features
- IBLT pre-flight: O(revision count) vs rsync's O(file size)
- W=16 sliding-window push: 13–15x speedup over stop-and-wait at WAN RTT
- Dataset-granular HDF5 sync: only modified datasets transferred
- CDC delta for any file type: insertion-stable chunk boundaries
- Full revision DAG: branch, merge, rollback, export, snapshot, GC
- QUIC transport: TLS 1.3, per-message streams via quinn 0.11

## Tests
~573 passing (default features); ~589 with --features simd-cdc

## Performance (Apple Silicon)
- Reconstruct rev=100: 68 µs (target ≤ 1 ms)
- BLAKE3 Rayon 1 MB: 10.3 GiB/s (target ≥ 5 GB/s)
- GC 500 revisions: 20.6 µs (target ≤ 2 s)
- W=16 vs W=1 at 5 ms RTT: 14.8x speedup
- No-op pre-flight at 16 MB: 4 ms vs rsync 35 ms (7.8x)

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
2026-04-04 18:41:22 -05:00

1.4 KiB
Raw Permalink Blame History

clawsync-core

Low-level primitives for ClawSync: checksums, content-defined chunking, delta computation, and compression.

Overview

clawsync-core provides the performance-critical building blocks used by the higher-level ClawSync crates. All hot paths are Rayon-parallelized and target ≥10 GB/s BLAKE3 throughput on AVX2 hardware.

Modules

checksum

  • blake3_hash(data) — single-call BLAKE3
  • blake3_hash_large(data) — Rayon tree hashing for inputs ≥128 KB (≥10 GiB/s)
  • blake3_hash_batch(chunks) — parallel batch hashing via par_iter
  • blake3_hex(data) / hash_to_hex(hash) — lookup-table hex encoding
  • xxh3_hash(data) — xxHash3 for fast block checksums

cdc

Content-defined chunking via fastcdc. Produces variable-length chunks with stable boundaries for efficient delta transfer.

delta

Parallel stripe delta computation (Rayon). Produces Vec<DeltaBlock> from two byte slices.

compress

Streaming zstd and lz4 frame wrappers for block-level compression.

Performance

Operation Throughput
BLAKE3 single-threaded ~3 GiB/s
BLAKE3 Rayon (1 MB input) ≥10 GiB/s
BLAKE3 batch (256 × 4 KB) ≥9 GiB/s
xxHash3 ≥20 GiB/s

Requires RUSTFLAGS="-C target-cpu=native" for full SIMD throughput (set in .cargo/config.toml of the workspace).

License

MIT — see repository root.