Files
clawhdf5/CLAUDE.md
T
osobhandClaude Opus 5.5 735db117a7 build: pure-Rust zlib-rs as the default deflate backend
The core crates (clawhdf5, -agent, -format, -io, -filters, -ann, -accel,
-netcdf4, -cli) now build no C by default: deflate defaults to zlib-rs,
a pure-Rust port of zlib-ng, and zlib-ng becomes the opt-in
`fast-deflate`, which overrides zlib-rs wherever it is enabled. A default
build no longer needs cmake or a C compiler.

Measured on tank, both builds run alternately, three rounds, medians:
zlib-rs is within 6% of zlib-ng on every HDF5 read and write (512x512
deflate-6 chunked write 1.458 vs 1.484 ms; 64 MB compressed read 64.4
vs 65.2 ms), and compressed output is byte-identical. Details in
BENCHMARKS.md, "Deflate backend".

Getting there took two fixes the first measurement exposed:

- zlib-rs needs `std` to detect SIMD at runtime. flate2 enables it via
  its default `runtime_detection`, which `default-features = false` had
  switched off, leaving zlib-rs 3.5x slower on inflate. The `zlib-rs`
  features now enable it.
- Both deflate paths streamed through flate2's 32 KiB read/write
  wrappers. They now hand the codec the whole chunk in one call, into a
  buffer sized up front (~5% on chunked writes). This also fixes a
  silent short read: the streaming reader returned a truncated stream's
  bytes without an error; a truncated chunk is now DecompressionError.
  In clawhdf5-filters, output longer than the stated size is now an
  error rather than silently cut off.

CI: ci-test.sh lints and tests the zlib-ng path, and fails if a
C-building crate (*-sys, cc, cmake) enters a core crate's default
dependency tree. The arm64 job no longer installs cmake.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-23 11:05:22 -05:00

8.5 KiB

clawhdf5

Purpose

Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated I/O. Used by ZeroClaw as its persistent memory and knowledge graph backend.

Architecture

Cargo workspace with 16 crates under crates/ (plus libaec-sys, an internal FFI bindings crate for the optional szip feature):

Crate Role
clawhdf5-format HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants
clawhdf5-io Read/write implementation
clawhdf5-filters Compression filters (gzip, LZ4, Zstd, Blosc)
clawhdf5-derive Proc-macro derive for HDF5-serializable structs
clawhdf5 Main facade crate
clawhdf5-netcdf4 NetCDF-4 compatibility layer
clawhdf5-ann HNSW approximate nearest-neighbor vector index
clawhdf5-agent Agent memory, session history, knowledge graph storage
clawhdf5-gpu GPU-accelerated I/O via wgpu (hand-written WGSL compute shaders)
clawhdf5-accel CPU SIMD acceleration path
clawhdf5-migrate SQLite → HDF5 agent-memory migration
clawhdf5-android Android JNI bindings
clawhdf5-cli Command-line interface
clawhdf5-napi Node.js native addon bindings
clawhdf5-py PyO3 Python bindings
clawhdf5-bench Benchmark suite

Key Features

  • Zero-C-dependency HDF5 read/write: no libhdf5, and deflate defaults to pure-Rust zlib-rs (fast-deflate opts into zlib-ng, which needs cmake). ci-test.sh fails if a C-building crate enters the core crates' default tree. flate2 must keep runtime_detection with zlib-rs — without it zlib-rs loses SIMD and inflates 3.5x slower.
  • HNSW vector index for semantic similarity search over agent memories — the clawhdf5-agent hnsw feature is on by default, so hybrid_search uses the approximate clawhdf5-ann index for the vector stage (the index mirrors the cache and self-heals on drift). Build the agent with --no-default-features --features float16 to force the exact linear cosine scan. The agent's parallel feature (also default) builds the index on a thread pool; the graph is identical with or without it. The index uses the HNSW paper's diversity heuristic for neighbour selection (plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its graph is saved to <store>.h5.ann at each checkpoint and reloaded by open() (tied to the checkpoint by a generation id; stale/damaged sidecars are ignored and the index rebuilt). MemoryConfig::quantized_index (on by default for new stores, persisted; stores predating the setting load as false and keep their f32 index — guarded by tests/fixtures/store_v2_5_0.h5; CLI opt-out is create --f32-index) stores the index's own copy of the embeddings as i8, which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors at 100K); because quantised distances are approximate and ef cannot compensate, the query path then re-scores the candidate pool against the exact embeddings, which holds recall at the f32 index's level. It is also faster at equal recall: 1.63x the QPS on x86-64 (AVX2) and 1.18x on a Raspberry Pi 5 (clawhdf5_accel::dot_i8, NEON SDOT via inline asm since the intrinsic is unstable; plain NEON on pre-dotprod cores). The aarch64 code is cfg'd out on x86, so x86 CI never compiles or lints it — test it on real ARM (rpivision02, 10.0.2.3, is a Pi 5). hybrid_search keeps one incremental BM25 index for the life of the store and never writes the store: Hebbian activation boosts are persisted by the next checkpoint (or on drop), not per query. Measure any search-path change with cargo run --release -p clawhdf5-bench --bin search_harness (baselines in BENCHMARKS.md).
  • WAL (write-ahead log) for crash-safe persistence, with a chained CRC32 trailer per entry (each entry's CRC folds in the previous entry's CRC) so a corrupted, reordered, duplicated, or spliced entry stops replay cleanly instead of loading bad or tampered data. The pre-chaining per-entry-CRC format (v2) is still fully readable; the oldest no-CRC format (v1) is only reachable through the one-time migration path in HDF5Memory::open, not through the public WalFile::read_entries. What the WAL guarantees: integrity, ordering, and recovery from a process crash at any point — including between a checkpoint and the WAL truncate (each checkpoint records a WalMark in /meta, and open() skips the WAL prefix the .h5 already contains, so entries are never applied twice). Checkpoints and snapshots are made durable as a unit (temp file synced, renamed, directory synced). What it does not guarantee: individual WAL appends are not fsynced (a deliberate latency trade-off), so saves made since the last checkpoint can be lost on power failure or kernel panic. Current header version is 4 (adds the Update record used by save_or_update); v3 files are read and upgraded in place.
  • A store has a single writer: HDF5Memory::create/open hold an exclusive advisory lock on <store>.h5.lock and a second opener gets MemoryError::Locked. Use HDF5Memory::open_read_only for a lock-free, never-writing point-in-time view (the CLI's recall/stats/agents-md/ export do). An unreadable WAL (torn header, bad magic) is quarantined to <store>.h5.wal.corrupt-<ts> rather than blocking open(); a WAL with an unknown newer version still fails and is left untouched.
  • MemoryConfig::compression is off by default; when on, embeddings are deflate-compressed, or Zstd with the agent's zstd feature (links libzstd).
  • Dataset::verify_provenance() (clawhdf5 facade, provenance feature, on by default) recomputes a dataset's SHA-256 and compares it against the _provenance_sha256 attribute written automatically on save when DatasetBuilder::with_provenance is used. It's opt-in per call, not run automatically on open — it decodes and hashes the whole dataset. The hash is unkeyed (tamper-evident, not tamper-proof): it detects accidental corruption, not a deliberate actor able to modify both the data and the stored hash.
  • clawhdf5-agent's HDF5Memory::save/save_batch/save_or_update run every write through an in-memory (session-scoped, not persisted to disk) provenance ledger and write-anomaly detector: a content hash per record (provenance.rs) for detecting accidental mid-session corruption, plus rate-limit/injection-pattern/source-distribution checks (anomaly.rs). Alerts never block a save — drain them with HDF5Memory::take_anomaly_alerts. MemorySource for this bookkeeping is inferred from the caller-supplied source_channel string (a heuristic, not an authenticated trust boundary).
  • GPU-accelerated batch I/O for large dataset processing
  • Python and Node.js bindings for cross-language use
  • NetCDF-4 compatibility for scientific data interop

Workflows

Build

cargo build --release

Test

cargo test --workspace

CI

.gitea/workflows/ci.yml has two jobs, both green as of 2026-09-22:

  • test (ubuntu-latest, in rust:latest) runs scripts/ci-test.sh with the h5py/netCDF4 interop suites required (CLAWHDF5_REQUIRE_INTEROP=1). Served by the tank and architect runners.
  • test-arm64 (linux_arm64) lints and tests the aarch64 code — the NEON kernels are cfg'd out on x86, so this is the only place they are built. Served by vision-01 (host mode) and vision-02 (Docker), so steps must work in both.

Keep workflows free of JavaScript actions (actions/checkout, actions/cache, …): rust:latest has no node, and not every runner reaches GitHub, where they are fetched from. Check out with plain git instead. The test job installs cmake for the opt-in fast-deflate (zlib-ng) steps; the default build needs no C toolchain, so test-arm64 does not. All runners are on gitea-runner 3.5.0, from docker.gitea.com/act_runnergitea/act_runner:latest on Docker Hub is frozen at 0.6.1.

CLI

cargo run -p clawhdf5-cli -- --help
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands

Python bindings

cd crates/clawhdf5-py
maturin develop
python -c "import clawhdf5; print(clawhdf5.__version__)"

Integration

ZeroClaw imports this as a Cargo feature (clawhdf5 feature flag) to persist agent memory with HNSW vector search for context retrieval.