CI / test (push) Failing after 3s
README.md: - Fix badly stale LongMemEval numbers (badge said Hit@5 46%, table showed fabricated ~46%/~0.34/~72% figures that never matched BENCHMARKS.md's actual results of Hit@5 100% session / 84.4% turn-level, MRR 1.0/0.6597) - Remove clawhdf5-types from the Crate Map — that crate was removed in an earlier cleanup pass but the README diagram was never updated; fix the crate count (16, not 17) and stale line-of-code figures (72,087/84K -> ~92K) - Fix a dead #benchmarks badge anchor (no such heading exists) -> #performance - Document the new clawhdf5-ann `parallel` feature (had no Feature Flags entry) - Note WAL's CRC32 per-entry check, link the new tank LongMemEval/SIMD/ vector-search reproduction section, update stale test-count comment (417+ -> 1,650+) and Phase 2 roadmap blurb (LongMemEval is now done) ROADMAP.md: - Check off "Academic benchmark cross-validation" (done via the tank LongMemEval re-run) and add a new "Recently closed out" section summarizing the Tier 3-4 hardening pass (Android JNI validation, pyo3 bump, WAL CRC32, bounds-check audit + fuzz harness that found 3 real bugs, HNSW optional parallel feature, workspace.dependencies) - Update stale test count (1,546 -> 1,650+) and last-updated date CLAUDE.md: mention WAL's per-entry CRC32 check CHANGELOG.md: add Security/Performance/Architecture/Documentation entries under Unreleased summarizing all of Tiers 1-4 (this had not been touched since 2026-06-04, predating the entire hardening pass)
2.6 KiB
2.6 KiB
clawhdf5
Purpose
Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated I/O. Used by ZeroClaw as its persistent memory and knowledge graph backend.
Architecture
Cargo workspace with 16 crates under crates/ (plus libaec-sys, an internal FFI bindings crate for the optional szip feature):
| Crate | Role |
|---|---|
clawhdf5-format |
HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants |
clawhdf5-io |
Read/write implementation |
clawhdf5-filters |
Compression filters (gzip, LZ4, Zstd, Blosc) |
clawhdf5-derive |
Proc-macro derive for HDF5-serializable structs |
clawhdf5 |
Main facade crate |
clawhdf5-netcdf4 |
NetCDF-4 compatibility layer |
clawhdf5-ann |
HNSW approximate nearest-neighbor vector index |
clawhdf5-agent |
Agent memory, session history, knowledge graph storage |
clawhdf5-gpu |
GPU-accelerated I/O via wgpu (hand-written WGSL compute shaders) |
clawhdf5-accel |
CPU SIMD acceleration path |
clawhdf5-migrate |
Schema migration engine |
clawhdf5-android |
Android JNI bindings |
clawhdf5-cli |
Command-line interface |
clawhdf5-napi |
Node.js native addon bindings |
clawhdf5-py |
PyO3 Python bindings |
clawhdf5-bench |
Benchmark suite |
Key Features
- Zero-dependency HDF5 read/write (no libhdf5 C library required)
- HNSW vector index for semantic similarity search over agent memories — the
clawhdf5-agenthnswfeature is on by default, sohybrid_searchuses the approximateclawhdf5-annindex for the vector stage (the index mirrors the cache and self-heals on drift). Build the agent with--no-default-features --features float16to force the exact linear cosine scan. - WAL (write-ahead log) for crash-safe persistence, with a CRC32 trailer per entry so a corrupted entry stops replay cleanly instead of loading bad data
- GPU-accelerated batch I/O for large dataset processing
- Python and Node.js bindings for cross-language use
- NetCDF-4 compatibility for scientific data interop
Workflows
Build
cargo build --release
Test
cargo test --workspace
CLI
cargo run -p clawhdf5-cli -- --help
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands
Python bindings
cd crates/clawhdf5-py
maturin develop
python -c "import clawhdf5; print(clawhdf5.__version__)"
Integration
ZeroClaw imports this as a Cargo feature (clawhdf5 feature flag) to persist agent memory with HNSW vector search for context retrieval.