Files
clawhdf5/CLAUDE.md
T
osobhandClaude Opus 5.5 c470244a6f feat(agent): HDF5Memory::search with source filters, re-ranking, confidence
`HDF5Memory::search(query_embedding, query_text, &SearchOptions)` is the
store's full search path. `SearchOptions::new(k)` is plain hybrid search
with the tuned default fusion; each further stage is opt-in:

- `with_sources([..])`: only records from these source channels. The
  filter applies before ranking, so a filtered search still returns up
  to k results, normalised over what it can return. The HNSW pool is
  over-fetched in proportion to what the filter removes, and the allowed
  records are scanned exactly whenever that costs fewer distance
  evaluations than the index would (~pool x M) — and as the fallback if
  the pool comes back short. Keyword matches are filtered too.
- `with_rerank(ReRankConfig)` re-ranks a max(3k, 10) candidate pool by
  relevance, recency, source authority and activation;
  `with_confidence(ConfidenceConfig)` drops low-confidence results;
  `at_time(now)` pins the recency clock.

These were reachable only through the OpenClaw backend, which is now
`search` with both on. Its Hebbian boost now goes to the k results it
returns rather than the whole 3k candidate pool. `hybrid_search` and
`hybrid_search_with` are wrappers and unchanged (tested bit for bit).

Measured on tank (search_harness --options-study --full, 3 runs): at
100K every filter — 50%, 10%, 1% of the store, and records far from the
query — returns the exact filtered top 10, and none is slower than an
unfiltered search (1%: 2.3 ms vs 4.6 ms). Re-rank + confidence costs
about 3%. A first version decided between index and exact scan by pool
size vs store size; it measured 0.976 recall at 12.3 ms on the
far-from-query filter, which is why the rule compares costs instead.

Tests: tests/search_options.rs (filter correctness and full pages via
both paths, far-from-query fallback, edge cases, equality with
hybrid_search_with, re-rank recency, confidence, boost scope).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-24 16:35:42 -05:00

9.7 KiB
Raw Blame History

clawhdf5

Purpose

Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated I/O. Used by ZeroClaw as its persistent memory and knowledge graph backend.

Architecture

Cargo workspace with 16 crates under crates/ (plus libaec-sys, an internal FFI bindings crate for the optional szip feature):

Crate Role
clawhdf5-format HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants
clawhdf5-io Read/write implementation
clawhdf5-filters Compression filters (gzip, LZ4, Zstd, Blosc)
clawhdf5-derive Proc-macro derive for HDF5-serializable structs
clawhdf5 Main facade crate
clawhdf5-netcdf4 NetCDF-4 compatibility layer
clawhdf5-ann HNSW approximate nearest-neighbor vector index
clawhdf5-agent Agent memory, session history, knowledge graph storage
clawhdf5-gpu GPU-accelerated I/O via wgpu (hand-written WGSL compute shaders)
clawhdf5-accel CPU SIMD acceleration path
clawhdf5-migrate SQLite → HDF5 agent-memory migration
clawhdf5-android Android JNI bindings
clawhdf5-cli Command-line interface
clawhdf5-napi Node.js native addon bindings
clawhdf5-py PyO3 Python bindings
clawhdf5-bench Benchmark suite

Key Features

  • Zero-C-dependency HDF5 read/write: no libhdf5, and deflate defaults to pure-Rust zlib-rs (fast-deflate opts into zlib-ng, which needs cmake). ci-test.sh fails if a C-building crate enters the core crates' default tree. flate2 must keep runtime_detection with zlib-rs — without it zlib-rs loses SIMD and inflates 3.5x slower. MSRV is 1.92 (rust-version, checked in CI).
  • HNSW vector index for semantic similarity search over agent memories — the clawhdf5-agent hnsw feature is on by default, so hybrid_search uses the approximate clawhdf5-ann index for the vector stage (the index mirrors the cache and self-heals on drift). Build the agent with --no-default-features --features float16 to force the exact linear cosine scan. The agent's parallel feature (also default) builds the index on a thread pool; the graph is identical with or without it. The index uses the HNSW paper's diversity heuristic for neighbour selection (plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its graph is saved to <store>.h5.ann at each checkpoint and reloaded by open() (tied to the checkpoint by a generation id; stale/damaged sidecars are ignored and the index rebuilt). MemoryConfig::quantized_index (on by default for new stores, persisted; stores predating the setting load as false and keep their f32 index — guarded by tests/fixtures/store_v2_5_0.h5; CLI opt-out is create --f32-index) stores the index's own copy of the embeddings as i8, which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors at 100K); because quantised distances are approximate and ef cannot compensate, the query path then re-scores the candidate pool against the exact embeddings, which holds recall at the f32 index's level. It is also faster at equal recall: 1.63x the QPS on x86-64 (AVX2) and 1.18x on a Raspberry Pi 5 (clawhdf5_accel::dot_i8, NEON SDOT via inline asm since the intrinsic is unstable; plain NEON on pre-dotprod cores). The aarch64 code is cfg'd out on x86, so x86 CI never compiles or lints it — test it on real ARM (rpivision02, 10.0.2.3, is a Pi 5). hybrid_search keeps one incremental BM25 index for the life of the store and never writes the store: Hebbian activation boosts are persisted by the next checkpoint (or on drop), not per query. Measure any search-path change with cargo run --release -p clawhdf5-bench --bin search_harness (baselines in BENCHMARKS.md).
  • WAL (write-ahead log) for crash-safe persistence, with a chained CRC32 trailer per entry (each entry's CRC folds in the previous entry's CRC) so a corrupted, reordered, duplicated, or spliced entry stops replay cleanly instead of loading bad or tampered data. The pre-chaining per-entry-CRC format (v2) is still fully readable; the oldest no-CRC format (v1) is only reachable through the one-time migration path in HDF5Memory::open, not through the public WalFile::read_entries. What the WAL guarantees: integrity, ordering, and recovery from a process crash at any point — including between a checkpoint and the WAL truncate (each checkpoint records a WalMark in /meta, and open() skips the WAL prefix the .h5 already contains, so entries are never applied twice). Checkpoints and snapshots are made durable as a unit (temp file synced, renamed, directory synced). What it does not guarantee: individual WAL appends are not fsynced (a deliberate latency trade-off), so saves made since the last checkpoint can be lost on power failure or kernel panic. Current header version is 4 (adds the Update record used by save_or_update); v3 files are read and upgraded in place.
  • A store has a single writer: HDF5Memory::create/open hold an exclusive advisory lock on <store>.h5.lock and a second opener gets MemoryError::Locked. Use HDF5Memory::open_read_only for a lock-free, never-writing point-in-time view (the CLI's recall/stats/agents-md/ export do). An unreadable WAL (torn header, bad magic) is quarantined to <store>.h5.wal.corrupt-<ts> rather than blocking open(); a WAL with an unknown newer version still fails and is left untouched.
  • MemoryConfig::float16 (off by default, persisted; CLI create --float16) writes /memory/embeddings as IEEE half precision (48% smaller file at 100K, same recall). MemoryCache::half_precision rounds each embedding as it enters the cache (push, update, WAL replay, and on load of a store still f32 on disk), so memory and file agree bit for bit; the conversions live in clawhdf5_format::float16 and must stay the single implementation. Values beyond ±65504 are MemoryError::InvalidEntry. Interop: every file must open in h5py — f32 datasets and empty datasets did not until 2026-09-23 (see docs/known-issues.md); the agent's h5py_interop test guards a whole store.
  • HDF5Memory::search(query_emb, text, &SearchOptions) is the full search path: optional source-channel filter (applied before ranking; exact scan of the allowed records whenever cheaper than pool × M index distance evaluations, and as the fallback when the pool comes back short), fusion, activation scaling, optional re-ranking and confidence rejection. hybrid_search/hybrid_search_with are thin wrappers; the OpenClaw backend is search with re-rank + confidence on. Measure changes with search_harness --options-study.
  • MemoryConfig::compression is off by default; when on, embeddings are deflate-compressed, or Zstd with the agent's zstd feature (links libzstd).
  • Dataset::verify_provenance() (clawhdf5 facade, provenance feature, on by default) recomputes a dataset's SHA-256 and compares it against the _provenance_sha256 attribute written automatically on save when DatasetBuilder::with_provenance is used. It's opt-in per call, not run automatically on open — it decodes and hashes the whole dataset. The hash is unkeyed (tamper-evident, not tamper-proof): it detects accidental corruption, not a deliberate actor able to modify both the data and the stored hash.
  • clawhdf5-agent's HDF5Memory::save/save_batch/save_or_update run every write through an in-memory (session-scoped, not persisted to disk) provenance ledger and write-anomaly detector: a content hash per record (provenance.rs) for detecting accidental mid-session corruption, plus rate-limit/injection-pattern/source-distribution checks (anomaly.rs). Alerts never block a save — drain them with HDF5Memory::take_anomaly_alerts. MemorySource for this bookkeeping is inferred from the caller-supplied source_channel string (a heuristic, not an authenticated trust boundary).
  • GPU-accelerated batch I/O for large dataset processing
  • Python and Node.js bindings for cross-language use
  • NetCDF-4 compatibility for scientific data interop

Workflows

Build

cargo build --release

Test

cargo test --workspace

CI

.gitea/workflows/ci.yml has two jobs, both green as of 2026-09-22:

  • test (ubuntu-latest, in rust:latest) runs scripts/ci-test.sh with the h5py/netCDF4 interop suites required (CLAWHDF5_REQUIRE_INTEROP=1). Served by the tank and architect runners.
  • test-arm64 (linux_arm64) lints and tests the aarch64 code — the NEON kernels are cfg'd out on x86, so this is the only place they are built. Served by vision-01 (host mode) and vision-02 (Docker), so steps must work in both.

Keep workflows free of JavaScript actions (actions/checkout, actions/cache, …): rust:latest has no node, and not every runner reaches GitHub, where they are fetched from. Check out with plain git instead. The test job installs cmake for the opt-in fast-deflate (zlib-ng) steps; the default build needs no C toolchain, so test-arm64 does not. All runners are on gitea-runner 3.5.0, from docker.gitea.com/act_runner — gitea/act_runner:latest on Docker Hub is frozen at 0.6.1.

CLI

cargo run -p clawhdf5-cli -- --help
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands

Python bindings

cd crates/clawhdf5-py
maturin develop
python -c "import clawhdf5; print(clawhdf5.__version__)"

Integration

ZeroClaw imports this as a Cargo feature (clawhdf5 feature flag) to persist agent memory with HNSW vector search for context retrieval.