The setting was persisted in /meta and otherwise ignored: embeddings
were always written as f32. It now does what it says.
clawhdf5-format:
- `DatasetBuilder::with_f16_data` writes IEEE binary16 (numpy float16),
rounding to nearest-even, and `make_f16_type`.
- `clawhdf5_format::float16` holds the f32 <-> f16 conversions, the one
implementation the writer, the reader and the agent all use. Checked
against the `half` crate on 16.7M f32 values and round-trips all 65536
half values; the h5py interop tests confirm the rounding matches
numpy's bit for bit (4020 values incl. ties, subnormals, overflow).
- Reading little-endian float16 as f32 has a fast path.
clawhdf5-agent:
- A float16 store writes /memory/embeddings as half precision, and
`MemoryCache::half_precision` rounds each embedding as it enters the
cache (save, update, WAL replay, and on load of a store still f32 on
disk), so memory and file agree bit for bit and a store searches the
same before and after a reopen (tested).
- Values beyond +-65504 are refused with the new
`MemoryError::InvalidEntry` rather than stored as infinity, on every
save path; batches are all or nothing, and a rejected ephemeral entry
stays in the ephemeral tier. Breaking for exhaustive matches.
- CLI: `create --float16`. Off by default.
Measured on tank, 384-dim, six runs alternating order, medians
(search_harness --float16-study --full): at 100K the file goes from
154.0 to 80.8 MiB (-48%), checkpoint 752 -> 512 ms, open 300 -> 252 ms;
vector recall@10 against an exact scan and hybrid_search latency do not
change. At 10K open is 3 ms slower. Also a test that h5py opens a whole
agent store, f32 and float16, and decodes every dataset.
Docs: README, BENCHMARKS.md ("float16 embedding storage"), CHANGELOG
(including the h5py interop fixes in the previous commit), CLAUDE.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
167 lines
9.2 KiB
Markdown
167 lines
9.2 KiB
Markdown
# clawhdf5
|
|
|
|
## Purpose
|
|
Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated I/O. Used by ZeroClaw as its persistent memory and knowledge graph backend.
|
|
|
|
## Architecture
|
|
|
|
Cargo workspace with 16 crates under `crates/` (plus `libaec-sys`, an internal FFI bindings crate for the optional `szip` feature):
|
|
|
|
| Crate | Role |
|
|
|-------|------|
|
|
| `clawhdf5-format` | HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants |
|
|
| `clawhdf5-io` | Read/write implementation |
|
|
| `clawhdf5-filters` | Compression filters (gzip, LZ4, Zstd, Blosc) |
|
|
| `clawhdf5-derive` | Proc-macro derive for HDF5-serializable structs |
|
|
| `clawhdf5` | Main facade crate |
|
|
| `clawhdf5-netcdf4` | NetCDF-4 compatibility layer |
|
|
| `clawhdf5-ann` | HNSW approximate nearest-neighbor vector index |
|
|
| `clawhdf5-agent` | Agent memory, session history, knowledge graph storage |
|
|
| `clawhdf5-gpu` | GPU-accelerated I/O via wgpu (hand-written WGSL compute shaders) |
|
|
| `clawhdf5-accel` | CPU SIMD acceleration path |
|
|
| `clawhdf5-migrate` | SQLite → HDF5 agent-memory migration |
|
|
| `clawhdf5-android` | Android JNI bindings |
|
|
| `clawhdf5-cli` | Command-line interface |
|
|
| `clawhdf5-napi` | Node.js native addon bindings |
|
|
| `clawhdf5-py` | PyO3 Python bindings |
|
|
| `clawhdf5-bench` | Benchmark suite |
|
|
|
|
## Key Features
|
|
- Zero-C-dependency HDF5 read/write: no libhdf5, and deflate defaults to
|
|
pure-Rust zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake).
|
|
`ci-test.sh` fails if a C-building crate enters the core crates' default
|
|
tree. flate2 must keep `runtime_detection` with zlib-rs — without it zlib-rs
|
|
loses SIMD and inflates 3.5x slower. MSRV is 1.92 (`rust-version`, checked
|
|
in CI).
|
|
- HNSW vector index for semantic similarity search over agent memories — the
|
|
`clawhdf5-agent` `hnsw` feature is **on by default**, so `hybrid_search` uses
|
|
the approximate `clawhdf5-ann` index for the vector stage (the index mirrors
|
|
the cache and self-heals on drift). Build the agent with
|
|
`--no-default-features --features float16` to force the exact linear cosine scan.
|
|
The agent's `parallel` feature (also default) builds the index on a thread
|
|
pool; the graph is identical with or without it.
|
|
The index uses the HNSW paper's diversity heuristic for neighbour selection
|
|
(plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its
|
|
graph is saved to `<store>.h5.ann` at each checkpoint and reloaded by `open()`
|
|
(tied to the checkpoint by a generation id; stale/damaged sidecars are
|
|
ignored and the index rebuilt). `MemoryConfig::quantized_index` (**on by
|
|
default** for new stores, persisted; stores predating the setting load as
|
|
`false` and keep their f32 index — guarded by
|
|
`tests/fixtures/store_v2_5_0.h5`; CLI opt-out is `create --f32-index`)
|
|
stores the index's own copy of the embeddings as `i8`,
|
|
which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors
|
|
at 100K); because quantised distances are approximate and `ef` cannot
|
|
compensate, the query path then re-scores the candidate pool against the
|
|
exact embeddings, which holds recall at the f32 index's level. It is also
|
|
faster at equal recall: 1.63x the QPS on x86-64 (AVX2) and 1.18x on a
|
|
Raspberry Pi 5 (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm since
|
|
the intrinsic is unstable; plain NEON on pre-dotprod cores). The aarch64
|
|
code is `cfg`'d out on x86, so x86 CI never compiles or lints it — test it
|
|
on real ARM (`rpivision02`, 10.0.2.3, is a Pi 5). `hybrid_search` keeps one incremental BM25
|
|
index for the life of the store and never writes the store: Hebbian
|
|
activation boosts are persisted by the next checkpoint (or on drop), not per
|
|
query. Measure any search-path change with
|
|
`cargo run --release -p clawhdf5-bench --bin search_harness` (baselines in
|
|
`BENCHMARKS.md`).
|
|
- WAL (write-ahead log) for crash-safe persistence, with a chained CRC32
|
|
trailer per entry (each entry's CRC folds in the previous entry's CRC) so a
|
|
corrupted, reordered, duplicated, or spliced entry stops replay cleanly
|
|
instead of loading bad or tampered data. The pre-chaining per-entry-CRC
|
|
format (v2) is still fully readable; the oldest no-CRC format (v1) is only
|
|
reachable through the one-time migration path in `HDF5Memory::open`, not
|
|
through the public `WalFile::read_entries`.
|
|
**What the WAL guarantees:** integrity, ordering, and recovery from a
|
|
*process* crash at any point — including between a checkpoint and the WAL
|
|
truncate (each checkpoint records a `WalMark` in `/meta`, and `open()` skips
|
|
the WAL prefix the `.h5` already contains, so entries are never applied
|
|
twice). Checkpoints and snapshots are made durable as a unit (temp file
|
|
synced, renamed, directory synced). **What it does not guarantee:**
|
|
individual WAL appends are *not* fsynced (a deliberate latency trade-off), so
|
|
saves made since the last checkpoint can be lost on power failure or kernel
|
|
panic. Current header version is 4 (adds the `Update` record used by
|
|
`save_or_update`); v3 files are read and upgraded in place.
|
|
- A store has a **single writer**: `HDF5Memory::create`/`open` hold an exclusive
|
|
advisory lock on `<store>.h5.lock` and a second opener gets
|
|
`MemoryError::Locked`. Use `HDF5Memory::open_read_only` for a lock-free,
|
|
never-writing point-in-time view (the CLI's `recall`/`stats`/`agents-md`/
|
|
`export` do). An unreadable WAL (torn header, bad magic) is quarantined to
|
|
`<store>.h5.wal.corrupt-<ts>` rather than blocking `open()`; a WAL with an
|
|
unknown *newer* version still fails and is left untouched.
|
|
- `MemoryConfig::float16` (off by default, persisted; CLI `create --float16`)
|
|
writes `/memory/embeddings` as IEEE half precision (48% smaller file at
|
|
100K, same recall). `MemoryCache::half_precision` rounds each embedding as
|
|
it enters the cache (push, update, WAL replay, and on load of a store still
|
|
`f32` on disk), so memory and file agree bit for bit; the conversions live
|
|
in `clawhdf5_format::float16` and must stay the single implementation.
|
|
Values beyond ±65504 are `MemoryError::InvalidEntry`. Interop: every file
|
|
must open in h5py — `f32` datasets and empty datasets did not until
|
|
2026-09-23 (see `docs/known-issues.md`); the agent's `h5py_interop` test
|
|
guards a whole store.
|
|
- `MemoryConfig::compression` is off by default; when on, embeddings are
|
|
deflate-compressed, or Zstd with the agent's `zstd` feature (links libzstd).
|
|
- `Dataset::verify_provenance()` (clawhdf5 facade, `provenance` feature, on by
|
|
default) recomputes a dataset's SHA-256 and compares it against the
|
|
`_provenance_sha256` attribute written automatically on save when
|
|
`DatasetBuilder::with_provenance` is used. It's opt-in per call, not run
|
|
automatically on open — it decodes and hashes the whole dataset. The hash
|
|
is unkeyed (tamper-*evident*, not tamper-*proof*): it detects accidental
|
|
corruption, not a deliberate actor able to modify both the data and the
|
|
stored hash.
|
|
- `clawhdf5-agent`'s `HDF5Memory::save`/`save_batch`/`save_or_update` run every
|
|
write through an in-memory (session-scoped, not persisted to disk)
|
|
provenance ledger and write-anomaly detector: a content hash per record
|
|
(`provenance.rs`) for detecting accidental mid-session corruption, plus
|
|
rate-limit/injection-pattern/source-distribution checks (`anomaly.rs`).
|
|
Alerts never block a save — drain them with `HDF5Memory::take_anomaly_alerts`.
|
|
`MemorySource` for this bookkeeping is inferred from the caller-supplied
|
|
`source_channel` string (a heuristic, not an authenticated trust boundary).
|
|
- GPU-accelerated batch I/O for large dataset processing
|
|
- Python and Node.js bindings for cross-language use
|
|
- NetCDF-4 compatibility for scientific data interop
|
|
|
|
## Workflows
|
|
|
|
### Build
|
|
```bash
|
|
cargo build --release
|
|
```
|
|
|
|
### Test
|
|
```bash
|
|
cargo test --workspace
|
|
```
|
|
|
|
### CI
|
|
`.gitea/workflows/ci.yml` has two jobs, both green as of 2026-09-22:
|
|
- **`test`** (`ubuntu-latest`, in `rust:latest`) runs `scripts/ci-test.sh` with
|
|
the h5py/netCDF4 interop suites required (`CLAWHDF5_REQUIRE_INTEROP=1`).
|
|
Served by the `tank` and `architect` runners.
|
|
- **`test-arm64`** (`linux_arm64`) lints and tests the aarch64 code — the NEON
|
|
kernels are `cfg`'d out on x86, so this is the only place they are built.
|
|
Served by `vision-01` (host mode) and `vision-02` (Docker), so steps must
|
|
work in both.
|
|
|
|
Keep workflows free of JavaScript actions (`actions/checkout`, `actions/cache`,
|
|
…): `rust:latest` has no `node`, and not every runner reaches GitHub, where
|
|
they are fetched from. Check out with plain `git` instead. The `test` job
|
|
installs `cmake` for the opt-in `fast-deflate` (zlib-ng) steps; the default
|
|
build needs no C toolchain, so `test-arm64` does not.
|
|
All runners are on `gitea-runner` 3.5.0, from `docker.gitea.com/act_runner`
|
|
— `gitea/act_runner:latest` on Docker Hub is frozen at 0.6.1.
|
|
|
|
### CLI
|
|
```bash
|
|
cargo run -p clawhdf5-cli -- --help
|
|
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands
|
|
```
|
|
|
|
### Python bindings
|
|
```bash
|
|
cd crates/clawhdf5-py
|
|
maturin develop
|
|
python -c "import clawhdf5; print(clawhdf5.__version__)"
|
|
```
|
|
|
|
## Integration
|
|
ZeroClaw imports this as a Cargo feature (`clawhdf5` feature flag) to persist agent memory with HNSW vector search for context retrieval.
|