Merge branch 'docs/refresh-reference' into docs/readme-refresh

This commit is contained in:
osobh
2026-09-28 11:14:06 -05:00
6 changed files with 955 additions and 1243 deletions
+100 -12
View File
@@ -51,6 +51,31 @@ target: Criterion stretched it where 5 s could not hold the samples it needed
---
## Current headline numbers
The newest dated measurement of each headline figure, as of 2026-09-28.
Everything below this section is the dated record behind them; sections whose
figures a later run replaced are marked *Superseded*. Machine "tank" is an AMD
Ryzen 7 7800X3D (8C/16T); rows marked idle were run with the 1-minute load
average below 2.
| Figure | Value | Measured | Command | Details |
|---|---|---|---|---|
| Agent memory search, `HDF5Memory::hybrid_search` p50 | 0.49 ms at 10K, 4.69 ms at 100K records | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` | [Current: search harness](#current-search-harness-2026-09-24) |
| LongMemEval `longmemeval_s` (full haystack), default hybrid 0.4/0.6, turn-level retrieval Hit@5 (not QA accuracy) | 81.4% | 2026-09-27, tank, search code of `7a8fae0` | `longmemeval_bench … --embeddings weights/all-minilm-l6-v2` | [Re-run with real embeddings](#re-run-with-real-embeddings-2026-09-27-tank), [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500) |
| Loaded store memory, 100K × 384 | 399 MiB (2.72x raw) with the `f32` index; 256 MiB (1.74x) with the int8 index (int8 side not re-run since it was first measured) | 2026-09-24, tank, `5c8323c` | `search_harness -- --footprint --full [--int8]` | [Memory footprint](#memory-footprint), [Quantising the index copy](#quantising-the-index-copy-quantized_index) |
| int8 index vs `f32` index, QPS at equal recall | 1.63x (x86-64 AVX2), 1.18x (Raspberry Pi 5, `SDOT`) | undated; not re-checked against the 2026-09-24 `f32` figure | `search_harness -- --full` | [Quantising the index copy](#quantising-the-index-copy-quantized_index), [On ARM](#on-arm-raspberry-pi-5-cortex-a76) |
| `float16` store file size, 100K × 384 | 80.8 MiB vs 154.0 MiB `f32` (48% smaller) | 2026-09-23, tank | `search_harness -- --float16-study --full` | [float16 embedding storage](#float16-embedding-storage-memoryconfigfloat16) |
| Full reads of chunked deflate data, 16 threads on one `File` | 4944 MB/s, 1.58x 16 h5py processes (noisy run: compare ratios, not MB/s) | 2026-09-26, tank, `c5334b1` | `concurrent_read` + `concurrent_read_h5py.py` | [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) |
| Same, clawhdf5 only, against the build before range-read M2/M3 | 8525 MB/s vs 6258 (+36%); contiguous and metadata reads at parity | 2026-09-27, tank (idle), `7a8fae0` vs `8f59b2e` | `concurrent_read --decode-threads 1 --reps 3` | [Local metadata and data reads after range-read M2/M3](#local-metadata-and-data-reads-after-range-read-m2m3-2026-09-27-tank) |
| `ObjectHeader::parse` (401 headers) | 23.5–23.6 µs, 1.0–2.6% below `8f59b2e` | 2026-09-27, tank (idle), `96086ad` | `cargo bench -p clawhdf5 --bench local_metadata_bench` | [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) |
| Selection reads, 64 MB chunked + deflate `f64` | full 63.2 ms; one 64 × 64 window 0.18 ms | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin read_harness` | [Current: read harness](#current-read-harness-2026-09-24) |
| Deflate backend, zlib-rs (default) vs zlib-ng | within 6% on every HDF5 read/write path | 2026-09-23, tank | `cargo bench -p clawhdf5-filters --bench deflate_bench` (and the two commands with it) | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng) |
| vs libhdf5 1.14.6: chunked deflate-6 write 512×512 / 128 attributes / 64 groups | 45.3x / 10.3x / 10.6x | 2026-08-03, tank | `cargo bench -p clawhdf5-bench --features libhdf5-compare` | [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) |
| Signed checkpoints | about 20% of a checkpoint (598 vs 495 ms at 100K) | 2026-09-25, tank | `search_harness -- --signing-study --full` | [Signed checkpoints](#signed-checkpoints) |
---
## Memory footprint
`cargo run --release -p clawhdf5-bench --bin search_harness -- --footprint --full`,
@@ -64,6 +89,10 @@ change at all. Measured that way a store holding the corpus twice and one
holding it once came out *identical* (1.00x both), which is how the first
attempt at this measurement went.
> *Superseded* by the current figures below (2026-09-24): this table is the
> record of the double-copy fix (commit 2e7e045, undated); the store measured
> 2.72x, not 2.43x, by the time the int8 index landed.
| N | vectors (raw) | reopened, before | reopened, after |
|---:|---:|---:|---:|
| 1 000 | 1 MiB | 5 MiB (3.41x) | 4 MiB (2.39x) |
@@ -379,6 +408,9 @@ point: does a selection cost what the *selection* costs?
### Baseline (v2.4.0): every selection decodes the whole dataset
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
> (2026-09-24). Kept as the before picture.
4096 x 2048 f64 (64 MB per dataset), chunks 256 x 256, file 129 MB
| layout | read | selected | time ms | MB/s of selection | vs full read |
@@ -404,6 +436,9 @@ point: does a selection cost what the *selection* costs?
### After: partial reads
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
> (2026-09-24).
Only the rows of a contiguous dataset, or the chunks, that overlap the
selection's bounding box are read/decoded. A 64 x 64 window of the compressed
dataset: **105 -> 0.39 ms**; one row: **106 -> 2.7 ms**; one column:
@@ -435,6 +470,9 @@ because the machine's speed drifted; compare the *vs full read* column.)
### After: parallel cached decode, fewer copies (full reads)
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
> (2026-09-24).
Full-read times, old and new binaries run alternately at the same moment (this
machine's absolute speed drifts over a long session, so only same-moment
comparisons mean anything):
@@ -542,6 +580,10 @@ rounds; run 2 also alternated `main` `425585e`.
### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank)
> The `object_header_parse_x401` row (+4.2%) is *superseded* by
> [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank)
> (2026-09-27, `96086ad`); the other rows are current.
`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main`
`7a8fae0` (PRs #18 and #19), each built in its own worktree and run as
separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen
@@ -585,8 +627,8 @@ What this shows:
- **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header;
the base and candidate ranges do not overlap). It is the cost of reading
continuation chunks from a bounded queue (the fix for unbounded reads on
crafted headers) and does not show in the listing. Kept open in
`docs/known-issues.md`.
crafted headers) and does not show in the listing. (Fixed later the
same day; see the section above and `docs/known-issues.md`.)
- **Full reads of deflate data got faster** after #18 (in-place chunk
decoding into the typed output and per-thread scratch buffers): +1.7% on
one thread, +36% at 16.
@@ -639,6 +681,10 @@ saturate memory bandwidth (1.05x).
### Results after the read fixes (2026-09-26, tank, `408f69e`)
> *Superseded* by [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)
> (2026-09-26, `c5334b1`), which closed the 16-thread gap listed at the end
> of this section.
Same machine, files and commands as the first run below, re-run on an idle
tank (load average 1.60 at the start; the 1-minute figure rose to about 5
during the clawhdf5 runs, mostly their own threads) after two fixes:
@@ -679,11 +725,17 @@ Read with care:
- At 16 threads every tool dropped in this run (h5py threads on contiguous
data from 8002 to 2285 MB/s, processes from 12846 to 6942), so the
16-thread rows are noisier than the others.
- Still behind: full reads of chunked data at 16 threads (0.69x-0.76x h5py
processes). See `docs/known-issues.md`.
- Still behind at this commit: full reads of chunked data at 16 threads
(0.69x-0.76x h5py processes); fixed by `c5334b1` (above), recorded as
fixed in `docs/known-issues.md`.
### First run, before the read fixes (2026-09-26, tank, `91644d8`)
> *Superseded* results: the tables and "What this shows" are the before
> picture for [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)
> (2026-09-26). The workload description and the **Run** box below are
> still how every `concurrent_read` figure in this file is produced.
Measured on tank (AMD Ryzen 7 7800X3D, 8 cores / 16 threads, 61 GiB, Linux
7.0) at commit `91644d8`, load average 1.84 when the run started (the
1-minute figure rose to 3.7 during the runs; that is mostly the benchmark's
@@ -720,8 +772,8 @@ What this shows:
- **clawhdf5 threads on one `File` do, for hyperslab reads of compressed
data:** 1244 MB/s at 16 threads, 9.7x h5py threads and 0.89x h5py
processes, without a process pool.
- **Where clawhdf5 is behind** (open performance bugs, see
`docs/known-issues.md`):
- **Where clawhdf5 was behind** at `91644d8` (both since fixed; see
`docs/known-issues.md`, "Concurrent and contiguous read performance"):
- *Full reads of chunked datasets stop scaling at about 4 threads*
(about 880 MB/s) while h5py processes reach 4424 MB/s. Hyperslab
reads, which bypass the `File`'s chunk cache, keep scaling, so the
@@ -811,6 +863,11 @@ Other flags (both harnesses): `--threads`, `--reps`, `--slab`, `--slabs`,
## Search harness baseline (v2.3.0)
> *Historical.* This baseline and the "After: …" subsections that follow
> record each step of the search work; they are *superseded* by
> [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24),
> the last subsection of this part.
Produced by `cargo run --release -p clawhdf5-bench --bin search_harness -- --full`
on deterministic **clustered** synthetic data (384-dim, unit-normalised; points =
cluster centre + noise — uniform random vectors are nearly equidistant in high
@@ -873,10 +930,11 @@ build: 9752.6 ms (10254 vectors/s) · exact scan: 40 QPS, p50 24648 µs
| 1000 | 11 | 3.9 | 0.9 | 68.1 | 5.48 | 5.57 | 182.5 |
| 10000 | 114 | 32.2 | 10.9 | 845.0 | 48.56 | 78.65 | 19.8 |
| 100000 | 1486 | 713.0 | 354.5 | 10486.5 | 883.51 | 975.23 | 1.1 |
wrote /tmp/claude-1000/-home-osobh-projects-clawhdf5/422f755e-dd25-4c35-8613-5439087e3aaa/scratchpad/baseline_full.json
### After: HNSW neighbour-selection heuristic
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
Same harness, same data, after replacing closest-M neighbour selection with the
HNSW paper's diversity heuristic (Algorithm 4, keeping pruned connections) for
both new links and back-link pruning. Recall@10 at `ef = 64`: **0.87 → 1.00**
@@ -922,6 +980,8 @@ build: 36472.8 ms (2742 vectors/s) · exact scan: 40 QPS, p50 24644 µs
### After: persistent keyword index, no store rewrite per query
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
`hybrid_search` used to rebuild the BM25 index from scratch (re-tokenising every
record) and rewrite the whole `.h5` file on **every query**. The index is now
kept for the life of the store and updated incrementally, and activation boosts
@@ -942,6 +1002,8 @@ index removes that.
### After: vector index persisted with the checkpoint
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
The HNSW graph (not the vectors, which the store already holds) is saved to
`<store>.h5.ann` at each checkpoint and reloaded by `open()`, tied to that
checkpoint by a generation id. The index is now built once per store (the *cold
@@ -959,6 +1021,8 @@ index incrementally.
### After: unit-vector dot product, reusable visited set
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
Cosine distance recomputed both vector norms on every evaluation; the index now
stores unit vectors and uses a plain dot product. The per-call `HashSet` of
visited nodes became a reusable epoch-stamped array. Recall is unchanged.
@@ -1004,6 +1068,8 @@ build: 21084.6 ms (4743 vectors/s) · exact scan: 39 QPS, p50 24739 µs
### After: unranked keyword scores, top-k merge (rankings unchanged)
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
A fusion study (`search_harness --fusion-study`) showed that capping the
keyword candidate pool is **not** a safe optimisation: against the current
full-corpus normalisation the final top-10 overlap is only 0.83-0.92 and the
@@ -1026,6 +1092,8 @@ results.
### After: batched bulk build (optionally parallel); deletions handled in search
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
Profiling showed **90% of a build's distance evaluations are in back-link
pruning**. The bulk build now inserts in batches: plan each node's neighbours
against the graph as it stood at the start of the batch, link, then prune every
@@ -1588,6 +1656,11 @@ MRR, or a one-question change in recency, is within this variation.
### Full haystack — `longmemeval_s`, n=500 (the number to cite)
This table is **BM25-only** (zero embeddings). With real embeddings and the
default hybrid 0.4/0.6 the same corpus gives turn Hit@5 **81.4%** (2026-09-27;
see [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500)),
which is the headline figure.
47.7 sessions and 493.5 turns per question; 4.0% of haystack sessions are evidence
sessions, so retrieval has to actually discriminate.
@@ -1695,7 +1768,9 @@ over rank-1 precision.
activation. Until now its combined score contained **no relevance term at
all** — `RerankInput` did not carry the retrieval score — so a caller that
re-ranked its candidates threw the retriever's ordering away and returned them
ordered by age. The OpenClaw backend did exactly that on every search.
ordered by age. `ClawhdfBackend` (the `openclaw` module) did exactly that on
every search. (OpenClaw itself never integrated clawhdf5; see
`docs/openclaw.md`.)
Measuring that is unambiguous. "Recency" below is the share of
`knowledge-update` questions where the newest gold session outranked the stale
@@ -1793,7 +1868,7 @@ worth stating plainly rather than hiding: LongMemEval questions share substantia
vocabulary with their evidence turns, which is close to the best case for lexical
matching, and MiniLM at 384 dimensions is a small embedding model.
> **Run:** `cargo run --release --bin longmemeval_bench --features embeddings -- \
> **Run:** `cargo run --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings -- \
> benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings weights/all-minilm-l6-v2`
> For the GPU path use `--features embeddings-cuda`. That requires `nvcc` on
> `PATH` at *build* time — cudarc's build script shells out to it. The toolkit
@@ -2212,7 +2287,8 @@ The tank row was measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D), commit
### Reproducibility
```bash
rustup override set nightly
# Any stable toolchain at or above the MSRV (1.92) works; the original
# 2026-07-01 run used a nightly, later runs stable.
# Latency benchmarks (Criterion)
cargo bench -p clawhdf5-agent
@@ -2315,7 +2391,9 @@ libhdf5 reads from a temp file including `open` + `read` + `close` overhead.
| clawhdf5 hyperslab (f64, 10% slice) | — | 4.09 µs / **1.8 GiB/s** | 50.1 µs / **1.5 GiB/s** |
libhdf5 f64 comparison excluded — clawhdf5's datatype encoding differs from libhdf5's (known
gap), making cross-format reads unreliable for comparison.
gap), making cross-format reads unreliable for comparison. (That gap was the float sign-bit
bug, fixed 2026-09-23: `docs/known-issues.md`, "Every `f32` dataset we wrote was unreadable by
h5py / libhdf5". The comparison has not been re-run since.)
### Chunked Read Throughput
@@ -2423,6 +2501,12 @@ global file mutex and flushes to disk on every attribute write or group creation
## vs libhdf5 Summary
> Measured on the original i7-12650H (clawhdf5 2026-07-01, libhdf5
> 2026-06-30). The newest run of this table is
> [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03)
> (2026-08-03), which reproduces every row within ~15% except chunked
> write (45.3x on tank).
| Workload | clawhdf5 | libhdf5 | Speedup |
|----------|----------|---------|---------|
| Sequential read, 1K f32 | 634 ns | 45.2 µs | **71×** |
@@ -2454,7 +2538,7 @@ to the page cache. There is no algorithmic headroom above ~1.7 GiB/s on this har
### Caveats
- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap). f64 results are clawhdf5-only.
- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap at the time; fixed 2026-09-23, see [Sequential Read Throughput](#sequential-read-throughput)). f64 results are clawhdf5-only.
- Serial benchmarks. clawhdf5 uses Rayon for chunk compression when > 2 chunks; that parallelism is already reflected in the chunked write numbers.
- clawhdf5 reads from `Vec<u8>` (zero-copy from mmap in production); libhdf5 reads from a temp file. This gives clawhdf5 a structural read advantage that reflects realistic API usage.
@@ -2651,6 +2735,10 @@ Same not-like-for-like caveat as the "Comparison to MemX" section at the top of
file applies — MemX's figure is end-to-end, these are a single component. Ratios are
an order-of-magnitude indication, not a benchmark result.
> The Ratio column below was retracted afterwards: see
> [Comparison to MemX](#comparison-to-memx-arxiv260316171). Kept as recorded
> on 2026-08-05; do not cite it.
| Metric | MemX (claimed, end-to-end) | ClawhDF5 (tank, component only) | Ratio |
|--------|----------------------------|----------------------------------|-------|
| 100K flat search | <90 ms | 6.60 ms | ~14x |
+232 -220
View File
@@ -1,253 +1,266 @@
# clawhdf5
## Purpose
Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated vector search. A standalone library. Its one verified consumer is ClawBrainHub (`.brain` files); no agent framework integrates it (OpenClaw and ZeroClaw claims were withdrawn on 2026-09-25 — neither was ever true).
Pure-Rust HDF5 implementation (read, write, in-place edit, remote and browser
reads) plus agent memory on top of it: HNSW vector search, a WAL-backed store,
and GPU vector distances. A standalone library. Its one verified consumer is
ClawBrainHub (`.brain` files); no agent framework integrates it (see
*Standing rules*).
## Architecture
Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal FFI bindings crate for the optional `szip` feature):
Cargo workspace, 19 crates under `crates/` (plus `libaec-sys`, the FFI crate
behind the optional `szip` feature). MSRV 1.92 (`rust-version`, checked in CI).
| Crate | Role |
|-------|------|
| `clawhdf5-format` | HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants |
| `clawhdf5-io` | Read/write implementation |
| `clawhdf5-filters` | Deflate backends (zlib-rs, zlib-ng, Apple Compression); the HDF5 filter pipeline, the filter registry (`clawhdf5_format::filter_registry`) and the other codecs (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec, and the pure-Rust plugin filters LZF, bitshuffle, bzip2, Blosc 1, and Blosc2 and ZFP read-only) live in `clawhdf5-format`. |
| `clawhdf5-derive` | Proc-macro derive for HDF5-serializable structs |
| `clawhdf5` | Main facade crate |
| `clawhdf5-netcdf4` | NetCDF-4 compatibility layer |
| `clawhdf5-ann` | HNSW approximate nearest-neighbor vector index |
| `clawhdf5-agent` | Agent memory, session history, knowledge graph storage |
| `clawhdf5-gpu` | GPU vector distance computation via wgpu (hand-written WGSL compute shaders) — not dataset I/O |
| `clawhdf5-accel` | CPU SIMD acceleration path |
| `clawhdf5-migrate` | SQLite → HDF5 agent-memory migration |
| `clawhdf5-format` | The HDF5 format: parsers and writer (superblock, headers, B-trees, heaps, chunk indexes), the `Storage` trait, the filter pipeline and registry (`filter_registry`), every codec except deflate (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec; pure-Rust LZF, bitshuffle, bzip2, Blosc 1; Blosc2 and ZFP read-only), `float16`, `checksum` |
| `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) |
| `clawhdf5-io` | I/O adapters (buffers, mmap, prefetch) |
| `clawhdf5-derive` | `#[derive(H5Type)]` for compound types |
| `clawhdf5` | Facade: `File`, `FileBuilder`, `Dataset`, `FileEditor` (`src/edit/`), SWMR reading (`src/swmr.rs`) |
| `clawhdf5-netcdf4` | NetCDF-4 read support |
| `clawhdf5-remote` | `open_url`: HTTP(S) range requests and object stores (S3, GCS, Azure) through `BlockCache` |
| `clawhdf5-tools` | `h5rs`: `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` |
| `clawhdf5-py` | PyO3 bindings (h5py-like API, remote files, `'r+'` editing) |
| `clawhdf5-wasm` | wasm-bindgen browser reader (`open(bytes)`, `openUrl(url)`); demo in `examples/wasm-viewer/` |
| `clawhdf5-ann` | HNSW index |
| `clawhdf5-agent` | Agent memory store (`HDF5Memory`), sessions, knowledge graph, BM25 |
| `clawhdf5-accel` | CPU SIMD kernels (AVX2, NEON) |
| `clawhdf5-gpu` | wgpu vector distances (WGSL) — not dataset I/O; HDF5 I/O is CPU-only |
| `clawhdf5-migrate` | SQLite → agent store migration |
| `clawhdf5-cli` | Agent-memory CLI |
| `clawhdf5-napi` | Node.js addon (the `packages/clawhdf5-node` wrapper is broken; `docs/known-issues.md`) |
| `clawhdf5-android` | Android JNI bindings |
| `clawhdf5-cli` | Command-line interface (agent memory) |
| `clawhdf5-tools` | `h5rs`: pure-Rust HDF5 tools — `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` (structural + checksum validator) |
| `clawhdf5-napi` | Node.js native addon bindings |
| `clawhdf5-py` | PyO3 Python bindings |
| `clawhdf5-wasm` | WebAssembly (wasm-bindgen) reader for the browser; demo in `examples/wasm-viewer/` |
| `clawhdf5-remote` | Remote files: `open_url` over HTTP(S) range requests and object stores (`object_store`: S3, GCS, Azure) through a mandatory block cache (`BlockCache`) |
| `clawhdf5-bench` | Benchmark suite |
| `clawhdf5-bench` | Benchmarks and harnesses (`search_harness`, `read_harness`, `concurrent_read`, `longmemeval_bench`, …) |
## Key Features
- Zero-C-dependency HDF5 read/write: no libhdf5, and deflate defaults to
pure-Rust zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake).
`ci-test.sh` fails if a C-building crate enters the core crates' default
tree. flate2 must keep `runtime_detection` with zlib-rs — without it zlib-rs
loses SIMD and inflates 3.5x slower. MSRV is 1.92 (`rust-version`, checked
in CI).
- HNSW vector index for semantic similarity search over agent memories — the
`clawhdf5-agent` `hnsw` feature is **on by default**, so `hybrid_search` uses
the approximate `clawhdf5-ann` index for the vector stage (the index mirrors
the cache and self-heals on drift). Build the agent with
`--no-default-features --features float16` to force the exact linear cosine scan.
The agent's `parallel` feature (also default) builds the index on a thread
pool; the graph is identical with or without it.
The index uses the HNSW paper's diversity heuristic for neighbour selection
(plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its
graph is saved to `<store>.h5.ann` at each checkpoint and reloaded by `open()`
(tied to the checkpoint by a generation id; stale/damaged sidecars are
ignored and the index rebuilt). `MemoryConfig::quantized_index` (**on by
default** for new stores, persisted; stores predating the setting load as
`false` and keep their f32 index — guarded by
`tests/fixtures/store_v2_5_0.h5`; CLI opt-out is `create --f32-index`)
stores the index's own copy of the embeddings as `i8`,
which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors
at 100K); because quantised distances are approximate and `ef` cannot
compensate, the query path then re-scores the candidate pool against the
exact embeddings, which holds recall at the f32 index's level. It is also
faster at equal recall: 1.63x the QPS on x86-64 (AVX2) and 1.18x on a
Raspberry Pi 5 (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm since
the intrinsic is unstable; plain NEON on pre-dotprod cores). The aarch64
code is `cfg`'d out on x86, so x86 CI never compiles or lints it — test it
on real ARM (`rpivision02`, 10.0.2.3, is a Pi 5). `hybrid_search` keeps one incremental BM25
index for the life of the store and never writes the store: Hebbian
activation boosts are persisted by the next checkpoint (or on drop), not per
query. Measure any search-path change with
`cargo run --release -p clawhdf5-bench --bin search_harness` (baselines in
`BENCHMARKS.md`).
- WAL (write-ahead log) for crash-safe persistence, with a chained CRC32
trailer per entry (each entry's CRC folds in the previous entry's CRC) so a
corrupted, reordered, duplicated, or spliced entry stops replay cleanly
instead of loading bad or tampered data. The pre-chaining per-entry-CRC
format (v2) is still fully readable; the oldest no-CRC format (v1) is only
reachable through the one-time migration path in `HDF5Memory::open`, not
through the public `WalFile::read_entries`.
**What the WAL guarantees:** integrity, ordering, and recovery from a
*process* crash at any point — including between a checkpoint and the WAL
truncate (each checkpoint records a `WalMark` in `/meta`, and `open()` skips
the WAL prefix the `.h5` already contains, so entries are never applied
twice). Checkpoints and snapshots are made durable as a unit (temp file
synced, renamed, directory synced). **What it does not guarantee:**
individual WAL appends are *not* fsynced (a deliberate latency trade-off), so
saves made since the last checkpoint can be lost on power failure or kernel
panic. Current header version is 4 (adds the `Update` record used by
`save_or_update`); v3 files are read and upgraded in place.
- A store has a **single writer**: `HDF5Memory::create`/`open` hold an exclusive
advisory lock on `<store>.h5.lock` and a second opener gets
`MemoryError::Locked`. Use `HDF5Memory::open_read_only` for a lock-free,
never-writing point-in-time view (the CLI's `recall`/`stats`/`agents-md`/
`export` do). An unreadable WAL (torn header, bad magic) is quarantined to
`<store>.h5.wal.corrupt-<ts>` rather than blocking `open()`; a WAL with an
unknown *newer* version still fails and is left untouched.
- `MemoryConfig::float16` (**on by default** for new stores, persisted;
existing stores keep their recorded `false` — guarded by the v2.5.0
fixture in `tests/float16_store.rs`; CLI opt-out is `create --f32`) writes
`/memory/embeddings` as IEEE half precision (48% smaller file at 100K;
LongMemEval with real MiniLM embeddings identical to f32).
`MemoryCache::half_precision` rounds each embedding as it enters the cache (push, update, WAL replay, and on load of a store still
`f32` on disk), so memory and file agree bit for bit; the conversions live
in `clawhdf5_format::float16` and must stay the single implementation.
Values beyond ±65504 are `MemoryError::InvalidEntry`. Interop: every file
must open in h5py — `f32` datasets and empty datasets did not until
2026-09-23 (see `docs/known-issues.md`); the agent's `h5py_interop` test
guards a whole store.
- `HDF5Memory::search(query_emb, text, &SearchOptions)` is the full search
path: optional source-channel filter (applied before ranking; exact scan of
the allowed records whenever cheaper than `pool × M` index distance
Reference docs: `docs/known-issues.md` (open issues table first — check it
before calling something a bug or a feature), `BENCHMARKS.md` (headline
numbers first), `CONFORMANCE.md` (generated), `docs/design/range-reads.md`
and `docs/design/swmr.md`, `CHANGELOG.md` (full detail of every fix).
## Standing rules
- **No C in the default build.** No libhdf5; deflate defaults to pure-Rust
zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake). `ci-test.sh`
fails if a C-building crate enters the core crates' default tree. Zstd,
SZIP, `https` (ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are opt-in. flate2
must keep `runtime_detection` with zlib-rs — without it zlib-rs loses SIMD
and inflates 3.5x slower.
- **Every file we write must open in h5py/libhdf5.** Interop tests compare
against h5py and h5dump; `f32` and empty datasets did not open until
2026-09-23.
- **float16 has one implementation:** `clawhdf5_format::float16`.
- **Claims need evidence.** Performance and integration claims in docs must
be measured, dated (with machine and command), or withdrawn. Benchmark
numbers are dated records: never edit a measured value, add a new dated
section and mark the old one superseded.
- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not and
never was an OpenClaw memory plugin; the old `memory.backend = "clawhdf5"`
config was never valid. `docs/openclaw.md` records what a real plugin would
need. The `openclaw` module's `ClawhdfBackend` is just `search` with
re-rank + confidence on.
- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream
v0.8.5 and the `osobh/zeroclaw` fork and their history): its memory
backends are its own; `clawhdf5-migrate`'s default SQLite layout is not
ZeroClaw's schema. Don't reintroduce integration claims without an
integration and a test against the real consumer.
- **known-issues.md:** one entry per bug; when fixed, record it in
`CHANGELOG.md` and move the entry to *Fixed (history)* with date, PR,
affected releases and what users must do — never delete it.
## HDF5 library: invariants and gotchas
- **Remote/range reads** (`docs/design/range-reads.md`, M0-M5 merged in PRs
#17-#21): every format-crate read path goes through `Storage`
(`read_at`/`read_ranges`/`hint`). `File::open_storage` takes any
`Storage`; `clawhdf5_remote::open_url` wraps HTTP (`HttpStorage`, ureq) or
`ObjectStoreStorage` in `BlockCache` (1 MiB blocks, LRU budget, in-flight
dedup, coalesced runs). Remote files are pinned by ETag/Last-Modified and
length (`RemoteError::FileChanged`). Zero-copy APIs and `File::as_bytes`
need an in-memory file. Parse through `File::storage()` and the `*_in`
functions, not `as_bytes`, in new code (the Python bindings do).
`ObjectStoreStorage` runs reads on its own small tokio runtime, so it
works from any thread.
- **SWMR** (`docs/design/swmr.md`): `File::open_swmr` reads a file a libhdf5
SWMR writer is appending to — positioned reads, no chunk cache, bounded
retries (100), `Dataset::refresh()`. clawhdf5 has no SWMR writer; remote
SWMR is out of scope.
- **Browser** (`clawhdf5-wasm`, read-only, no Zstd/SZIP): `openUrl` reads
through the restartable "NeedBytes" cache (`src/lazy.rs`: a call is re-run
after each wave of misses; no block is evicted while a call runs); the HTTP
is JavaScript (`js/remote.js`).
- **In-place editing** (`clawhdf5::FileEditor`): overwrites values, grows and
shrinks chunked datasets (every chunk index) and sets attributes (compact
and dense) without rewriting the file, changing indexes and heaps as
libhdf5 does; freed space is reused within one editor. Anything it cannot
do safely is `Error::Unsupported` before any write (limits in
`docs/known-issues.md`). The algorithms follow libhdf5 `hdf5_1_14_6`
(github.com/HDFGroup/hdf5). Test changes with `cargo test -p
clawhdf5-tools --test edit_interop --test edit_coverage_interop`.
- **Provenance:** `Dataset::verify_provenance()` (facade `provenance`
feature, default) re-hashes a dataset against its `_provenance_sha256`
attribute (`DatasetBuilder::with_provenance`). Opt-in per call; unkeyed
hash — tamper-evident, not tamper-proof.
## Agent memory: invariants and gotchas
- **Search.** `HDF5Memory::search(query_emb, text, &SearchOptions)` is the
full path: optional source-channel filter (before ranking; exact scan of
the allowed records when cheaper than `pool × M` index distance
evaluations, and as the fallback when the pool comes back short), fusion,
activation scaling, optional re-ranking and confidence rejection.
`hybrid_search`/`hybrid_search_with` are thin wrappers; `ClawhdfBackend`
(the `openclaw` module) is `search` with re-rank + confidence on.
- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not an
OpenClaw memory plugin and never was — the old `memory.backend = "clawhdf5"`
config was never valid. Don't reintroduce OpenClaw claims; `docs/openclaw.md`
records what a real plugin would need.
- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream
v0.8.5 and the `osobh/zeroclaw` fork, and their full history): no
`clawhdf5` feature or backend exists; ZeroClaw's memory backends are
sqlite/lucid/postgres/qdrant/markdown/none behind its own `Memory` trait.
`clawhdf5-migrate`'s default SQLite layout (`memory_chunks`, `sessions`,
`entities`, `relations`) is not ZeroClaw's schema either (ZeroClaw's is a
`memories` table). Don't reintroduce integration claims without an
integration and a test against the real consumer. Measure changes with
`search_harness --options-study`.
- `MemoryConfig::compression` is off by default; when on, embeddings are
deflate-compressed, or Zstd with the agent's `zstd` feature (links libzstd).
- Signed checkpoints (`clawhdf5-agent` `signing` module): with
`HDF5Memory::set_signing_key` every checkpoint stores an Ed25519-signed
manifest (SHA-256 per record in a Merkle tree + settings/sessions/graph
hashes; per-record hashes in `/integrity/record_hashes`);
`HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover exactly
what the file persists in the form the loader returns it (strings lose
trailing NULs; an empty WAL mark is not written) or untouched stores stop
verifying — `tests/signed_store.rs` round-trips awkward strings. The key is
never persisted; a signed store refuses to checkpoint without it
(`MemoryError::SigningKeyRequired`, and `MemoryError` is `#[non_exhaustive]`).
`hybrid_search`/`hybrid_search_with` are thin wrappers. It keeps one
incremental BM25 index for the life of the store and never writes the
store: Hebbian activation boosts are persisted by the next checkpoint (or
on drop).
- **HNSW** (`hnsw` feature, default): the approximate `clawhdf5-ann` index
mirrors the cache and self-heals on drift; build the agent with
`--no-default-features --features float16` for the exact linear scan.
`parallel` (default) builds it on a thread pool with an identical graph.
Neighbour selection uses the HNSW paper's diversity heuristic (closest-M
capped recall at 0.31 recall@10 at 100K on clustered data). The graph is
saved to `<store>.h5.ann` at each checkpoint, tied to it by a generation
id; a stale or damaged sidecar is ignored and the index rebuilt.
- **`MemoryConfig::quantized_index`** (default on for new stores, persisted;
older stores load as `false` — guarded by `tests/fixtures/store_v2_5_0.h5`;
CLI `create --f32-index`): the index's copy of the embeddings is `i8`, and
the query path re-scores candidates against the exact embeddings. The
aarch64 kernels (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm) are
`cfg`'d out on x86, so x86 CI never compiles them — test on real ARM
(`rpivision02`, 10.0.2.3, a Pi 5) or rely on the `test-arm64` job.
- **`MemoryConfig::float16`** (default on for new stores, persisted; older
stores keep `false` — guarded in `tests/float16_store.rs`; CLI `create
--f32`): `/memory/embeddings` is IEEE half. `MemoryCache::half_precision`
rounds each embedding as it enters the cache (push, update, WAL replay, and
load of a store still `f32` on disk) so memory and file agree bit for bit.
Values beyond ±65504 are `MemoryError::InvalidEntry`. The agent's
`h5py_interop` test guards that a whole store opens in h5py.
- **WAL.** Chained CRC32 per entry (a corrupted, reordered, duplicated or
spliced entry stops replay cleanly). Header version 4 (`Update` record for
`save_or_update`); v3 is upgraded in place, v2 read, v1 only through the
one-time migration in `HDF5Memory::open`. Each checkpoint records a
`WalMark` in `/meta` so `open()` never applies an entry twice; checkpoints
and snapshots are durable as a unit (temp file synced, renamed, directory
synced). Individual WAL appends are **not** fsynced (deliberate): saves
since the last checkpoint can be lost on power failure or kernel panic.
- **Single writer.** `create`/`open` hold an exclusive lock on
`<store>.h5.lock` (`MemoryError::Locked` for a second opener);
`open_read_only` is a lock-free point-in-time view (CLI `recall`/`stats`/
`agents-md`/`export`). An unreadable WAL is quarantined to
`<store>.h5.wal.corrupt-<ts>`; a WAL of an unknown newer version fails and
is left untouched.
- **Signed checkpoints** (`signing` module): with `set_signing_key` each
checkpoint stores an Ed25519-signed manifest (per-record SHA-256 in a
Merkle tree plus settings/sessions/graph hashes; `/integrity/record_hashes`);
`HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover
exactly what the file persists in the form the loader returns it (strings
lose trailing NULs; an empty WAL mark is not written) —
`tests/signed_store.rs` round-trips awkward strings. The key is never
persisted; a signed store refuses to checkpoint without it
(`MemoryError::SigningKeyRequired`; `MemoryError` is `#[non_exhaustive]`).
WAL entries after the checkpoint are not covered.
- `Dataset::verify_provenance()` (clawhdf5 facade, `provenance` feature, on by
default) recomputes a dataset's SHA-256 and compares it against the
`_provenance_sha256` attribute written automatically on save when
`DatasetBuilder::with_provenance` is used. It's opt-in per call, not run
automatically on open — it decodes and hashes the whole dataset. The hash
is unkeyed (tamper-*evident*, not tamper-*proof*): it detects accidental
corruption, not a deliberate actor able to modify both the data and the
stored hash.
- `clawhdf5-agent`'s `HDF5Memory::save`/`save_batch`/`save_or_update` run every
write through an in-memory (session-scoped, not persisted to disk)
provenance ledger and write-anomaly detector: a content hash per record
(`provenance.rs`) for detecting accidental mid-session corruption, plus
rate-limit/injection-pattern/source-distribution checks (`anomaly.rs`).
Alerts never block a save — drain them with `HDF5Memory::take_anomaly_alerts`.
`MemorySource` for this bookkeeping is inferred from the caller-supplied
`source_channel` string (a heuristic, not an authenticated trust boundary).
- In-place modification: `clawhdf5::FileEditor` (`crates/clawhdf5/src/edit/`)
overwrites values, grows and shrinks chunked datasets (every chunk index,
version-2 B-trees included) and sets attributes (compact and dense
storage) in existing files (h5py- or clawhdf5-written) without rewriting
them, changing indexes and heaps as libhdf5 does (index shapes and heap
bookkeeping are compared with libhdf5's in the tests); space an edit
frees is reused by later edits of the same editor. Anything it cannot do
safely is `Error::Unsupported` before any write (limits in
`docs/known-issues.md`). Test changes with
`cargo test -p clawhdf5-tools --test edit_interop --test
edit_coverage_interop` (h5py, h5dump, `h5rs check`, structure comparisons
with libhdf5; libhdf5 sources for the algorithms are at
github.com/HDFGroup/hdf5, tag `hdf5_1_14_6`).
- Remote files (`clawhdf5-remote`, range-read milestone M3 of
`docs/design/range-reads.md`): `open_url("http://…")` gives a
`clawhdf5::File` over `File::open_storage`, read through `BlockCache`
(1 MiB blocks, LRU byte budget, per-block in-flight dedup across threads,
runs coalesced into parallel requests). `HttpStorage` pins the file by
ETag/Last-Modified and length (a change is `RemoteError::FileChanged`),
refuses servers that ignore `Range` unless a full download is allowed,
and retries transient failures. `ObjectStoreStorage` (feature
`object-store`, pure Rust) runs each read on a small owned tokio
runtime and waits on a channel, so it works from any thread, including
inside `spawn_blocking` or another runtime. Default build is plain HTTP with
no C; `https` (rustls + ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are
opt-in. Tests run a std-only HTTP server
(`tests/common/server.rs`, also the `range_server` example);
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
file over HTTP with `File::open`.
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
call is re-run after each wave of misses; no block evicted while a call
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
version) and tests it under Node and headless Chromium (a Playwright
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
(range server with request counts, 200 MB budget file); the CI container
has neither, so CI runs the native `h5py_interop` and `lazy` tests
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
numbers in the example's README predate `openUrl`.
- Python and Node.js bindings for cross-language use
- NetCDF-4 compatibility for scientific data interop
- **Write bookkeeping.** `save`/`save_batch`/`save_or_update` feed an
in-memory, session-scoped provenance ledger and anomaly detector
(`provenance.rs`, `anomaly.rs`); alerts never block a save
(`take_anomaly_alerts`). `MemorySource` is inferred from the caller's
`source_channel` string — a heuristic, not a trust boundary.
- `MemoryConfig::compression` is off by default (deflate, or Zstd with the
agent's `zstd` feature, which links libzstd).
## Workflows
### Build
Put `$HOME/.cargo/bin` on `PATH`. The h5py/netCDF4 interop tests find their
Python through `CLAWHDF5_PYTHON` (or `.venv/bin/python`); create it with
`python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 hdf5plugin`.
Set `CLAWHDF5_REQUIRE_INTEROP=1` to make a missing interpreter a failure.
```bash
cargo build --release
```
### Test
```bash
cargo test --workspace
bash scripts/ci-test.sh # everything CI runs (see below)
```
### CI
`.gitea/workflows/ci.yml` has two jobs, both green as of 2026-09-22:
- **`test`** (`ubuntu-latest`, in `rust:latest`) runs `scripts/ci-test.sh` with
the h5py/netCDF4 interop suites required (`CLAWHDF5_REQUIRE_INTEROP=1`).
Served by the `tank` and `architect` runners.
- **`test-arm64`** (`linux_arm64`) lints and tests the aarch64 code — the NEON
kernels are `cfg`'d out on x86, so this is the only place they are built.
Served by `vision-01` (host mode) and `vision-02` (Docker), so steps must
work in both.
### CI (`.gitea/workflows/`)
- **`ci.yml` `test`** (`ubuntu-latest`, `rust:latest` container; runners
`tank`, `architect`): installs h5py/netCDF4/xarray/hdf5plugin/maturin/pytest,
`hdf5-tools` and `cmake`, then runs `scripts/ci-test.sh` with
`CLAWHDF5_REQUIRE_INTEROP=1`. The script runs: fmt; clippy (workspace, the
format feature matrix, each plugin filter alone, parallel, fast-deflate,
remote with all backends, h5rs remote); "no C in the default build";
wasm32 build and clippy; `check-32bit-casts.sh`; the wasm package under
Node when `node` and `wasm-bindgen` exist (not in CI); the MSRV check;
`cargo test` (workspace plus feature variants: format matrix, parallel,
remote/object_store, h5rs URLs, ann parallel, fast-deflate); the h5py
interop suites (`writer_h5py_tests --include-ignored`, plugin filters,
ZFP); the Python package (clippy, `maturin build`, pytest vs h5py);
`cargo bench --no-run`; `check-nostd.sh`; an optional fuzz smoke run
(`CLAWHDF5_FUZZ_SECONDS`).
- **`ci.yml` `test-arm64`** (`linux_arm64`; `vision-01` host mode,
`vision-02` Docker — steps must work in both): clippy and tests of
`clawhdf5-accel`, `-ann`, `-format`; the only place the NEON kernels build.
- **`conformance.yml`** (nightly 03:17 UTC and manual): probe unit tests,
`conformance/test_ref.py`, then `conformance/run.sh` (gate:
`conformance/check.py` against `baseline.json`).
Keep workflows free of JavaScript actions (`actions/checkout`, `actions/cache`,
…): `rust:latest` has no `node`, and not every runner reaches GitHub, where
they are fetched from. Check out with plain `git` instead. The `test` job
installs `cmake` for the opt-in `fast-deflate` (zlib-ng) steps; the default
build needs no C toolchain, so `test-arm64` does not.
All runners are on `gitea-runner` 3.5.0, from `docker.gitea.com/act_runner`
— `gitea/act_runner:latest` on Docker Hub is frozen at 0.6.1.
Keep workflows free of JavaScript actions (`actions/checkout`,
`actions/cache`, …): `rust:latest` has no `node` and not every runner reaches
GitHub. Check out with plain `git`. Runners are `gitea-runner` 3.5.0 from
`docker.gitea.com/act_runner` (`gitea/act_runner:latest` on Docker Hub is
frozen at 0.6.1).
### CLI
### Conformance
```bash
cargo run -p clawhdf5-cli -- --help
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands
CLAWHDF5_PYTHON=.venv/bin/python bash conformance/run.sh --no-fetch # writes CONFORMANCE.md
```
Reads 697 files of eight pinned corpora with clawhdf5 and h5py and compares
them object by object (602 ok as of 2026-09-27). `CONFORMANCE.md` is
generated — never hand-edit it (its wording lives in `conformance/report.py`).
Use `--update-baseline` only after an intended change in results.
`CONFORMANCE_CACHE` points at an existing corpus cache (`conformance/.cache`,
about 450 MB). See `conformance/README.md`.
### HDF5 tools (`h5rs`, crate `clawhdf5-tools`)
### HDF5 tools (`h5rs`)
```bash
cargo run -p clawhdf5-tools -- ls -r file.h5 # also dump [--json], stat, diff, check
bash scripts/h5rs-fuzz.sh # every subcommand over the CVE corpus: no panic/crash/hang
bash scripts/h5rs-check-ok-files.sh --data # check passes every fully-read conformance file
```
Its interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian
`hdf5-tools`, installed in CI); `dump` must stay byte-identical to h5dump on
the test files.
Interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian `hdf5-tools`);
`dump` must stay byte-identical to h5dump on the test files.
### Remote and browser tests
- `clawhdf5-remote` tests run a std-only HTTP server
(`tests/common/server.rs`, also the `range_server` example);
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
file over HTTP with `File::open`.
- wasm: `bash examples/wasm-viewer/test/run.sh` builds the package (needs the
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node and
headless Chromium (Playwright's download in `~/.cache/ms-playwright` on
tank) against `test/serve.py` (range server with request counts). CI has
neither, so it runs the native `h5py_interop` and `lazy` tests
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus).
### Python bindings
```bash
cd crates/clawhdf5-py
maturin develop
python -c "import clawhdf5; print(clawhdf5.__version__)"
cd crates/clawhdf5-py && maturin develop
python -m pytest crates/clawhdf5-py/tests # compares with h5py; editing tests want CLAWHDF5_H5RS=<path to h5rs>
```
### Benchmarks
- Search path: `cargo run --release -p clawhdf5-bench --bin search_harness`
(`--full`, `--options-study`, `--footprint`, …); reads: `read_harness`,
`concurrent_read`; criterion benches with `cargo bench -p <crate>`.
- Run on an idle machine (1-minute load average below 2; wait otherwise),
alternate base and candidate binaries for A/B comparisons, and record date,
machine, commit and command with every number in `BENCHMARKS.md`.
- `scripts/run-benchmarks.sh` is stale (it benchmarks `rustyhdf5-format` and
overwrites `BENCHMARKS.md`) — do not run it.
### CLI
```bash
cargo run -p clawhdf5-cli -- --help
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot
```
## Integration
@@ -255,8 +268,7 @@ python -c "import clawhdf5; print(clawhdf5.__version__)"
verified consumer: `cbh-core` reads and writes `.brain` files through the
facade (`File`, `FileBuilder`, `AttrValue`, `Selection`), `cbh-scanner`
uses the facade, and `cbh-cli` uses `clawhdf5_agent::bm25::BM25Index`. It
depends on this repo by path (`../clawhdf5`), so it builds against whatever
is checked out — changes to those APIs reach it directly. Verified
2026-09-25 against main: builds, and its 204 tests pass.
- OpenClaw and ZeroClaw were both described as consumers; neither integrates
clawhdf5 (see Key Features and `docs/openclaw.md`).
depends on this repo by path (`../clawhdf5`), so changes to those APIs
reach it directly. Verified 2026-09-25 against main: builds, and its 204
tests pass.
- OpenClaw and ZeroClaw integrate nothing (see *Standing rules*).
+40 -36
View File
@@ -1,40 +1,33 @@
# Design: range reads (reading HDF5 without holding the whole file)
Status: proposal, 2026-09-26; the plan for Phase 3's largest architectural
change. Progress: M0 and M1 are done, and so is M2 (branch
`feat/p3-m2-raw-data`): every read path of the format crate works through
`Storage`, v2 B-trees, dense groups and raw data included, and
`File::open_storage` gives the facade's read API over any `Storage` (see
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
object stores) and URLs in `h5rs` (see the M3 status below). M5 (SWMR) is
done on branch `feat/p3-m5-swmr-reader`, with its own design in
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
next. Every count in §1–§2 was
Status (updated 2026-09-28): **implemented and merged.** Proposed
2026-09-26 as the plan for Phase 3's largest architectural change; every
milestone below is on `main`:
object stores) and URLs in `h5rs` (see the M3 status below); the Python
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
| Milestone | What | Merged |
|---|---|---|
| M0 | indexed name lookups, checked address conversion | PR #17 (`8f59b2e`) |
| M1 | metadata parsers over the `Storage` trait | PR #17 (`8f59b2e`) |
| M2 | raw data over `Storage`, `File::open_storage` | PR #18 (`a4c2ace`) |
| M3 | `clawhdf5-remote` (block cache, HTTP(S), object stores), URLs in `h5rs`; Python `clawhdf5.File(url)` | PR #18 (`a4c2ace`); Python in PR #19 (`7a8fae0`) |
| M4 | wasm `openUrl` through the restartable `NeedBytes` mode | PR #19 (`7a8fae0`); fewer round trips in PR #21 (`9b5803f`) |
| M5 | SWMR reader (`File::open_swmr`), design in [`swmr.md`](swmr.md) | PR #19 (`7a8fae0`) |
change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet.
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
Each milestone's own *Status* note in §4 records what was built and how it
differs from the plan. What is still missing is tracked in
[`docs/known-issues.md`](../known-issues.md) ("Range reads", "Remote
files" and "`clawhdf5-wasm`" limits); the main gaps are a paged file's page
size as the block size, remote SWMR, and a SWMR writer.
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
reader, through the restartable `NeedBytes` mode (see the M4 status
below). M5 (SWMR) is not started.
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
converted code without changing the plan: object-header continuation chunks
are followed without recursion (still one bounded `read_at` per chunk), and
implicit chunk indexes are addressed over the maximum chunk grid (in
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
working on the whole file in memory; it is not part of this design. Every
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
every count in a milestone's status on the date it gives, with the
commands given next to it. No timing numbers appear here on purpose: the machine was shared
with other build jobs when this was written.
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) is not part
of this design. Every count in §1–§2 was taken on `tank` on 2026-09-26 at
commit `de2a53f`, and every count in a milestone's status on the date it
gives, with the commands given next to it. No timing numbers appear here on
purpose: the machine was shared with other build jobs when this was written.
## The problem
@@ -398,7 +391,8 @@ Rejected. It is how one would retrofit a C library that cannot change; we can.
Adopt **(a)**, with a block cache as a required part of every non-local
backend, **(c)** as a cache policy, and the wasm path through the restartable
`NeedBytes` mode. Every milestone keeps `main` green: `cargo test
--workspace`, clippy, the conformance gate at 575/697 unchanged, and the mmap
--workspace`, clippy, the conformance gate unchanged (575/697 when this was written; 602/697
after PR #21), and the mmap
fast path within benchmark noise.
**M0 — prerequisites (≈1 week).**
@@ -414,7 +408,8 @@ fast path within benchmark noise.
n children decodes its links O(n) times. Look names up through the index
(above) and let a listing hand out its entries, so the cache has less to
absorb.
- *Status 2026-09-26:* done on branch `perf/p3-indexed-lookups` — link and
- *Status 2026-09-26:* done on branch `perf/p3-indexed-lookups` (merged in
PR #17) — link and
attribute names through the name indexes (`group_v2::resolve_child`,
`attribute::find_attribute_in_file`; creation-order lookups by name do
not exist in the API, so the creation-order index is still only listed),
@@ -438,6 +433,11 @@ fast path within benchmark noise.
(`fn parse(data: &[u8], ..) { parse_in(data, ..) }`, generic core), so
callers and the other crates don't move yet.
- Replace the 5 open-ended slices and 38 `len()` checks with bounded reads.
- *Status 2026-09-26:* done on branch `feat/p3-storage-trait` (merged in
PR #17) for the `Storage` trait and the metadata parsers listed in
`CHANGELOG.md` under "Range reads, milestone M1"; group B-tree v2 lookups,
dense groups and the facade were converted in M2. Both error enums are
`#[non_exhaustive]`; storage failures are `FormatError::Storage`.
**M2 — raw data over the trait (1–2 weeks).**
- `data_read`, `chunked_read`, `parallel_read`, `partial_read`, `vds`,
@@ -449,7 +449,8 @@ fast path within benchmark noise.
on other backends (they already return `Option`/`Result`).
- Facade: `File::open_storage(Box<dyn Storage + Send + Sync>)`; `File::open`
keeps mmap and `from_bytes` keeps `Vec`, both through `impl Storage for [u8]`.
- *Status 2026-09-26:* done on branch `feat/p3-m2-raw-data`. As planned,
- *Status 2026-09-26:* done on branch `feat/p3-m2-raw-data` (merged in PR
#18). As planned,
with these choices:
- `File::open_storage` takes an `Arc<dyn Storage + Send + Sync>` (the
file handle is shared by its datasets and may be sent across threads).
@@ -487,7 +488,8 @@ fast path within benchmark noise.
§2; the page size for paged files; the first block prefetched on open) and
a request counter exposed for tests and users.
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote` (merged in PR
#18), except the
Python bindings (done 2026-09-27, below), with these choices:
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
@@ -526,7 +528,7 @@ fast path within benchmark noise.
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
6.4 MB: 35 001 object headers spread over the file).
- *Status 2026-09-27, Python bindings:* done on branch
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
`feat/p3-python-remote-edit` (merged in PR #19). `clawhdf5.File(url)` and
`File.open_url(url, **options)` (cache and HTTP options) go through
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
@@ -543,7 +545,8 @@ fast path within benchmark noise.
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
to a whole download when the server does not answer 206.
- `examples/wasm-viewer`: open by URL.
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy` (merged in PR
#19), as planned,
with these choices:
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
`Storage` over the blocks fetched so far. A call (open, list, read)
@@ -607,7 +610,7 @@ fast path within benchmark noise.
missing blocks: listing 3000 datasets went from 185 passes to 6.
The 32-bit risk below is covered by a Node test that reads data at
3 GiB from a mock server and is refused a 4 GiB file.
- Fewer round trips (2026-09-27, later): the walks descend into every
- Fewer round trips (2026-09-27, later; merged in PR #21): the walks descend into every
child after a failure (not only read the siblings), and parsers
call `Storage::hint` for what they read next (node bodies, object
header chunks, a dense group's heap blocks, a listing's child
@@ -621,7 +624,8 @@ fast path within benchmark noise.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
blocks past the old end. Needs libhdf5 SWMR semantics research first.
- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader`; design and
- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader` (merged in
PR #19); design and
libhdf5 research in [`swmr.md`](swmr.md). Differences from the sketch
above: the refresh is per dataset (`Dataset::refresh`, as libhdf5's
`H5Drefresh`), not per file — a SWMR writer only grows datasets, and the
+9 -4
View File
@@ -1,7 +1,8 @@
# Design: reading files a SWMR writer is still appending to (range-read M5)
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
(see "Status" at the end). This is milestone M5 of
Status: design 2026-09-27; the reader is implemented and merged (branch
`feat/p3-m5-swmr-reader`, PR #19, `7a8fae0`; see "Status" at the end).
clawhdf5 has no SWMR writer. This is milestone M5 of
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
@@ -165,7 +166,7 @@ writer cannot add them), and `MmapFile`/`LazyFile`.
## Status
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
above, merged to `main` in PR #19 (`7a8fae0`) (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
build): no read returned a value the writer had not written at that
@@ -175,4 +176,8 @@ variant of the test with the chunk cache left on in live mode fails it
(stale chunk index / edge chunk), which is why live files do not use it.
Also found: `File::open` of such a file had been failing since the
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).
end-of-file check of 2026-09-26 (item 1; fixed before any release, see
[`docs/known-issues.md`](../known-issues.md#files-a-swmr-writer-had-open-could-not-be-read-past-a-stale-end-of-file)).
Not done (tracked in `docs/known-issues.md`, "Range reads" limits): SWMR
writing, remote SWMR, and live reading through `MmapFile`/`LazyFile`.
+571 -970
View File
File diff suppressed because it is too large Load Diff
+3 -1
View File
@@ -71,4 +71,6 @@ Building blocks, usable as a library today, but not an OpenClaw plugin:
and export rewrites every heading as `##`.
- `crates/clawhdf5-napi` and `packages/clawhdf5-node` — Node bindings and a
TypeScript wrapper. **Not published, not built or tested in CI, and known to
be broken**; see `docs/known-issues.md`.
be broken**; see
[`docs/known-issues.md`](known-issues.md#the-nodejs-package-packagesclawhdf5-node-does-not-work)
(re-checked 2026-09-28: unchanged).