Files
osobhandClaude Opus 5.5 b5a5041655 Write files HDF5 1.8 can read: FileWriter/FileBuilder::libver_bounds
New `LibVer` (V18, V110, V112, V114, V200, Latest) and
`libver_bounds(low, high)` on the format crate's `FileWriter` and the
facade's `FileBuilder`, as libhdf5's H5Pset_libver_bounds / h5py's
libver=(low, high). The default stays (V110, Latest), byte for byte what
was written before.

With a low bound of 1.8: superblock version 2, layout message version 3
(contiguous, compact, chunked) and a version-1 B-tree chunk index for
every chunked dataset, resizable ones included -- what libhdf5 2.x writes
under libver=('v108', 'latest'). The new chunk B-tree writer
(btree_v1_write.rs) replays H5B_insert with the H5Dbtree.c callbacks for
row-major insertion (split ratios 0.1/0.5/0.9, right keys moved as
H5D__btree_cmp3 moves them, root kept in place): its trees equal
libhdf5's node for node for 1-D/2-D/3-D, 2- and 3-level, filtered and
unfiltered datasets (libhdf5 writing without a chunk cache).

The high bound refuses, with FormatError::LibverBound before anything is
written, what needs a newer format: virtual datasets and the paged
file-space strategy (1.10), the 1.12 reference types (datatype v4),
native complex (datatype v5, HDF5 2.0), and a low bound above the high.

Tests: tools/tests/libver_v18.rs writes every writer feature under
(V18, V18), and HDF5 1.8.23's h5dump (scripts/build-hdf5-1.8.sh; skipped
when absent) dumps it exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset; h5py, clawhdf5 and h5rs check --data agree;
then FileEditor grows/appends/annotates it and h5py appends, and every
reader checks again. read_harness gains --v18 and --chunk N.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00
..

clawhdf5-bench

The measurement harnesses behind BENCHMARKS.md: HDF5 read and write speed (against libhdf5 and h5py where noted) and the agent store's search, footprint and retrieval quality. Not meant for publishing; nothing else in the workspace depends on it. Run everything with --release, and quote numbers with the machine, date and command, as BENCHMARKS.md does.

Binaries

Binary Measures
read_harness full reads vs hyperslab selections of a chunked 2-D dataset (compressed and not) and a contiguous one: does a selection cost scale with the selection or the dataset? (-- --large for 512 MB)
concurrent_read decoded read throughput vs threads on one open File; scripts/concurrent_read_h5py.py runs the same workload with h5py (threads and processes) and scripts/compare_concurrent_read.py tabulates both
search_harness HNSW recall@10 vs exact search, QPS and latency per ef, and end-to-end HDF5Memory ingest/checkpoint/open/search at 1K–100K (--full); studies: --float16-study, --options-study, --signing-study, --ann-only --uniform
longmemeval_bench LongMemEval retrieval recall (turn and session Hit@k, MRR) — retrieval, not QA accuracy. Oracle or full longmemeval_s haystack; --features embeddings (or embeddings-cuda) embeds with MiniLM, otherwise the vector stage is inert and the run is BM25-only
memory_arena a deterministic multi-session retrieval benchmark (BM25-only)
footprint_bench file size and bytes per record at 100–100K records, float16 or --f32, WAL on/off, compressed or not
consolidation_efficiency retrieval before and after consolidation on signal + noise records
ephemeral_perf the in-memory ephemeral tier's set/get latency
mpi_io_bench clawhdf5-io's MpiVol (root-read + broadcast, not collective I/O); needs --features mpi-io and mpirun
cargo run --release -p clawhdf5-bench --bin search_harness -- --full
cargo run --release -p clawhdf5-bench --bin read_harness

Criterion benches and example

  • cargo bench -p clawhdf5-bench runs h5bench_write, h5bench_read and h5bench_meta (h5bench-style sequential, chunked, strided and metadata workloads). --features libhdf5-compare adds the same workloads through libhdf5 (the hdf5-metno crate; needs a system libhdf5 1.14).
  • examples/worldmodel_sampling.rs: shuffled per-frame reads of a (N, H, W, C) uint8 dataset, clawhdf5 against h5py on the same file.

Features

Feature What Builds C
libhdf5-compare libhdf5 variants of the Criterion benches links the system libhdf5
mpi-io mpi_io_bench yes (mpi-sys; needs an MPI installation)
embeddings MiniLM embeddings for longmemeval_bench (candle) yes (a cc build dependency in the candle/tokenizers tree)
embeddings-cuda the same on a CUDA GPU (minutes instead of hours on the full haystack) yes (CUDA)

License

MIT