Every crate under crates/ now has a README (android, bench, cli, napi and wasm had none), each saying what the crate is, its main types and functions (names checked against the code), its cargo features with defaults and which ones build C (checked with `cargo tree`), and links to the top-level docs. Corrections to the old stubs: - clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs clawhdf5-format as a dependency. - clawhdf5-filters: deflate backends only, and no library crate depends on it; the filter pipeline and every other codec are in -format. - clawhdf5-gpu: vector distance compute, not I/O; not used by HDF5Memory::search. - clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO. - clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and search(q, k, ef). - clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm backends are reported but run the scalar kernels. - clawhdf5-gpu: the old example called l2_distances, which does not exist (l2_search). - clawhdf5-agent: it described a "vector store" with "GPU acceleration"; it now covers HDF5Memory, search options, WAL, signing, the graph. - crates.io/docs.rs badges removed and `cargo install <crate>` replaced: nothing is published; depend on git. - fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh. - tools: the FileEditor interop tests that live in this crate. - remote, py: license, other front ends, limits, File.mode/flush/chunks. The Rust examples of the facade, format, filters, accel, ann, derive and agent READMEs were compiled and run as tests (netcdf4, gpu and remote compiled only) in a scratch crate; the CLI example was run. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
3.0 KiB
3.0 KiB
clawhdf5-bench
The measurement harnesses behind BENCHMARKS.md:
HDF5 read and write speed (against libhdf5 and h5py where noted) and the
agent store's search, footprint and retrieval quality. Not meant for
publishing; nothing else in the workspace depends on it. Run everything with
--release, and quote numbers with the machine, date and command, as
BENCHMARKS.md does.
Binaries
| Binary | Measures |
|---|---|
read_harness |
full reads vs hyperslab selections of a chunked 2-D dataset (compressed and not) and a contiguous one: does a selection cost scale with the selection or the dataset? (-- --large for 512 MB) |
concurrent_read |
decoded read throughput vs threads on one open File; scripts/concurrent_read_h5py.py runs the same workload with h5py (threads and processes) and scripts/compare_concurrent_read.py tabulates both |
search_harness |
HNSW recall@10 vs exact search, QPS and latency per ef, and end-to-end HDF5Memory ingest/checkpoint/open/search at 1K–100K (--full); studies: --float16-study, --options-study, --signing-study, --ann-only --uniform |
longmemeval_bench |
LongMemEval retrieval recall (turn and session Hit@k, MRR) — retrieval, not QA accuracy. Oracle or full longmemeval_s haystack; --features embeddings (or embeddings-cuda) embeds with MiniLM, otherwise the vector stage is inert and the run is BM25-only |
memory_arena |
a deterministic multi-session retrieval benchmark (BM25-only) |
footprint_bench |
file size and bytes per record at 100–100K records, float16 or --f32, WAL on/off, compressed or not |
consolidation_efficiency |
retrieval before and after consolidation on signal + noise records |
ephemeral_perf |
the in-memory ephemeral tier's set/get latency |
mpi_io_bench |
clawhdf5-io's MpiVol (root-read + broadcast, not collective I/O); needs --features mpi-io and mpirun |
cargo run --release -p clawhdf5-bench --bin search_harness -- --full
cargo run --release -p clawhdf5-bench --bin read_harness
Criterion benches and example
cargo bench -p clawhdf5-benchrunsh5bench_write,h5bench_readandh5bench_meta(h5bench-style sequential, chunked, strided and metadata workloads).--features libhdf5-compareadds the same workloads through libhdf5 (thehdf5-metnocrate; needs a system libhdf5 1.14).examples/worldmodel_sampling.rs: shuffled per-frame reads of a(N, H, W, C)uint8dataset, clawhdf5 against h5py on the same file.
Features
| Feature | What | Builds C |
|---|---|---|
libhdf5-compare |
libhdf5 variants of the Criterion benches | links the system libhdf5 |
mpi-io |
mpi_io_bench |
yes (mpi-sys; needs an MPI installation) |
embeddings |
MiniLM embeddings for longmemeval_bench (candle) |
yes (a cc build dependency in the candle/tokenizers tree) |
embeddings-cuda |
the same on a CUDA GPU (minutes instead of hours on the full haystack) | yes (CUDA) |
License
MIT