osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00

clawhdf5

A pure-Rust HDF5 reader, writer and in-place editor — no libhdf5, and no C by default — with an agent-memory store built on it.

License: MIT Rust Conformance LongMemEval

clawhdf5 implements the HDF5 file format from the specification, in Rust. It reads superblocks v0–3, every group and chunk-index structure libhdf5 writes, the standard filters and the common plugin filters, variable-length data and virtual datasets, and follows files a SWMR writer is appending to. It reads files from libhdf5, h5py and netCDF-4 and writes files they read. The same library opens files over HTTP and in object stores by range requests, runs in the browser as WebAssembly, and has Python bindings with an h5py-shaped API.

Two things live in this repository:

  • The HDF5 library — the clawhdf5 crate and its parts, h5rs command-line tools, Python, WebAssembly and NetCDF-4 layers.
  • Agent memory (clawhdf5-agent) — a single-file store for AI agents (HNSW + BM25 hybrid search, write-ahead log, signed checkpoints) whose files are ordinary HDF5. See docs/agent-memory.md.

Nothing is published to crates.io or PyPI yet: use it from git or a checkout.

Contents

Evidence

Conformance. Every file of eight public corpora — the libhdf5 source tree's test files, the HDF Group's CVE reproducer corpus, netcdf-c, netcdf4-python, pyfive, h5wasm, h5py's and xarray's data files, 697 files in all, pinned by commit — is read by clawhdf5 and by h5py/libhdf5 and compared object by object (object set, shapes, a SHA-256 of every dataset's and attribute's values). Run of 2026-09-28 on tank, h5py 3.16 / HDF5 2.0 (CONFORMANCE.md):

files ok (identical to h5py) mismatch libhdf5 cannot open ref-bug¹ our-error¹ panic / hang / crash / OOM
697 602 0 92 2 1 0

¹ The three remaining objects are corrupt data (scale-offset codes past the end of a chunk, short unfiltered chunks, an N-Bit parameter list one value short) that HDF5 2.0 returns only by reading past a buffer; clawhdf5 refuses them, as libhdf5's development branch and its own test_filter_bad_params do. Details and evidence in CONFORMANCE.md § Reference bugs; the N-Bit file is counted as ref-bug or our-error depending on the run (known-issues). The run is a nightly CI job (.gitea/workflows/conformance.yml) that fails on any panic, hang or crash, or on an ok file that stops being ok.

Robustness on hostile files. On the 147 CVE and fuzzer files (CONFORMANCE.md § CVE corpus):

tool panic crash hang OOM
clawhdf5 0 0 0 0
h5dump 1.14.6 0 2 0 0
h5py 3.16.0 / HDF5 2.0.0 0 1 0 0

Sizes and addresses read from a file are checked before use (overflow-checked arithmetic, fallible allocation on the chunked read paths, bounded recursion in B-trees and object-header chains), and scripts/h5rs-fuzz.sh runs every h5rs subcommand over the corpus looking for panics, crashes and hangs.

Reads from many threads. A File is Send + Sync and there is no library-wide lock, so one open file serves many threads. Full reads of 64 deflate-compressed 64 MiB datasets, each read decoding on its calling thread (concurrent_read --decode-threads 1), tank (Ryzen 7 7800X3D, 16 threads), 2026-09-26, commit c5334b1 (BENCHMARKS.md):

threads clawhdf5, one File h5py, threads h5py, processes clawhdf5 / h5py processes
1 670 MB/s 410 MB/s 397 MB/s 1.69x
16 4944 MB/s 390 MB/s 3135 MB/s 1.58x

That run was noisier than others on the same machine, so compare ratios within it rather than MB/s across runs. A contiguous (uncompressed) full read on one thread ran at 6718 MB/s against h5py's 5545 in the same run.

Against libhdf5 1.14.6 from Rust, tank, 2026-08-03 (BENCHMARKS.md § Independent Validation): sequential read of 100K f32 23.3 µs vs 63.6 µs (2.7x); 128 attribute writes 85.2 µs vs 877 µs (10.3x); 64 group creates 130 µs vs 1.37 ms (10.6x); a 512×512 f32 chunked deflate-6 write 1.44 ms vs 65.0 ms (re-measured 2026-09-23 with the pure-Rust deflate: 1.46 ms vs 51.4 ms, 35x); a 100K f32 sequential write is a tie. The writer (FileBuilder) assembles a file in memory and writes it once, which is part of that difference; read the caveats in BENCHMARKS.md before quoting these.

What is supported

Limits and open issues, with dates, are in docs/known-issues.md.

Area Supported Read only Not supported
File format Superblock v0–v3, user blocks, v1/v2 object headers Metadata cache images Writing files HDF5 1.8 can read
Groups and links Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links Following external links (explicit error); user-defined links are skipped
Datatypes Integers and IEEE floats of every width and byte order (incl. f16), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex (HDF5 2.0 class 11) Variable-length strings and sequences, object references Writing variable-length data; decoding region and attribute references; x87 long double and binary128
Layouts and chunk indexes Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does); fill values; resizable datasets; virtual datasets (read limits in known-issues) Chunk indexes v1 B-tree and implicit (the editor also changes them) External raw data files (explicit error)
Filters deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP Other filter IDs, unless you register a codec (filter_registry::register_filter)
Editing in place FileEditor: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; filters this build cannot encode (refused before any write)
Access Local files (mmap or buffered), bytes in memory, any Storage backend, HTTP(S) and S3/GCS/Azure via clawhdf5-remote, SWMR reading (File::open_swmr, Dataset::refresh) Remote files and the browser are read-only SWMR writing; remote SWMR; MPI collective I/O (clawhdf5-io's mpi-io reads on one rank and broadcasts)
Bindings Python (read, 'w' for numeric arrays, 'r+' editing, URLs), NetCDF-4 (CF scale/offset/fill) WebAssembly (open(bytes), openUrl); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets Node.js (the package does not work; see known-issues)

Plugin filters other than LZF are cargo features (bitshuffle, bzip2, blosc, blosc2, zfp, or plugin-filters for all of them), all pure Rust; h5py + hdf5plugin read what clawhdf5 writes with them, and ZFP decodes bit-exact against hdf5plugin 7.1. pcodec (opt-in) uses a private filter ID that only clawhdf5 reads.

C dependencies, precisely. The core crates build no C by default: no libhdf5, and deflate is zlib-rs, whose output was byte-identical to zlib-ng's at levels 1, 6 and 9 on the benchmark inputs and which matched its HDF5 read and write speed within 6% (tank, 2026-09-23, BENCHMARKS.md). CI fails if a C-building crate enters their default dependency tree. C comes in only when you ask: fast-deflate (zlib-ng, needs cmake), zstd, szip, https and the cloud stores (ring / aws-lc-rs), the BLAS backends, clawhdf5-migrate (bundled SQLite) and the Node.js bindings. One exception links rather than builds C: on macOS the default system-zlib-decompress feature inflates with the system libz first.

Install

The crates are not on crates.io; depend on the repository (MSRV 1.92):

[dependencies]
clawhdf5        = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
# optional parts
clawhdf5-remote = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }   # HTTP / object stores
clawhdf5-agent  = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }   # agent memory

or, with a checkout, clawhdf5 = { path = "../clawhdf5/crates/clawhdf5" }. Add features = ["plugin-filters"] for every plugin filter.

The Python package is not on PyPI; build it with maturin into a virtualenv:

python -m venv .venv && . .venv/bin/activate
pip install maturin numpy
maturin develop --release -m crates/clawhdf5-py/Cargo.toml   # add --features https for https://
python -c "import clawhdf5; print(clawhdf5.__version__)"

h5rs: cargo install --path crates/clawhdf5-tools (add --features remote for URLs).

Quick start: Rust

use clawhdf5::{AttrValue, File, FileBuilder, Selection};

// Write: a chunked, deflate-compressed 2-D dataset that can grow along axis 0.
let data: Vec<f64> = (0..1000 * 64).map(|i| i as f64).collect();
let mut b = FileBuilder::new();
b.create_dataset("run/temps")          // intermediate groups are created, as in h5py
    .with_f64_data(&data)
    .with_shape(&[1000, 64])
    .with_maxshape(&[u64::MAX, 64])    // u64::MAX = unlimited
    .with_chunks(&[100, 64])
    .with_deflate(6)
    .set_attr("units", AttrValue::String("K".into()));
b.write("data.h5")?;

// Read: whole datasets, or a hyperslab (only the chunks it touches are decoded).
let file = File::open("data.h5")?;
let ds = file.dataset("run/temps")?;
assert_eq!(ds.shape()?, vec![1000, 64]);
let all = ds.read_f64()?;
let rows = ds.read_f64_selection(&Selection::Hyperslab {
    start: vec![10, 0], stride: vec![1, 1], count: vec![2, 64], block: vec![1, 1],
})?;
assert_eq!(rows.len(), 128);
println!("{:?} {:?}", ds.attr("units")?, file.root().groups()?);

Edit that file in place — no rewrite; each call is written and synced before it returns, and anything the editor cannot do safely is refused before a byte is written:

use clawhdf5::{AttrValue, FileEditor, Selection};

let mut ed = FileEditor::open("data.h5")?;          // exclusive lock, as libhdf5 takes
ed.resize("run/temps", &[1100, 64])?;              // h5py: ds.resize((1100, 64))
let sel = Selection::Hyperslab {
    start: vec![1000, 0], stride: vec![1, 1], count: vec![100, 64], block: vec![1, 1],
};
ed.write_values("run/temps", &sel, &vec![0.5f64; 100 * 64])?; // ds[1000:1100] = 0.5
ed.set_attr("run/temps", "calibrated", &AttrValue::I64(1))?;

The editor changes chunk indexes and heaps as libhdf5 does (the tests compare index shapes and heap bookkeeping with libhdf5's, and check every edited file with h5py, h5dump and h5rs check). More in docs/QUICKSTART.md: groups and links, filters, strings and variable-length data, NetCDF-4.

Quick start: Python

import numpy as np
import clawhdf5

with clawhdf5.File("data.h5", "r") as f:
    print(list(f.keys()))              # member names, like h5py
    ds = f["group/temperatures"]       # relative or absolute paths
    print(ds.shape, ds.dtype, ds.chunks)
    block = ds[100:200, ::4]           # a small selection decodes only its chunks
    row = ds[-1]                       # integers drop the axis
    picked = ds[[1, 5, 9], :]          # one increasing index list per key
    units = ds.attrs["units"]          # attributes come back as h5py returns them
    everything = np.asarray(ds)
    ids = f["table"]["id"]             # compound -> structured array; one field

with clawhdf5.File("data.h5", "r+") as f:          # edited in place, h5py semantics
    f["group/temperatures"][100:200, ::4] = 0.0
    f["series"].resize(5000, axis=0)   # chunked datasets, within maxshape
    f["series"][4000:] = np.ones(1000)
    f["group"].attrs["calibrated"] = True

with clawhdf5.File("http://data.example.org/run42.h5") as f:   # range requests, no download
    first = f["group/temperatures"][0]

Reads release the GIL, so Python threads read in parallel. The test suite (crates/clawhdf5-py/tests) compares every read and every edit with h5py, locally and over HTTP. Types, keys, writing ('w': numeric arrays) and limits: crates/clawhdf5-py/README.md.

Remote files, the browser, SWMR

Remote files (clawhdf5-remote, design: docs/design/range-reads.md). The same clawhdf5::File, over HTTP range requests or an object store, through a block cache (1 MiB blocks, LRU budget, concurrent requests deduplicated, runs coalesced into parallel requests). The file is pinned by ETag / Last-Modified and length: a file that changes on the server is an error, never a mix of old and new bytes.

let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/tall.h5")?;
let values = file.dataset("/g2/dset2.1")?.read_f64()?;

Try it with the crate's test server:

cargo run -p clawhdf5-remote --example range_server -- crates/clawhdf5/tests/fixtures 127.0.0.1:8000
cargo run -p clawhdf5-remote --example read_url -- http://127.0.0.1:8000/tall.h5 /g2/dset2.1

Plain HTTP builds no C; https (rustls + ring) and s3 / gcs / azure are opt-in features. The cloud backends are built and their URL handling tested, but have not been run against a real bucket.

In the browser (clawhdf5-wasm, demo and API in examples/wasm-viewer): open(bytes) reads a file held in memory; openUrl(url) reads a file on a web server by range requests, fetching only what each call needs, on the main thread (no worker, no synchronous XHR). In a 200 MB h5py file, listing the root, reading two small datasets, a group's attributes, the large dataset's shape and a 10-value window of it took 5 requests and 6 MiB (tank, 2026-09-27, the viewer's Node + Chromium test suite).

SWMR reading (design: docs/design/swmr.md). File::open_swmr follows a file a libhdf5 SWMR writer (h5py f.swmr_mode = True) is still appending to, as h5py's File(path, "r", swmr=True) does: Dataset::refresh() picks up the new extent, and a read that races the writer (a checksum failing mid-flush) is retried, up to 100 attempts as in libhdf5, never returned torn. Tested live against an h5py writer appending to Extensible-Array and v2-B-tree indexed datasets for 2 500 steps (20 000 in a release build), beside h5py's own SWMR reader. clawhdf5 does not write SWMR files.

let file = clawhdf5::File::open_swmr("live.h5")?;
let mut ds = file.dataset("samples")?;
while file.swmr_writer_active()? {   // add your own timeout: a writer that died keeps the flag set
    ds.refresh()?;
    let n = ds.shape()?[0];
    // read the new rows ...
    std::thread::sleep(std::time::Duration::from_millis(100));
}

h5rs tools

h5rs (crate clawhdf5-tools) is a pure-Rust counterpart of the HDF5 command-line tools:

h5rs ls -r file.h5          # like h5ls
h5rs dump file.h5           # like h5dump: DDL, or --json (hdf5-json)
h5rs stat file.h5           # like h5stat
h5rs diff a.h5 b.h5         # like h5diff
h5rs check --data file.h5   # structural and checksum validator

dump output is byte-identical to h5dump's on the interop test files, and the ls/stat/diff tests compare with h5ls, h5stat and h5diff. check walks the file's structures, verifies their checksums (superblock, object headers, v2 B-trees, fractal heaps, chunk indexes) and with --data decodes every dataset; it validates with the library's own parsers, so it accepts what they accept. With --features remote every subcommand takes a URL. Details: crates/clawhdf5-tools/README.md.

Agent memory

clawhdf5-agent stores an agent's memories — text, embeddings, sessions, a knowledge graph — in one HDF5 file (readable by h5py), with:

  • Hybrid search: HNSW (clawhdf5-ann) vector + BM25 keyword, weighted 0.4 / 0.6, optional source filter, re-ranking and confidence rejection. On the full LongMemEval longmemeval_s haystack (500 questions, real MiniLM embeddings) turn-level Hit@5 is 81.4% — retrieval recall, not the official QA-accuracy metric (tank, re-run 2026-09-27, BENCHMARKS.md).
  • Compact by default: float16 embeddings on disk (48% smaller at 100K; tank, 2026-09-23) and an int8 index copy with exact re-scoring — 1.74x the raw vectors in memory at 100K instead of 2.72x, and 1.63x the QPS at equal recall on AVX2 (int8 side measured 2026-09-19/20, machine not recorded, not re-run since; see BENCHMARKS.md).
  • Durability: a write-ahead log with a chained CRC per entry, crash-safe checkpoints, a single-writer lock and a read-only open. WAL appends are not fsynced: saves since the last checkpoint can be lost on power failure.
  • Signed checkpoints: Ed25519 over a SHA-256 Merkle tree of the records, settings, sessions and graph; HDF5Memory::verify names the edited records.
use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions};

let mut memory = HDF5Memory::create(MemoryConfig::new("agent.h5".into(), "my-agent", 384))?;
memory.save(MemoryEntry {
    chunk: "User prefers dark mode and vim keybindings.".into(),
    embedding: embed("User prefers dark mode and vim keybindings."), // your embedder
    source_channel: "chat".into(),
    timestamp: now,
    session_id: "session-001".into(),
    tags: "preference".into(),
})?;
for r in memory.search(&embed("what editor?"), "editor preferences", &SearchOptions::new(5)) {
    println!("[{:.3}] {}", r.score, r.chunk);
}

Architecture, every module, performance tables, feature flags, file schema, CLI and SQLite migration: docs/agent-memory.md.

Crate map

19 crates under crates/, plus libaec-sys (FFI for the optional SZIP filter).

Crate Role
HDF5
clawhdf5 The facade: File, FileBuilder, FileEditor, Dataset, Group, SWMR reading
clawhdf5-format The format itself (superblock, headers, B-trees, heaps, datatypes), the filter pipeline and registry, every codec but the deflate backends; no_std-capable
clawhdf5-filters Deflate backends (zlib-rs default, zlib-ng, Apple Compression)
clawhdf5-io I/O helpers: mmap, async, an HSDS client, mpi-io (not collective I/O)
clawhdf5-remote HTTP(S) and object-store files through a block cache
clawhdf5-netcdf4 NetCDF-4 dimensions, variables, CF attributes
clawhdf5-derive Derive macros for HDF5-serialisable structs
clawhdf5-tools h5rs: ls, dump, stat, diff, check
Bindings
clawhdf5-py Python (PyO3 + numpy)
clawhdf5-wasm Browser (wasm-bindgen), read-only
clawhdf5-napi Node.js (unpublished; does not work, see known-issues)
clawhdf5-android Android JNI bindings for the agent store
Agent memory
clawhdf5-agent The memory store
clawhdf5-ann HNSW index (f32 or i8 storage)
clawhdf5-accel SIMD kernels (AVX2, NEON incl. SDOT; AVX-512 behind a feature)
clawhdf5-gpu Vector distance computation on the GPU (wgpu, WGSL); HDF5 I/O is CPU-only
clawhdf5-migrate SQLite → agent store migration
clawhdf5-cli The clawhdf5 agent-memory CLI
clawhdf5-bench Benchmarks and harnesses

Building and testing

cargo build --workspace              # pure Rust: no cmake or C compiler needed
cargo test --workspace
scripts/ci-test.sh                   # what CI runs: fmt, clippy matrix, tests, interop, no_std, no-C check
conformance/run.sh                   # the conformance report (needs h5py, hdf5plugin, h5dump)

The interop suites need a Python with h5py (and netCDF4, xarray); on a PEP 668 system that has to be a virtualenv, which ci-test.sh finds as .venv or through CLAWHDF5_PYTHON. Without one they skip; set CLAWHDF5_REQUIRE_INTEROP=1 to make that a failure, as CI does:

python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 xarray

CI (.gitea/workflows/) runs ci-test.sh on x86-64, lints and tests the NEON code on aarch64, and runs the conformance corpus nightly.

Documentation

docs/QUICKSTART.md Longer quick starts: HDF5 in Rust and Python, NetCDF-4, agent memory, CLI
docs/USE_CASES.md Where clawhdf5 fits, and where it does not
docs/agent-memory.md The agent-memory store in full
CONFORMANCE.md The conformance report, generated by conformance/run.sh
BENCHMARKS.md Every measurement, with date, machine and command
docs/known-issues.md Open limits and fixed bugs, with dates
CHANGELOG.md Changes, including everything since v2.7.0
docs/README.md Index of every document

Who uses it

ClawBrainHub is the one verified consumer: its .brain files are HDF5 files it reads and writes through the facade (File, FileBuilder, AttrValue, Selection), and its CLI uses clawhdf5_agent::bm25::BM25Index (builds and passes its tests against main, checked 2026-09-25). clawhdf5 is not an OpenClaw memory plugin (docs/openclaw.md), and ZeroClaw does not use it.

License

MIT — see LICENSE.

S
Description
No description provided
Readme MIT
5.9 MiB
v2.1.0
Latest
2026-06-03 10:39:34 +00:00
Languages
Rust 96%
Python 2.8%
Shell 0.7%
TypeScript 0.4%
JavaScript 0.1%