Numbers, API names, feature defaults and PR references checked against CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history. - int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives them (2026-09-19/20, machine not recorded, not re-run) instead of none; the Pi 5 1.18x carries 2026-09-21. - BENCHMARKS headline: the libhdf5 chunked-write figure is the newest measurement (35x, 2026-09-23), not 45.3x (2026-08-03). - Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug) in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to the bad_nbit_parms_walk.h5 flip. - README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/ time datasets too; zlib-rs byte-identity scoped to what was measured; macOS default links the system libz for inflate. - Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4 unlimited-dimension size warning. - agent-memory.md: string-dataset compression threshold, agents-md prints Markdown, float16 file sizes linked to their study. - known-issues.md: contiguous selection reads, 1.21x vs h5py threads. - docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify, dated figures, fast-math is not BLAS. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
5.0 KiB
5.0 KiB
clawhdf5-format
The HDF5 file format in pure Rust: parsers and writers for every on-disk
structure, the filter pipeline and its codecs, and the shared type
definitions the other crates use. Most users want the
clawhdf5 facade, which wraps this crate in an
h5py-like API; use this one directly for low-level access or in no_std
code.
Not on crates.io yet; depend on it from git:
[dependencies]
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
What is in it
- Parsing: superblock v0–v3 (
superblock, with the superblock extension and metadata cache images,superblock_ext), object headers v1 and v2 (object_header), every header message the readers use (datatype,dataspace,data_layoutv1–v4 including virtual datasets,fill_value,attribute,link_message,shared_message, ...), groups old and new (group_v1symbol tables with local heaps,group_v2with fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single chunk, implicit, fixed array, extensible array, v2 B-tree). - Reading data:
data_read(contiguous, compact, chunked),partial_readandselection(hyperslabs and points),vl_data(variable-length strings and sequences through the global heap),chunk_cache. - Storage: the
storage::Storagetrait (read_at,read_ranges,len,hint) that every read path goes through, so a file can be read from memory, a file handle or a remote backend (clawhdf5-remote). - Writing:
file_writer::FileWriterand the builders intype_builders(datasets, groups, attributes, compound and enum types, links, virtual datasets, creation-order tracking); chunk indexes and dense-storage B-trees of any size (chunked_write,btree_v2_write,ea_writer). Output is read by h5py and h5dump. - Filters:
filter_pipelineandfilter_registry(look up by ID; other IDs can be registered at run time withregister_filter). Built in: deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4, Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle, bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only). - Shared pieces:
float16(the one IEEE half-precision conversion the workspace uses),provenance(SHA-256 dataset hashes),checksum(Jenkins lookup3 for v2+ structures).
Example
use clawhdf5_format::file_writer::{AttrValue, FileWriter};
use clawhdf5_format::{group_v2, object_header, signature, superblock};
// Write a file to memory
let mut fw = FileWriter::new();
fw.create_dataset("data")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_shape(&[3])
.set_attr("unit", AttrValue::String("m/s".into()));
let bytes = fw.finish().unwrap();
// Parse it back: superblock -> path -> object header
let (_user_block, file) = signature::split_user_block(&bytes).unwrap();
let sb = superblock::Superblock::parse(file, 0).unwrap();
let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap();
let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size)
.unwrap();
assert!(!hdr.messages.is_empty());
Features
| Feature | Default | What | Builds C |
|---|---|---|---|
std |
yes | standard library; without it the crate is no_std + alloc (CI builds it for thumbv7em-none-eabihf) |
no |
checksum |
yes | verify Jenkins lookup3 checksums | no |
deflate |
yes | deflate through flate2 | no |
zlib-rs |
yes | flate2's pure-Rust zlib-rs backend, with runtime_detection (without it zlib-rs loses SIMD and inflates 3.5x slower) |
no |
system-zlib-decompress |
yes | macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere | no (links the system libz on macOS) |
provenance |
yes | SHA-256 provenance hashes | no |
lzf |
yes | LZF (32000) | no |
parallel |
no | rayon-parallel chunk decoding | no |
fast-checksum |
no | hardware CRC32 through crc32fast |
no |
lz4 |
no | LZ4 (32004) | no |
pcodec |
no | pcodec | no |
bitshuffle, bzip2, blosc |
no | 32008, 307, 32001, read and write | no |
blosc2, zfp |
no | 32026, 32013, read only | no |
plugin-filters |
no | all six plugin filters above | no |
lookup-stats |
no | counters for name-lookup benchmarks | no |
zstd |
no | Zstandard (32015) | yes (libzstd) |
szip |
no | SZIP (4) decoding | links the system libaec (libaec-dev) |
fast-deflate |
no | zlib-ng | yes (cmake) |
system-zlib |
no | the system zlib | yes (libz-sys) |
blake3_hash |
no | provenance::blake3_hash |
yes (cc) |
Robustness
Every parser is meant to return an error, never panic, on hostile input:
nine cargo-fuzz targets live in fuzz/, the conformance
sweep includes the HDF Group's CVE corpus
(CONFORMANCE.md), and header checks follow
libhdf5's. Open gaps are in docs/known-issues.md.
License
MIT