Files
clawhdf5/crates/clawhdf5-format/README.md
T
osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00

5.0 KiB
Raw Blame History

clawhdf5-format

The HDF5 file format in pure Rust: parsers and writers for every on-disk structure, the filter pipeline and its codecs, and the shared type definitions the other crates use. Most users want the clawhdf5 facade, which wraps this crate in an h5py-like API; use this one directly for low-level access or in no_std code.

Not on crates.io yet; depend on it from git:

[dependencies]
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }

What is in it

  • Parsing: superblock v0–v3 (superblock, with the superblock extension and metadata cache images, superblock_ext), object headers v1 and v2 (object_header), every header message the readers use (datatype, dataspace, data_layout v1–v4 including virtual datasets, fill_value, attribute, link_message, shared_message, ...), groups old and new (group_v1 symbol tables with local heaps, group_v2 with fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single chunk, implicit, fixed array, extensible array, v2 B-tree).
  • Reading data: data_read (contiguous, compact, chunked), partial_read and selection (hyperslabs and points), vl_data (variable-length strings and sequences through the global heap), chunk_cache.
  • Storage: the storage::Storage trait (read_at, read_ranges, len, hint) that every read path goes through, so a file can be read from memory, a file handle or a remote backend (clawhdf5-remote).
  • Writing: file_writer::FileWriter and the builders in type_builders (datasets, groups, attributes, compound and enum types, links, virtual datasets, creation-order tracking); chunk indexes and dense-storage B-trees of any size (chunked_write, btree_v2_write, ea_writer). Output is read by h5py and h5dump.
  • Filters: filter_pipeline and filter_registry (look up by ID; other IDs can be registered at run time with register_filter). Built in: deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4, Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle, bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only).
  • Shared pieces: float16 (the one IEEE half-precision conversion the workspace uses), provenance (SHA-256 dataset hashes), checksum (Jenkins lookup3 for v2+ structures).

Example

use clawhdf5_format::file_writer::{AttrValue, FileWriter};
use clawhdf5_format::{group_v2, object_header, signature, superblock};

// Write a file to memory
let mut fw = FileWriter::new();
fw.create_dataset("data")
    .with_f64_data(&[1.0, 2.0, 3.0])
    .with_shape(&[3])
    .set_attr("unit", AttrValue::String("m/s".into()));
let bytes = fw.finish().unwrap();

// Parse it back: superblock -> path -> object header
let (_user_block, file) = signature::split_user_block(&bytes).unwrap();
let sb = superblock::Superblock::parse(file, 0).unwrap();
let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap();
let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size)
    .unwrap();
assert!(!hdr.messages.is_empty());

Features

Feature Default What Builds C
std yes standard library; without it the crate is no_std + alloc (CI builds it for thumbv7em-none-eabihf) no
checksum yes verify Jenkins lookup3 checksums no
deflate yes deflate through flate2 no
zlib-rs yes flate2's pure-Rust zlib-rs backend, with runtime_detection (without it zlib-rs loses SIMD and inflates 3.5x slower) no
system-zlib-decompress yes macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere no (links the system libz on macOS)
provenance yes SHA-256 provenance hashes no
lzf yes LZF (32000) no
parallel no rayon-parallel chunk decoding no
fast-checksum no hardware CRC32 through crc32fast no
lz4 no LZ4 (32004) no
pcodec no pcodec no
bitshuffle, bzip2, blosc no 32008, 307, 32001, read and write no
blosc2, zfp no 32026, 32013, read only no
plugin-filters no all six plugin filters above no
lookup-stats no counters for name-lookup benchmarks no
zstd no Zstandard (32015) yes (libzstd)
szip no SZIP (4) decoding links the system libaec (libaec-dev)
fast-deflate no zlib-ng yes (cmake)
system-zlib no the system zlib yes (libz-sys)
blake3_hash no provenance::blake3_hash yes (cc)

Robustness

Every parser is meant to return an error, never panic, on hostile input: nine cargo-fuzz targets live in fuzz/, the conformance sweep includes the HDF Group's CVE corpus (CONFORMANCE.md), and header checks follow libhdf5's. Open gaps are in docs/known-issues.md.

License

MIT