A version-4 layout stores every chunk dimension in the fewest bytes that hold the largest one (H5D__chunk_set_sizes: (log2(dim) + 8) / 8), so a chunk dimension of 65 536 to 16 777 215 takes 3 bytes. Only widths 1, 2, 4 and 8 were decoded; an h5py file with chunks=(70000,) and libver='latest' failed with UnexpectedEof. Widths 1-8 are decoded now; 0 and more than 8 are refused with libhdf5's "encoded chunk dimension size is too large", and a dimension past u32 is refused, not truncated. The review asked for libhdf5's check that the stored width matches the one computed from the dimensions. HDF5 2.0.0 (h5py 3.16) refuses any mismatch, but HDFGroup/hdf5@e124c36 ("Allow reading of files with chunk dimensions encoded using more bytes than necessary", 2026-06-05) relaxed it to refusing only a width too small for the dimensions, which cannot happen once the dimensions have been decoded from that width. Follow current libhdf5: a wider-than-needed encoding is read. clawhdf5's own writer produces such layouts (the next commit fixes that). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5-format
Pure-Rust HDF5 binary format parsing and writing — no C dependencies.
Features
- Zero-copy superblock, object header, and B-tree parsing
- Chunked dataset read/write with filter pipelines
no_stdsupport (disablestdfeature)- Optional parallel reads via Rayon
- SHA-256 provenance tracking
Usage
use clawhdf5_format::Superblock;
let data = std::fs::read("data.h5").unwrap();
let sb = Superblock::from_bytes(&data).unwrap();
println!("HDF5 version {}.{}", sb.version_major(), sb.version_minor());
License
MIT