Both indexes place each chunk at a linear index computed from the dataset's maximum dimensions (libhdf5's max_down_chunks), and the Extensible Array first swizzles its unlimited dimension to the slowest position. We linearised by the current dimensions, so any dataset whose shape was smaller than its maxshape, or whose unlimited dimension was not the first, read back scrambled without an error: h5py libver="latest" files with maxshape (10, None) or (20, 10), and the libhdf5 test files h5fc_ext*.h5 and test_ld.h5. The linearisation now lives in chunk_grid (shared with the writers), and slots beyond the current extent are ignored as the library does. read_fixed_array_chunks / read_extensible_array_chunks take the dataspace's max dimensions. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5-format
Pure-Rust HDF5 binary format parsing and writing — no C dependencies.
Features
- Zero-copy superblock, object header, and B-tree parsing
- Chunked dataset read/write with filter pipelines
no_stdsupport (disablestdfeature)- Optional parallel reads via Rayon
- SHA-256 provenance tracking
Usage
use clawhdf5_format::Superblock;
let data = std::fs::read("data.h5").unwrap();
let sb = Superblock::from_bytes(&data).unwrap();
println!("HDF5 version {}.{}", sb.version_major(), sb.version_minor());
License
MIT