read_f32/read_f64/read_i32/read_i64/read_u64 of a chunked dataset that stores exactly that type in native byte order now decode every chunk straight into the Vec<T> they return (data_read::read_chunked_native), on File (through its chunk cache), MmapFile and LazyFile. Before, the chunks went into a byte buffer that read_as_* then copied into a second, typed one: two dataset-sized allocations and a full extra copy per read. The output is zeroed pages from the allocator, backed by transparent huge pages when large, like the byte reader's. Other types and byte orders, and datasets with no storage or external data, keep converting through the byte readers; unallocated chunks read as the fill value as before. tests/chunked_read_paths_interop.rs checks every chunked read path (File twice, so cached; from_bytes; MmapFile; LazyFile; small, strided and point selections; with and without the parallel feature) against h5py for 1-8 byte integers and 2-8 byte floats in both byte orders, through deflate, shuffle, Fletcher32, LZF, SZIP and Blosc, with partial edge chunks, sparse datasets with default and non-default fill values, and datasets larger than the chunk cache. A filter this build lacks must be an error (or, when an optional filter declined every chunk, the right data). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5-format
Pure-Rust HDF5 binary format parsing and writing — no C dependencies.
Features
- Zero-copy superblock, object header, and B-tree parsing
- Chunked dataset read/write with filter pipelines
no_stdsupport (disablestdfeature)- Optional parallel reads via Rayon
- SHA-256 provenance tracking
Usage
use clawhdf5_format::Superblock;
let data = std::fs::read("data.h5").unwrap();
let sb = Superblock::from_bytes(&data).unwrap();
println!("HDF5 version {}.{}", sb.version_major(), sb.version_minor());
License
MIT