h5py's built-in compression="lzf" failed with UnsupportedFilter(32000). The new `lzf` feature (no dependencies, on by default in clawhdf5-format and the facade) decodes the raw liblzf stream h5py's filter stores, bounded by the chunk size, and encodes it: DatasetBuilder::with_lzf() (or with_plugin_filter(PluginFilter::Lzf)) writes the filter with h5py's cd_values (filter version 4, liblzf 0x0105, chunk size in bytes), flagged optional as h5py does. ChunkOptions gains a `plugin` field for the plugin filters; build_pipeline_for_chunk passes the chunk size to filters that record it. tests/plugin_filters_interop.rs: h5py writes LZF (alone, with shuffle, with shuffle+fletcher32) over 12 dtype/shape/chunk/data cases with partial edge chunks and incompressible data, and every dataset reads byte for byte equal to its unfiltered twin; our LZF output (1-D and 2-D, edge chunks, with and without shuffle) reads back in h5py. Both fail with the decoder removed. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5-format
Pure-Rust HDF5 binary format parsing and writing — no C dependencies.
Features
- Zero-copy superblock, object header, and B-tree parsing
- Chunked dataset read/write with filter pipelines
no_stdsupport (disablestdfeature)- Optional parallel reads via Rayon
- SHA-256 provenance tracking
Usage
use clawhdf5_format::Superblock;
let data = std::fs::read("data.h5").unwrap();
let sb = Superblock::from_bytes(&data).unwrap();
println!("HDF5 version {}.{}", sb.version_major(), sb.version_minor());
License
MIT