Files
clawhdf5/crates/clawhdf5-format
osobhandClaude Opus 5.5 5a20cf04e8 Write chunks of 4 GiB or more; FileEditor refuses to rewrite them
The writer gives a chunk of more than u32::MAX bytes layout message
version 5 and, filtered, index elements whose stored size takes the file's
size of lengths, as libhdf5 2.x does (the Fixed and Extensible Array
structures match libhdf5's byte for byte). Chunk dimensions of 2^32 or more
and filters that cannot take such a chunk (LZF, bitshuffle, bzip2, Blosc,
pcodec) are refused instead of truncated. Chunks are extracted row by row
and one at a time; deflate no longer cuts input at 4 GiB - 1 bytes, nor
holds the worst-case bound of a large chunk; an LZ4 chunk of 4 GiB or more
is read as the registered framing.

FileEditor refuses writing values into, or pruning/allocating, chunks of
4 GiB or more before anything is written; growing the extent and
attributes still work.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 23:44:48 -05:00
..

clawhdf5-format

The HDF5 file format in pure Rust: parsers and writers for every on-disk structure, the filter pipeline and its codecs, and the shared type definitions the other crates use. Most users want the clawhdf5 facade, which wraps this crate in an h5py-like API; use this one directly for low-level access or in no_std code.

Not on crates.io yet; depend on it from git:

[dependencies]
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }

What is in it

  • Parsing: superblock v0–v3 (superblock, with the superblock extension and metadata cache images, superblock_ext), object headers v1 and v2 (object_header), every header message the readers use (datatype, dataspace, data_layout v1–v4 including virtual datasets, fill_value, attribute, link_message, shared_message, ...), groups old and new (group_v1 symbol tables with local heaps, group_v2 with fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single chunk, implicit, fixed array, extensible array, v2 B-tree).
  • Reading data: data_read (contiguous, compact, chunked), partial_read and selection (hyperslabs and points), vl_data (variable-length strings and sequences through the global heap), chunk_cache.
  • Storage: the storage::Storage trait (read_at, read_ranges, len, hint) that every read path goes through, so a file can be read from memory, a file handle or a remote backend (clawhdf5-remote).
  • Writing: file_writer::FileWriter and the builders in type_builders (datasets, groups, attributes, compound and enum types, links, virtual datasets, creation-order tracking); chunk indexes and dense-storage B-trees of any size (chunked_write, btree_v2_write, ea_writer, and the version-1 chunk B-tree of btree_v1_write). Output is read by h5py and h5dump; FileWriter::libver_bounds (libver) picks the format: HDF5 1.10 by default, or one HDF5 1.8 reads.
  • Filters: filter_pipeline and filter_registry (look up by ID; other IDs can be registered at run time with register_filter). Built in: deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4, Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle, bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only).
  • Shared pieces: float16 (the one IEEE half-precision conversion the workspace uses), provenance (SHA-256 dataset hashes), checksum (Jenkins lookup3 for v2+ structures).

Example

use clawhdf5_format::file_writer::{AttrValue, FileWriter};
use clawhdf5_format::{group_v2, object_header, signature, superblock};

// Write a file to memory
let mut fw = FileWriter::new();
fw.create_dataset("data")
    .with_f64_data(&[1.0, 2.0, 3.0])
    .with_shape(&[3])
    .set_attr("unit", AttrValue::String("m/s".into()));
let bytes = fw.finish().unwrap();

// Parse it back: superblock -> path -> object header
let (_user_block, file) = signature::split_user_block(&bytes).unwrap();
let sb = superblock::Superblock::parse(file, 0).unwrap();
let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap();
let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size)
    .unwrap();
assert!(!hdr.messages.is_empty());

Features

Feature Default What Builds C
std yes standard library; without it the crate is no_std + alloc (CI builds it for thumbv7em-none-eabihf) no
checksum yes verify Jenkins lookup3 checksums no
deflate yes deflate through flate2 no
zlib-rs yes flate2's pure-Rust zlib-rs backend, with runtime_detection (without it zlib-rs loses SIMD and inflates 3.5x slower) no
system-zlib-decompress yes macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere no (links the system libz on macOS)
provenance yes SHA-256 provenance hashes no
lzf yes LZF (32000) no
parallel no rayon-parallel chunk decoding no
fast-checksum no hardware CRC32 through crc32fast no
lz4 no LZ4 (32004) no
pcodec no pcodec no
bitshuffle, bzip2, blosc no 32008, 307, 32001, read and write no
blosc2, zfp no 32026, 32013, read only no
plugin-filters no all six plugin filters above no
lookup-stats no counters for name-lookup benchmarks no
zstd no Zstandard (32015) yes (libzstd)
szip no SZIP (4) decoding links the system libaec (libaec-dev)
fast-deflate no zlib-ng yes (cmake)
system-zlib no the system zlib yes (libz-sys)
blake3_hash no provenance::blake3_hash yes (cc)

Robustness

Every parser is meant to return an error, never panic, on hostile input: nine cargo-fuzz targets live in fuzz/, the conformance sweep includes the HDF Group's CVE corpus (CONFORMANCE.md), and header checks follow libhdf5's. Open gaps are in docs/known-issues.md.

License

MIT