Files
osobhandClaude Opus 5.5 e4b3a76dc7 Chunk dimensions of 2^32 or more; unfiltered 4 GiB chunks written without copies
- DataLayout::Chunked::chunk_dimensions is Vec<u64> (was Vec<u32>), and
  the chunk index readers, writers and serializers take &[u64]: layout
  messages of version 4/5 store each dimension in up to 8 bytes, and
  libhdf5 2.x writes dimensions of 2^32 or more (layout version 5). Such
  dimensions were refused on read (InvalidChunkDimensions) and write. A
  version-3 layout (4-byte dimensions) is never written for them; a chunk
  whose size overflows 64 bits is refused when opened.
- The file writer lays chunked datasets out as pieces referring to the
  chunks instead of copying them into one buffer per pass, and an
  unfiltered chunk that is a contiguous run of the dataset's data (a
  dataset stored as one chunk of its shape, row blocks) borrows it; a
  filtered one is compressed straight from it. Contiguous datasets are not
  copied either. FileWriter::finish_with streams the file to a callback;
  FileBuilder::write uses it, so the file is never assembled in memory.
  DatasetBuilder::with_u8_data_owned takes the data without a copy.
  Peak RSS writing one unfiltered 1 GiB chunk (with_u8_data_owned +
  write): 5.0 GiB before, 1.0 GiB after; with_u8_data: 6.0 -> 2.0 GiB.
- extract_chunk no longer panics on data shorter than the shape.
- Tests: huge_chunk_dims.h5 fixture (libhdf5 2.0.0 via h5py 3.16, u8
  chunks of 2^32 + 7), and opt-in end-to-end tests of chunk dims >= 2^32
  (filtered and unfiltered, read and written, h5py and h5dump 2.2.0), of
  an unfiltered 4 GiB+ chunk written by clawhdf5, and of LZ4/Zstd chunks
  of that size; example write_one_chunk for memory measurements.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:33 -05:00
..

clawhdf5-format

The HDF5 file format in pure Rust: parsers and writers for every on-disk structure, the filter pipeline and its codecs, and the shared type definitions the other crates use. Most users want the clawhdf5 facade, which wraps this crate in an h5py-like API; use this one directly for low-level access or in no_std code.

Not on crates.io yet; depend on it from git:

[dependencies]
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }

What is in it

  • Parsing: superblock v0–v3 (superblock, with the superblock extension and metadata cache images, superblock_ext), object headers v1 and v2 (object_header), every header message the readers use (datatype, dataspace, data_layout v1–v4 including virtual datasets, fill_value, attribute, link_message, shared_message, ...), groups old and new (group_v1 symbol tables with local heaps, group_v2 with fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single chunk, implicit, fixed array, extensible array, v2 B-tree).
  • Reading data: data_read (contiguous, compact, chunked), partial_read and selection (hyperslabs and points), vl_data (variable-length strings and sequences through the global heap), chunk_cache.
  • Storage: the storage::Storage trait (read_at, read_ranges, len, hint) that every read path goes through, so a file can be read from memory, a file handle or a remote backend (clawhdf5-remote).
  • Writing: file_writer::FileWriter and the builders in type_builders (datasets, groups, attributes, compound and enum types, links, virtual datasets, creation-order tracking); chunk indexes and dense-storage B-trees of any size (chunked_write, btree_v2_write, ea_writer, and the version-1 chunk B-tree of btree_v1_write). Output is read by h5py and h5dump; FileWriter::libver_bounds (libver) picks the format: HDF5 1.10 by default, or one HDF5 1.8 reads.
  • Filters: filter_pipeline and filter_registry (look up by ID; other IDs can be registered at run time with register_filter). Built in: deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4, Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle, bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only).
  • Shared pieces: float16 (the one IEEE half-precision conversion the workspace uses), provenance (SHA-256 dataset hashes), checksum (Jenkins lookup3 for v2+ structures).

Example

use clawhdf5_format::file_writer::{AttrValue, FileWriter};
use clawhdf5_format::{group_v2, object_header, signature, superblock};

// Write a file to memory
let mut fw = FileWriter::new();
fw.create_dataset("data")
    .with_f64_data(&[1.0, 2.0, 3.0])
    .with_shape(&[3])
    .set_attr("unit", AttrValue::String("m/s".into()));
let bytes = fw.finish().unwrap();

// Parse it back: superblock -> path -> object header
let (_user_block, file) = signature::split_user_block(&bytes).unwrap();
let sb = superblock::Superblock::parse(file, 0).unwrap();
let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap();
let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size)
    .unwrap();
assert!(!hdr.messages.is_empty());

Features

Feature Default What Builds C
std yes standard library; without it the crate is no_std + alloc (CI builds it for thumbv7em-none-eabihf) no
checksum yes verify Jenkins lookup3 checksums no
deflate yes deflate through flate2 no
zlib-rs yes flate2's pure-Rust zlib-rs backend, with runtime_detection (without it zlib-rs loses SIMD and inflates 3.5x slower) no
system-zlib-decompress yes macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere no (links the system libz on macOS)
provenance yes SHA-256 provenance hashes no
lzf yes LZF (32000) no
parallel no rayon-parallel chunk decoding no
fast-checksum no hardware CRC32 through crc32fast no
lz4 no LZ4 (32004) no
pcodec no pcodec no
bitshuffle, bzip2, blosc no 32008, 307, 32001, read and write no
blosc2, zfp no 32026, 32013, read only no
plugin-filters no all six plugin filters above no
lookup-stats no counters for name-lookup benchmarks no
zstd no Zstandard (32015) yes (libzstd)
szip no SZIP (4) decoding links the system libaec (libaec-dev)
fast-deflate no zlib-ng yes (cmake)
system-zlib no the system zlib yes (libz-sys)
blake3_hash no provenance::blake3_hash yes (cc)

Robustness

Every parser is meant to return an error, never panic, on hostile input: nine cargo-fuzz targets live in fuzz/, the conformance sweep includes the HDF Group's CVE corpus (CONFORMANCE.md), and header checks follow libhdf5's. Open gaps are in docs/known-issues.md.

License

MIT