Files
osobhandClaude Opus 5.5 e4b3a76dc7 Chunk dimensions of 2^32 or more; unfiltered 4 GiB chunks written without copies
- DataLayout::Chunked::chunk_dimensions is Vec<u64> (was Vec<u32>), and
  the chunk index readers, writers and serializers take &[u64]: layout
  messages of version 4/5 store each dimension in up to 8 bytes, and
  libhdf5 2.x writes dimensions of 2^32 or more (layout version 5). Such
  dimensions were refused on read (InvalidChunkDimensions) and write. A
  version-3 layout (4-byte dimensions) is never written for them; a chunk
  whose size overflows 64 bits is refused when opened.
- The file writer lays chunked datasets out as pieces referring to the
  chunks instead of copying them into one buffer per pass, and an
  unfiltered chunk that is a contiguous run of the dataset's data (a
  dataset stored as one chunk of its shape, row blocks) borrows it; a
  filtered one is compressed straight from it. Contiguous datasets are not
  copied either. FileWriter::finish_with streams the file to a callback;
  FileBuilder::write uses it, so the file is never assembled in memory.
  DatasetBuilder::with_u8_data_owned takes the data without a copy.
  Peak RSS writing one unfiltered 1 GiB chunk (with_u8_data_owned +
  write): 5.0 GiB before, 1.0 GiB after; with_u8_data: 6.0 -> 2.0 GiB.
- extract_chunk no longer panics on data shorter than the shape.
- Tests: huge_chunk_dims.h5 fixture (libhdf5 2.0.0 via h5py 3.16, u8
  chunks of 2^32 + 7), and opt-in end-to-end tests of chunk dims >= 2^32
  (filtered and unfiltered, read and written, h5py and h5dump 2.2.0), of
  an unfiltered 4 GiB+ chunk written by clawhdf5, and of LZ4/Zstd chunks
  of that size; example write_one_chunk for memory measurements.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-29 20:47:33 -05:00
..

clawhdf5-py

Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is clawhdf5 (import clawhdf5); it needs numpy and no libhdf5.

Install

Not on PyPI yet. Build it into a virtualenv with maturin:

pip install maturin numpy
cd crates/clawhdf5-py
maturin develop --release
python -c "import clawhdf5; print(clawhdf5.__version__)"

Reading

The read API follows h5py:

import numpy as np
import clawhdf5

with clawhdf5.File("data.h5", "r") as f:
    f.keys(), f["group"].items(), "group/data" in f
    ds = f["group/data"]           # or f["/group/data"], f["group"]["data"]
    ds.shape, ds.dtype, ds.attrs["units"]
    ds[10:20, ::2]                 # a small selection reads only its chunks
    ds[-1], ds[..., 0], ds[[1, 4, 7]]
    np.asarray(ds)
    f["table"]["id"]               # a compound field
  • Dataset.dtype is the numpy dtype h5py reports: integers and IEEE floats of every width in either byte order, bool, enums (with dtype.metadata['enum']), complex, S<n> fixed strings, object for variable-length strings (bytes values) and sequences (array values), V<n> opaque, array types, and compounds as structured dtypes. Non-IEEE floats (HDF5 2.x's bfloat16, FP8, FP6 and FP4) read as h5py 3.16 reads them: as the narrowest IEEE float that holds them (float32 for bfloat16, float16 for the 1-byte formats) in the file's byte order, with libhdf5's values (tests/test_small_floats.py); inside a compound or array type they raise TypeError, and writing them raises NotImplementedError. Other types raise TypeError.
  • Keys are h5py's: integers, slices with a positive step, ..., one increasing list of integers, compound field names. Each maps onto a hyperslab selection. None and negative steps are refused with h5py's errors; boolean masks (which h5py supports) raise NotImplementedError, for reads and writes.
  • What is read from the file: a selection whose bounding box covers at most half the dataset decodes only the chunks (or contiguous rows) the box overlaps. The library decodes the whole dataset for a larger box (including a strided slice such as ds[::100] across a chunked dataset), and for compact, virtual and unwritten datasets and chunked ones with a non-default fill value. An index list is read one group of neighbouring chunks at a time (a new group only past a chunk with no selected index), so each chunk is decoded once. ds[()], ds[...] and np.asarray(ds) use the file's chunk cache; other selections do not.
  • The bytes the library reads become the numpy array's buffer without a copy, and the read runs with the GIL released, so threads read in parallel. A bug in the library (a Rust panic) raises clawhdf5.InternalError, a RuntimeError.
  • Attributes return what h5py returns; clawhdf5.Empty stands for a null dataspace (h5py's Empty).
  • Also as in h5py: File.mode ('r', or 'r+' for a writable file), File.flush() (a no-op: edits are already synced), Dataset.chunks.

Remote files

A URL instead of a path reads the file where it is, through clawhdf5-remote: HTTP range requests through a block cache (1 MiB blocks, 64 MiB budget by default), fetching only the blocks a read needs. The whole read API works the same, and the GIL is released while waiting on the network.

f = clawhdf5.File("http://host/data.h5")          # default options
f = clawhdf5.File.open_url(
    "http://host/data.h5",
    block_size=256 * 1024, cache_size=128 << 20,  # the block cache
    headers={"Authorization": "Bearer ..."},      # sent to this origin only
    retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
    allow_full_download=False,  # a server without Range support: refuse
    require_validator=False,    # refuse servers without ETag/Last-Modified
)
f.remote_stats   # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
  • The file is pinned when opened (ETag or Last-Modified, and length): if it changes on the server, reads raise OSError instead of mixing versions. Network failures are OSError too.
  • Remote files are read-only.
  • Schemes: the default build (no C) reads http://. https:// needs maturin develop --release --features https (rustls with ring, which compiles C); s3://, gs:// and az:// need the s3, gcs and azure features (credentials from the environment; aws-lc-rs, C).

Writing

clawhdf5.File(path, "w") with create_dataset(name, data=array, chunks=..., compression="gzip"), create_group and attrs[...] = ... writes float64, float32, int64, int32, uint8, complex64 and complex128 arrays; the file is written on close(). Complex arrays are stored as h5py stores them, a compound {r, i} that every libhdf5 reads (not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the Rust API writes that on request). Both forms read back as numpy complex64/complex128.

libver= sets the library version bounds as in h5py: 'v108', 'v110', 'v112', 'v114', 'v200' or 'latest' (the low bound; the high bound is then 'latest'), or a (low, high) tuple. With 'v108' the file is written in the HDF5 1.8 format (version-2 superblock, version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the HDF5 1.10 format. 'earliest' as the low bound writes the 1.8 format too, with a UserWarning: clawhdf5 cannot write the pre-1.8 format. The argument is ignored when reading and refused for 'r+'/'a'.

with clawhdf5.File("old.h5", "w", libver="v108") as f:   # HDF5 1.8 reads it
    f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip")

Editing a file in place

clawhdf5.File(path, "r+") (or "a" on an existing file) edits the file where it is, through clawhdf5's FileEditor; the file is locked until close(), and every edit is written and synced before the statement returns.

with clawhdf5.File("data.h5", "r+") as f:
    ds = f["grid"]
    ds[10:20, ::2] = 0                  # h5py keys and broadcasting
    ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5]  # one index list: exact shape
    f["series"].resize((5000, 3))       # or .resize(5000, axis=0)
    f["series"].attrs["units"] = "K"
    f.attrs.create("version", 2, dtype="u1")
  • Values: a numpy array is converted to the dataset's dtype as libhdf5 converts it (integers saturate at the target's limits; floats are truncated toward zero and clipped); anything else goes through numpy.asarray(value, dtype=ds.dtype), as in h5py. Writing NaN into an integer dataset raises ValueError (libhdf5 would store an arbitrary value). A few libhdf5 edge cases differ on purpose; see docs/known-issues.md.
  • Shapes: ds.resize grows or shrinks chunked datasets within their maxshape, as h5py; datasets and attrs objects taken before an edit see its result.
  • Attributes: numeric, bool, complex, bytes and str data of any shape. str is stored as a fixed-length UTF-8 string (h5py stores a variable-length one), so h5py reads it back as bytes.
  • Not supported (NotImplementedError, nothing written): creating or deleting datasets and groups, deleting attributes (creating and replacing them works, compact or dense), writing compound fields by name, variable-length data, HDF5 array types, and whatever FileEditor refuses (listed in docs/known-issues.md).

Tests

pip install pytest h5py
pytest crates/clawhdf5-py/tests

tests/test_read_vs_h5py.py compares every read with h5py on a file h5py writes, opened locally and over HTTP (an in-process range server, tests/conftest.py); tests/test_remote.py checks remote reads (requests, failures, the GIL); tests/test_edit.py applies every edit through h5py and clawhdf5 to copies of a file and compares them through h5py (and h5dump, and h5rs check when CLAWHDF5_H5RS names it). scripts/ci-test.sh builds the wheel and runs these in CI.

License

MIT