FileEditor::resize (on main since PR #18, a4c2ace) scrambled the values of
a chunked dataset whose dataspace records no maximum dimensions when it
shrank it: clawhdf5's writer stores such a dataspace for every chunked
dataset created without a maxshape, the maximum is then the current
dimensions, and the Fixed Array index linearises chunks by the maximum, so
patching only the current dimensions moved every chunk after the first
row. h5py, h5dump and our reader all read the wrong values; the dataset
could not grow back either.
libhdf5 never writes such a dataspace (H5S_set_extent_simple records the
maximum, equal to the dimensions when none is given); reading one,
H5S_extent_get_dims reports the current dimensions as the maximum and
H5S_set_extent checks against none, so its own H5Dset_extent scrambles
such a file the same way. The editor now records the maximum libhdf5 would
have written (the dimensions the index was built with) before changing the
current ones, moving the grown dataspace message in the header when it
must. The writer records the maximum of every chunked dataset too, so
h5py can resize what clawhdf5 writes (the pinned file hashes of three
no-maxshape cases in plugin_filters_interop change by 8 bytes a dimension).
Tests: edit_resize_interop.rs (a 2.7.0-written fixture, new FileBuilder
files and h5py files through shrinks, zero extents and growth, against a
model with our reader and h5py; h5py resizing a FileBuilder file), and in
test_edit.py resizes checked against a numpy model, independently of h5py.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5-py
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
clawhdf5 (import clawhdf5); it needs numpy and no libhdf5.
Install
Not on PyPI yet. Build it into a virtualenv with maturin:
pip install maturin numpy
cd crates/clawhdf5-py
maturin develop --release
python -c "import clawhdf5; print(clawhdf5.__version__)"
Reading
The read API follows h5py:
import numpy as np
import clawhdf5
with clawhdf5.File("data.h5", "r") as f:
f.keys(), f["group"].items(), "group/data" in f
ds = f["group/data"] # or f["/group/data"], f["group"]["data"]
ds.shape, ds.dtype, ds.attrs["units"]
ds[10:20, ::2] # a small selection reads only its chunks
ds[-1], ds[..., 0], ds[[1, 4, 7]]
np.asarray(ds)
f["table"]["id"] # a compound field
Dataset.dtypeis the numpy dtype h5py reports: integers and IEEE floats of every width in either byte order,bool, enums (withdtype.metadata['enum']), complex,S<n>fixed strings,objectfor variable-length strings (bytesvalues) and sequences (array values),V<n>opaque, array types, and compounds as structured dtypes. Other types raiseTypeError.- Keys are h5py's: integers, slices with a positive step,
..., one increasing list of integers, compound field names. Each maps onto a hyperslab selection.None, negative steps and boolean masks are refused with h5py's errors. - What is read from the file: a selection whose bounding box covers at
most half the dataset decodes only the chunks (or contiguous rows) the box
overlaps. The library decodes the whole dataset for a larger box
(including a strided slice such as
ds[::100]across a chunked dataset), and for compact, virtual and unwritten datasets and chunked ones with a non-default fill value. An index list is read one group of neighbouring chunks at a time (a new group only past a chunk with no selected index), so each chunk is decoded once.ds[()],ds[...]andnp.asarray(ds)use the file's chunk cache; other selections do not. - The bytes the library reads become the numpy array's buffer without a
copy, and the read runs with the GIL released, so threads read in
parallel. A bug in the library (a Rust panic) raises
clawhdf5.InternalError, aRuntimeError. - Attributes return what h5py returns;
clawhdf5.Emptystands for a null dataspace (h5py'sEmpty).
Remote files
A URL instead of a path reads the file where it is, through
clawhdf5-remote: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise
OSErrorinstead of mixing versions. Network failures areOSErrortoo. - Remote files are read-only.
- Schemes: the default build (no C) reads
http://.https://needsmaturin develop --release --features https(rustls with ring, which compiles C);s3://,gs://andaz://need thes3,gcsandazurefeatures (credentials from the environment; aws-lc-rs, C).
Writing
clawhdf5.File(path, "w") with create_dataset(name, data=array, chunks=..., compression="gzip"), create_group and attrs[...] = ...
writes float64, float32, int64, int32 and uint8 arrays; the file is
written on close().
Editing a file in place
clawhdf5.File(path, "r+") (or "a" on an existing file) edits the file
where it is, through clawhdf5's FileEditor; the file is locked until
close(), and every edit is written and synced before the statement
returns.
with clawhdf5.File("data.h5", "r+") as f:
ds = f["grid"]
ds[10:20, ::2] = 0 # h5py keys and broadcasting
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
f["series"].attrs["units"] = "K"
f.attrs.create("version", 2, dtype="u1")
- Values: a numpy array is converted to the dataset's dtype as libhdf5
converts it (integers saturate at the target's limits; floats are
truncated toward zero and clipped); anything else goes through
numpy.asarray(value, dtype=ds.dtype), as in h5py. Writing NaN into an integer dataset raisesValueError(libhdf5 would store an arbitrary value). A few libhdf5 edge cases differ on purpose; seedocs/known-issues.md. - Shapes:
ds.resizegrows or shrinks chunked datasets within theirmaxshape, as h5py; datasets andattrsobjects taken before an edit see its result. - Attributes: numeric, bool, complex, bytes and
strdata of any shape.stris stored as a fixed-length UTF-8 string (h5py stores a variable-length one), so h5py reads it back asbytes. - Not supported (
NotImplementedError, nothing written): creating or deleting datasets, groups and attributes, writing compound fields by name, variable-length data, HDF5 array types, and whateverFileEditorrefuses (listed indocs/known-issues.md).
Tests
pip install pytest h5py
pytest crates/clawhdf5-py/tests
tests/test_read_vs_h5py.py compares every read with h5py on a file h5py
writes, opened locally and over HTTP (an in-process range server,
tests/conftest.py); tests/test_remote.py checks remote reads (requests,
failures, the GIL); tests/test_edit.py applies every edit through h5py and
clawhdf5 to copies of a file and compares them through h5py (and h5dump,
and h5rs check when CLAWHDF5_H5RS names it). scripts/ci-test.sh builds
the wheel and runs these in CI.
License
MIT