- docs/README.md links the improvement logs and June plans where the refresh archived them (docs/archive/). - clawhdf5-py README: 'r+' creates and replaces attributes (compact or dense); only deleting them is unsupported. - scripts/run-benchmarks.sh benchmarked the pre-rename rustyhdf5-format and overwrote BENCHMARKS.md; nothing referenced it. Removed. - Cargo.toml descriptions no longer name rustyhdf5/edgehdf5; clawhdf5-gpu says it is not HDF5 I/O. - benchmarks/cross_platform.sh pointed at a ROADMAP section that no longer exists. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
6.5 KiB
clawhdf5-py
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
clawhdf5 (import clawhdf5); it needs numpy and no libhdf5.
Install
Not on PyPI yet. Build it into a virtualenv with maturin:
pip install maturin numpy
cd crates/clawhdf5-py
maturin develop --release
python -c "import clawhdf5; print(clawhdf5.__version__)"
Reading
The read API follows h5py:
import numpy as np
import clawhdf5
with clawhdf5.File("data.h5", "r") as f:
f.keys(), f["group"].items(), "group/data" in f
ds = f["group/data"] # or f["/group/data"], f["group"]["data"]
ds.shape, ds.dtype, ds.attrs["units"]
ds[10:20, ::2] # a small selection reads only its chunks
ds[-1], ds[..., 0], ds[[1, 4, 7]]
np.asarray(ds)
f["table"]["id"] # a compound field
Dataset.dtypeis the numpy dtype h5py reports: integers and IEEE floats of every width in either byte order,bool, enums (withdtype.metadata['enum']), complex,S<n>fixed strings,objectfor variable-length strings (bytesvalues) and sequences (array values),V<n>opaque, array types, and compounds as structured dtypes. Other types raiseTypeError.- Keys are h5py's: integers, slices with a positive step,
..., one increasing list of integers, compound field names. Each maps onto a hyperslab selection.Noneand negative steps are refused with h5py's errors; boolean masks (which h5py supports) raiseNotImplementedError, for reads and writes. - What is read from the file: a selection whose bounding box covers at
most half the dataset decodes only the chunks (or contiguous rows) the box
overlaps. The library decodes the whole dataset for a larger box
(including a strided slice such as
ds[::100]across a chunked dataset), and for compact, virtual and unwritten datasets and chunked ones with a non-default fill value. An index list is read one group of neighbouring chunks at a time (a new group only past a chunk with no selected index), so each chunk is decoded once.ds[()],ds[...]andnp.asarray(ds)use the file's chunk cache; other selections do not. - The bytes the library reads become the numpy array's buffer without a
copy, and the read runs with the GIL released, so threads read in
parallel. A bug in the library (a Rust panic) raises
clawhdf5.InternalError, aRuntimeError. - Attributes return what h5py returns;
clawhdf5.Emptystands for a null dataspace (h5py'sEmpty). - Also as in h5py:
File.mode('r', or'r+'for a writable file),File.flush()(a no-op: edits are already synced),Dataset.chunks.
Remote files
A URL instead of a path reads the file where it is, through
clawhdf5-remote: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise
OSErrorinstead of mixing versions. Network failures areOSErrortoo. - Remote files are read-only.
- Schemes: the default build (no C) reads
http://.https://needsmaturin develop --release --features https(rustls with ring, which compiles C);s3://,gs://andaz://need thes3,gcsandazurefeatures (credentials from the environment; aws-lc-rs, C).
Writing
clawhdf5.File(path, "w") with create_dataset(name, data=array, chunks=..., compression="gzip"), create_group and attrs[...] = ...
writes float64, float32, int64, int32 and uint8 arrays; the file is
written on close().
Editing a file in place
clawhdf5.File(path, "r+") (or "a" on an existing file) edits the file
where it is, through clawhdf5's FileEditor; the file is locked until
close(), and every edit is written and synced before the statement
returns.
with clawhdf5.File("data.h5", "r+") as f:
ds = f["grid"]
ds[10:20, ::2] = 0 # h5py keys and broadcasting
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
f["series"].attrs["units"] = "K"
f.attrs.create("version", 2, dtype="u1")
- Values: a numpy array is converted to the dataset's dtype as libhdf5
converts it (integers saturate at the target's limits; floats are
truncated toward zero and clipped); anything else goes through
numpy.asarray(value, dtype=ds.dtype), as in h5py. Writing NaN into an integer dataset raisesValueError(libhdf5 would store an arbitrary value). A few libhdf5 edge cases differ on purpose; seedocs/known-issues.md. - Shapes:
ds.resizegrows or shrinks chunked datasets within theirmaxshape, as h5py; datasets andattrsobjects taken before an edit see its result. - Attributes: numeric, bool, complex, bytes and
strdata of any shape.stris stored as a fixed-length UTF-8 string (h5py stores a variable-length one), so h5py reads it back asbytes. - Not supported (
NotImplementedError, nothing written): creating or deleting datasets and groups, deleting attributes (creating and replacing them works, compact or dense), writing compound fields by name, variable-length data, HDF5 array types, and whateverFileEditorrefuses (listed indocs/known-issues.md).
Tests
pip install pytest h5py
pytest crates/clawhdf5-py/tests
tests/test_read_vs_h5py.py compares every read with h5py on a file h5py
writes, opened locally and over HTTP (an in-process range server,
tests/conftest.py); tests/test_remote.py checks remote reads (requests,
failures, the GIL); tests/test_edit.py applies every edit through h5py and
clawhdf5 to copies of a file and compares them through h5py (and h5dump,
and h5rs check when CLAWHDF5_H5RS names it). scripts/ci-test.sh builds
the wheel and runs these in CI.
License
MIT