The Python bindings could not open a remote file: they parsed through File::as_bytes() in eight places (path lookups, object headers, dataspaces, attributes, group listings, the global heap of variable-length data), which a storage-backed file does not have. - Every object of a File now shares one handle (src/handle.rs) that runs all file access, metadata included, with the GIL released and parses through File::storage() and the clawhdf5_format *_in functions. Local files take the same path (their storage is the mmap). - clawhdf5.File(url) opens any scheme://... through clawhdf5_remote::storage_for_url (read-only; another mode is a ValueError). File.open_url(url, **options) takes the cache and HTTP options (block_size, cache_size, headers, retries, timeout, allow_full_download, max_full_download, require_validator, max_redirects, max_parallel); File.remote_stats gives the block cache's counters. - Default build: plain HTTP only, no C. https (rustls/ring) and s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and ci-test.sh's no-C check now covers the crate. - A failed storage read (network error, file changed on the server) is an OSError, never KeyError/ValueError and never data; `key in group` raises it instead of answering False. Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and 1 KiB blocks) against a range-capable http.server in the test process (conftest.RangeServer); test_remote.py covers request counts, cache hits, a server without Range support, a changed file, a server that hangs up, 16 threads, and a spinning thread that keeps running while a read waits on 0.2 s requests. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
4.6 KiB
clawhdf5-py
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
clawhdf5 (import clawhdf5); it needs numpy and no libhdf5.
Install
Not on PyPI yet. Build it into a virtualenv with maturin:
pip install maturin numpy
cd crates/clawhdf5-py
maturin develop --release
python -c "import clawhdf5; print(clawhdf5.__version__)"
Reading
The read API follows h5py:
import numpy as np
import clawhdf5
with clawhdf5.File("data.h5", "r") as f:
f.keys(), f["group"].items(), "group/data" in f
ds = f["group/data"] # or f["/group/data"], f["group"]["data"]
ds.shape, ds.dtype, ds.attrs["units"]
ds[10:20, ::2] # a small selection reads only its chunks
ds[-1], ds[..., 0], ds[[1, 4, 7]]
np.asarray(ds)
f["table"]["id"] # a compound field
Dataset.dtypeis the numpy dtype h5py reports: integers and IEEE floats of every width in either byte order,bool, enums (withdtype.metadata['enum']), complex,S<n>fixed strings,objectfor variable-length strings (bytesvalues) and sequences (array values),V<n>opaque, array types, and compounds as structured dtypes. Other types raiseTypeError.- Keys are h5py's: integers, slices with a positive step,
..., one increasing list of integers, compound field names. Each maps onto a hyperslab selection.None, negative steps and boolean masks are refused with h5py's errors. - What is read from the file: a selection whose bounding box covers at
most half the dataset decodes only the chunks (or contiguous rows) the box
overlaps. The library decodes the whole dataset for a larger box
(including a strided slice such as
ds[::100]across a chunked dataset), and for compact, virtual and unwritten datasets and chunked ones with a non-default fill value. An index list is read one group of neighbouring chunks at a time (a new group only past a chunk with no selected index), so each chunk is decoded once.ds[()],ds[...]andnp.asarray(ds)use the file's chunk cache; other selections do not. - The bytes the library reads become the numpy array's buffer without a
copy, and the read runs with the GIL released, so threads read in
parallel. A bug in the library (a Rust panic) raises
clawhdf5.InternalError, aRuntimeError. - Attributes return what h5py returns;
clawhdf5.Emptystands for a null dataspace (h5py'sEmpty).
Remote files
A URL instead of a path reads the file where it is, through
clawhdf5-remote: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise
OSErrorinstead of mixing versions. Network failures areOSErrortoo. - Remote files are read-only.
- Schemes: the default build (no C) reads
http://.https://needsmaturin develop --release --features https(rustls with ring, which compiles C);s3://,gs://andaz://need thes3,gcsandazurefeatures (credentials from the environment; aws-lc-rs, C).
Writing
clawhdf5.File(path, "w") with create_dataset(name, data=array, chunks=..., compression="gzip"), create_group and attrs[...] = ...
writes float64, float32, int64, int32 and uint8 arrays; the file is
written on close().
Tests
pip install pytest h5py
pytest crates/clawhdf5-py/tests
tests/test_read_vs_h5py.py compares every read with h5py on a file h5py
writes, opened locally and over HTTP (an in-process range server,
tests/conftest.py); tests/test_remote.py checks remote reads (requests,
failures, the GIL). scripts/ci-test.sh builds the wheel and runs these
in CI.
License
MIT