# clawhdf5-py [![crates.io](https://img.shields.io/crates/v/clawhdf5-py.svg)](https://crates.io/crates/clawhdf5-py) [![docs.rs](https://docs.rs/clawhdf5-py/badge.svg)](https://docs.rs/clawhdf5-py) Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is `clawhdf5` (`import clawhdf5`); it needs numpy and no libhdf5. ## Install Not on PyPI yet. Build it into a virtualenv with [maturin](https://www.maturin.rs): ```bash pip install maturin numpy cd crates/clawhdf5-py maturin develop --release python -c "import clawhdf5; print(clawhdf5.__version__)" ``` ## Reading The read API follows h5py: ```python import numpy as np import clawhdf5 with clawhdf5.File("data.h5", "r") as f: f.keys(), f["group"].items(), "group/data" in f ds = f["group/data"] # or f["/group/data"], f["group"]["data"] ds.shape, ds.dtype, ds.attrs["units"] ds[10:20, ::2] # a small selection reads only its chunks ds[-1], ds[..., 0], ds[[1, 4, 7]] np.asarray(ds) f["table"]["id"] # a compound field ``` - `Dataset.dtype` is the numpy dtype h5py reports: integers and IEEE floats of every width in either byte order, `bool`, enums (with `dtype.metadata['enum']`), complex, `S` fixed strings, `object` for variable-length strings (`bytes` values) and sequences (array values), `V` opaque, array types, and compounds as structured dtypes. Other types raise `TypeError`. - Keys are h5py's: integers, slices with a positive step, `...`, one increasing list of integers, compound field names. Each maps onto a hyperslab selection. `None` and negative steps are refused with h5py's errors; boolean masks (which h5py supports) raise `NotImplementedError`, for reads and writes. - What is read from the file: a selection whose bounding box covers at most half the dataset decodes only the chunks (or contiguous rows) the box overlaps. The library decodes the whole dataset for a larger box (including a strided slice such as `ds[::100]` across a chunked dataset), and for compact, virtual and unwritten datasets and chunked ones with a non-default fill value. An index list is read one group of neighbouring chunks at a time (a new group only past a chunk with no selected index), so each chunk is decoded once. `ds[()]`, `ds[...]` and `np.asarray(ds)` use the file's chunk cache; other selections do not. - The bytes the library reads become the numpy array's buffer without a copy, and the read runs with the GIL released, so threads read in parallel. A bug in the library (a Rust panic) raises `clawhdf5.InternalError`, a `RuntimeError`. - Attributes return what h5py returns; `clawhdf5.Empty` stands for a null dataspace (h5py's `Empty`). ## Remote files A URL instead of a path reads the file where it is, through `clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks, 64 MiB budget by default), fetching only the blocks a read needs. The whole read API works the same, and the GIL is released while waiting on the network. ```python f = clawhdf5.File("http://host/data.h5") # default options f = clawhdf5.File.open_url( "http://host/data.h5", block_size=256 * 1024, cache_size=128 << 20, # the block cache headers={"Authorization": "Bearer ..."}, # sent to this origin only retries=3, timeout=30.0, max_redirects=5, max_parallel=8, allow_full_download=False, # a server without Range support: refuse require_validator=False, # refuse servers without ETag/Last-Modified ) f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...} ``` - The file is pinned when opened (ETag or Last-Modified, and length): if it changes on the server, reads raise `OSError` instead of mixing versions. Network failures are `OSError` too. - Remote files are read-only. - Schemes: the default build (no C) reads `http://`. `https://` needs `maturin develop --release --features https` (rustls with ring, which compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and `azure` features (credentials from the environment; aws-lc-rs, C). ## Writing `clawhdf5.File(path, "w")` with `create_dataset(name, data=array, chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...` writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is written on `close()`. ## Editing a file in place `clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file where it is, through clawhdf5's `FileEditor`; the file is locked until `close()`, and every edit is written and synced before the statement returns. ```python with clawhdf5.File("data.h5", "r+") as f: ds = f["grid"] ds[10:20, ::2] = 0 # h5py keys and broadcasting ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape f["series"].resize((5000, 3)) # or .resize(5000, axis=0) f["series"].attrs["units"] = "K" f.attrs.create("version", 2, dtype="u1") ``` - Values: a numpy array is converted to the dataset's dtype as libhdf5 converts it (integers saturate at the target's limits; floats are truncated toward zero and clipped); anything else goes through `numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an integer dataset raises `ValueError` (libhdf5 would store an arbitrary value). A few libhdf5 edge cases differ on purpose; see `docs/known-issues.md`. - Shapes: `ds.resize` grows or shrinks chunked datasets within their `maxshape`, as h5py; datasets and `attrs` objects taken before an edit see its result. - Attributes: numeric, bool, complex, bytes and `str` data of any shape. `str` is stored as a fixed-length UTF-8 string (h5py stores a variable-length one), so h5py reads it back as `bytes`. - Not supported (`NotImplementedError`, nothing written): creating or deleting datasets, groups and attributes, writing compound fields by name, variable-length data, HDF5 array types, and whatever `FileEditor` refuses (listed in `docs/known-issues.md`). ## Tests ```bash pip install pytest h5py pytest crates/clawhdf5-py/tests ``` `tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py writes, opened locally and over HTTP (an in-process range server, `tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests, failures, the GIL); `tests/test_edit.py` applies every edit through h5py and clawhdf5 to copies of a file and compares them through h5py (and `h5dump`, and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds the wheel and runs these in CI. ## License MIT