The READMEs said ds[...] reads only the selected elements, and the facade's read_selection docs that only intersecting chunks are decompressed. The bounding-box path runs only when the box covers at most half the dataset; larger boxes (any strided slice across the dataset), compact, virtual and unwritten datasets and chunked ones with a non-default fill value decode the whole dataset. The READMEs, the facade and format docs, the bindings' docstrings and known-issues now say so, and how index lists are read. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
84 lines
3.1 KiB
Markdown
84 lines
3.1 KiB
Markdown
# clawhdf5-py
|
|
|
|
[](https://crates.io/crates/clawhdf5-py)
|
|
[](https://docs.rs/clawhdf5-py)
|
|
|
|
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
|
|
`clawhdf5` (`import clawhdf5`); it needs numpy and no libhdf5.
|
|
|
|
## Install
|
|
|
|
Not on PyPI yet. Build it into a virtualenv with [maturin](https://www.maturin.rs):
|
|
|
|
```bash
|
|
pip install maturin numpy
|
|
cd crates/clawhdf5-py
|
|
maturin develop --release
|
|
python -c "import clawhdf5; print(clawhdf5.__version__)"
|
|
```
|
|
|
|
## Reading
|
|
|
|
The read API follows h5py:
|
|
|
|
```python
|
|
import numpy as np
|
|
import clawhdf5
|
|
|
|
with clawhdf5.File("data.h5", "r") as f:
|
|
f.keys(), f["group"].items(), "group/data" in f
|
|
ds = f["group/data"] # or f["/group/data"], f["group"]["data"]
|
|
ds.shape, ds.dtype, ds.attrs["units"]
|
|
ds[10:20, ::2] # a small selection reads only its chunks
|
|
ds[-1], ds[..., 0], ds[[1, 4, 7]]
|
|
np.asarray(ds)
|
|
f["table"]["id"] # a compound field
|
|
```
|
|
|
|
- `Dataset.dtype` is the numpy dtype h5py reports: integers and IEEE floats
|
|
of every width in either byte order, `bool`, enums (with
|
|
`dtype.metadata['enum']`), complex, `S<n>` fixed strings, `object` for
|
|
variable-length strings (`bytes` values) and sequences (array values),
|
|
`V<n>` opaque, array types, and compounds as structured dtypes.
|
|
Other types raise `TypeError`.
|
|
- Keys are h5py's: integers, slices with a positive step, `...`, one
|
|
increasing list of integers, compound field names. Each maps onto a
|
|
hyperslab selection. `None`, negative steps and boolean masks are refused
|
|
with h5py's errors.
|
|
- What is read from the file: a selection whose bounding box covers at
|
|
most half the dataset decodes only the chunks (or contiguous rows) the box
|
|
overlaps. The library decodes the whole dataset for a larger box
|
|
(including a strided slice such as `ds[::100]` across a chunked dataset),
|
|
and for compact, virtual and unwritten datasets and chunked ones with a
|
|
non-default fill value. An index list is read one group of neighbouring
|
|
chunks at a time (a new group only past a chunk with no selected index),
|
|
so each chunk is decoded once. `ds[()]`, `ds[...]` and `np.asarray(ds)`
|
|
use the file's chunk cache; other selections do not.
|
|
- The bytes the library reads become the numpy array's buffer without a
|
|
copy, and the read runs with the GIL released, so threads read in
|
|
parallel. A bug in the library (a Rust panic) raises
|
|
`clawhdf5.InternalError`, a `RuntimeError`.
|
|
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
|
|
dataspace (h5py's `Empty`).
|
|
|
|
## Writing
|
|
|
|
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
|
|
chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
|
|
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
|
|
written on `close()`.
|
|
|
|
## Tests
|
|
|
|
```bash
|
|
pip install pytest h5py
|
|
pytest crates/clawhdf5-py/tests
|
|
```
|
|
|
|
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
|
|
writes. `scripts/ci-test.sh` builds the wheel and runs these in CI.
|
|
|
|
## License
|
|
|
|
MIT
|