clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a FileEditor, and with it the file's exclusive lock, until close(): - ds[key] = value: h5py's keys and broadcasting (numpy's rules for slices and integers with extra leading 1-axes allowed; the exact shape for an index list, a scalar only where h5py expands it). Arrays are converted as libhdf5 converts them in native byte order (integers saturate, floats truncate toward zero and clip, integers go into h5py's bool enum by value); other values through numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer dataset is a ValueError instead of libhdf5's arbitrary value. The value preparation is a small Python module compiled into the extension (src/edit_helpers.py). - ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules. - attrs[name] = value, attrs.create(name, data, shape, dtype), attrs.modify: numeric, bool, complex, bytes and str data of any shape, with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor cannot write variable-length strings). - File.mode, File.flush(), Dataset.chunks. Each edit runs with the GIL released under the file handle's write lock (no read sees a half-written edit), then the file is reopened; datasets and attrs objects re-read their shape and attributes when the handle's edit generation moved. What the editor cannot do is NotImplementedError before anything is written: deleting attributes or objects, creating datasets or groups, compound fields by name, variable-length data, and FileEditor's own limits. Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft conversions in non-native byte order (a float in (-1, 0) becomes the integer minimum, same-size unsigned->signed wraps) and native casts that are undefined in C (half floats into unsigned, float(max) rounded up) -- clawhdf5 saturates as libhdf5's native path does; listed in docs/known-issues.md. Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to copies of the same file and both read back through h5py after each edit, on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed sequence over every chunk index kind, compact/contiguous/gzip layouts and numeric, bool, enum, complex, string and compound types, 16 random sequences of 40 edits, and a numeric conversion matrix; a refused edit must be refused by both and leave the file unchanged. Also dense attributes, locking, objects seeing edits, readers racing a writer, and h5dump (plus h5rs check in ci-test.sh) on every edited file. The read-vs-h5py suite also runs on a file opened 'r+'. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
154 lines
6.4 KiB
Markdown
154 lines
6.4 KiB
Markdown
# clawhdf5-py
|
|
|
|
[](https://crates.io/crates/clawhdf5-py)
|
|
[](https://docs.rs/clawhdf5-py)
|
|
|
|
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
|
|
`clawhdf5` (`import clawhdf5`); it needs numpy and no libhdf5.
|
|
|
|
## Install
|
|
|
|
Not on PyPI yet. Build it into a virtualenv with [maturin](https://www.maturin.rs):
|
|
|
|
```bash
|
|
pip install maturin numpy
|
|
cd crates/clawhdf5-py
|
|
maturin develop --release
|
|
python -c "import clawhdf5; print(clawhdf5.__version__)"
|
|
```
|
|
|
|
## Reading
|
|
|
|
The read API follows h5py:
|
|
|
|
```python
|
|
import numpy as np
|
|
import clawhdf5
|
|
|
|
with clawhdf5.File("data.h5", "r") as f:
|
|
f.keys(), f["group"].items(), "group/data" in f
|
|
ds = f["group/data"] # or f["/group/data"], f["group"]["data"]
|
|
ds.shape, ds.dtype, ds.attrs["units"]
|
|
ds[10:20, ::2] # a small selection reads only its chunks
|
|
ds[-1], ds[..., 0], ds[[1, 4, 7]]
|
|
np.asarray(ds)
|
|
f["table"]["id"] # a compound field
|
|
```
|
|
|
|
- `Dataset.dtype` is the numpy dtype h5py reports: integers and IEEE floats
|
|
of every width in either byte order, `bool`, enums (with
|
|
`dtype.metadata['enum']`), complex, `S<n>` fixed strings, `object` for
|
|
variable-length strings (`bytes` values) and sequences (array values),
|
|
`V<n>` opaque, array types, and compounds as structured dtypes.
|
|
Other types raise `TypeError`.
|
|
- Keys are h5py's: integers, slices with a positive step, `...`, one
|
|
increasing list of integers, compound field names. Each maps onto a
|
|
hyperslab selection. `None`, negative steps and boolean masks are refused
|
|
with h5py's errors.
|
|
- What is read from the file: a selection whose bounding box covers at
|
|
most half the dataset decodes only the chunks (or contiguous rows) the box
|
|
overlaps. The library decodes the whole dataset for a larger box
|
|
(including a strided slice such as `ds[::100]` across a chunked dataset),
|
|
and for compact, virtual and unwritten datasets and chunked ones with a
|
|
non-default fill value. An index list is read one group of neighbouring
|
|
chunks at a time (a new group only past a chunk with no selected index),
|
|
so each chunk is decoded once. `ds[()]`, `ds[...]` and `np.asarray(ds)`
|
|
use the file's chunk cache; other selections do not.
|
|
- The bytes the library reads become the numpy array's buffer without a
|
|
copy, and the read runs with the GIL released, so threads read in
|
|
parallel. A bug in the library (a Rust panic) raises
|
|
`clawhdf5.InternalError`, a `RuntimeError`.
|
|
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
|
|
dataspace (h5py's `Empty`).
|
|
|
|
## Remote files
|
|
|
|
A URL instead of a path reads the file where it is, through
|
|
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
|
|
64 MiB budget by default), fetching only the blocks a read needs. The whole
|
|
read API works the same, and the GIL is released while waiting on the
|
|
network.
|
|
|
|
```python
|
|
f = clawhdf5.File("http://host/data.h5") # default options
|
|
f = clawhdf5.File.open_url(
|
|
"http://host/data.h5",
|
|
block_size=256 * 1024, cache_size=128 << 20, # the block cache
|
|
headers={"Authorization": "Bearer ..."}, # sent to this origin only
|
|
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
|
|
allow_full_download=False, # a server without Range support: refuse
|
|
require_validator=False, # refuse servers without ETag/Last-Modified
|
|
)
|
|
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
|
|
```
|
|
|
|
- The file is pinned when opened (ETag or Last-Modified, and length): if it
|
|
changes on the server, reads raise `OSError` instead of mixing versions.
|
|
Network failures are `OSError` too.
|
|
- Remote files are read-only.
|
|
- Schemes: the default build (no C) reads `http://`. `https://` needs
|
|
`maturin develop --release --features https` (rustls with ring, which
|
|
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
|
|
`azure` features (credentials from the environment; aws-lc-rs, C).
|
|
|
|
## Writing
|
|
|
|
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
|
|
chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
|
|
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
|
|
written on `close()`.
|
|
|
|
## Editing a file in place
|
|
|
|
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
|
|
where it is, through clawhdf5's `FileEditor`; the file is locked until
|
|
`close()`, and every edit is written and synced before the statement
|
|
returns.
|
|
|
|
```python
|
|
with clawhdf5.File("data.h5", "r+") as f:
|
|
ds = f["grid"]
|
|
ds[10:20, ::2] = 0 # h5py keys and broadcasting
|
|
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
|
|
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
|
|
f["series"].attrs["units"] = "K"
|
|
f.attrs.create("version", 2, dtype="u1")
|
|
```
|
|
|
|
- Values: a numpy array is converted to the dataset's dtype as libhdf5
|
|
converts it (integers saturate at the target's limits; floats are
|
|
truncated toward zero and clipped); anything else goes through
|
|
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
|
|
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
|
|
value). A few libhdf5 edge cases differ on purpose; see
|
|
`docs/known-issues.md`.
|
|
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
|
|
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
|
|
see its result.
|
|
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
|
|
`str` is stored as a fixed-length UTF-8 string (h5py stores a
|
|
variable-length one), so h5py reads it back as `bytes`.
|
|
- Not supported (`NotImplementedError`, nothing written): creating or
|
|
deleting datasets, groups and attributes, writing compound fields by
|
|
name, variable-length data, HDF5 array types, and whatever
|
|
`FileEditor` refuses (listed in `docs/known-issues.md`).
|
|
|
|
## Tests
|
|
|
|
```bash
|
|
pip install pytest h5py
|
|
pytest crates/clawhdf5-py/tests
|
|
```
|
|
|
|
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
|
|
writes, opened locally and over HTTP (an in-process range server,
|
|
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
|
|
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
|
|
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
|
|
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
|
|
the wheel and runs these in CI.
|
|
|
|
## License
|
|
|
|
MIT
|