Files
clawhdf5/crates/clawhdf5-py/README.md
T
osobhandClaude Opus 5.5 bce07e9cb9
CI / test-arm64 (pull_request) Successful in 1m38s
CI / test (pull_request) Successful in 31m22s
feat: check HDF5 2.x small floats against libhdf5 2.2.0
Fixture written by libhdf5 2.2.0 (built from tag 2.2.0) through ctypes:
every bit pattern of FP4 E2M1, FP6 E2M3/E3M2, FP8 E4M3/E5M2 and a
bfloat16 LE/BE set, as datasets and attributes, with what H5Dread/H5Aread
return into double and float and the conversion exceptions libhdf5
raises. clawhdf5 already decoded every value as libhdf5 does, including
an all-ones exponent as inf/NaN in the OCP formats that have none
(documented as a deliberate match in known-issues).

- data_read: NaNs of non-native float layouts get libhdf5's bits (sign
  kept, every mantissa bit set) in f64 and f32.
- h5rs dump/ls name these types as h5dump/h5ls 2.x do
  (H5T_FLOAT_F4E2M1, "FP4 E2M1 4-bit float", float4-e2m1 ...), checked
  against h5dump 2.2.0's output of the fixture.
- Python bindings read them as h5py 3.16 does (float32 for bfloat16,
  float16 for the 1-byte formats, file byte order, same bytes as h5py);
  writing them in 'r+' is refused.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 21:25:03 -05:00

160 lines
6.9 KiB
Markdown

# clawhdf5-py
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
`clawhdf5` (`import clawhdf5`); it needs numpy and no libhdf5.
## Install
Not on PyPI yet. Build it into a virtualenv with [maturin](https://www.maturin.rs):
```bash
pip install maturin numpy
cd crates/clawhdf5-py
maturin develop --release
python -c "import clawhdf5; print(clawhdf5.__version__)"
```
## Reading
The read API follows h5py:
```python
import numpy as np
import clawhdf5
with clawhdf5.File("data.h5", "r") as f:
f.keys(), f["group"].items(), "group/data" in f
ds = f["group/data"] # or f["/group/data"], f["group"]["data"]
ds.shape, ds.dtype, ds.attrs["units"]
ds[10:20, ::2] # a small selection reads only its chunks
ds[-1], ds[..., 0], ds[[1, 4, 7]]
np.asarray(ds)
f["table"]["id"] # a compound field
```
- `Dataset.dtype` is the numpy dtype h5py reports: integers and IEEE floats
of every width in either byte order, `bool`, enums (with
`dtype.metadata['enum']`), complex, `S<n>` fixed strings, `object` for
variable-length strings (`bytes` values) and sequences (array values),
`V<n>` opaque, array types, and compounds as structured dtypes.
Non-IEEE floats (HDF5 2.x's bfloat16, FP8, FP6 and FP4) read as h5py
3.16 reads them: as the narrowest IEEE float that holds them (`float32`
for bfloat16, `float16` for the 1-byte formats) in the file's byte order,
with libhdf5's values (`tests/test_small_floats.py`); inside a compound or
array type they raise `TypeError`, and writing them raises
`NotImplementedError`. Other types raise `TypeError`.
- Keys are h5py's: integers, slices with a positive step, `...`, one
increasing list of integers, compound field names. Each maps onto a
hyperslab selection. `None` and negative steps are refused
with h5py's errors; boolean masks (which h5py supports) raise
`NotImplementedError`, for reads and writes.
- What is read from the file: a selection whose bounding box covers at
most half the dataset decodes only the chunks (or contiguous rows) the box
overlaps. The library decodes the whole dataset for a larger box
(including a strided slice such as `ds[::100]` across a chunked dataset),
and for compact, virtual and unwritten datasets and chunked ones with a
non-default fill value. An index list is read one group of neighbouring
chunks at a time (a new group only past a chunk with no selected index),
so each chunk is decoded once. `ds[()]`, `ds[...]` and `np.asarray(ds)`
use the file's chunk cache; other selections do not.
- The bytes the library reads become the numpy array's buffer without a
copy, and the read runs with the GIL released, so threads read in
parallel. A bug in the library (a Rust panic) raises
`clawhdf5.InternalError`, a `RuntimeError`.
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
dataspace (h5py's `Empty`).
- Also as in h5py: `File.mode` (`'r'`, or `'r+'` for a writable file), `File.flush()` (a no-op:
edits are already synced), `Dataset.chunks`.
## Remote files
A URL instead of a path reads the file where it is, through
`clawhdf5-remote`: HTTP range requests through a block cache (1 MiB blocks,
64 MiB budget by default), fetching only the blocks a read needs. The whole
read API works the same, and the GIL is released while waiting on the
network.
```python
f = clawhdf5.File("http://host/data.h5") # default options
f = clawhdf5.File.open_url(
"http://host/data.h5",
block_size=256 * 1024, cache_size=128 << 20, # the block cache
headers={"Authorization": "Bearer ..."}, # sent to this origin only
retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
allow_full_download=False, # a server without Range support: refuse
require_validator=False, # refuse servers without ETag/Last-Modified
)
f.remote_stats # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
```
- The file is pinned when opened (ETag or Last-Modified, and length): if it
changes on the server, reads raise `OSError` instead of mixing versions.
Network failures are `OSError` too.
- Remote files are read-only.
- Schemes: the default build (no C) reads `http://`. `https://` needs
`maturin develop --release --features https` (rustls with ring, which
compiles C); `s3://`, `gs://` and `az://` need the `s3`, `gcs` and
`azure` features (credentials from the environment; aws-lc-rs, C).
## Writing
`clawhdf5.File(path, "w")` with `create_dataset(name, data=array,
chunks=..., compression="gzip")`, `create_group` and `attrs[...] = ...`
writes `float64`, `float32`, `int64`, `int32` and `uint8` arrays; the file is
written on `close()`.
## Editing a file in place
`clawhdf5.File(path, "r+")` (or `"a"` on an existing file) edits the file
where it is, through clawhdf5's `FileEditor`; the file is locked until
`close()`, and every edit is written and synced before the statement
returns.
```python
with clawhdf5.File("data.h5", "r+") as f:
ds = f["grid"]
ds[10:20, ::2] = 0 # h5py keys and broadcasting
ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5] # one index list: exact shape
f["series"].resize((5000, 3)) # or .resize(5000, axis=0)
f["series"].attrs["units"] = "K"
f.attrs.create("version", 2, dtype="u1")
```
- Values: a numpy array is converted to the dataset's dtype as libhdf5
converts it (integers saturate at the target's limits; floats are
truncated toward zero and clipped); anything else goes through
`numpy.asarray(value, dtype=ds.dtype)`, as in h5py. Writing NaN into an
integer dataset raises `ValueError` (libhdf5 would store an arbitrary
value). A few libhdf5 edge cases differ on purpose; see
`docs/known-issues.md`.
- Shapes: `ds.resize` grows or shrinks chunked datasets within their
`maxshape`, as h5py; datasets and `attrs` objects taken before an edit
see its result.
- Attributes: numeric, bool, complex, bytes and `str` data of any shape.
`str` is stored as a fixed-length UTF-8 string (h5py stores a
variable-length one), so h5py reads it back as `bytes`.
- Not supported (`NotImplementedError`, nothing written): creating or
deleting datasets and groups, deleting attributes (creating and
replacing them works, compact or dense), writing compound fields by
name, variable-length data, HDF5 array types, and whatever
`FileEditor` refuses (listed in `docs/known-issues.md`).
## Tests
```bash
pip install pytest h5py
pytest crates/clawhdf5-py/tests
```
`tests/test_read_vs_h5py.py` compares every read with h5py on a file h5py
writes, opened locally and over HTTP (an in-process range server,
`tests/conftest.py`); `tests/test_remote.py` checks remote reads (requests,
failures, the GIL); `tests/test_edit.py` applies every edit through h5py and
clawhdf5 to copies of a file and compares them through h5py (and `h5dump`,
and `h5rs check` when `CLAWHDF5_H5RS` names it). `scripts/ci-test.sh` builds
the wheel and runs these in CI.
## License
MIT