Files
clawhdf5/crates/clawhdf5-py/README.md
T
osobhandClaude Opus 5.5 b55b24b7ba docs: crate READMEs describe each crate as it is today
Every crate under crates/ now has a README (android, bench, cli, napi and
wasm had none), each saying what the crate is, its main types and
functions (names checked against the code), its cargo features with
defaults and which ones build C (checked with `cargo tree`), and links to
the top-level docs.

Corrections to the old stubs:
- clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs
  clawhdf5-format as a dependency.
- clawhdf5-filters: deflate backends only, and no library crate depends
  on it; the filter pipeline and every other codec are in -format.
- clawhdf5-gpu: vector distance compute, not I/O; not used by
  HDF5Memory::search.
- clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO.
- clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and
  search(q, k, ef).
- clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm
  backends are reported but run the scalar kernels.
- clawhdf5-gpu: the old example called l2_distances, which does not
  exist (l2_search).
- clawhdf5-agent: it described a "vector store" with "GPU acceleration";
  it now covers HDF5Memory, search options, WAL, signing, the graph.
- crates.io/docs.rs badges removed and `cargo install <crate>` replaced:
  nothing is published; depend on git.
- fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh.
- tools: the FileEditor interop tests that live in this crate.
- remote, py: license, other front ends, limits, File.mode/flush/chunks.

The Rust examples of the facade, format, filters, accel, ann, derive and
agent READMEs were compiled and run as tests (netcdf4, gpu and remote
compiled only) in a scratch crate; the CLI example was run.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:13:30 -05:00

6.4 KiB

clawhdf5-py

Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is clawhdf5 (import clawhdf5); it needs numpy and no libhdf5.

Install

Not on PyPI yet. Build it into a virtualenv with maturin:

pip install maturin numpy
cd crates/clawhdf5-py
maturin develop --release
python -c "import clawhdf5; print(clawhdf5.__version__)"

Reading

The read API follows h5py:

import numpy as np
import clawhdf5

with clawhdf5.File("data.h5", "r") as f:
    f.keys(), f["group"].items(), "group/data" in f
    ds = f["group/data"]           # or f["/group/data"], f["group"]["data"]
    ds.shape, ds.dtype, ds.attrs["units"]
    ds[10:20, ::2]                 # a small selection reads only its chunks
    ds[-1], ds[..., 0], ds[[1, 4, 7]]
    np.asarray(ds)
    f["table"]["id"]               # a compound field
  • Dataset.dtype is the numpy dtype h5py reports: integers and IEEE floats of every width in either byte order, bool, enums (with dtype.metadata['enum']), complex, S<n> fixed strings, object for variable-length strings (bytes values) and sequences (array values), V<n> opaque, array types, and compounds as structured dtypes. Other types raise TypeError.
  • Keys are h5py's: integers, slices with a positive step, ..., one increasing list of integers, compound field names. Each maps onto a hyperslab selection. None and negative steps are refused with h5py's errors; boolean masks (which h5py supports) raise NotImplementedError, for reads and writes.
  • What is read from the file: a selection whose bounding box covers at most half the dataset decodes only the chunks (or contiguous rows) the box overlaps. The library decodes the whole dataset for a larger box (including a strided slice such as ds[::100] across a chunked dataset), and for compact, virtual and unwritten datasets and chunked ones with a non-default fill value. An index list is read one group of neighbouring chunks at a time (a new group only past a chunk with no selected index), so each chunk is decoded once. ds[()], ds[...] and np.asarray(ds) use the file's chunk cache; other selections do not.
  • The bytes the library reads become the numpy array's buffer without a copy, and the read runs with the GIL released, so threads read in parallel. A bug in the library (a Rust panic) raises clawhdf5.InternalError, a RuntimeError.
  • Attributes return what h5py returns; clawhdf5.Empty stands for a null dataspace (h5py's Empty).
  • Also as in h5py: File.mode ('r', or 'r+' for a writable file), File.flush() (a no-op: edits are already synced), Dataset.chunks.

Remote files

A URL instead of a path reads the file where it is, through clawhdf5-remote: HTTP range requests through a block cache (1 MiB blocks, 64 MiB budget by default), fetching only the blocks a read needs. The whole read API works the same, and the GIL is released while waiting on the network.

f = clawhdf5.File("http://host/data.h5")          # default options
f = clawhdf5.File.open_url(
    "http://host/data.h5",
    block_size=256 * 1024, cache_size=128 << 20,  # the block cache
    headers={"Authorization": "Bearer ..."},      # sent to this origin only
    retries=3, timeout=30.0, max_redirects=5, max_parallel=8,
    allow_full_download=False,  # a server without Range support: refuse
    require_validator=False,    # refuse servers without ETag/Last-Modified
)
f.remote_stats   # {'requests': ..., 'bytes_fetched': ..., 'hits': ..., ...}
  • The file is pinned when opened (ETag or Last-Modified, and length): if it changes on the server, reads raise OSError instead of mixing versions. Network failures are OSError too.
  • Remote files are read-only.
  • Schemes: the default build (no C) reads http://. https:// needs maturin develop --release --features https (rustls with ring, which compiles C); s3://, gs:// and az:// need the s3, gcs and azure features (credentials from the environment; aws-lc-rs, C).

Writing

clawhdf5.File(path, "w") with create_dataset(name, data=array, chunks=..., compression="gzip"), create_group and attrs[...] = ... writes float64, float32, int64, int32 and uint8 arrays; the file is written on close().

Editing a file in place

clawhdf5.File(path, "r+") (or "a" on an existing file) edits the file where it is, through clawhdf5's FileEditor; the file is locked until close(), and every edit is written and synced before the statement returns.

with clawhdf5.File("data.h5", "r+") as f:
    ds = f["grid"]
    ds[10:20, ::2] = 0                  # h5py keys and broadcasting
    ds[[1, 4, 7], 3] = [1.5, 2.5, 3.5]  # one index list: exact shape
    f["series"].resize((5000, 3))       # or .resize(5000, axis=0)
    f["series"].attrs["units"] = "K"
    f.attrs.create("version", 2, dtype="u1")
  • Values: a numpy array is converted to the dataset's dtype as libhdf5 converts it (integers saturate at the target's limits; floats are truncated toward zero and clipped); anything else goes through numpy.asarray(value, dtype=ds.dtype), as in h5py. Writing NaN into an integer dataset raises ValueError (libhdf5 would store an arbitrary value). A few libhdf5 edge cases differ on purpose; see docs/known-issues.md.
  • Shapes: ds.resize grows or shrinks chunked datasets within their maxshape, as h5py; datasets and attrs objects taken before an edit see its result.
  • Attributes: numeric, bool, complex, bytes and str data of any shape. str is stored as a fixed-length UTF-8 string (h5py stores a variable-length one), so h5py reads it back as bytes.
  • Not supported (NotImplementedError, nothing written): creating or deleting datasets, groups and attributes, writing compound fields by name, variable-length data, HDF5 array types, and whatever FileEditor refuses (listed in docs/known-issues.md).

Tests

pip install pytest h5py
pytest crates/clawhdf5-py/tests

tests/test_read_vs_h5py.py compares every read with h5py on a file h5py writes, opened locally and over HTTP (an in-process range server, tests/conftest.py); tests/test_remote.py checks remote reads (requests, failures, the GIL); tests/test_edit.py applies every edit through h5py and clawhdf5 to copies of a file and compares them through h5py (and h5dump, and h5rs check when CLAWHDF5_H5RS names it). scripts/ci-test.sh builds the wheel and runs these in CI.

License

MIT