h5py supports boolean masks for reads and writes; clawhdf5 supports
neither, so a mask is an unsupported operation (NotImplementedError, as
for every other edit the bindings cannot do), not an invalid key.
Tests: test_unsupported_edits_are_clear_errors (1-D, N-D and per-axis
mask writes, file unchanged) and test_boolean_masks_are_refused (reads).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5.File(path, 'r+') (and 'a' on an existing file) holds a
FileEditor, and with it the file's exclusive lock, until close():
- ds[key] = value: h5py's keys and broadcasting (numpy's rules for
slices and integers with extra leading 1-axes allowed; the exact shape
for an index list, a scalar only where h5py expands it). Arrays are
converted as libhdf5 converts them in native byte order (integers
saturate, floats truncate toward zero and clip, integers go into h5py's
bool enum by value); other values through
numpy.asarray(value, dtype=ds.dtype), as h5py does. NaN into an integer
dataset is a ValueError instead of libhdf5's arbitrary value. The value
preparation is a small Python module compiled into the extension
(src/edit_helpers.py).
- ds.resize(shape) / ds.resize(n, axis=k) with h5py's argument rules.
- attrs[name] = value, attrs.create(name, data, shape, dtype),
attrs.modify: numeric, bool, complex, bytes and str data of any shape,
with h5py's HDF5 types; str is stored as fixed-length UTF-8 (the editor
cannot write variable-length strings).
- File.mode, File.flush(), Dataset.chunks.
Each edit runs with the GIL released under the file handle's write lock
(no read sees a half-written edit), then the file is reopened;
datasets and attrs objects re-read their shape and attributes when the
handle's edit generation moved. What the editor cannot do is
NotImplementedError before anything is written: deleting attributes or
objects, creating datasets or groups, compound fields by name,
variable-length data, and FileEditor's own limits.
Where libhdf5 2.0 (h5py 3.16) converts inconsistently -- its soft
conversions in non-native byte order (a float in (-1, 0) becomes the
integer minimum, same-size unsigned->signed wraps) and native casts that
are undefined in C (half floats into unsigned, float(max) rounded up) --
clawhdf5 saturates as libhdf5's native path does; listed in
docs/known-issues.md.
Tests (tests/test_edit.py): every edit applied by h5py and by clawhdf5 to
copies of the same file and both read back through h5py after each edit,
on h5py files (libver earliest, v114, latest) and a clawhdf5 file: a fixed
sequence over every chunk index kind, compact/contiguous/gzip layouts and
numeric, bool, enum, complex, string and compound types, 16 random
sequences of 40 edits, and a numeric conversion matrix; a refused edit
must be refused by both and leave the file unchanged. Also dense
attributes, locking, objects seeing edits, readers racing a writer, and
h5dump (plus h5rs check in ci-test.sh) on every edited file. The
read-vs-h5py suite also runs on a file opened 'r+'.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.
- Every object of a File now shares one handle (src/handle.rs) that
runs all file access, metadata included, with the GIL released and
parses through File::storage() and the clawhdf5_format *_in functions.
Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
clawhdf5_remote::storage_for_url (read-only; another mode is a
ValueError). File.open_url(url, **options) takes the cache and HTTP
options (block_size, cache_size, headers, retries, timeout,
allow_full_download, max_full_download, require_validator,
max_redirects, max_parallel); File.remote_stats gives the block
cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
OSError, never KeyError/ValueError and never data; `key in group`
raises it instead of answering False.
Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
test_threads_read_the_same_file passed with the GIL held. The new
test_reads_release_the_gil measures the longest stall of a spinning
Python thread while another reads: with py.detach removed from the read
it stalled 0.062 s of a 0.064 s read and failed; with it, about 3 ms.
test_errors_match_h5py now compares the result whenever h5py reads the
key, instead of only checking that we raise when h5py raises, over a
longer key list.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ds[np.array(1)] went down the index-list path, where tolist() returns a
scalar and extracting a list of indices raised a confusing TypeError.
h5py treats it as an integer index; so do we now. The h5py comparison
keys include 0-d arrays (signed and unsigned) on each axis; they failed
before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every ds[...] and g[k] resolved the path from the root again, two or
three times per open, and resolving a name in a large group scans its
links: visiting a group was O(n^2). 4000 scalar datasets in one group
took 39 s (v1 group) and 131 s (dense) to list, read and re-read; now
0.3 s each. A Dataset keeps its object address, a Group (and the file's
root) its address and, after the first lookup, its link table.
New facade API File::dataset_at(address), tested in integration_tests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each run of consecutive indices was its own uncached hyperslab read, so
a list over a compressed chunked dataset decoded the same chunk once per
run (d[range(0, 200000, 40)] over 20 gzip chunks: 8 s, h5py 0.014 s).
Plan::reads now groups the indices — a group ends only where a whole
chunk holds no selected index, or, unchunked, at a gap over 64 KiB — and
the selected rows are gathered from each group's block in Rust. Now
3.8 ms (h5py 4.1 ms, release, tank). The new test (1-D, 2-D and
contiguous, compared with h5py, 2 s bound) took 5.8 s before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
np.concatenate copies structured dtypes field by field into np.empty, so
the padding of ds[[0, 3, 6]] held process memory. The runs' bytes are
joined in Rust, whole elements at a time, before anything becomes numpy:
the padding is the file's bytes (h5py's) and the result is still a view
of the Rust buffer. The h5py comparisons now compare every byte of
structured values; the new test failed on the padding before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
PanicException derives from BaseException, so `except Exception` let a
library bug through. Every call from the bindings into the library now
runs under catch_unwind and a panic becomes InternalError (RuntimeError)
naming the object. Tests: a hidden hook panics inside the guard; and the
v4 chunk indexes are compared with h5py from Python — with the library
fix reverted, ds[0:30] of the implicit-index dataset now raises
InternalError instead of aborting the test run.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ds[key] read the whole dataset and sliced it in numpy, and knew six
dtypes. Keys (ints, positive-step slices, Ellipsis, one increasing index
list, compound field names) now map onto hyperslab selections, and the
facade's read_selection bytes become the numpy buffer without a copy
(PyArray::from_vec viewed as the dtype). dtype mapping follows h5py for
all integer/IEEE float widths and byte orders, bool, enum, complex, fixed
and variable-length strings, vlen sequences, opaque, array types and
(nested, padded) compounds; anything it cannot describe exactly is a
TypeError. Attributes return what h5py returns; groups and files gain
the rest of the h5py mapping interface. Reads run under py.detach.
tests/test_read_vs_h5py.py compares >500 reads with h5py 3.16 on an
h5py-written file, checks errors match, that a damaged chunk outside the
selection is never touched, and 8 threads reading at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>