The Python bindings could not open a remote file: they parsed through
File::as_bytes() in eight places (path lookups, object headers,
dataspaces, attributes, group listings, the global heap of
variable-length data), which a storage-backed file does not have.
- Every object of a File now shares one handle (src/handle.rs) that
runs all file access, metadata included, with the GIL released and
parses through File::storage() and the clawhdf5_format *_in functions.
Local files take the same path (their storage is the mmap).
- clawhdf5.File(url) opens any scheme://... through
clawhdf5_remote::storage_for_url (read-only; another mode is a
ValueError). File.open_url(url, **options) takes the cache and HTTP
options (block_size, cache_size, headers, retries, timeout,
allow_full_download, max_full_download, require_validator,
max_redirects, max_parallel); File.remote_stats gives the block
cache's counters.
- Default build: plain HTTP only, no C. https (rustls/ring) and
s3/gcs/azure (aws-lc-rs) are opt-in features of clawhdf5-py, and
ci-test.sh's no-C check now covers the crate.
- A failed storage read (network error, file changed on the server) is an
OSError, never KeyError/ValueError and never data; `key in group`
raises it instead of answering False.
Tests: the read-vs-h5py suite runs locally and over HTTP (1 MiB and
1 KiB blocks) against a range-capable http.server in the test process
(conftest.RangeServer); test_remote.py covers request counts, cache
hits, a server without Range support, a changed file, a server that
hangs up, 16 threads, and a spinning thread that keeps running while a
read waits on 0.2 s requests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
test_threads_read_the_same_file passed with the GIL held. The new
test_reads_release_the_gil measures the longest stall of a spinning
Python thread while another reads: with py.detach removed from the read
it stalled 0.062 s of a 0.064 s read and failed; with it, about 3 ms.
test_errors_match_h5py now compares the result whenever h5py reads the
key, instead of only checking that we raise when h5py raises, over a
longer key list.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ds[np.array(1)] went down the index-list path, where tolist() returns a
scalar and extracting a list of indices raised a confusing TypeError.
h5py treats it as an integer index; so do we now. The h5py comparison
keys include 0-d arrays (signed and unsigned) on each axis; they failed
before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Every ds[...] and g[k] resolved the path from the root again, two or
three times per open, and resolving a name in a large group scans its
links: visiting a group was O(n^2). 4000 scalar datasets in one group
took 39 s (v1 group) and 131 s (dense) to list, read and re-read; now
0.3 s each. A Dataset keeps its object address, a Group (and the file's
root) its address and, after the first lookup, its link table.
New facade API File::dataset_at(address), tested in integration_tests.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each run of consecutive indices was its own uncached hyperslab read, so
a list over a compressed chunked dataset decoded the same chunk once per
run (d[range(0, 200000, 40)] over 20 gzip chunks: 8 s, h5py 0.014 s).
Plan::reads now groups the indices — a group ends only where a whole
chunk holds no selected index, or, unchunked, at a gap over 64 KiB — and
the selected rows are gathered from each group's block in Rust. Now
3.8 ms (h5py 4.1 ms, release, tank). The new test (1-D, 2-D and
contiguous, compared with h5py, 2 s bound) took 5.8 s before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
np.concatenate copies structured dtypes field by field into np.empty, so
the padding of ds[[0, 3, 6]] held process memory. The runs' bytes are
joined in Rust, whole elements at a time, before anything becomes numpy:
the padding is the file's bytes (h5py's) and the result is still a view
of the Rust buffer. The h5py comparisons now compare every byte of
structured values; the new test failed on the padding before.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
PanicException derives from BaseException, so `except Exception` let a
library bug through. Every call from the bindings into the library now
runs under catch_unwind and a panic becomes InternalError (RuntimeError)
naming the object. Tests: a hidden hook panics inside the guard; and the
v4 chunk indexes are compared with h5py from Python — with the library
fix reverted, ds[0:30] of the implicit-index dataset now raises
InternalError instead of aborting the test run.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
ds[key] read the whole dataset and sliced it in numpy, and knew six
dtypes. Keys (ints, positive-step slices, Ellipsis, one increasing index
list, compound field names) now map onto hyperslab selections, and the
facade's read_selection bytes become the numpy buffer without a copy
(PyArray::from_vec viewed as the dtype). dtype mapping follows h5py for
all integer/IEEE float widths and byte orders, bool, enum, complex, fixed
and variable-length strings, vlen sequences, opaque, array types and
(nested, padded) compounds; anything it cannot describe exactly is a
TypeError. Attributes return what h5py returns; groups and files gain
the rest of the h5py mapping interface. Reads run under py.detach.
tests/test_read_vs_h5py.py compares >500 reads with h5py 3.16 on an
h5py-written file, checks errors match, that a damaged chunk outside the
selection is never touched, and 8 threads reading at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>