Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python

# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
This commit is contained in:
osobh
2026-09-27 07:59:14 -05:00
33 changed files with 4057 additions and 393 deletions
+142
View File
@@ -72,6 +72,148 @@ Design: `docs/design/swmr.md`.
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
groups or attributes (a SWMR writer cannot add objects or attributes).
### Correctness: edits planned from another file after a rename or `chdir` (2026-09-27)
- **`FileEditor` planned each edit by re-opening its path but wrote
through the file it held open** (fixed 2026-09-27; on main since PR #18,
no release). When the path came to name another file between edits — a
rename or replacement, or, for a relative path, a change of working
directory — an edit was laid out from the other file's metadata and
written into the held one, corrupting it (h5py: "invalid dataset size,
likely file corruption"). The Python `'r+'` handle had the same flaw in
its reads: it reopened the path after every edit, so reads came from the
other file. The editor now plans every edit from the file it holds, and
its path is canonicalised at open. New `FileEditor::reader()` opens the
held file anew for reading (on Linux through `/proc/self/fd`, so it
follows a renamed file; elsewhere by the path, refused when the path no
longer names the held file), without sharing the editor's lock; the
Python handle reads through it, and a `'w'` file is written at the
absolute path it was opened with. Tests: `edit_tests.rs`'s
`edits_go_to_the_file_held_not_the_path`; `test_edit.py`'s
`test_relative_path_and_chdir` and `test_path_replaced_between_edits`
(the review's repro).
### Correctness: zero extents in Fixed/Extensible Array chunk indexes (2026-09-27)
- **A chunked dataset whose maximum (or, with none recorded, current)
extent is 0 along a dimension made the reader divide by zero** (fixed
2026-09-27): `h5rs check` panicked ("attempt to divide by zero",
`chunk_grid.rs`) and the next `FileEditor::resize` failed with an
internal error. The unfixed editor produced such files by resizing a
clawhdf5-written dataset to a zero extent (12 of 30 extra random-edit
seeds on clawhdf5-written files). Such an index has no slot for any
chunk of the dataset, and `ChunkGrid::offsets` now says so instead of
dividing by the zero stride. Tests: `chunk_grid`'s
`zero_extent_has_no_chunks`, `edit_interop.rs`'s
`zero_extent_resizes_without_a_recorded_maximum` (a file the unfixed
editor left checks clean and resizes on; a 2.7.0-written file through
zero extents checks clean at each step), and `test_edit.py`'s random
edits on seeds 10 to 39 of a clawhdf5-written file.
### Correctness: resizing chunked datasets with no recorded maximum (2026-09-27)
- **`FileEditor::resize` scrambled the values of a chunked dataset whose
dataspace records no maximum dimensions when it shrank it** (fixed
2026-09-27). The editor shipped on main in PR #18 (a4c2ace) and reached
Python as `Dataset.resize` in `'r+'` files; no release has it. clawhdf5's
writer stores such a dataspace for every chunked dataset created without
a `maxshape`, with a Fixed Array (or Single Chunk) chunk index. The
editor patched only the current dimensions, and with no maximum
recorded the maximum is the current dimensions — which is also what the
Fixed Array linearises chunks by — so a shrink moved every chunk after
the first row: h5py, h5dump and our reader all read wrong values
without complaint (20x20, chunks 6x6, resized to 15x15: row 6 read
`0 0 0 0 0 0 120 ...`). After a shrink the dataset could not grow back
either (`3 exceeds the maximum 0`). libhdf5 itself never writes such a
dataspace (`H5S_set_extent_simple` records the maximum, equal to the
dimensions when none is given); reading one, `H5S_extent_get_dims`
reports the current dimensions as the maximum and `H5S_set_extent`
checks against no maximum at all, so libhdf5's own `H5Dset_extent`
scrambles such a file the same way (and lets it grow past its Fixed
Array). The editor now records the maximum libhdf5 would have written
— the dimensions before the first resize, the ones the index was built
with — then changes the current ones (the dataspace message grows by one
length per dimension and moves in the header when it has to). The
dataset then shrinks, grows back to that extent and refuses more, as
one libhdf5 wrote would. The writer (`FileBuilder`) now records the
maximum of every chunked dataset too, as libhdf5 does, so h5py can
resize what it writes (8 more bytes per dimension). Tests:
`crates/clawhdf5/tests/edit_resize_interop.rs` (a 2.7.0-written fixture,
new `FileBuilder` files and h5py files through shrinks, zero extents
and growth, checked against a model with our reader and h5py; h5py
resizing a `FileBuilder` file) and `test_edit.py`'s numpy-model checks.
### Python bindings: in-place editing (2026-09-27)
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
file for editing through `clawhdf5::FileEditor`, holding its exclusive
lock until `close()`:
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
one increasing index list) and broadcasting (numpy's rules for slices
and integers, allowing extra leading length-1 axes; the exact shape for
an index list, a scalar only where h5py expands it). A numpy array is
converted to the dataset's dtype as libhdf5 converts it in native byte
order (integers saturate; floats are truncated toward zero and clipped;
integers go into h5py's bool enum by value, as libhdf5 stores them);
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
does. NaN into an integer dataset is a `ValueError`.
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
and errors (`TypeError` for a dataset that is not chunked).
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
enum, complex as the `r`/`i` compound), except that `str` becomes
fixed-length UTF-8.
- Every edit is written and synced before it returns, then the file is
reopened: datasets and attrs objects taken earlier see new shapes and
attributes, and reads on other threads wait while an edit is written.
- What the editor cannot do raises `NotImplementedError` and writes
nothing (deleting attributes or objects, creating datasets or groups,
compound fields by name, variable-length data, ...;
`docs/known-issues.md`).
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
- Boolean-mask keys (`ds[mask]`, `ds[mask] = v`), which h5py supports,
raise `NotImplementedError` (they raised `TypeError`).
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
by h5py and by clawhdf5, and both read back through h5py after every
edit, on files h5py writes with `libver` earliest, v114 and latest and on
a clawhdf5-written one: a fixed sequence over every chunk index kind,
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
and compound types, and 16 random sequences of 40 edits (writes,
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
and leave its file unchanged. A matrix of every numeric source dtype into
every numeric dataset dtype at the edge values, dense attribute storage,
locking, objects seeing each other's edits, readers racing a writer
(never a partly written dataset). Every edited file must pass `h5dump`
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
a file opened `'r+'`.
### Python bindings: remote files (2026-09-27)
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
through the block cache, the whole read API (groups, attributes, every
dataset type and index the local reader handles). A URL is any
`scheme://…`; a remote file is read-only (another mode is a
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
`allow_full_download`, `max_full_download`, `require_validator`,
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
cache's counters. The default wheel builds plain HTTP only (no C: rustls
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
check now covers `clawhdf5-py`.
- **Every read parses through `File::storage()`** instead of
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
group listings, variable-length data through the global heap), inside one
shared file handle that releases the GIL for all file access, not only
dataset reads: a read waiting on the network lets other Python threads
run. A failed read of the storage (a network error, a file changed on
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
data; `key in group` raises it instead of answering `False`.
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
blocks and 1 KiB blocks) against a range-capable `http.server` in the
test process, plus `tests/test_remote.py`: request counts of a small
read, cache hits, a server without `Range` support (refused, or a
whole download when allowed), a file changed on the server, a server
that hangs up, 16 threads on one remote file, and a thread that keeps
running while a read waits on 0.2 s requests.
### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by