Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python
# Conflicts: # CHANGELOG.md # docs/design/range-reads.md # docs/known-issues.md
This commit is contained in:
+142
@@ -72,6 +72,148 @@ Design: `docs/design/swmr.md`.
|
||||
`HttpStorage` pins the length), `MmapFile`/`LazyFile`, and refreshing
|
||||
groups or attributes (a SWMR writer cannot add objects or attributes).
|
||||
|
||||
### Correctness: edits planned from another file after a rename or `chdir` (2026-09-27)
|
||||
- **`FileEditor` planned each edit by re-opening its path but wrote
|
||||
through the file it held open** (fixed 2026-09-27; on main since PR #18,
|
||||
no release). When the path came to name another file between edits — a
|
||||
rename or replacement, or, for a relative path, a change of working
|
||||
directory — an edit was laid out from the other file's metadata and
|
||||
written into the held one, corrupting it (h5py: "invalid dataset size,
|
||||
likely file corruption"). The Python `'r+'` handle had the same flaw in
|
||||
its reads: it reopened the path after every edit, so reads came from the
|
||||
other file. The editor now plans every edit from the file it holds, and
|
||||
its path is canonicalised at open. New `FileEditor::reader()` opens the
|
||||
held file anew for reading (on Linux through `/proc/self/fd`, so it
|
||||
follows a renamed file; elsewhere by the path, refused when the path no
|
||||
longer names the held file), without sharing the editor's lock; the
|
||||
Python handle reads through it, and a `'w'` file is written at the
|
||||
absolute path it was opened with. Tests: `edit_tests.rs`'s
|
||||
`edits_go_to_the_file_held_not_the_path`; `test_edit.py`'s
|
||||
`test_relative_path_and_chdir` and `test_path_replaced_between_edits`
|
||||
(the review's repro).
|
||||
|
||||
### Correctness: zero extents in Fixed/Extensible Array chunk indexes (2026-09-27)
|
||||
- **A chunked dataset whose maximum (or, with none recorded, current)
|
||||
extent is 0 along a dimension made the reader divide by zero** (fixed
|
||||
2026-09-27): `h5rs check` panicked ("attempt to divide by zero",
|
||||
`chunk_grid.rs`) and the next `FileEditor::resize` failed with an
|
||||
internal error. The unfixed editor produced such files by resizing a
|
||||
clawhdf5-written dataset to a zero extent (12 of 30 extra random-edit
|
||||
seeds on clawhdf5-written files). Such an index has no slot for any
|
||||
chunk of the dataset, and `ChunkGrid::offsets` now says so instead of
|
||||
dividing by the zero stride. Tests: `chunk_grid`'s
|
||||
`zero_extent_has_no_chunks`, `edit_interop.rs`'s
|
||||
`zero_extent_resizes_without_a_recorded_maximum` (a file the unfixed
|
||||
editor left checks clean and resizes on; a 2.7.0-written file through
|
||||
zero extents checks clean at each step), and `test_edit.py`'s random
|
||||
edits on seeds 10 to 39 of a clawhdf5-written file.
|
||||
|
||||
### Correctness: resizing chunked datasets with no recorded maximum (2026-09-27)
|
||||
- **`FileEditor::resize` scrambled the values of a chunked dataset whose
|
||||
dataspace records no maximum dimensions when it shrank it** (fixed
|
||||
2026-09-27). The editor shipped on main in PR #18 (a4c2ace) and reached
|
||||
Python as `Dataset.resize` in `'r+'` files; no release has it. clawhdf5's
|
||||
writer stores such a dataspace for every chunked dataset created without
|
||||
a `maxshape`, with a Fixed Array (or Single Chunk) chunk index. The
|
||||
editor patched only the current dimensions, and with no maximum
|
||||
recorded the maximum is the current dimensions — which is also what the
|
||||
Fixed Array linearises chunks by — so a shrink moved every chunk after
|
||||
the first row: h5py, h5dump and our reader all read wrong values
|
||||
without complaint (20x20, chunks 6x6, resized to 15x15: row 6 read
|
||||
`0 0 0 0 0 0 120 ...`). After a shrink the dataset could not grow back
|
||||
either (`3 exceeds the maximum 0`). libhdf5 itself never writes such a
|
||||
dataspace (`H5S_set_extent_simple` records the maximum, equal to the
|
||||
dimensions when none is given); reading one, `H5S_extent_get_dims`
|
||||
reports the current dimensions as the maximum and `H5S_set_extent`
|
||||
checks against no maximum at all, so libhdf5's own `H5Dset_extent`
|
||||
scrambles such a file the same way (and lets it grow past its Fixed
|
||||
Array). The editor now records the maximum libhdf5 would have written
|
||||
— the dimensions before the first resize, the ones the index was built
|
||||
with — then changes the current ones (the dataspace message grows by one
|
||||
length per dimension and moves in the header when it has to). The
|
||||
dataset then shrinks, grows back to that extent and refuses more, as
|
||||
one libhdf5 wrote would. The writer (`FileBuilder`) now records the
|
||||
maximum of every chunked dataset too, as libhdf5 does, so h5py can
|
||||
resize what it writes (8 more bytes per dimension). Tests:
|
||||
`crates/clawhdf5/tests/edit_resize_interop.rs` (a 2.7.0-written fixture,
|
||||
new `FileBuilder` files and h5py files through shrinks, zero extents
|
||||
and growth, checked against a model with our reader and h5py; h5py
|
||||
resizing a `FileBuilder` file) and `test_edit.py`'s numpy-model checks.
|
||||
|
||||
### Python bindings: in-place editing (2026-09-27)
|
||||
- **`clawhdf5.File(path, 'r+')`** (and `'a'` on an existing file) opens a
|
||||
file for editing through `clawhdf5::FileEditor`, holding its exclusive
|
||||
lock until `close()`:
|
||||
- `ds[key] = value` with h5py's keys (integers, slices with steps, `...`,
|
||||
one increasing index list) and broadcasting (numpy's rules for slices
|
||||
and integers, allowing extra leading length-1 axes; the exact shape for
|
||||
an index list, a scalar only where h5py expands it). A numpy array is
|
||||
converted to the dataset's dtype as libhdf5 converts it in native byte
|
||||
order (integers saturate; floats are truncated toward zero and clipped;
|
||||
integers go into h5py's bool enum by value, as libhdf5 stores them);
|
||||
other values through `numpy.asarray(value, dtype=ds.dtype)`, as h5py
|
||||
does. NaN into an integer dataset is a `ValueError`.
|
||||
- `ds.resize(shape)` / `ds.resize(n, axis=k)` with h5py's argument rules
|
||||
and errors (`TypeError` for a dataset that is not chunked).
|
||||
- `obj.attrs[name] = value`, `attrs.create(name, data, shape, dtype)`,
|
||||
`attrs.modify`: numeric, bool, complex, bytes and `str` data of any
|
||||
shape, stored with the HDF5 types h5py uses (bool as the `FALSE`/`TRUE`
|
||||
enum, complex as the `r`/`i` compound), except that `str` becomes
|
||||
fixed-length UTF-8.
|
||||
- Every edit is written and synced before it returns, then the file is
|
||||
reopened: datasets and attrs objects taken earlier see new shapes and
|
||||
attributes, and reads on other threads wait while an edit is written.
|
||||
- What the editor cannot do raises `NotImplementedError` and writes
|
||||
nothing (deleting attributes or objects, creating datasets or groups,
|
||||
compound fields by name, variable-length data, ...;
|
||||
`docs/known-issues.md`).
|
||||
- `File.mode`, `File.flush()` (a no-op), `Dataset.chunks`.
|
||||
- Boolean-mask keys (`ds[mask]`, `ds[mask] = v`), which h5py supports,
|
||||
raise `NotImplementedError` (they raised `TypeError`).
|
||||
- Tests (`tests/test_edit.py`): each edit applied to two copies of a file,
|
||||
by h5py and by clawhdf5, and both read back through h5py after every
|
||||
edit, on files h5py writes with `libver` earliest, v114 and latest and on
|
||||
a clawhdf5-written one: a fixed sequence over every chunk index kind,
|
||||
compact/contiguous/gzip layouts and numeric, bool, enum, complex, string
|
||||
and compound types, and 16 random sequences of 40 edits (writes,
|
||||
resizes, attributes); where h5py refuses an edit clawhdf5 must refuse it
|
||||
and leave its file unchanged. A matrix of every numeric source dtype into
|
||||
every numeric dataset dtype at the edge values, dense attribute storage,
|
||||
locking, objects seeing each other's edits, readers racing a writer
|
||||
(never a partly written dataset). Every edited file must pass `h5dump`
|
||||
and, in `ci-test.sh`, `h5rs check`. The read-vs-h5py suite also runs on
|
||||
a file opened `'r+'`.
|
||||
|
||||
### Python bindings: remote files (2026-09-27)
|
||||
- **`clawhdf5.File(url)`** opens `http://` URLs (and `https://`, `s3://`,
|
||||
`gs://`, `az://` in a wheel built with the `https`, `s3`, `gcs`,
|
||||
`azure` features) through `clawhdf5-remote`'s `open_url`: range requests
|
||||
through the block cache, the whole read API (groups, attributes, every
|
||||
dataset type and index the local reader handles). A URL is any
|
||||
`scheme://…`; a remote file is read-only (another mode is a
|
||||
`ValueError`). **`clawhdf5.File.open_url(url, **options)`** takes
|
||||
`block_size`, `cache_size`, `headers`, `retries`, `timeout`,
|
||||
`allow_full_download`, `max_full_download`, `require_validator`,
|
||||
`max_redirects` and `max_parallel`; `File.remote_stats` gives the block
|
||||
cache's counters. The default wheel builds plain HTTP only (no C: rustls
|
||||
needs ring, and the cloud clients aws-lc-rs), and `ci-test.sh`'s no-C
|
||||
check now covers `clawhdf5-py`.
|
||||
- **Every read parses through `File::storage()`** instead of
|
||||
`File::as_bytes()` (path lookups, object headers, dataspaces, attributes,
|
||||
group listings, variable-length data through the global heap), inside one
|
||||
shared file handle that releases the GIL for all file access, not only
|
||||
dataset reads: a read waiting on the network lets other Python threads
|
||||
run. A failed read of the storage (a network error, a file changed on
|
||||
the server) is an `OSError`, never a `KeyError`/`ValueError` and never
|
||||
data; `key in group` raises it instead of answering `False`.
|
||||
- Tests: the whole read-vs-h5py suite also runs over HTTP (default 1 MiB
|
||||
blocks and 1 KiB blocks) against a range-capable `http.server` in the
|
||||
test process, plus `tests/test_remote.py`: request counts of a small
|
||||
read, cache hits, a server without `Range` support (refused, or a
|
||||
whole download when allowed), a file changed on the server, a server
|
||||
that hangs up, 16 threads on one remote file, and a thread that keeps
|
||||
running while a read waits on 0.2 s requests.
|
||||
|
||||
### Range reads, milestone M3: remote files (2026-09-26)
|
||||
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
||||
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
|
||||
|
||||
Reference in New Issue
Block a user