Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python

# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
This commit is contained in:
osobh
2026-09-27 07:59:14 -05:00
33 changed files with 4057 additions and 393 deletions
+87 -6
View File
@@ -30,6 +30,33 @@ h5py's SWMR reader). A file still being written is read with
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
at its length at open and is not meant for files that change while open.
## Shrinking a chunked dataset with no recorded maximum scrambled it
**Status:** fixed 2026-09-27, before any release (`FileEditor::resize`
shipped on main in PR #18, a4c2ace). Files the writer produced before the
fix still lack the maximum; see *Existing files*.
clawhdf5's writer stored no maximum dimensions for a chunked dataset
created without a `maxshape` (libhdf5 always stores one, equal to the
dimensions when none is given). With none recorded, the maximum is the
current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index
places chunks by the maximum. `FileEditor::resize` changed only the
current dimensions, so a shrink moved every existing chunk and every
reader returned wrong values; after a shrink the dataset could not grow
back. Found by the review of the Python editing work.
**Fix:** before the first resize of such a dataset the editor records the
maximum libhdf5 would have written (the dimensions the index was built
with); the writer now records it for every chunked dataset. **Test:**
`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:**
datasets written before the fix have no recorded maximum. The fixed
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
does not** — it scrambles them the same way, and lets them grow past
their Fixed Array. Resize them once with a fixed `FileEditor` (a resize
to the same shape changes nothing; shrink and grow back) before letting
libhdf5 resize them. A dataset already shrunk by the unfixed
editor holds misplaced chunks; rewrite it from a good copy.
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to
@@ -139,6 +166,57 @@ space it leaves is too small for its next, larger version.
**No journal.** A crash while an edit patches existing structures can leave
the file inconsistent; see the `FileEditor` documentation.
**Renamed files outside Linux.** Edits always go to the file the editor
opened. `FileEditor::reader()` (which the Python `'r+'` handle reads
through) reopens it through `/proc/self/fd` on Linux; elsewhere it reopens
the path, and after the path was renamed or replaced it fails (on Unix;
Windows cannot tell and would read whatever the path names).
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
**Status:** open (added 2026-09-27). The Python bindings edit through
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
before anything is written. On top of them:
- **No new or deleted objects:** `create_dataset`/`create_group` in an
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
`NotImplementedError` (the editor changes values, shapes and
attributes only). Mode `'a'` works on an existing file only.
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
whole elements of the same structured dtype are), HDF5 array-type
elements, variable-length data, strings padded with spaces or
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
ones, h5py's, are writable), compounds containing such strings, null
dataspaces, index-list writes of more than 2²² elements (write them
in slices), and boolean-mask keys (`ds[mask] = v`, and mask reads;
h5py supports both).
- **`str` attributes are fixed-length UTF-8**, where h5py writes
variable-length strings: h5py reads them back as `bytes`
(`numpy.bytes_`), not `str`.
- **Numeric conversion follows libhdf5's native-order results, not its
bugs.** Arrays are converted as libhdf5 converts them (integers
saturate, floats are truncated toward zero and clipped), checked value by
value against h5py 3.16 / HDF5 2.0 in
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
differs from h5py on purpose:
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
minimum or 2⁶³ depending on the type);
- when the dataset or the array is not in native byte order, libhdf5's
"soft" conversions store a float in (-1, 0) as the integer minimum and
wrap an unsigned value too large for the signed type of the same size
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
native order;
- libhdf5's native casts that are undefined in C: half floats into
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
integers (→ minimum), a float equal to the integer maximum rounded up
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
- **Each edit reopens the file** (a new memory map) so that reads see it;
reads from other threads wait while an edit is written.
## Selection reads that decode more than the selection
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
@@ -867,9 +945,9 @@ cache, but:
`read_*_zerocopy`) need the file in memory and answer
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
panics for such a file (`File::contiguous_bytes()` is the fallible form).
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
`remote` feature).
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
(`h5rs` and the Python bindings read through `File::storage`, and take
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -885,9 +963,12 @@ cache, but:
**Status:** open (added 2026-09-26, milestone M3 of
`docs/design/range-reads.md`).
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
parses through `File::as_bytes`, which a remote file does not have; the
wasm reader's `openUrl` is milestone M4.
- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
the default wheel reads plain `http://` only: `https://` needs a wheel
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.