Merge branch 'feat/p3-python-remote-edit' into feat/p3-wasm-swmr-python
# Conflicts: # CHANGELOG.md # docs/design/range-reads.md # docs/known-issues.md
This commit is contained in:
@@ -12,6 +12,10 @@ done on branch `feat/p3-m5-swmr-reader`, with its own design in
|
||||
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
|
||||
next. Every count in §1–§2 was
|
||||
|
||||
object stores) and URLs in `h5rs` (see the M3 status below); the Python
|
||||
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
|
||||
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
|
||||
|
||||
change. Progress: M1, first part (the `Storage` trait and the metadata
|
||||
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
||||
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
||||
@@ -477,7 +481,7 @@ fast path within benchmark noise.
|
||||
a request counter exposed for tests and users.
|
||||
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
|
||||
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
|
||||
Python bindings, with these choices:
|
||||
Python bindings (done 2026-09-27, below), with these choices:
|
||||
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
|
||||
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
|
||||
sits below the facade.
|
||||
@@ -514,6 +518,17 @@ fast path within benchmark noise.
|
||||
same work without a cache: 141 936 requests. Per file: A lists in 2
|
||||
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
|
||||
6.4 MB: 35 001 object headers spread over the file).
|
||||
- *Status 2026-09-27, Python bindings:* done on branch
|
||||
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
|
||||
`File.open_url(url, **options)` (cache and HTTP options) go through
|
||||
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
|
||||
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
|
||||
parsing (path lookups, headers, attributes, listings, the global heap)
|
||||
moved from `File::as_bytes` to `File::storage()` and the `*_in`
|
||||
functions, and every file access, metadata included, runs with the GIL
|
||||
released. Checked by running the whole read-vs-h5py suite over an
|
||||
in-process range server (1 MiB and 1 KiB blocks) and by request counts
|
||||
in `crates/clawhdf5-py/tests/test_remote.py`.
|
||||
|
||||
**M4 — wasm lazy loading (1–2 weeks).**
|
||||
- `clawhdf5-wasm`: `openUrl(url) -> Promise<H5File>` backed by `fetch` with a
|
||||
|
||||
+87
-6
@@ -30,6 +30,33 @@ h5py's SWMR reader). A file still being written is read with
|
||||
`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file
|
||||
at its length at open and is not meant for files that change while open.
|
||||
|
||||
## Shrinking a chunked dataset with no recorded maximum scrambled it
|
||||
|
||||
**Status:** fixed 2026-09-27, before any release (`FileEditor::resize`
|
||||
shipped on main in PR #18, a4c2ace). Files the writer produced before the
|
||||
fix still lack the maximum; see *Existing files*.
|
||||
|
||||
clawhdf5's writer stored no maximum dimensions for a chunked dataset
|
||||
created without a `maxshape` (libhdf5 always stores one, equal to the
|
||||
dimensions when none is given). With none recorded, the maximum is the
|
||||
current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index
|
||||
places chunks by the maximum. `FileEditor::resize` changed only the
|
||||
current dimensions, so a shrink moved every existing chunk and every
|
||||
reader returned wrong values; after a shrink the dataset could not grow
|
||||
back. Found by the review of the Python editing work.
|
||||
|
||||
**Fix:** before the first resize of such a dataset the editor records the
|
||||
maximum libhdf5 would have written (the dimensions the index was built
|
||||
with); the writer now records it for every chunked dataset. **Test:**
|
||||
`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:**
|
||||
datasets written before the fix have no recorded maximum. The fixed
|
||||
editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`)
|
||||
does not** — it scrambles them the same way, and lets them grow past
|
||||
their Fixed Array. Resize them once with a fixed `FileEditor` (a resize
|
||||
to the same shape changes nothing; shrink and grow back) before letting
|
||||
libhdf5 resize them. A dataset already shrunk by the unfixed
|
||||
editor holds misplaced chunks; rewrite it from a good copy.
|
||||
|
||||
## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768
|
||||
|
||||
**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to
|
||||
@@ -139,6 +166,57 @@ space it leaves is too small for its next, larger version.
|
||||
**No journal.** A crash while an edit patches existing structures can leave
|
||||
the file inconsistent; see the `FileEditor` documentation.
|
||||
|
||||
**Renamed files outside Linux.** Edits always go to the file the editor
|
||||
opened. `FileEditor::reader()` (which the Python `'r+'` handle reads
|
||||
through) reopens it through `/proc/self/fd` on Linux; elsewhere it reopens
|
||||
the path, and after the path was renamed or replaced it fails (on Unix;
|
||||
Windows cannot tell and would read whatever the path names).
|
||||
|
||||
## Python in-place editing (`clawhdf5.File(path, 'r+')`) limits
|
||||
|
||||
**Status:** open (added 2026-09-27). The Python bindings edit through
|
||||
`FileEditor`, so its limits above apply, raised as `NotImplementedError`
|
||||
before anything is written. On top of them:
|
||||
|
||||
- **No new or deleted objects:** `create_dataset`/`create_group` in an
|
||||
`'r+'` file, `del f[name]` and `del obj.attrs[name]` raise
|
||||
`NotImplementedError` (the editor changes values, shapes and
|
||||
attributes only). Mode `'a'` works on an existing file only.
|
||||
- **Not writable from Python:** compound fields by name (`ds['x'] = …`;
|
||||
whole elements of the same structured dtype are), HDF5 array-type
|
||||
elements, variable-length data, strings padded with spaces or
|
||||
NUL-terminated (libhdf5 converts those differently from numpy; NUL-padded
|
||||
ones, h5py's, are writable), compounds containing such strings, null
|
||||
dataspaces, index-list writes of more than 2²² elements (write them
|
||||
in slices), and boolean-mask keys (`ds[mask] = v`, and mask reads;
|
||||
h5py supports both).
|
||||
- **`str` attributes are fixed-length UTF-8**, where h5py writes
|
||||
variable-length strings: h5py reads them back as `bytes`
|
||||
(`numpy.bytes_`), not `str`.
|
||||
- **Numeric conversion follows libhdf5's native-order results, not its
|
||||
bugs.** Arrays are converted as libhdf5 converts them (integers
|
||||
saturate, floats are truncated toward zero and clipped), checked value by
|
||||
value against h5py 3.16 / HDF5 2.0 in
|
||||
`crates/clawhdf5-py/tests/test_edit.py::test_numeric_conversions_match_h5py`
|
||||
(2026-09-27, tank). Where libhdf5 itself is inconsistent, clawhdf5
|
||||
differs from h5py on purpose:
|
||||
- NaN into an integer dataset is a `ValueError` (libhdf5 stores 0, the
|
||||
minimum or 2⁶³ depending on the type);
|
||||
- when the dataset or the array is not in native byte order, libhdf5's
|
||||
"soft" conversions store a float in (-1, 0) as the integer minimum and
|
||||
wrap an unsigned value too large for the signed type of the same size
|
||||
(65535 → -1); clawhdf5 gives 0 and the maximum, as libhdf5 does in
|
||||
native order;
|
||||
- libhdf5's native casts that are undefined in C: half floats into
|
||||
unsigned integers (-1 → 65535, +inf → 0), half-float ±inf into signed
|
||||
integers (→ minimum), a float equal to the integer maximum rounded up
|
||||
in its precision (`float32(2**31 - 1)` into `int32`, `float64(2**64 -
|
||||
1)` into `uint64`: → minimum or 0); clawhdf5 saturates;
|
||||
- a double in (65504, 65520) into a half float: libhdf5 stores infinity,
|
||||
clawhdf5 (numpy) rounds to 65504 as IEEE 754 does.
|
||||
- **Each edit reopens the file** (a new memory map) so that reads see it;
|
||||
reads from other threads wait while an edit is written.
|
||||
|
||||
## Selection reads that decode more than the selection
|
||||
|
||||
**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so
|
||||
@@ -867,9 +945,9 @@ cache, but:
|
||||
`read_*_zerocopy`) need the file in memory and answer
|
||||
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
|
||||
panics for such a file (`File::contiguous_bytes()` is the fallible form).
|
||||
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
|
||||
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
|
||||
`remote` feature).
|
||||
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
|
||||
(`h5rs` and the Python bindings read through `File::storage`, and take
|
||||
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
|
||||
- The file's length is read once, at open: a growing file (SWMR) is not
|
||||
followed (milestone M5). A remote file is pinned at open, so one that
|
||||
grows is `RemoteError::FileChanged`.
|
||||
@@ -885,9 +963,12 @@ cache, but:
|
||||
**Status:** open (added 2026-09-26, milestone M3 of
|
||||
`docs/design/range-reads.md`).
|
||||
|
||||
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
|
||||
parses through `File::as_bytes`, which a remote file does not have; the
|
||||
wasm reader's `openUrl` is milestone M4.
|
||||
- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is
|
||||
milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but
|
||||
the default wheel reads plain `http://` only: `https://` needs a wheel
|
||||
built with `--features https` (rustls with ring, which compiles C), and
|
||||
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
|
||||
The Python tests run against an in-process `http.server` only.
|
||||
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
||||
The design's policy of using a paged file's page size as the block size
|
||||
is not implemented, and only the first block is read ahead.
|
||||
|
||||
Reference in New Issue
Block a user