Merge branch 'feat/p3-m4-wasm-lazy' into feat/p3-wasm-swmr-python

# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
This commit is contained in:
osobh
2026-09-27 07:59:14 -05:00
24 changed files with 3752 additions and 293 deletions
+68 -3
View File
@@ -948,6 +948,11 @@ cache, but:
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
(`h5rs` and the Python bindings read through `File::storage`, and take
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
feature; the wasm reader's `openUrl` reads by range requests since
2026-09-27, its `open(bytes)` takes a whole file).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -969,6 +974,10 @@ cache, but:
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
`File::as_bytes`, which a remote file does not have. (The browser can
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.
@@ -1007,10 +1016,66 @@ cache, but:
## `clawhdf5-wasm` (browser) limits
**Status:** open (by design for now; added 2026-09-26).
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
- The whole file is held in memory: `open()` takes its bytes. There are no
HTTP range reads, so a multi-GB file does not fit a browser tab.
- `open()` holds the whole file in memory (it takes its bytes), so a
multi-GB local file does not fit a browser tab. A file on a web server
can be opened with `openUrl()` instead, which fetches only the byte
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
with these limits:
- **Round trips:** a call runs as passes over the blocks fetched so far
and is re-run after each wave of misses, so a call costs one round
trip per wave, not one for everything: a chunk index is walked a
level (or a node) per round trip, while the chunks of a read are
fetched together. Listing a group asks for every child's object
header, and every node of a level of the group's index, in one pass
(since 2026-09-27; it was one round trip per header block): 3000
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
`libver="latest"` file (dense links). Each pass re-parses what the
call reads (CPU, not network). With headers spread through the file
(h5py writes each next to its data) a listing still fetches most of
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
- **Memory:** a call keeps every block it reads until it finishes (the
cache budget applies between calls). It may fetch at most `maxFetch`
bytes (512 MiB by default, at most 1 GiB), and a single read longer
than that fails before anything is fetched: a hostile server cannot
make the page fetch or allocate what a file's lengths claim. `read()`
of a dataset that would take more than 1 GiB while it is decoded
(stored bytes, the values widened to 64 bits, the result) fails,
naming `readHyperslab`; read windows of large datasets. On wasm32 a
buffer past 2 GiB cannot exist and a failed allocation aborts the
whole module (every open file on the page), which these limits keep
from happening; before 2026-09-27 both did abort it.
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
open. The format code turns file offsets into `usize` to use them
(with a clean error past it), which is 32 bits on wasm32, so nothing
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
tested (with a mock server); files above 200 MB have not been served
for real.
- **Cross-origin servers** must allow CORS for the page's origin and
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
answer `HEAD` with `Content-Length`. The file is pinned at open by its
`ETag` or `Last-Modified` (and its length): when the page cannot see
either header, only a change of length is detected.
- **A server without range support** (it answers `200`) costs a whole
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
with `fallback: "error"`. Every body, this one and each `206`, is read
as it arrives and cut off past its limit (the range asked for, or
`maxDownload`): a server cannot make the page buffer more.
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
size is not used. No retries: a failed request fails the call (calling
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included (a page on 127.0.0.1 reading a
file from localhost, with and without exposed headers); not in Firefox
or Safari.
- The native corpus comparison (`tests/lazy.rs` with
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
and which chunk's error is reported depends on the iteration order of
the chunk index (a `HashMap`, seeded per process), so the lazy and
the range-storage reads can name different errors. Both are errors;
not specific to `openUrl` (it predates it).
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`.