Merge branch 'feat/p3-m4-wasm-lazy' into feat/p3-wasm-swmr-python

# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
This commit is contained in:
osobh
2026-09-27 07:59:14 -05:00
24 changed files with 3752 additions and 293 deletions
+81
View File
@@ -214,6 +214,87 @@ Design: `docs/design/swmr.md`.
that hangs up, 16 threads on one remote file, and a thread that keeps
running while a read waits on 0.2 s requests.
### Range reads, milestone M4: remote files in the browser (2026-09-27)
- **`openUrl(url, opts)`** in `clawhdf5-wasm` opens an HDF5/NetCDF-4 file
on a web server without downloading it: the returned `RemoteFile` has
`H5File`'s methods (`kind`, `list`, `info`, `attrs`, `attrErrors`,
`read`, `readHyperslab`), each returning a promise, and `stats()`
(requests, bytes fetched, file size). Only the byte ranges a call needs
are fetched, with `fetch` and `Range` headers, in 1 MiB blocks
(`blockSize`) kept in a 64 MiB cache (`cacheSize`). `open(bytes)` is
unchanged.
- **How, on the browser's main thread:** the design's restartable
"NeedBytes" mode (`clawhdf5_wasm::lazy::LazyStorage`). A call runs as a
pass over the blocks fetched so far; a read that misses records the
missing blocks and fails, the pass's result is dropped (even if a parser
caught the error and carried on), the blocks are fetched and the pass is
re-run. No Web Worker and no synchronous XHR (what h5wasm's lazy files
need). No block is evicted while a call is in flight, so every call ends;
the cache is trimmed between calls, raw-data blocks before metadata.
`clawhdf5-remote`'s `BlockCache` is not reused: it fetches by blocking,
and it evicts, and does not keep large reads, during a read, where a
restartable pass needs every block it read to stay until it finishes.
- **Every answer is checked** (`js/remote.js`): a `206` with exactly the
bytes asked for (`Content-Range`, when the page can see it, and the body
length), and the same `ETag`/`Last-Modified` and length as at open, or
the call fails: never data from another file or offset. A server that
answers `200` to a `Range` request is downloaded whole (up to
`maxDownload`, 512 MiB) unless `fallback: "error"`. Other options:
`headers`, `credentials`, `parallel` (6 requests at a time), `fetch`.
- **The viewer** (`examples/wasm-viewer`) has a URL box, opens
`?file=<url>` lazily, and shows the requests and bytes fetched.
- Counted on tank, 2026-09-27 (`WASM_BIG_MB=200 bash
examples/wasm-viewer/test/run.sh`, requests as `test/serve.py` counted
them): in a 200 MB h5py file, listing the root, reading two small
datasets, a group's attributes, the large dataset's shape and a
10-value window of it (25 million values) took 5 requests and 6 MiB.
With
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, the 622 corpus files
up to 16 MiB that open read over HTTP as from bytes: the same listings,
attributes and values (datasets up to 2^20 values), and an error
wherever bytes give one. Natively, `cargo test -p clawhdf5-wasm --test
lazy` with the same variable compares 656 files (up to 64 MiB, error
messages included, against the facade's range-storage path).
- `Reader::open_storage` (the wasm crate's core over any `Storage`), and
variable-length strings resolve through the file's storage rather than
`File::as_bytes`.
- The package is larger: the facade's `Storage` read path is now
reachable from JavaScript (it was compiled out before), and the promise
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
build machine was shared); the viewer README's size table predates M4.
- **Hardened after review (2026-09-27):**
- Sizes a server or a dataset names are errors, never an abort of the
wasm module (which took every open file on the page with it): a read
past 2 GiB aborted in the lazy cache, reachable by a hostile server
claiming a large file and a 2 GiB heap collection, and `read()` of a
256 MiB `u8` dataset aborted widening it to 64 bits. New option
`maxFetch` (512 MiB, at most 1 GiB): what one call may fetch, and the
longest single read, refused before fetching. `read()` refuses a
dataset that would take more than 1 GiB to decode, naming
`readHyperslab`. A file of 4 GiB or more is refused at open (wasm32
reads offsets as 32-bit); `maxDownload` is at most 1 GiB.
- Response bodies are read as they arrive and cut off at the length
asked for (`maxDownload` for a `200`): a `206` with a gigabyte body
was buffered whole before its length was checked.
- Listing a group reads every child's header, and every node of each
level of the group's index, in one pass: 3000 datasets (h5py, 198 MB)
listed in 6 passes and 73 requests at 1 MiB blocks instead of 185
passes and 184 serial requests (`libver="latest"`: 9 passes instead of
189). In `clawhdf5-format`, the B-tree v1/v2 collectors, the symbol
table node loop and the dense-link loop read (without using) the
siblings after the first that fails, then return that error: same
results and errors, more reads only on failure (free in memory). The
lazy cache no longer re-fetches a cached block to merge two requests.
- `headers` may be a `Headers` instance or `[name, value]` pairs (a
`Headers` was silently dropped); when one range request fails the
others in flight are aborted; `parallel` must be an integer from 1 to
1024.
- Tests: `test/serve.py` serves ranges without exposed `Content-Range`/
`ETag` (`/noexpose/`, and `/unexposed/` for a real cross-origin page
in Chromium), so the HEAD-length path runs end to end; hostile and
oversized files (`make_fixture.py`'s `write_limits`), flooding
bodies, aborted siblings.
### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by