Merge branch 'feat/p3-m4-wasm-lazy' into feat/p3-wasm-swmr-python

# Conflicts:
#	CHANGELOG.md
#	docs/design/range-reads.md
#	docs/known-issues.md
This commit is contained in:
osobh
2026-09-27 07:59:14 -05:00
24 changed files with 3752 additions and 293 deletions
+74 -3
View File
@@ -20,13 +20,20 @@ change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet.
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
reader, through the restartable `NeedBytes` mode (see the M4 status
below). M5 (SWMR) is not started.
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
converted code without changing the plan: object-header continuation chunks
are followed without recursion (still one bounded `read_at` per chunk), and
implicit chunk indexes are addressed over the maximum chunk grid (in
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
working on the whole file in memory; it is not part of this design. Every count below was
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
next to it. No timing numbers appear here on purpose: the machine was shared
working on the whole file in memory; it is not part of this design. Every
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
every count in a milestone's status on the date it gives, with the
commands given next to it. No timing numbers appear here on purpose: the machine was shared
with other build jobs when this was written.
## The problem
@@ -536,6 +543,70 @@ fast path within benchmark noise.
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
to a whole download when the server does not answer 206.
- `examples/wasm-viewer`: open by URL.
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
with these choices:
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
`Storage` over the blocks fetched so far. A call (open, list, read)
runs as a pass; a read that misses records its blocks and fails with
a storage error. The error itself is not the signal: parsers catch
errors and carry on (a listing leaves out a link it cannot resolve),
so *any* pass that recorded a miss is thrown away, whatever it
returned, and re-run once the missing ranges have been fetched
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
synchronous XHR would have kept the parsers' single pass, but it
needs the page to run the reader off the main thread (a second
module, messages for every call and every typed array copied back),
since synchronous XHR can return binary data only there; the
restartable loop costs a
re-parse per wave of misses instead, which is CPU, not network.
- **Progress:** no block is evicted while a call is in flight
(`LazyStorage::operation`), so each pass that does not finish asks for
at least one new block, and a call ends after at most one pass per
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
evicted before metadata. The price: a call holds everything it reads
until it finishes.
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
backend by blocking, evicts during a read, and does not keep reads
that miss more than half its budget; a restartable pass needs every
block it has read to be there when it is re-run. The block
arithmetic and coalescing (runs of consecutive blocks, a one-block
hole filled, at most 8 MiB per request) follow it; the cache is
about 450 lines (with its documentation) in the wasm crate, with no
in-flight tracking or HTTP. (A hole already cached is not filled:
it would be fetched again.)
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
requests at a time, every answer checked (206, `Content-Range` when
visible, body length, the ETag/Last-Modified and length of the
first answer). A `200` to the first request is kept as the whole
file up to `maxDownload` (512 MiB), or refused with
`fallback: "error"`. The first block comes with the request that
learns the length, as in M3.
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
--test lazy` compares what the viewer shows of each file (kinds,
listings, attributes, info, whole reads, hyperslabs) read lazily at
512 B to 1 MiB blocks with the in-memory reader and the facade's
range-storage path, for built files, the h5py/netCDF4 fixture and,
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
files. The built package, under Node and headless Chromium
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
`test/serve.py`, which counts requests): every fixture check over
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
attributes and a 10-value window of the 25-million-value dataset
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
with `open(bytes)`.
- After review (2026-09-27): a call may fetch at most `maxFetch`
(512 MiB, at most 1 GiB) and a longer single read is refused before
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
more is refused at open on wasm32 (offsets become `usize`); every
response body is cut off at the length asked for. Parsers that walk
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
the wasm listing's child headers) read the siblings after the first
failure before returning it, so a pass asks for a whole level's
missing blocks: listing 3000 datasets went from 185 passes to 6.
The 32-bit risk below is covered by a Node test that reads data at
3 GiB from a mock server and is refused a 4 GiB file.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
+68 -3
View File
@@ -948,6 +948,11 @@ cache, but:
`LazyFile`, `MmapFile` and the wasm bindings still read a whole file
(`h5rs` and the Python bindings read through `File::storage`, and take
URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`).
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
feature; the wasm reader's `openUrl` reads by range requests since
2026-09-27, its `open(bytes)` takes a whole file).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -969,6 +974,10 @@ cache, but:
built with `--features https` (rustls with ring, which compiles C), and
`s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs).
The Python tests run against an in-process `http.server` only.
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
`File::as_bytes`, which a remote file does not have. (The browser can
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.
@@ -1007,10 +1016,66 @@ cache, but:
## `clawhdf5-wasm` (browser) limits
**Status:** open (by design for now; added 2026-09-26).
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
- The whole file is held in memory: `open()` takes its bytes. There are no
HTTP range reads, so a multi-GB file does not fit a browser tab.
- `open()` holds the whole file in memory (it takes its bytes), so a
multi-GB local file does not fit a browser tab. A file on a web server
can be opened with `openUrl()` instead, which fetches only the byte
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
with these limits:
- **Round trips:** a call runs as passes over the blocks fetched so far
and is re-run after each wave of misses, so a call costs one round
trip per wave, not one for everything: a chunk index is walked a
level (or a node) per round trip, while the chunks of a read are
fetched together. Listing a group asks for every child's object
header, and every node of a level of the group's index, in one pass
(since 2026-09-27; it was one round trip per header block): 3000
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
`libver="latest"` file (dense links). Each pass re-parses what the
call reads (CPU, not network). With headers spread through the file
(h5py writes each next to its data) a listing still fetches most of
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
- **Memory:** a call keeps every block it reads until it finishes (the
cache budget applies between calls). It may fetch at most `maxFetch`
bytes (512 MiB by default, at most 1 GiB), and a single read longer
than that fails before anything is fetched: a hostile server cannot
make the page fetch or allocate what a file's lengths claim. `read()`
of a dataset that would take more than 1 GiB while it is decoded
(stored bytes, the values widened to 64 bits, the result) fails,
naming `readHyperslab`; read windows of large datasets. On wasm32 a
buffer past 2 GiB cannot exist and a failed allocation aborts the
whole module (every open file on the page), which these limits keep
from happening; before 2026-09-27 both did abort it.
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
open. The format code turns file offsets into `usize` to use them
(with a clean error past it), which is 32 bits on wasm32, so nothing
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
tested (with a mock server); files above 200 MB have not been served
for real.
- **Cross-origin servers** must allow CORS for the page's origin and
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
answer `HEAD` with `Content-Length`. The file is pinned at open by its
`ETag` or `Last-Modified` (and its length): when the page cannot see
either header, only a change of length is detected.
- **A server without range support** (it answers `200`) costs a whole
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
with `fallback: "error"`. Every body, this one and each `206`, is read
as it arrives and cut off past its limit (the range asked for, or
`maxDownload`): a server cannot make the page buffer more.
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
size is not used. No retries: a failed request fails the call (calling
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included (a page on 127.0.0.1 reading a
file from localhost, with and without exposed headers); not in Firefox
or Safari.
- The native corpus comparison (`tests/lazy.rs` with
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
and which chunk's error is reported depends on the iteration order of
the chunk index (a `HashMap`, seeded per process), so the lazy and
the range-storage reads can name different errors. Both are errors;
not specific to `openUrl` (it predates it).
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`.