docs: range-read M4 (openUrl in the browser): changelog, limits, design status

CHANGELOG (Unreleased), known-issues (the browser can open URLs; the
limits of openUrl: round trips per wave of misses, a call holds what it
reads, CORS and validator visibility, the download fallback), the M4
status in docs/design/range-reads.md (why NeedBytes rather than a
Worker, why not clawhdf5-remote's BlockCache, what was tested; the
status paragraph at the top lost a garbled duplicate), the viewer's
README (API, options, how it works, tests; the size table is marked as
predating openUrl), CLAUDE.md and the README crate list.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:47:29 -05:00
co-authored by Claude Opus 5.5
parent b8c85f7627
commit f825a89e23
7 changed files with 245 additions and 49 deletions
+61 -10
View File
@@ -7,20 +7,19 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
`File::open_storage` gives the facade's read API over any `Storage` (see
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is
next. Every count in §1–§2 was
change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet.
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
reader, through the restartable `NeedBytes` mode (see the M4 status
below). M5 (SWMR) is not started.
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
converted code without changing the plan: object-header continuation chunks
are followed without recursion (still one bounded `read_at` per chunk), and
implicit chunk indexes are addressed over the maximum chunk grid (in
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
working on the whole file in memory; it is not part of this design. Every count below was
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
next to it. No timing numbers appear here on purpose: the machine was shared
working on the whole file in memory; it is not part of this design. Every
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
every count in a milestone's status on the date it gives, with the
commands given next to it. No timing numbers appear here on purpose: the machine was shared
with other build jobs when this was written.
## The problem
@@ -519,6 +518,58 @@ fast path within benchmark noise.
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
to a whole download when the server does not answer 206.
- `examples/wasm-viewer`: open by URL.
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
with these choices:
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
`Storage` over the blocks fetched so far. A call (open, list, read)
runs as a pass; a read that misses records its blocks and fails with
a storage error. The error itself is not the signal: parsers catch
errors and carry on (a listing leaves out a link it cannot resolve),
so *any* pass that recorded a miss is thrown away, whatever it
returned, and re-run once the missing ranges have been fetched
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
synchronous XHR would have kept the parsers' single pass, but it
needs the page to run the reader off the main thread (a second
module, messages for every call and every typed array copied back),
since synchronous XHR can return binary data only there; the
restartable loop costs a
re-parse per wave of misses instead, which is CPU, not network.
- **Progress:** no block is evicted while a call is in flight
(`LazyStorage::operation`), so each pass that does not finish asks for
at least one new block, and a call ends after at most one pass per
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
evicted before metadata. The price: a call holds everything it reads
until it finishes.
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
backend by blocking, evicts during a read, and does not keep reads
that miss more than half its budget; a restartable pass needs every
block it has read to be there when it is re-run. The block
arithmetic and coalescing (runs of consecutive blocks, a one-block
hole filled, at most 8 MiB per request) follow it; the cache is
about 450 lines (with its documentation) in the wasm crate, with no
in-flight tracking or HTTP.
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
requests at a time, every answer checked (206, `Content-Range` when
visible, body length, the ETag/Last-Modified and length of the
first answer). A `200` to the first request is kept as the whole
file up to `maxDownload` (512 MiB), or refused with
`fallback: "error"`. The first block comes with the request that
learns the length, as in M3.
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
--test lazy` compares what the viewer shows of each file (kinds,
listings, attributes, info, whole reads, hyperslabs) read lazily at
512 B to 1 MiB blocks with the in-memory reader and the facade's
range-storage path, for built files, the h5py/netCDF4 fixture and,
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
files. The built package, under Node and headless Chromium
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
`test/serve.py`, which counts requests): every fixture check over
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
attributes and a 10-value window of the 25-million-value dataset
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
with `open(bytes)`.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
+37 -9
View File
@@ -844,9 +844,10 @@ cache, but:
`read_*_zerocopy`) need the file in memory and answer
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
panics for such a file (`File::contiguous_bytes()` is the fallible form).
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
`remote` feature).
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
feature; the wasm reader's `openUrl` reads by range requests since
2026-09-27, its `open(bytes)` takes a whole file).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -862,9 +863,9 @@ cache, but:
**Status:** open (added 2026-09-26, milestone M3 of
`docs/design/range-reads.md`).
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
parses through `File::as_bytes`, which a remote file does not have; the
wasm reader's `openUrl` is milestone M4.
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
`File::as_bytes`, which a remote file does not have. (The browser can
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.
@@ -903,10 +904,37 @@ cache, but:
## `clawhdf5-wasm` (browser) limits
**Status:** open (by design for now; added 2026-09-26).
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
- The whole file is held in memory: `open()` takes its bytes. There are no
HTTP range reads, so a multi-GB file does not fit a browser tab.
- `open()` holds the whole file in memory (it takes its bytes), so a
multi-GB local file does not fit a browser tab. A file on a web server
can be opened with `openUrl()` instead, which fetches only the byte
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
with these limits:
- **Round trips:** a call runs as passes over the blocks fetched so far
and is re-run after each wave of misses, so a call costs one round
trip per wave, not one for everything: a chunk index is walked a
level (or a node) per round trip, while the chunks of a read are
fetched together. Each pass re-parses what the call reads (CPU, not
network).
- **Memory:** a call keeps every block it reads until it finishes (the
cache budget applies between calls), so reading a large dataset whole
needs its stored bytes in memory next to the result; read windows
(`readHyperslab`) of large datasets. Files above 200 MB have not been
tested, nor offsets above 4 GiB on wasm32.
- **Cross-origin servers** must allow CORS for the page's origin and
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
answer `HEAD` with `Content-Length`. The file is pinned at open by its
`ETag` or `Last-Modified` (and its length): when the page cannot see
either header, only a change of length is detected.
- **A server without range support** (it answers `200`) costs a whole
download, up to `maxDownload` (512 MiB), or an error with
`fallback: "error"`.
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
size is not used. No retries: a failed request fails the call (calling
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server; not in Firefox or Safari.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`.