docs: range-read M4 (openUrl in the browser): changelog, limits, design status
CHANGELOG (Unreleased), known-issues (the browser can open URLs; the limits of openUrl: round trips per wave of misses, a call holds what it reads, CORS and validator visibility, the download fallback), the M4 status in docs/design/range-reads.md (why NeedBytes rather than a Worker, why not clawhdf5-remote's BlockCache, what was tested; the status paragraph at the top lost a garbled duplicate), the viewer's README (API, options, how it works, tests; the size table is marked as predating openUrl), CLAUDE.md and the README crate list. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
+61
-10
@@ -7,20 +7,19 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
|
||||
`File::open_storage` gives the facade's read API over any `Storage` (see
|
||||
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
|
||||
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
|
||||
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is
|
||||
next. Every count in §1–§2 was
|
||||
|
||||
change. Progress: M1, first part (the `Storage` trait and the metadata
|
||||
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
||||
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
||||
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
|
||||
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
|
||||
reader, through the restartable `NeedBytes` mode (see the M4 status
|
||||
below). M5 (SWMR) is not started.
|
||||
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||
converted code without changing the plan: object-header continuation chunks
|
||||
are followed without recursion (still one bounded `read_at` per chunk), and
|
||||
implicit chunk indexes are addressed over the maximum chunk grid (in
|
||||
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
|
||||
working on the whole file in memory; it is not part of this design. Every count below was
|
||||
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
|
||||
next to it. No timing numbers appear here on purpose: the machine was shared
|
||||
working on the whole file in memory; it is not part of this design. Every
|
||||
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
|
||||
every count in a milestone's status on the date it gives, with the
|
||||
commands given next to it. No timing numbers appear here on purpose: the machine was shared
|
||||
with other build jobs when this was written.
|
||||
|
||||
## The problem
|
||||
@@ -519,6 +518,58 @@ fast path within benchmark noise.
|
||||
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
|
||||
to a whole download when the server does not answer 206.
|
||||
- `examples/wasm-viewer`: open by URL.
|
||||
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
|
||||
with these choices:
|
||||
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
|
||||
`Storage` over the blocks fetched so far. A call (open, list, read)
|
||||
runs as a pass; a read that misses records its blocks and fails with
|
||||
a storage error. The error itself is not the signal: parsers catch
|
||||
errors and carry on (a listing leaves out a link it cannot resolve),
|
||||
so *any* pass that recorded a miss is thrown away, whatever it
|
||||
returned, and re-run once the missing ranges have been fetched
|
||||
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
|
||||
synchronous XHR would have kept the parsers' single pass, but it
|
||||
needs the page to run the reader off the main thread (a second
|
||||
module, messages for every call and every typed array copied back),
|
||||
since synchronous XHR can return binary data only there; the
|
||||
restartable loop costs a
|
||||
re-parse per wave of misses instead, which is CPU, not network.
|
||||
- **Progress:** no block is evicted while a call is in flight
|
||||
(`LazyStorage::operation`), so each pass that does not finish asks for
|
||||
at least one new block, and a call ends after at most one pass per
|
||||
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
|
||||
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
|
||||
evicted before metadata. The price: a call holds everything it reads
|
||||
until it finishes.
|
||||
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
|
||||
backend by blocking, evicts during a read, and does not keep reads
|
||||
that miss more than half its budget; a restartable pass needs every
|
||||
block it has read to be there when it is re-run. The block
|
||||
arithmetic and coalescing (runs of consecutive blocks, a one-block
|
||||
hole filled, at most 8 MiB per request) follow it; the cache is
|
||||
about 450 lines (with its documentation) in the wasm crate, with no
|
||||
in-flight tracking or HTTP.
|
||||
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
|
||||
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
|
||||
requests at a time, every answer checked (206, `Content-Range` when
|
||||
visible, body length, the ETag/Last-Modified and length of the
|
||||
first answer). A `200` to the first request is kept as the whole
|
||||
file up to `maxDownload` (512 MiB), or refused with
|
||||
`fallback: "error"`. The first block comes with the request that
|
||||
learns the length, as in M3.
|
||||
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
|
||||
--test lazy` compares what the viewer shows of each file (kinds,
|
||||
listings, attributes, info, whole reads, hyperslabs) read lazily at
|
||||
512 B to 1 MiB blocks with the in-memory reader and the facade's
|
||||
range-storage path, for built files, the h5py/netCDF4 fixture and,
|
||||
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
|
||||
files. The built package, under Node and headless Chromium
|
||||
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
|
||||
`test/serve.py`, which counts requests): every fixture check over
|
||||
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
|
||||
attributes and a 10-value window of the 25-million-value dataset
|
||||
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
|
||||
with `open(bytes)`.
|
||||
|
||||
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
||||
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
||||
|
||||
+37
-9
@@ -844,9 +844,10 @@ cache, but:
|
||||
`read_*_zerocopy`) need the file in memory and answer
|
||||
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
|
||||
panics for such a file (`File::contiguous_bytes()` is the fallible form).
|
||||
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
|
||||
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
|
||||
`remote` feature).
|
||||
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
|
||||
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
|
||||
feature; the wasm reader's `openUrl` reads by range requests since
|
||||
2026-09-27, its `open(bytes)` takes a whole file).
|
||||
- The file's length is read once, at open: a growing file (SWMR) is not
|
||||
followed (milestone M5). A remote file is pinned at open, so one that
|
||||
grows is `RemoteError::FileChanged`.
|
||||
@@ -862,9 +863,9 @@ cache, but:
|
||||
**Status:** open (added 2026-09-26, milestone M3 of
|
||||
`docs/design/range-reads.md`).
|
||||
|
||||
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
|
||||
parses through `File::as_bytes`, which a remote file does not have; the
|
||||
wasm reader's `openUrl` is milestone M4.
|
||||
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
|
||||
`File::as_bytes`, which a remote file does not have. (The browser can
|
||||
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
|
||||
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
|
||||
The design's policy of using a paged file's page size as the block size
|
||||
is not implemented, and only the first block is read ahead.
|
||||
@@ -903,10 +904,37 @@ cache, but:
|
||||
|
||||
## `clawhdf5-wasm` (browser) limits
|
||||
|
||||
**Status:** open (by design for now; added 2026-09-26).
|
||||
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
|
||||
|
||||
- The whole file is held in memory: `open()` takes its bytes. There are no
|
||||
HTTP range reads, so a multi-GB file does not fit a browser tab.
|
||||
- `open()` holds the whole file in memory (it takes its bytes), so a
|
||||
multi-GB local file does not fit a browser tab. A file on a web server
|
||||
can be opened with `openUrl()` instead, which fetches only the byte
|
||||
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
|
||||
with these limits:
|
||||
- **Round trips:** a call runs as passes over the blocks fetched so far
|
||||
and is re-run after each wave of misses, so a call costs one round
|
||||
trip per wave, not one for everything: a chunk index is walked a
|
||||
level (or a node) per round trip, while the chunks of a read are
|
||||
fetched together. Each pass re-parses what the call reads (CPU, not
|
||||
network).
|
||||
- **Memory:** a call keeps every block it reads until it finishes (the
|
||||
cache budget applies between calls), so reading a large dataset whole
|
||||
needs its stored bytes in memory next to the result; read windows
|
||||
(`readHyperslab`) of large datasets. Files above 200 MB have not been
|
||||
tested, nor offsets above 4 GiB on wasm32.
|
||||
- **Cross-origin servers** must allow CORS for the page's origin and
|
||||
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
|
||||
answer `HEAD` with `Content-Length`. The file is pinned at open by its
|
||||
`ETag` or `Last-Modified` (and its length): when the page cannot see
|
||||
either header, only a change of length is detected.
|
||||
- **A server without range support** (it answers `200`) costs a whole
|
||||
download, up to `maxDownload` (512 MiB), or an error with
|
||||
`fallback: "error"`.
|
||||
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
|
||||
size is not used. No retries: a failed request fails the call (calling
|
||||
again retries it; what was fetched stays cached).
|
||||
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
||||
a local server; not in Firefox or Safari.
|
||||
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
||||
refused with an error naming the type; attributes of those types come back
|
||||
as `value: null` with their `dtype`.
|
||||
|
||||
Reference in New Issue
Block a user