diff --git a/CHANGELOG.md b/CHANGELOG.md index 664a831..bb41c2b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,55 @@ ## Unreleased +### Range reads, milestone M4: remote files in the browser (2026-09-27) +- **`openUrl(url, opts)`** in `clawhdf5-wasm` opens an HDF5/NetCDF-4 file + on a web server without downloading it: the returned `RemoteFile` has + `H5File`'s methods (`kind`, `list`, `info`, `attrs`, `attrErrors`, + `read`, `readHyperslab`), each returning a promise, and `stats()` + (requests, bytes fetched, file size). Only the byte ranges a call needs + are fetched, with `fetch` and `Range` headers, in 1 MiB blocks + (`blockSize`) kept in a 64 MiB cache (`cacheSize`). `open(bytes)` is + unchanged. +- **How, on the browser's main thread:** the design's restartable + "NeedBytes" mode (`clawhdf5_wasm::lazy::LazyStorage`). A call runs as a + pass over the blocks fetched so far; a read that misses records the + missing blocks and fails, the pass's result is dropped (even if a parser + caught the error and carried on), the blocks are fetched and the pass is + re-run. No Web Worker and no synchronous XHR (what h5wasm's lazy files + need). No block is evicted while a call is in flight, so every call ends; + the cache is trimmed between calls, raw-data blocks before metadata. + `clawhdf5-remote`'s `BlockCache` is not reused: it fetches by blocking, + and it evicts, and does not keep large reads, during a read, where a + restartable pass needs every block it read to stay until it finishes. +- **Every answer is checked** (`js/remote.js`): a `206` with exactly the + bytes asked for (`Content-Range`, when the page can see it, and the body + length), and the same `ETag`/`Last-Modified` and length as at open, or + the call fails: never data from another file or offset. A server that + answers `200` to a `Range` request is downloaded whole (up to + `maxDownload`, 512 MiB) unless `fallback: "error"`. Other options: + `headers`, `credentials`, `parallel` (6 requests at a time), `fetch`. +- **The viewer** (`examples/wasm-viewer`) has a URL box, opens + `?file=` lazily, and shows the requests and bytes fetched. +- Counted on tank, 2026-09-27 (`WASM_BIG_MB=200 bash + examples/wasm-viewer/test/run.sh`, requests as `test/serve.py` counted + them): in a 200 MB h5py file, listing the root, reading two small + datasets, a group's attributes, the large dataset's shape and a + 10-value window of it (25 million values) took 5 requests and 6 MiB. + With + `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, the 622 corpus files + up to 16 MiB that open read over HTTP as from bytes: the same listings, + attributes and values (datasets up to 2^20 values), and an error + wherever bytes give one. Natively, `cargo test -p clawhdf5-wasm --test + lazy` with the same variable compares 656 files (up to 64 MiB, error + messages included, against the facade's range-storage path). +- `Reader::open_storage` (the wasm crate's core over any `Storage`), and + variable-length strings resolve through the file's storage rather than + `File::as_bytes`. +- The package is larger: the facade's `Storage` read path is now + reachable from JavaScript (it was compiled out before), and the promise + glue and `remote.js` add JavaScript. Not measured for the docs yet (the + build machine was shared); the viewer README's size table predates M4. + ### Range reads, milestone M3: remote files (2026-09-26) - **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives a `clawhdf5::File` (through `File::open_storage`) that reads the file by diff --git a/CLAUDE.md b/CLAUDE.md index b382f26..8c95cc5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -181,14 +181,19 @@ Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal F `CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus file over HTTP with `File::open`. - GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only -- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only, file held in memory; - no Zstd/SZIP since they link C) and the `examples/wasm-viewer/` page. - `examples/wasm-viewer/test/run.sh` builds the package (needs the - `wasm-bindgen` CLI at the crate's exact version) and tests it under Node - and headless Chromium (a Playwright download in `~/.cache/ms-playwright` - on tank); the CI container has neither, so CI runs the native - `clawhdf5-wasm` `h5py_interop` test on the same fixture. Size numbers are - in the example's README. +- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since + they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds + the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range + requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a + call is re-run after each wave of misses; no block evicted while a call + runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh` + builds the package (needs the `wasm-bindgen` CLI at the crate's exact + version) and tests it under Node and headless Chromium (a Playwright + download in `~/.cache/ms-playwright` on tank) against `test/serve.py` + (range server with request counts, 200 MB budget file); the CI container + has neither, so CI runs the native `h5py_interop` and `lazy` tests + (`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size + numbers in the example's README predate `openUrl`. - Python and Node.js bindings for cross-language use - NetCDF-4 compatibility for scientific data interop diff --git a/README.md b/README.md index 3301b09..621e0dd 100644 --- a/README.md +++ b/README.md @@ -757,7 +757,7 @@ clawhdf5 workspace (19 crates, ~86K lines of Rust in src/, ~104K with tests ├── Bindings │ ├── clawhdf5-py — Python (PyO3) │ ├── clawhdf5-napi — Node.js (napi-rs) -│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only) +│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only; remote files by HTTP range requests) │ └── Tooling ├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check diff --git a/docs/design/range-reads.md b/docs/design/range-reads.md index af44ae2..40170c1 100644 --- a/docs/design/range-reads.md +++ b/docs/design/range-reads.md @@ -7,20 +7,19 @@ change. Progress: M0 and M1 are done, and so is M2 (branch `File::open_storage` gives the facade's read API over any `Storage` (see `CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch `feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S), -object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is -next. Every count in §1–§2 was - -change. Progress: M1, first part (the `Storage` trait and the metadata -parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done; -group B-tree v2 lookups, dense groups and the facade are not converted yet. -Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched +object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on +branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser +reader, through the restartable `NeedBytes` mode (see the M4 status +below). M5 (SWMR) is not started. +Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched converted code without changing the plan: object-header continuation chunks are followed without recursion (still one bounded `read_at` per chunk), and implicit chunk indexes are addressed over the maximum chunk grid (in `chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps -working on the whole file in memory; it is not part of this design. Every count below was -taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given -next to it. No timing numbers appear here on purpose: the machine was shared +working on the whole file in memory; it is not part of this design. Every +count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and +every count in a milestone's status on the date it gives, with the +commands given next to it. No timing numbers appear here on purpose: the machine was shared with other build jobs when this was written. ## The problem @@ -519,6 +518,58 @@ fast path within benchmark noise. Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back to a whole download when the server does not answer 206. - `examples/wasm-viewer`: open by URL. +- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned, + with these choices: + - **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a + `Storage` over the blocks fetched so far. A call (open, list, read) + runs as a pass; a read that misses records its blocks and fails with + a storage error. The error itself is not the signal: parsers catch + errors and carry on (a listing leaves out a link it cannot resolve), + so *any* pass that recorded a miss is thrown away, whatever it + returned, and re-run once the missing ranges have been fetched + (`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with + synchronous XHR would have kept the parsers' single pass, but it + needs the page to run the reader off the main thread (a second + module, messages for every call and every typed array copied back), + since synchronous XHR can return binary data only there; the + restartable loop costs a + re-parse per wave of misses instead, which is CPU, not network. + - **Progress:** no block is evicted while a call is in flight + (`LazyStorage::operation`), so each pass that does not finish asks for + at least one new block, and a call ends after at most one pass per + block it reads. The budget (`cacheSize`, 64 MiB) is applied between + calls, raw-data blocks (`read_ranges`, or reads longer than a block) + evicted before metadata. The price: a call holds everything it reads + until it finishes. + - **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its + backend by blocking, evicts during a read, and does not keep reads + that miss more than half its budget; a restartable pass needs every + block it has read to be there when it is re-run. The block + arithmetic and coalescing (runs of consecutive blocks, a one-block + hole filled, at most 8 MiB per request) follow it; the cache is + about 450 lines (with its documentation) in the wasm crate, with no + in-flight tracking or HTTP. + - **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`, + shipped as a wasm-bindgen snippet): `fetch` with `Range`, six + requests at a time, every answer checked (206, `Content-Range` when + visible, body length, the ETag/Last-Modified and length of the + first answer). A `200` to the first request is kept as the whole + file up to `maxDownload` (512 MiB), or refused with + `fallback: "error"`. The first block comes with the request that + learns the length, as in M3. + - Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm + --test lazy` compares what the viewer shows of each file (kinds, + listings, attributes, info, whole reads, hyperslabs) read lazily at + 512 B to 1 MiB blocks with the in-memory reader and the facade's + range-storage path, for built files, the h5py/netCDF4 fixture and, + with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus + files. The built package, under Node and headless Chromium + (`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against + `test/serve.py`, which counts requests): every fixture check over + HTTP; in a 200 MB h5py file, listing, two small reads, a group's + attributes and a 10-value window of the 25-million-value dataset + took 5 requests, 6 MiB; with the corpus variable, 622 files agree + with `open(bytes)`. **M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow; add `File::refresh()` that re-reads the superblock/EOF and invalidates cached diff --git a/docs/known-issues.md b/docs/known-issues.md index e97f602..1f9a79b 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -844,9 +844,10 @@ cache, but: `read_*_zerocopy`) need the file in memory and answer `FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()` panics for such a file (`File::contiguous_bytes()` is the fallible form). - `LazyFile`, `MmapFile` and the Python and wasm bindings still read a - whole file (`h5rs` reads through `File::storage`, and takes URLs with its - `remote` feature). + `LazyFile`, `MmapFile` and the Python bindings still read a whole file + (`h5rs` reads through `File::storage`, and takes URLs with its `remote` + feature; the wasm reader's `openUrl` reads by range requests since + 2026-09-27, its `open(bytes)` takes a whole file). - The file's length is read once, at open: a growing file (SWMR) is not followed (milestone M5). A remote file is pinned at open, so one that grows is `RemoteError::FileChanged`. @@ -862,9 +863,9 @@ cache, but: **Status:** open (added 2026-09-26, milestone M3 of `docs/design/range-reads.md`). -- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3) - parses through `File::as_bytes`, which a remote file does not have; the - wasm reader's `openUrl` is milestone M4. +- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through + `File::as_bytes`, which a remote file does not have. (The browser can + since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.) - **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise). The design's policy of using a paged file's page size as the block size is not implemented, and only the first block is read ahead. @@ -903,10 +904,37 @@ cache, but: ## `clawhdf5-wasm` (browser) limits -**Status:** open (by design for now; added 2026-09-26). +**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27). -- The whole file is held in memory: `open()` takes its bytes. There are no - HTTP range reads, so a multi-GB file does not fit a browser tab. +- `open()` holds the whole file in memory (it takes its bytes), so a + multi-GB local file does not fit a browser tab. A file on a web server + can be opened with `openUrl()` instead, which fetches only the byte + ranges each call needs (milestone M4 of `docs/design/range-reads.md`), + with these limits: + - **Round trips:** a call runs as passes over the blocks fetched so far + and is re-run after each wave of misses, so a call costs one round + trip per wave, not one for everything: a chunk index is walked a + level (or a node) per round trip, while the chunks of a read are + fetched together. Each pass re-parses what the call reads (CPU, not + network). + - **Memory:** a call keeps every block it reads until it finishes (the + cache budget applies between calls), so reading a large dataset whole + needs its stored bytes in memory next to the result; read windows + (`readHyperslab`) of large datasets. Files above 200 MB have not been + tested, nor offsets above 4 GiB on wasm32. + - **Cross-origin servers** must allow CORS for the page's origin and + either expose `Content-Range` (`Access-Control-Expose-Headers`) or + answer `HEAD` with `Content-Length`. The file is pinned at open by its + `ETag` or `Last-Modified` (and its length): when the page cannot see + either header, only a change of length is detected. + - **A server without range support** (it answers `200`) costs a whole + download, up to `maxDownload` (512 MiB), or an error with + `fallback: "error"`. + - Fixed block size (`blockSize`, 1 MiB by default); a paged file's page + size is not used. No retries: a failed request fails the call (calling + again retries it; what was fetched stays cached). + - Tested under Node 22 and headless Chromium (Playwright's build) against + a local server; not in Firefox or Safari. - Compound, reference, opaque, bitfield, time and VL-sequence datasets are refused with an error naming the type; attributes of those types come back as `value: null` with their `dtype`. diff --git a/examples/wasm-viewer/README.md b/examples/wasm-viewer/README.md index 1a5c47d..39ad3f1 100644 --- a/examples/wasm-viewer/README.md +++ b/examples/wasm-viewer/README.md @@ -2,10 +2,12 @@ A single page that opens an HDF5 or NetCDF-4 file entirely in the browser with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a -file, browse its groups, and look at a dataset's type, shape, attributes -and values (a 50 x 12 window at a time, read as a hyperslab, with the -leading dimensions of a 3-D+ dataset held at chosen indices). The file never -leaves the page. +file or give a URL, browse its groups, and look at a dataset's type, shape, +attributes and values (a 50 x 12 window at a time, read as a hyperslab, +with the leading dimensions of a 3-D+ dataset held at chosen indices). A +local file never leaves the page. A file given by URL is not downloaded: +only the byte ranges each view needs are fetched (HTTP range requests), +and the header shows the requests and bytes that has cost so far. ## Build and open @@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file:// ``` -Then open . `?file=&path=` opens a -file from a URL (same origin, or one serving CORS headers) and selects an -object in it, e.g. `?file=data/run1.h5&path=/results/energy`. +Then open . The URL box, or +`?file=&path=`, opens a file by URL (same origin, or a server +sending CORS headers, see Limits) and selects an object in it, e.g. +`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not +answer range requests, so a file served by it is downloaded whole (the +header says so); `test/serve.py` does: + +```bash +python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer +# prints its port; open http://127.0.0.1:/?file=data/run1.h5 +``` ## JavaScript API ```js -import init, { open } from "./pkg/clawhdf5_wasm.js"; +import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js"; await init(); const f = open(new Uint8Array(await blob.arrayBuffer())); f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first @@ -32,8 +42,36 @@ f.attrs("/grid"); // [{ name, value, dtype }] f.read("/grid"); // { shape, dtype, data } f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block? f.free(); + +// A file on a web server, read by HTTP range requests as needed. The same +// methods, each returning a promise. +const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 }); +await r.list("/"); +await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches +r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes } +r.free(); ``` +`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block +fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept +between calls, default 64 MiB), `fallback` (`"download"`, the default, +reads the whole file when the server ignores `Range`, up to `maxDownload` +bytes, default 512 MiB; `"error"` refuses such a server), `headers` and +`credentials` (passed to every request), `parallel` (requests in flight, +default 6), `fetch` (a `fetch`-compatible function to use). + +How it works: the reader is synchronous and a page cannot block on the +network, so each call runs as a *pass* over the blocks fetched so far. A +pass that needs a block not yet fetched is abandoned, the missing blocks +are fetched (in parallel, adjacent blocks in one request), and the pass is +run again, until one completes (`docs/design/range-reads.md`, M4). Opening +costs one request (the first block, which also gives the file's size); +listing a group whose metadata is in blocks already fetched costs none; +reading a chunked dataset costs a round trip for its chunk index (a few +for a deep one) and one batch of requests for its chunks. Every answer is +checked — a `206` with exactly the bytes asked for, from the same file +(ETag or Last-Modified, and length) — or the call fails. + `data` is the typed array of the stored width (`Float64Array`, `Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`, `BigUint64Array`), or an array of strings for fixed- and variable-length @@ -43,7 +81,14 @@ else throws an `Error` naming the type. ## Limits -- Read-only, and the whole file is held in memory (no range requests). +- Read-only. `open(bytes)` holds the whole file in memory. +- `openUrl`: a cross-origin server must allow CORS and expose + `Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer + `HEAD` with `Content-Length`; if the page cannot see `ETag` or + `Last-Modified` either, a file replaced on the server is detected only by + a change of length. A call keeps what it reads until it finishes, so a + whole read of a large dataset needs its stored bytes in memory; read + windows of large datasets. More in `docs/known-issues.md`. - Compound, reference, opaque and variable-length-sequence datasets are refused with an error. Attributes of those types are listed with `value: null` and their `dtype`. @@ -56,28 +101,46 @@ else throws an `Error` naming the type. ## Tests `test/run.sh` builds the package, writes `fixture.h5` (h5py) and -`fixture.nc` (netCDF4) with `test/make_fixture.py`, then: +`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB +h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under +`TMPDIR`), serves them with `test/serve.py` (range requests, a request +counter, and `/norange/...` for a server without range support), then: - runs `test/test.mjs` under Node: every dataset (whole and a strided hyperslab), listing and attribute is compared with what libhdf5 reads back, error paths are checked, and so are the page's DOM-free helpers - (`viewer-lib.js`); + (`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB + and 512 B blocks), calls in flight at once, the request budget of + `big.h5` (listing it and reading small things of it must take at most 8 + requests and under 5% of the file), the download fallback and its + limit, and the errors: HTTP status, a file that changes, a server that + sends the wrong bytes or stops honouring `Range`. + `CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up + to 16 MiB) read by URL with the same file read from bytes; - runs `test/browser.sh`: loads the page in headless Chromium with - `?file=fixture.h5&path=...` for eight objects and checks the rendered tree, - types, shapes, attribute and value cells, and the error shown for an - unsupported type. Skipped when no Chromium is found (`CHROME` names one; - a Playwright download under `~/.cache/ms-playwright` is picked up). - Drag-and-drop and the file picker are not driven by it; they share - `load()` with the `?file=` path. + `?file=fix/fixture.h5&path=...` for eight objects and checks the rendered + tree, types, shapes, attribute and value cells, the request counter, and + the error shown for an unsupported type; then a server without range + support, and `big.h5` (a small dataset and a window of the large one, + with a single-digit percentage of the file fetched). Skipped when no + Chromium is found (`CHROME` names one; a Playwright download under + `~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker + and the URL box are not driven by it; they share `setFile()` with the + `?file=` path. The same expectations are checked natively, without Node, by -`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI -container has no Node or browser). +`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against +the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the +corpus too), which is what CI runs (the CI container has no Node or +browser). ## Size Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14, -`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`: +`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is +larger now and the table has not been re-measured: the reader has grown +since, and `openUrl` (2026-09-27) made the facade's range-read path +reachable from JavaScript and added promise glue and `remote.js`. | | raw | gzip -9 | |---|---:|---:| diff --git a/examples/wasm-viewer/test/test.mjs b/examples/wasm-viewer/test/test.mjs index 7a264dc..c0575e2 100644 --- a/examples/wasm-viewer/test/test.mjs +++ b/examples/wasm-viewer/test/test.mjs @@ -247,7 +247,7 @@ async function remoteTests() { const server = await serverStats(); eq(st.requests, server.requests, "big: requests counted"); eq(st.bytesFetched, server.bytes, "big: bytes counted"); - console.log(`big.h5 (${big.size} bytes): listed, 3 small reads and a window in ` + + console.log(`big.h5 (${big.size} bytes): listed, 2 small reads, attributes, info and a window in ` + `${server.requests} requests, ${server.bytes} bytes (${(100 * server.bytes / big.size).toFixed(2)}%)`); assert.ok(server.requests <= 8, `big: ${server.requests} requests`); assert.ok(server.bytes * 20 < big.size, `big: ${server.bytes} bytes fetched`);