docs: range-read M4 (openUrl in the browser): changelog, limits, design status

CHANGELOG (Unreleased), known-issues (the browser can open URLs; the
limits of openUrl: round trips per wave of misses, a call holds what it
reads, CORS and validator visibility, the download fallback), the M4
status in docs/design/range-reads.md (why NeedBytes rather than a
Worker, why not clawhdf5-remote's BlockCache, what was tested; the
status paragraph at the top lost a garbled duplicate), the viewer's
README (API, options, how it works, tests; the size table is marked as
predating openUrl), CLAUDE.md and the README crate list.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 06:47:29 -05:00
co-authored by Claude Opus 5.5
parent b8c85f7627
commit f825a89e23
7 changed files with 245 additions and 49 deletions
+83 -20
View File
@@ -2,10 +2,12 @@
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
file, browse its groups, and look at a dataset's type, shape, attributes
and values (a 50 x 12 window at a time, read as a hyperslab, with the
leading dimensions of a 3-D+ dataset held at chosen indices). The file never
leaves the page.
file or give a URL, browse its groups, and look at a dataset's type, shape,
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
local file never leaves the page. A file given by URL is not downloaded:
only the byte ranges each view needs are fetched (HTTP range requests),
and the header shows the requests and bytes that has cost so far.
## Build and open
@@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
```
Then open <http://localhost:8000/>. `?file=<url>&path=<object>` opens a
file from a URL (same origin, or one serving CORS headers) and selects an
object in it, e.g. `?file=data/run1.h5&path=/results/energy`.
Then open <http://localhost:8000/>. The URL box, or
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
sending CORS headers, see Limits) and selects an object in it, e.g.
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
answer range requests, so a file served by it is downloaded whole (the
header says so); `test/serve.py` does:
```bash
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
```
## JavaScript API
```js
import init, { open } from "./pkg/clawhdf5_wasm.js";
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
await init();
const f = open(new Uint8Array(await blob.arrayBuffer()));
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
@@ -32,8 +42,36 @@ f.attrs("/grid"); // [{ name, value, dtype }]
f.read("/grid"); // { shape, dtype, data }
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
f.free();
// A file on a web server, read by HTTP range requests as needed. The same
// methods, each returning a promise.
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
await r.list("/");
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
r.free();
```
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
between calls, default 64 MiB), `fallback` (`"download"`, the default,
reads the whole file when the server ignores `Range`, up to `maxDownload`
bytes, default 512 MiB; `"error"` refuses such a server), `headers` and
`credentials` (passed to every request), `parallel` (requests in flight,
default 6), `fetch` (a `fetch`-compatible function to use).
How it works: the reader is synchronous and a page cannot block on the
network, so each call runs as a *pass* over the blocks fetched so far. A
pass that needs a block not yet fetched is abandoned, the missing blocks
are fetched (in parallel, adjacent blocks in one request), and the pass is
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
costs one request (the first block, which also gives the file's size);
listing a group whose metadata is in blocks already fetched costs none;
reading a chunked dataset costs a round trip for its chunk index (a few
for a deep one) and one batch of requests for its chunks. Every answer is
checked — a `206` with exactly the bytes asked for, from the same file
(ETag or Last-Modified, and length) — or the call fails.
`data` is the typed array of the stored width (`Float64Array`,
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
`BigUint64Array`), or an array of strings for fixed- and variable-length
@@ -43,7 +81,14 @@ else throws an `Error` naming the type.
## Limits
- Read-only, and the whole file is held in memory (no range requests).
- Read-only. `open(bytes)` holds the whole file in memory.
- `openUrl`: a cross-origin server must allow CORS and expose
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
`Last-Modified` either, a file replaced on the server is detected only by
a change of length. A call keeps what it reads until it finishes, so a
whole read of a large dataset needs its stored bytes in memory; read
windows of large datasets. More in `docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with
`value: null` and their `dtype`.
@@ -56,28 +101,46 @@ else throws an `Error` naming the type.
## Tests
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
`fixture.nc` (netCDF4) with `test/make_fixture.py`, then:
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
counter, and `/norange/...` for a server without range support), then:
- runs `test/test.mjs` under Node: every dataset (whole and a strided
hyperslab), listing and attribute is compared with what libhdf5 reads
back, error paths are checked, and so are the page's DOM-free helpers
(`viewer-lib.js`);
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
and 512 B blocks), calls in flight at once, the request budget of
`big.h5` (listing it and reading small things of it must take at most 8
requests and under 5% of the file), the download fallback and its
limit, and the errors: HTTP status, a file that changes, a server that
sends the wrong bytes or stops honouring `Range`.
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
to 16 MiB) read by URL with the same file read from bytes;
- runs `test/browser.sh`: loads the page in headless Chromium with
`?file=fixture.h5&path=...` for eight objects and checks the rendered tree,
types, shapes, attribute and value cells, and the error shown for an
unsupported type. Skipped when no Chromium is found (`CHROME` names one;
a Playwright download under `~/.cache/ms-playwright` is picked up).
Drag-and-drop and the file picker are not driven by it; they share
`load()` with the `?file=` path.
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
tree, types, shapes, attribute and value cells, the request counter, and
the error shown for an unsupported type; then a server without range
support, and `big.h5` (a small dataset and a window of the large one,
with a single-digit percentage of the file fetched). Skipped when no
Chromium is found (`CHROME` names one; a Playwright download under
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
and the URL box are not driven by it; they share `setFile()` with the
`?file=` path.
The same expectations are checked natively, without Node, by
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI
container has no Node or browser).
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
corpus too), which is what CI runs (the CI container has no Node or
browser).
## Size
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`:
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
larger now and the table has not been re-measured: the reader has grown
since, and `openUrl` (2026-09-27) made the facade's range-read path
reachable from JavaScript and added promise glue and `remote.js`.
| | raw | gzip -9 |
|---|---:|---:|