docs: range-read M4 (openUrl in the browser): changelog, limits, design status
CHANGELOG (Unreleased), known-issues (the browser can open URLs; the limits of openUrl: round trips per wave of misses, a call holds what it reads, CORS and validator visibility, the download fallback), the M4 status in docs/design/range-reads.md (why NeedBytes rather than a Worker, why not clawhdf5-remote's BlockCache, what was tested; the status paragraph at the top lost a garbled duplicate), the viewer's README (API, options, how it works, tests; the size table is marked as predating openUrl), CLAUDE.md and the README crate list. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -2,10 +2,12 @@
|
||||
|
||||
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
|
||||
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
|
||||
file, browse its groups, and look at a dataset's type, shape, attributes
|
||||
and values (a 50 x 12 window at a time, read as a hyperslab, with the
|
||||
leading dimensions of a 3-D+ dataset held at chosen indices). The file never
|
||||
leaves the page.
|
||||
file or give a URL, browse its groups, and look at a dataset's type, shape,
|
||||
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
|
||||
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
|
||||
local file never leaves the page. A file given by URL is not downloaded:
|
||||
only the byte ranges each view needs are fetched (HTTP range requests),
|
||||
and the header shows the requests and bytes that has cost so far.
|
||||
|
||||
## Build and open
|
||||
|
||||
@@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe
|
||||
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
|
||||
```
|
||||
|
||||
Then open <http://localhost:8000/>. `?file=<url>&path=<object>` opens a
|
||||
file from a URL (same origin, or one serving CORS headers) and selects an
|
||||
object in it, e.g. `?file=data/run1.h5&path=/results/energy`.
|
||||
Then open <http://localhost:8000/>. The URL box, or
|
||||
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
|
||||
sending CORS headers, see Limits) and selects an object in it, e.g.
|
||||
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
|
||||
answer range requests, so a file served by it is downloaded whole (the
|
||||
header says so); `test/serve.py` does:
|
||||
|
||||
```bash
|
||||
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
|
||||
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
|
||||
```
|
||||
|
||||
## JavaScript API
|
||||
|
||||
```js
|
||||
import init, { open } from "./pkg/clawhdf5_wasm.js";
|
||||
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
|
||||
await init();
|
||||
const f = open(new Uint8Array(await blob.arrayBuffer()));
|
||||
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
|
||||
@@ -32,8 +42,36 @@ f.attrs("/grid"); // [{ name, value, dtype }]
|
||||
f.read("/grid"); // { shape, dtype, data }
|
||||
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
|
||||
f.free();
|
||||
|
||||
// A file on a web server, read by HTTP range requests as needed. The same
|
||||
// methods, each returning a promise.
|
||||
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
|
||||
await r.list("/");
|
||||
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
|
||||
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
|
||||
r.free();
|
||||
```
|
||||
|
||||
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
|
||||
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
|
||||
between calls, default 64 MiB), `fallback` (`"download"`, the default,
|
||||
reads the whole file when the server ignores `Range`, up to `maxDownload`
|
||||
bytes, default 512 MiB; `"error"` refuses such a server), `headers` and
|
||||
`credentials` (passed to every request), `parallel` (requests in flight,
|
||||
default 6), `fetch` (a `fetch`-compatible function to use).
|
||||
|
||||
How it works: the reader is synchronous and a page cannot block on the
|
||||
network, so each call runs as a *pass* over the blocks fetched so far. A
|
||||
pass that needs a block not yet fetched is abandoned, the missing blocks
|
||||
are fetched (in parallel, adjacent blocks in one request), and the pass is
|
||||
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
||||
costs one request (the first block, which also gives the file's size);
|
||||
listing a group whose metadata is in blocks already fetched costs none;
|
||||
reading a chunked dataset costs a round trip for its chunk index (a few
|
||||
for a deep one) and one batch of requests for its chunks. Every answer is
|
||||
checked — a `206` with exactly the bytes asked for, from the same file
|
||||
(ETag or Last-Modified, and length) — or the call fails.
|
||||
|
||||
`data` is the typed array of the stored width (`Float64Array`,
|
||||
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
|
||||
`BigUint64Array`), or an array of strings for fixed- and variable-length
|
||||
@@ -43,7 +81,14 @@ else throws an `Error` naming the type.
|
||||
|
||||
## Limits
|
||||
|
||||
- Read-only, and the whole file is held in memory (no range requests).
|
||||
- Read-only. `open(bytes)` holds the whole file in memory.
|
||||
- `openUrl`: a cross-origin server must allow CORS and expose
|
||||
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
|
||||
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
|
||||
`Last-Modified` either, a file replaced on the server is detected only by
|
||||
a change of length. A call keeps what it reads until it finishes, so a
|
||||
whole read of a large dataset needs its stored bytes in memory; read
|
||||
windows of large datasets. More in `docs/known-issues.md`.
|
||||
- Compound, reference, opaque and variable-length-sequence datasets are
|
||||
refused with an error. Attributes of those types are listed with
|
||||
`value: null` and their `dtype`.
|
||||
@@ -56,28 +101,46 @@ else throws an `Error` naming the type.
|
||||
## Tests
|
||||
|
||||
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
|
||||
`fixture.nc` (netCDF4) with `test/make_fixture.py`, then:
|
||||
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
|
||||
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
|
||||
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
|
||||
counter, and `/norange/...` for a server without range support), then:
|
||||
|
||||
- runs `test/test.mjs` under Node: every dataset (whole and a strided
|
||||
hyperslab), listing and attribute is compared with what libhdf5 reads
|
||||
back, error paths are checked, and so are the page's DOM-free helpers
|
||||
(`viewer-lib.js`);
|
||||
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
|
||||
and 512 B blocks), calls in flight at once, the request budget of
|
||||
`big.h5` (listing it and reading small things of it must take at most 8
|
||||
requests and under 5% of the file), the download fallback and its
|
||||
limit, and the errors: HTTP status, a file that changes, a server that
|
||||
sends the wrong bytes or stops honouring `Range`.
|
||||
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
|
||||
to 16 MiB) read by URL with the same file read from bytes;
|
||||
- runs `test/browser.sh`: loads the page in headless Chromium with
|
||||
`?file=fixture.h5&path=...` for eight objects and checks the rendered tree,
|
||||
types, shapes, attribute and value cells, and the error shown for an
|
||||
unsupported type. Skipped when no Chromium is found (`CHROME` names one;
|
||||
a Playwright download under `~/.cache/ms-playwright` is picked up).
|
||||
Drag-and-drop and the file picker are not driven by it; they share
|
||||
`load()` with the `?file=` path.
|
||||
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
|
||||
tree, types, shapes, attribute and value cells, the request counter, and
|
||||
the error shown for an unsupported type; then a server without range
|
||||
support, and `big.h5` (a small dataset and a window of the large one,
|
||||
with a single-digit percentage of the file fetched). Skipped when no
|
||||
Chromium is found (`CHROME` names one; a Playwright download under
|
||||
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
|
||||
and the URL box are not driven by it; they share `setFile()` with the
|
||||
`?file=` path.
|
||||
|
||||
The same expectations are checked natively, without Node, by
|
||||
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI
|
||||
container has no Node or browser).
|
||||
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
|
||||
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
|
||||
corpus too), which is what CI runs (the CI container has no Node or
|
||||
browser).
|
||||
|
||||
## Size
|
||||
|
||||
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
|
||||
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`:
|
||||
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
|
||||
larger now and the table has not been re-measured: the reader has grown
|
||||
since, and `openUrl` (2026-09-27) made the facade's range-read path
|
||||
reachable from JavaScript and added promise glue and `remote.js`.
|
||||
|
||||
| | raw | gzip -9 |
|
||||
|---|---:|---:|
|
||||
|
||||
Reference in New Issue
Block a user