CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB decode limit, the 4 GiB file limit on wasm32, bodies cut off at their length, listing passes, the cross-origin tests, and a pre-existing nondeterministic error choice on cve-2025-2310.h5 that can fail the native corpus comparison), the viewer README (options, how listing costs, tests) and the M4 status in docs/design/range-reads.md. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
178 lines
9.4 KiB
Markdown
178 lines
9.4 KiB
Markdown
# HDF5 viewer in the browser
|
|
|
|
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
|
|
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
|
|
file or give a URL, browse its groups, and look at a dataset's type, shape,
|
|
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
|
|
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
|
|
local file never leaves the page. A file given by URL is not downloaded:
|
|
only the byte ranges each view needs are fetched (HTTP range requests),
|
|
and the header shows the requests and bytes that has cost so far.
|
|
|
|
## Build and open
|
|
|
|
```bash
|
|
rustup target add wasm32-unknown-unknown
|
|
cargo install wasm-bindgen-cli --version 0.2.129 # must equal the crate version; build.sh checks
|
|
bash examples/wasm-viewer/build.sh # writes examples/wasm-viewer/pkg/ (not committed)
|
|
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
|
|
```
|
|
|
|
Then open <http://localhost:8000/>. The URL box, or
|
|
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
|
|
sending CORS headers, see Limits) and selects an object in it, e.g.
|
|
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
|
|
answer range requests, so a file served by it is downloaded whole (the
|
|
header says so); `test/serve.py` does:
|
|
|
|
```bash
|
|
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
|
|
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
|
|
```
|
|
|
|
## JavaScript API
|
|
|
|
```js
|
|
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
|
|
await init();
|
|
const f = open(new Uint8Array(await blob.arrayBuffer()));
|
|
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
|
|
f.info("/grid"); // { shape, maxshape, dtype, elementShape }
|
|
f.attrs("/grid"); // [{ name, value, dtype }]
|
|
f.read("/grid"); // { shape, dtype, data }
|
|
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
|
|
f.free();
|
|
|
|
// A file on a web server, read by HTTP range requests as needed. The same
|
|
// methods, each returning a promise.
|
|
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
|
|
await r.list("/");
|
|
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
|
|
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
|
|
r.free();
|
|
```
|
|
|
|
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
|
|
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
|
|
between calls, default 64 MiB), `maxFetch` (bytes one call may fetch, and
|
|
so the longest single read, up to 1 GiB, default 512 MiB), `fallback`
|
|
(`"download"`, the default, reads the whole file when the server ignores
|
|
`Range`, up to `maxDownload` bytes, default 512 MiB, at most 1 GiB;
|
|
`"error"` refuses such a server), `headers` (a `Headers`, `[name, value]`
|
|
pairs or an object) and `credentials` (passed to every request),
|
|
`parallel` (requests in flight, 1 to 1024, default 6; when one fails the
|
|
others are aborted), `fetch` (a `fetch`-compatible function to use).
|
|
|
|
How it works: the reader is synchronous and a page cannot block on the
|
|
network, so each call runs as a *pass* over the blocks fetched so far. A
|
|
pass that needs a block not yet fetched is abandoned, the missing blocks
|
|
are fetched (in parallel, adjacent blocks in one request), and the pass is
|
|
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
|
costs one request (the first block, which also gives the file's size);
|
|
listing a group whose metadata is in blocks already fetched costs none,
|
|
and otherwise a round trip per level of the group's index plus one for
|
|
its children's headers, all fetched together;
|
|
reading a chunked dataset costs a round trip for its chunk index (a few
|
|
for a deep one) and one batch of requests for its chunks. Every answer is
|
|
checked — a `206` with exactly the bytes asked for, from the same file
|
|
(ETag or Last-Modified, and length) — or the call fails.
|
|
|
|
`data` is the typed array of the stored width (`Float64Array`,
|
|
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
|
|
`BigUint64Array`), or an array of strings for fixed- and variable-length
|
|
strings and enumerations (h5py booleans read as `"TRUE"`/`"FALSE"`). Array
|
|
datatypes are flattened, their dimensions appended to `shape`. Anything
|
|
else throws an `Error` naming the type.
|
|
|
|
## Limits
|
|
|
|
- Read-only. `open(bytes)` holds the whole file in memory.
|
|
- `openUrl`: a cross-origin server must allow CORS and expose
|
|
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
|
|
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
|
|
`Last-Modified` either, a file replaced on the server is detected only by
|
|
a change of length. A call keeps what it reads until it finishes, and
|
|
may fetch at most `maxFetch`; `read()` of a dataset that would take more
|
|
than 1 GiB to decode is refused (read windows of large datasets with
|
|
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
|
|
Every response body is cut off past the length asked for. More in
|
|
`docs/known-issues.md`.
|
|
- Compound, reference, opaque and variable-length-sequence datasets are
|
|
refused with an error. Attributes of those types are listed with
|
|
`value: null` and their `dtype`.
|
|
- No Zstd or SZIP filters (they link C): such a dataset fails with
|
|
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
|
|
scale-offset are read (within the limits in `docs/known-issues.md`).
|
|
- Virtual datasets whose sources are in other files, and external links,
|
|
cannot be followed: there is no file system.
|
|
|
|
## Tests
|
|
|
|
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
|
|
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
|
|
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
|
|
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
|
|
counter, `/norange/...` for a server without range support,
|
|
`/noexpose/...` and `/unexposed/...` for one that does not expose its
|
|
headers to CORS), then:
|
|
|
|
- runs `test/test.mjs` under Node: every dataset (whole and a strided
|
|
hyperslab), listing and attribute is compared with what libhdf5 reads
|
|
back, error paths are checked, and so are the page's DOM-free helpers
|
|
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
|
|
and 512 B blocks), calls in flight at once, the request budget of
|
|
`big.h5` (listing it and reading small things of it must take at most 8
|
|
requests and under 5% of the file), the download fallback and its
|
|
limit, and the errors: HTTP status, a file that changes, a server that
|
|
sends the wrong bytes or stops honouring `Range`. Also the cross-origin
|
|
path with no exposed headers (`/noexpose/`: length from `HEAD`), the
|
|
size limits on `make_fixture.py`'s `limits.h5`, `hostile_vl.h5` (a heap
|
|
collection claiming 2 GiB) and `far.h5` (data at 3 GiB, served by a
|
|
mock), bodies longer than asked for, `headers` forms, `parallel`, and
|
|
sibling requests aborted after a failure.
|
|
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
|
|
to 16 MiB) read by URL with the same file read from bytes;
|
|
- runs `test/browser.sh`: loads the page in headless Chromium with
|
|
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
|
|
tree, types, shapes, attribute and value cells, the request counter, and
|
|
the error shown for an unsupported type; then the file from another
|
|
origin (localhost), with CORS exposing `Content-Range` and exposing
|
|
nothing (`/unexposed/`, where the server must see a `HEAD`), a server
|
|
without range support, and `big.h5` (a small dataset and a window of the large one,
|
|
with a single-digit percentage of the file fetched). Skipped when no
|
|
Chromium is found (`CHROME` names one; a Playwright download under
|
|
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
|
|
and the URL box are not driven by it; they share `setFile()` with the
|
|
`?file=` path.
|
|
|
|
The same expectations are checked natively, without Node, by
|
|
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
|
|
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
|
|
corpus too), which is what CI runs (the CI container has no Node or
|
|
browser).
|
|
|
|
## Size
|
|
|
|
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
|
|
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
|
|
larger now and the table has not been re-measured: the reader has grown
|
|
since, and `openUrl` (2026-09-27) made the facade's range-read path
|
|
reachable from JavaScript and added promise glue and `remote.js`.
|
|
|
|
| | raw | gzip -9 |
|
|
|---|---:|---:|
|
|
| `pkg/clawhdf5_wasm_bg.wasm` (profile `wasm-release`, opt-level `s`) | 627,501 B | 191,639 B |
|
|
| `pkg/clawhdf5_wasm.js` (wasm-bindgen glue) | 21,826 B | 4,487 B |
|
|
| same wasm at opt-level `z` | 693,068 B | 192,550 B |
|
|
| same wasm at opt-level `3` | 544,035 B | 198,803 B |
|
|
| h5wasm 0.10.3: wasm embedded in `dist/esm/hdf5_util.js` | 3,544,184 B | 907,096 B |
|
|
| h5wasm 0.10.3: `dist/esm/hdf5_util.js` as shipped | 4,150,134 B | 986,699 B |
|
|
|
|
h5wasm figures: `npm pack [email protected]` (npm reports
|
|
`dist.unpackedSize` 14,731,385 B for the whole package), wasm extracted from
|
|
the `binaryDecode` literal in `hdf5_util.js`. h5wasm is the whole of libhdf5
|
|
(writing, every datatype, plugins), so this compares download size, not
|
|
equal functionality. No `wasm-opt` pass was applied (binaryen is not
|
|
installed on tank). opt-level `s` is used because it is the smallest
|
|
compressed.
|