Files
clawhdf5/examples/wasm-viewer/README.md
T
osobhandClaude Opus 5.5 2e5b059530 docs: fewer round trips for remote files in the browser, counted
CHANGELOG (Unreleased): the v1 B-tree lookup, `Storage::hint`, the walks
that go on past a missing node, and the counts before and after on an
h5py file like the reviewer's (3000 datasets, 198 MB, earliest and
latest libver, 1 MiB and 64 KiB blocks), the corpus read lazily at
512 B and 64 KiB blocks, and the Node/Chromium suite.
known-issues (browser limits, round trips): the new counts, why the
passes cannot go lower (the chain of addresses), why merging nearby
requests does not help such a file, and that a second listing refetches
at 1 MiB blocks when the file's metadata blocks exceed `cacheSize`.
range-reads.md M4 status and the viewer README follow.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:48:47 -05:00

9.6 KiB

HDF5 viewer in the browser

A single page that opens an HDF5 or NetCDF-4 file entirely in the browser with clawhdf5-wasm (clawhdf5's reader compiled to WebAssembly): drop a file or give a URL, browse its groups, and look at a dataset's type, shape, attributes and values (a 50 x 12 window at a time, read as a hyperslab, with the leading dimensions of a 3-D+ dataset held at chosen indices). A local file never leaves the page. A file given by URL is not downloaded: only the byte ranges each view needs are fetched (HTTP range requests), and the header shows the requests and bytes that has cost so far.

Build and open

rustup target add wasm32-unknown-unknown
cargo install wasm-bindgen-cli --version 0.2.129   # must equal the crate version; build.sh checks
bash examples/wasm-viewer/build.sh                  # writes examples/wasm-viewer/pkg/ (not committed)
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://

Then open http://localhost:8000/. The URL box, or ?file=<url>&path=<object>, opens a file by URL (same origin, or a server sending CORS headers, see Limits) and selects an object in it, e.g. ?file=data/run1.h5&path=/results/energy. Python's http.server does not answer range requests, so a file served by it is downloaded whole (the header says so); test/serve.py does:

python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5

JavaScript API

import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
await init();
const f = open(new Uint8Array(await blob.arrayBuffer()));
f.list("/");                          // [{ name, kind: "group" | "dataset" }], groups first
f.info("/grid");                      // { shape, maxshape, dtype, elementShape }
f.attrs("/grid");                     // [{ name, value, dtype }]
f.read("/grid");                      // { shape, dtype, data }
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]);  // start, count, stride?, block?
f.free();

// A file on a web server, read by HTTP range requests as needed. The same
// methods, each returning a promise.
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
await r.list("/");
await r.readHyperslab("/grid", [0, 0], [10, 5]);  // fetches only the chunks it touches
r.stats();    // { lazy, size, requests, bytesFetched, cachedBytes, passes }
r.free();

openUrl(url, opts) options, all optional: blockSize (bytes per block fetched, 512 B to 64 MiB, default 1 MiB), cacheSize (bytes of blocks kept between calls, default 64 MiB), maxFetch (bytes one call may fetch, and so the longest single read, up to 1 GiB, default 512 MiB), fallback ("download", the default, reads the whole file when the server ignores Range, up to maxDownload bytes, default 512 MiB, at most 1 GiB; "error" refuses such a server), headers (a Headers, [name, value] pairs or an object) and credentials (passed to every request), parallel (requests in flight, 1 to 1024, default 6; when one fails the others are aborted), fetch (a fetch-compatible function to use).

How it works: the reader is synchronous and a page cannot block on the network, so each call runs as a pass over the blocks fetched so far. A pass that needs a block not yet fetched is abandoned, the missing blocks are fetched (in parallel, adjacent blocks in one request), and the pass is run again, until one completes (docs/design/range-reads.md, M4). Opening costs one request (the first block, which also gives the file's size); listing a group whose metadata is in blocks already fetched costs none, and otherwise about a round trip per level of the group's index plus one for its children's headers, all fetched together (the reader fetches what it knows it reads next along with what a pass missed); opening one object looks its name up in the group's index, not the whole group; reading a chunked dataset costs a round trip for its chunk index (a few for a deep one) and one batch of requests for its chunks. Every answer is checked — a 206 with exactly the bytes asked for, from the same file (ETag or Last-Modified, and length) — or the call fails.

data is the typed array of the stored width (Float64Array, Float32Array also for f16, Int8Array ... BigInt64Array, BigUint64Array), or an array of strings for fixed- and variable-length strings and enumerations (h5py booleans read as "TRUE"/"FALSE"). Array datatypes are flattened, their dimensions appended to shape. Anything else throws an Error naming the type.

Limits

  • Read-only. open(bytes) holds the whole file in memory.
  • openUrl: a cross-origin server must allow CORS and expose Content-Range (Access-Control-Expose-Headers: Content-Range) or answer HEAD with Content-Length; if the page cannot see ETag or Last-Modified either, a file replaced on the server is detected only by a change of length. A call keeps what it reads until it finishes, and may fetch at most maxFetch; read() of a dataset that would take more than 1 GiB to decode is refused (read windows of large datasets with readHyperslab). Files of 4 GiB or more are refused at open (wasm32). Every response body is cut off past the length asked for. More in docs/known-issues.md.
  • Compound, reference, opaque and variable-length-sequence datasets are refused with an error. Attributes of those types are listed with value: null and their dtype.
  • No Zstd or SZIP filters (they link C): such a dataset fails with unsupported filter. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and scale-offset are read (within the limits in docs/known-issues.md).
  • Virtual datasets whose sources are in other files, and external links, cannot be followed: there is no file system.

Tests

test/run.sh builds the package, writes fixture.h5 (h5py) and fixture.nc (netCDF4) with test/make_fixture.py, and big.h5, a 200 MB h5py file (WASM_BIG_MB sets its size, 0 leaves it out; it goes under TMPDIR), serves them with test/serve.py (range requests, a request counter, /norange/... for a server without range support, /noexpose/... and /unexposed/... for one that does not expose its headers to CORS), then:

  • runs test/test.mjs under Node: every dataset (whole and a strided hyperslab), listing and attribute is compared with what libhdf5 reads back, error paths are checked, and so are the page's DOM-free helpers (viewer-lib.js). Then the same checks on the files opened by URL (1 MiB and 512 B blocks), calls in flight at once, the request budget of big.h5 (listing it and reading small things of it must take at most 8 requests and under 5% of the file), the download fallback and its limit, and the errors: HTTP status, a file that changes, a server that sends the wrong bytes or stops honouring Range. Also the cross-origin path with no exposed headers (/noexpose/: length from HEAD), the size limits on make_fixture.py's limits.h5, hostile_vl.h5 (a heap collection claiming 2 GiB) and far.h5 (data at 3 GiB, served by a mock), bodies longer than asked for, headers forms, parallel, and sibling requests aborted after a failure. CLAWHDF5_WASM_CORPUS=DIR also compares every HDF5 file under DIR (up to 16 MiB) read by URL with the same file read from bytes;
  • runs test/browser.sh: loads the page in headless Chromium with ?file=fix/fixture.h5&path=... for eight objects and checks the rendered tree, types, shapes, attribute and value cells, the request counter, and the error shown for an unsupported type; then the file from another origin (localhost), with CORS exposing Content-Range and exposing nothing (/unexposed/, where the server must see a HEAD), a server without range support, and big.h5 (a small dataset and a window of the large one, with a single-digit percentage of the file fetched). Skipped when no Chromium is found (CHROME names one; a Playwright download under ~/.cache/ms-playwright is picked up). Drag-and-drop, the file picker and the URL box are not driven by it; they share setFile() with the ?file= path.

The same expectations are checked natively, without Node, by crates/clawhdf5-wasm/tests/h5py_interop.rs, and the lazy reader against the in-memory one by tests/lazy.rs (with CLAWHDF5_WASM_CORPUS, over the corpus too), which is what CI runs (the CI container has no Node or browser).

Size

Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14, gzip -9 -n), after bash examples/wasm-viewer/build.sh. The package is larger now and the table has not been re-measured: the reader has grown since, and openUrl (2026-09-27) made the facade's range-read path reachable from JavaScript and added promise glue and remote.js.

raw gzip -9
pkg/clawhdf5_wasm_bg.wasm (profile wasm-release, opt-level s) 627,501 B 191,639 B
pkg/clawhdf5_wasm.js (wasm-bindgen glue) 21,826 B 4,487 B
same wasm at opt-level z 693,068 B 192,550 B
same wasm at opt-level 3 544,035 B 198,803 B
h5wasm 0.10.3: wasm embedded in dist/esm/hdf5_util.js 3,544,184 B 907,096 B
h5wasm 0.10.3: dist/esm/hdf5_util.js as shipped 4,150,134 B 986,699 B

h5wasm figures: npm pack [email protected] (npm reports dist.unpackedSize 14,731,385 B for the whole package), wasm extracted from the binaryDecode literal in hdf5_util.js. h5wasm is the whole of libhdf5 (writing, every datatype, plugins), so this compares download size, not equal functionality. No wasm-opt pass was applied (binaryen is not installed on tank). opt-level s is used because it is the smallest compressed.