Lazy remote files in the browser (M4), SWMR reader (M5), Python remote reads and editing #19

Merged
osobh merged 33 commits from feat/p3-wasm-swmr-python into main 2026-09-27 14:47:26 +00:00
7 changed files with 245 additions and 49 deletions
Showing only changes of commit f825a89e23 - Show all commits
+49
View File
@@ -2,6 +2,55 @@
## Unreleased
### Range reads, milestone M4: remote files in the browser (2026-09-27)
- **`openUrl(url, opts)`** in `clawhdf5-wasm` opens an HDF5/NetCDF-4 file
on a web server without downloading it: the returned `RemoteFile` has
`H5File`'s methods (`kind`, `list`, `info`, `attrs`, `attrErrors`,
`read`, `readHyperslab`), each returning a promise, and `stats()`
(requests, bytes fetched, file size). Only the byte ranges a call needs
are fetched, with `fetch` and `Range` headers, in 1 MiB blocks
(`blockSize`) kept in a 64 MiB cache (`cacheSize`). `open(bytes)` is
unchanged.
- **How, on the browser's main thread:** the design's restartable
"NeedBytes" mode (`clawhdf5_wasm::lazy::LazyStorage`). A call runs as a
pass over the blocks fetched so far; a read that misses records the
missing blocks and fails, the pass's result is dropped (even if a parser
caught the error and carried on), the blocks are fetched and the pass is
re-run. No Web Worker and no synchronous XHR (what h5wasm's lazy files
need). No block is evicted while a call is in flight, so every call ends;
the cache is trimmed between calls, raw-data blocks before metadata.
`clawhdf5-remote`'s `BlockCache` is not reused: it fetches by blocking,
and it evicts, and does not keep large reads, during a read, where a
restartable pass needs every block it read to stay until it finishes.
- **Every answer is checked** (`js/remote.js`): a `206` with exactly the
bytes asked for (`Content-Range`, when the page can see it, and the body
length), and the same `ETag`/`Last-Modified` and length as at open, or
the call fails: never data from another file or offset. A server that
answers `200` to a `Range` request is downloaded whole (up to
`maxDownload`, 512 MiB) unless `fallback: "error"`. Other options:
`headers`, `credentials`, `parallel` (6 requests at a time), `fetch`.
- **The viewer** (`examples/wasm-viewer`) has a URL box, opens
`?file=<url>` lazily, and shows the requests and bytes fetched.
- Counted on tank, 2026-09-27 (`WASM_BIG_MB=200 bash
examples/wasm-viewer/test/run.sh`, requests as `test/serve.py` counted
them): in a 200 MB h5py file, listing the root, reading two small
datasets, a group's attributes, the large dataset's shape and a
10-value window of it (25 million values) took 5 requests and 6 MiB.
With
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, the 622 corpus files
up to 16 MiB that open read over HTTP as from bytes: the same listings,
attributes and values (datasets up to 2^20 values), and an error
wherever bytes give one. Natively, `cargo test -p clawhdf5-wasm --test
lazy` with the same variable compares 656 files (up to 64 MiB, error
messages included, against the facade's range-storage path).
- `Reader::open_storage` (the wasm crate's core over any `Storage`), and
variable-length strings resolve through the file's storage rather than
`File::as_bytes`.
- The package is larger: the facade's `Storage` read path is now
reachable from JavaScript (it was compiled out before), and the promise
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
build machine was shared); the viewer README's size table predates M4.
### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
a `clawhdf5::File` (through `File::open_storage`) that reads the file by
+13 -8
View File
@@ -181,14 +181,19 @@ Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal F
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
file over HTTP with `File::open`.
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only, file held in memory;
no Zstd/SZIP since they link C) and the `examples/wasm-viewer/` page.
`examples/wasm-viewer/test/run.sh` builds the package (needs the
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node
and headless Chromium (a Playwright download in `~/.cache/ms-playwright`
on tank); the CI container has neither, so CI runs the native
`clawhdf5-wasm` `h5py_interop` test on the same fixture. Size numbers are
in the example's README.
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
call is re-run after each wave of misses; no block evicted while a call
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
version) and tests it under Node and headless Chromium (a Playwright
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
(range server with request counts, 200 MB budget file); the CI container
has neither, so CI runs the native `h5py_interop` and `lazy` tests
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
numbers in the example's README predate `openUrl`.
- Python and Node.js bindings for cross-language use
- NetCDF-4 compatibility for scientific data interop
+1 -1
View File
@@ -757,7 +757,7 @@ clawhdf5 workspace (19 crates, ~86K lines of Rust in src/, ~104K with tests
├── Bindings
│ ├── clawhdf5-py — Python (PyO3)
│ ├── clawhdf5-napi — Node.js (napi-rs)
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only)
│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only; remote files by HTTP range requests)
│
└── Tooling
├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check
+61 -10
View File
@@ -7,20 +7,19 @@ change. Progress: M0 and M1 are done, and so is M2 (branch
`File::open_storage` gives the facade's read API over any `Storage` (see
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
object stores) and URLs in `h5rs` (see the M3 status below). M4 (wasm) is
next. Every count in §1–§2 was
change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet.
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
reader, through the restartable `NeedBytes` mode (see the M4 status
below). M5 (SWMR) is not started.
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
converted code without changing the plan: object-header continuation chunks
are followed without recursion (still one bounded `read_at` per chunk), and
implicit chunk indexes are addressed over the maximum chunk grid (in
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
working on the whole file in memory; it is not part of this design. Every count below was
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
next to it. No timing numbers appear here on purpose: the machine was shared
working on the whole file in memory; it is not part of this design. Every
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
every count in a milestone's status on the date it gives, with the
commands given next to it. No timing numbers appear here on purpose: the machine was shared
with other build jobs when this was written.
## The problem
@@ -519,6 +518,58 @@ fast path within benchmark noise.
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
to a whole download when the server does not answer 206.
- `examples/wasm-viewer`: open by URL.
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
with these choices:
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
`Storage` over the blocks fetched so far. A call (open, list, read)
runs as a pass; a read that misses records its blocks and fails with
a storage error. The error itself is not the signal: parsers catch
errors and carry on (a listing leaves out a link it cannot resolve),
so *any* pass that recorded a miss is thrown away, whatever it
returned, and re-run once the missing ranges have been fetched
(`attempt` → `Step::Need(ranges)` → `supply` → again). A Worker with
synchronous XHR would have kept the parsers' single pass, but it
needs the page to run the reader off the main thread (a second
module, messages for every call and every typed array copied back),
since synchronous XHR can return binary data only there; the
restartable loop costs a
re-parse per wave of misses instead, which is CPU, not network.
- **Progress:** no block is evicted while a call is in flight
(`LazyStorage::operation`), so each pass that does not finish asks for
at least one new block, and a call ends after at most one pass per
block it reads. The budget (`cacheSize`, 64 MiB) is applied between
calls, raw-data blocks (`read_ranges`, or reads longer than a block)
evicted before metadata. The price: a call holds everything it reads
until it finishes.
- **Not `clawhdf5-remote`'s `BlockCache`.** It fetches through its
backend by blocking, evicts during a read, and does not keep reads
that miss more than half its budget; a restartable pass needs every
block it has read to be there when it is re-run. The block
arithmetic and coalescing (runs of consecutive blocks, a one-block
hole filled, at most 8 MiB per request) follow it; the cache is
about 450 lines (with its documentation) in the wasm crate, with no
in-flight tracking or HTTP.
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
requests at a time, every answer checked (206, `Content-Range` when
visible, body length, the ETag/Last-Modified and length of the
first answer). A `200` to the first request is kept as the whole
file up to `maxDownload` (512 MiB), or refused with
`fallback: "error"`. The first block comes with the request that
learns the length, as in M3.
- Tested (tank, 2026-09-27): natively, `cargo test -p clawhdf5-wasm
--test lazy` compares what the viewer shows of each file (kinds,
listings, attributes, info, whole reads, hyperslabs) read lazily at
512 B to 1 MiB blocks with the in-memory reader and the facade's
range-storage path, for built files, the h5py/netCDF4 fixture and,
with `CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus`, 656 corpus
files. The built package, under Node and headless Chromium
(`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`, against
`test/serve.py`, which counts requests): every fixture check over
HTTP; in a 200 MB h5py file, listing, two small reads, a group's
attributes and a 10-value window of the 25-million-value dataset
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
with `open(bytes)`.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
+37 -9
View File
@@ -844,9 +844,10 @@ cache, but:
`read_*_zerocopy`) need the file in memory and answer
`FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()`
panics for such a file (`File::contiguous_bytes()` is the fallible form).
`LazyFile`, `MmapFile` and the Python and wasm bindings still read a
whole file (`h5rs` reads through `File::storage`, and takes URLs with its
`remote` feature).
`LazyFile`, `MmapFile` and the Python bindings still read a whole file
(`h5rs` reads through `File::storage`, and takes URLs with its `remote`
feature; the wasm reader's `openUrl` reads by range requests since
2026-09-27, its `open(bytes)` takes a whole file).
- The file's length is read once, at open: a growing file (SWMR) is not
followed (milestone M5). A remote file is pinned at open, so one that
grows is `RemoteError::FileChanged`.
@@ -862,9 +863,9 @@ cache, but:
**Status:** open (added 2026-09-26, milestone M3 of
`docs/design/range-reads.md`).
- **Python and the browser cannot open URLs yet.** `clawhdf5.File` (PyO3)
parses through `File::as_bytes`, which a remote file does not have; the
wasm reader's `openUrl` is milestone M4.
- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through
`File::as_bytes`, which a remote file does not have. (The browser can
since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.)
- **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise).
The design's policy of using a paged file's page size as the block size
is not implemented, and only the first block is read ahead.
@@ -903,10 +904,37 @@ cache, but:
## `clawhdf5-wasm` (browser) limits
**Status:** open (by design for now; added 2026-09-26).
**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27).
- The whole file is held in memory: `open()` takes its bytes. There are no
HTTP range reads, so a multi-GB file does not fit a browser tab.
- `open()` holds the whole file in memory (it takes its bytes), so a
multi-GB local file does not fit a browser tab. A file on a web server
can be opened with `openUrl()` instead, which fetches only the byte
ranges each call needs (milestone M4 of `docs/design/range-reads.md`),
with these limits:
- **Round trips:** a call runs as passes over the blocks fetched so far
and is re-run after each wave of misses, so a call costs one round
trip per wave, not one for everything: a chunk index is walked a
level (or a node) per round trip, while the chunks of a read are
fetched together. Each pass re-parses what the call reads (CPU, not
network).
- **Memory:** a call keeps every block it reads until it finishes (the
cache budget applies between calls), so reading a large dataset whole
needs its stored bytes in memory next to the result; read windows
(`readHyperslab`) of large datasets. Files above 200 MB have not been
tested, nor offsets above 4 GiB on wasm32.
- **Cross-origin servers** must allow CORS for the page's origin and
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
answer `HEAD` with `Content-Length`. The file is pinned at open by its
`ETag` or `Last-Modified` (and its length): when the page cannot see
either header, only a change of length is detected.
- **A server without range support** (it answers `200`) costs a whole
download, up to `maxDownload` (512 MiB), or an error with
`fallback: "error"`.
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
size is not used. No retries: a failed request fails the call (calling
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server; not in Firefox or Safari.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`.
+83 -20
View File
@@ -2,10 +2,12 @@
A single page that opens an HDF5 or NetCDF-4 file entirely in the browser
with `clawhdf5-wasm` (clawhdf5's reader compiled to WebAssembly): drop a
file, browse its groups, and look at a dataset's type, shape, attributes
and values (a 50 x 12 window at a time, read as a hyperslab, with the
leading dimensions of a 3-D+ dataset held at chosen indices). The file never
leaves the page.
file or give a URL, browse its groups, and look at a dataset's type, shape,
attributes and values (a 50 x 12 window at a time, read as a hyperslab,
with the leading dimensions of a 3-D+ dataset held at chosen indices). A
local file never leaves the page. A file given by URL is not downloaded:
only the byte ranges each view needs are fetched (HTTP range requests),
and the header shows the requests and bytes that has cost so far.
## Build and open
@@ -16,14 +18,22 @@ bash examples/wasm-viewer/build.sh # writes examples/wasm-viewe
python3 -m http.server -d examples/wasm-viewer 8000 # wasm cannot load from file://
```
Then open <http://localhost:8000/>. `?file=<url>&path=<object>` opens a
file from a URL (same origin, or one serving CORS headers) and selects an
object in it, e.g. `?file=data/run1.h5&path=/results/energy`.
Then open <http://localhost:8000/>. The URL box, or
`?file=<url>&path=<object>`, opens a file by URL (same origin, or a server
sending CORS headers, see Limits) and selects an object in it, e.g.
`?file=data/run1.h5&path=/results/energy`. Python's `http.server` does not
answer range requests, so a file served by it is downloaded whole (the
header says so); `test/serve.py` does:
```bash
python3 examples/wasm-viewer/test/serve.py --root data=/path/to/files --root =examples/wasm-viewer
# prints its port; open http://127.0.0.1:<port>/?file=data/run1.h5
```
## JavaScript API
```js
import init, { open } from "./pkg/clawhdf5_wasm.js";
import init, { open, openUrl } from "./pkg/clawhdf5_wasm.js";
await init();
const f = open(new Uint8Array(await blob.arrayBuffer()));
f.list("/"); // [{ name, kind: "group" | "dataset" }], groups first
@@ -32,8 +42,36 @@ f.attrs("/grid"); // [{ name, value, dtype }]
f.read("/grid"); // { shape, dtype, data }
f.readHyperslab("/grid", [0, 0], [10, 5], [2, 1]); // start, count, stride?, block?
f.free();
// A file on a web server, read by HTTP range requests as needed. The same
// methods, each returning a promise.
const r = await openUrl("https://example.org/run1.h5", { blockSize: 1 << 20 });
await r.list("/");
await r.readHyperslab("/grid", [0, 0], [10, 5]); // fetches only the chunks it touches
r.stats(); // { lazy, size, requests, bytesFetched, cachedBytes, passes }
r.free();
```
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
between calls, default 64 MiB), `fallback` (`"download"`, the default,
reads the whole file when the server ignores `Range`, up to `maxDownload`
bytes, default 512 MiB; `"error"` refuses such a server), `headers` and
`credentials` (passed to every request), `parallel` (requests in flight,
default 6), `fetch` (a `fetch`-compatible function to use).
How it works: the reader is synchronous and a page cannot block on the
network, so each call runs as a *pass* over the blocks fetched so far. A
pass that needs a block not yet fetched is abandoned, the missing blocks
are fetched (in parallel, adjacent blocks in one request), and the pass is
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
costs one request (the first block, which also gives the file's size);
listing a group whose metadata is in blocks already fetched costs none;
reading a chunked dataset costs a round trip for its chunk index (a few
for a deep one) and one batch of requests for its chunks. Every answer is
checked — a `206` with exactly the bytes asked for, from the same file
(ETag or Last-Modified, and length) — or the call fails.
`data` is the typed array of the stored width (`Float64Array`,
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
`BigUint64Array`), or an array of strings for fixed- and variable-length
@@ -43,7 +81,14 @@ else throws an `Error` naming the type.
## Limits
- Read-only, and the whole file is held in memory (no range requests).
- Read-only. `open(bytes)` holds the whole file in memory.
- `openUrl`: a cross-origin server must allow CORS and expose
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
`Last-Modified` either, a file replaced on the server is detected only by
a change of length. A call keeps what it reads until it finishes, so a
whole read of a large dataset needs its stored bytes in memory; read
windows of large datasets. More in `docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with
`value: null` and their `dtype`.
@@ -56,28 +101,46 @@ else throws an `Error` naming the type.
## Tests
`test/run.sh` builds the package, writes `fixture.h5` (h5py) and
`fixture.nc` (netCDF4) with `test/make_fixture.py`, then:
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
counter, and `/norange/...` for a server without range support), then:
- runs `test/test.mjs` under Node: every dataset (whole and a strided
hyperslab), listing and attribute is compared with what libhdf5 reads
back, error paths are checked, and so are the page's DOM-free helpers
(`viewer-lib.js`);
(`viewer-lib.js`). Then the same checks on the files opened by URL (1 MiB
and 512 B blocks), calls in flight at once, the request budget of
`big.h5` (listing it and reading small things of it must take at most 8
requests and under 5% of the file), the download fallback and its
limit, and the errors: HTTP status, a file that changes, a server that
sends the wrong bytes or stops honouring `Range`.
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
to 16 MiB) read by URL with the same file read from bytes;
- runs `test/browser.sh`: loads the page in headless Chromium with
`?file=fixture.h5&path=...` for eight objects and checks the rendered tree,
types, shapes, attribute and value cells, and the error shown for an
unsupported type. Skipped when no Chromium is found (`CHROME` names one;
a Playwright download under `~/.cache/ms-playwright` is picked up).
Drag-and-drop and the file picker are not driven by it; they share
`load()` with the `?file=` path.
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
tree, types, shapes, attribute and value cells, the request counter, and
the error shown for an unsupported type; then a server without range
support, and `big.h5` (a small dataset and a window of the large one,
with a single-digit percentage of the file fetched). Skipped when no
Chromium is found (`CHROME` names one; a Playwright download under
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
and the URL box are not driven by it; they share `setFile()` with the
`?file=` path.
The same expectations are checked natively, without Node, by
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, which is what CI runs (the CI
container has no Node or browser).
`crates/clawhdf5-wasm/tests/h5py_interop.rs`, and the lazy reader against
the in-memory one by `tests/lazy.rs` (with `CLAWHDF5_WASM_CORPUS`, over the
corpus too), which is what CI runs (the CI container has no Node or
browser).
## Size
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`:
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
larger now and the table has not been re-measured: the reader has grown
since, and `openUrl` (2026-09-27) made the facade's range-read path
reachable from JavaScript and added promise glue and `remote.js`.
| | raw | gzip -9 |
|---|---:|---:|
+1 -1
View File
@@ -247,7 +247,7 @@ async function remoteTests() {
const server = await serverStats();
eq(st.requests, server.requests, "big: requests counted");
eq(st.bytesFetched, server.bytes, "big: bytes counted");
console.log(`big.h5 (${big.size} bytes): listed, 3 small reads and a window in ` +
console.log(`big.h5 (${big.size} bytes): listed, 2 small reads, attributes, info and a window in ` +
`${server.requests} requests, ${server.bytes} bytes (${(100 * server.bytes / big.size).toFixed(2)}%)`);
assert.ok(server.requests <= 8, `big: ${server.requests} requests`);
assert.ok(server.bytes * 20 < big.size, `big: ${server.bytes} bytes fetched`);