docs: openUrl hardening after review (limits, listing passes, CORS tests)

CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB
decode limit, the 4 GiB file limit on wasm32, bodies cut off at their
length, listing passes, the cross-origin tests, and a pre-existing
nondeterministic error choice on cve-2025-2310.h5 that can fail the
native corpus comparison), the viewer README (options, how listing
costs, tests) and the M4 status in docs/design/range-reads.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 07:54:29 -05:00
co-authored by Claude Opus 5.5
parent 2379e6163e
commit 7a4acbce12
4 changed files with 113 additions and 23 deletions
+32
View File
@@ -50,6 +50,38 @@
reachable from JavaScript (it was compiled out before), and the promise reachable from JavaScript (it was compiled out before), and the promise
glue and `remote.js` add JavaScript. Not measured for the docs yet (the glue and `remote.js` add JavaScript. Not measured for the docs yet (the
build machine was shared); the viewer README's size table predates M4. build machine was shared); the viewer README's size table predates M4.
- **Hardened after review (2026-09-27):**
- Sizes a server or a dataset names are errors, never an abort of the
wasm module (which took every open file on the page with it): a read
past 2 GiB aborted in the lazy cache, reachable by a hostile server
claiming a large file and a 2 GiB heap collection, and `read()` of a
256 MiB `u8` dataset aborted widening it to 64 bits. New option
`maxFetch` (512 MiB, at most 1 GiB): what one call may fetch, and the
longest single read, refused before fetching. `read()` refuses a
dataset that would take more than 1 GiB to decode, naming
`readHyperslab`. A file of 4 GiB or more is refused at open (wasm32
reads offsets as 32-bit); `maxDownload` is at most 1 GiB.
- Response bodies are read as they arrive and cut off at the length
asked for (`maxDownload` for a `200`): a `206` with a gigabyte body
was buffered whole before its length was checked.
- Listing a group reads every child's header, and every node of each
level of the group's index, in one pass: 3000 datasets (h5py, 198 MB)
listed in 6 passes and 73 requests at 1 MiB blocks instead of 185
passes and 184 serial requests (`libver="latest"`: 9 passes instead of
189). In `clawhdf5-format`, the B-tree v1/v2 collectors, the symbol
table node loop and the dense-link loop read (without using) the
siblings after the first that fails, then return that error: same
results and errors, more reads only on failure (free in memory). The
lazy cache no longer re-fetches a cached block to merge two requests.
- `headers` may be a `Headers` instance or `[name, value]` pairs (a
`Headers` was silently dropped); when one range request fails the
others in flight are aborted; `parallel` must be an integer from 1 to
1024.
- Tests: `test/serve.py` serves ranges without exposed `Content-Range`/
`ETag` (`/noexpose/`, and `/unexposed/` for a real cross-origin page
in Chromium), so the HEAD-length path runs end to end; hostile and
oversized files (`make_fixture.py`'s `write_limits`), flooding
bodies, aborted siblings.
### Range reads, milestone M3: remote files (2026-09-26) ### Range reads, milestone M3: remote files (2026-09-26)
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives - **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
+13 -1
View File
@@ -548,7 +548,8 @@ fast path within benchmark noise.
arithmetic and coalescing (runs of consecutive blocks, a one-block arithmetic and coalescing (runs of consecutive blocks, a one-block
hole filled, at most 8 MiB per request) follow it; the cache is hole filled, at most 8 MiB per request) follow it; the cache is
about 450 lines (with its documentation) in the wasm crate, with no about 450 lines (with its documentation) in the wasm crate, with no
in-flight tracking or HTTP. in-flight tracking or HTTP. (A hole already cached is not filled:
it would be fetched again.)
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`, - **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
requests at a time, every answer checked (206, `Content-Range` when requests at a time, every answer checked (206, `Content-Range` when
@@ -570,6 +571,17 @@ fast path within benchmark noise.
attributes and a 10-value window of the 25-million-value dataset attributes and a 10-value window of the 25-million-value dataset
took 5 requests, 6 MiB; with the corpus variable, 622 files agree took 5 requests, 6 MiB; with the corpus variable, 622 files agree
with `open(bytes)`. with `open(bytes)`.
- After review (2026-09-27): a call may fetch at most `maxFetch`
(512 MiB, at most 1 GiB) and a longer single read is refused before
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
more is refused at open on wasm32 (offsets become `usize`); every
response body is cut off at the length asked for. Parsers that walk
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
the wasm listing's child headers) read the siblings after the first
failure before returning it, so a pass asks for a whole level's
missing blocks: listing 3000 datasets went from 185 passes to 6.
The 32-bit risk below is covered by a Node test that reads data at
3 GiB from a mock server and is refused a 4 GiB file.
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow; **M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
+38 -9
View File
@@ -915,26 +915,55 @@ cache, but:
and is re-run after each wave of misses, so a call costs one round and is re-run after each wave of misses, so a call costs one round
trip per wave, not one for everything: a chunk index is walked a trip per wave, not one for everything: a chunk index is walked a
level (or a node) per round trip, while the chunks of a read are level (or a node) per round trip, while the chunks of a read are
fetched together. Each pass re-parses what the call reads (CPU, not fetched together. Listing a group asks for every child's object
network). header, and every node of a level of the group's index, in one pass
(since 2026-09-27; it was one round trip per header block): 3000
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
`libver="latest"` file (dense links). Each pass re-parses what the
call reads (CPU, not network). With headers spread through the file
(h5py writes each next to its data) a listing still fetches most of
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
- **Memory:** a call keeps every block it reads until it finishes (the - **Memory:** a call keeps every block it reads until it finishes (the
cache budget applies between calls), so reading a large dataset whole cache budget applies between calls). It may fetch at most `maxFetch`
needs its stored bytes in memory next to the result; read windows bytes (512 MiB by default, at most 1 GiB), and a single read longer
(`readHyperslab`) of large datasets. Files above 200 MB have not been than that fails before anything is fetched: a hostile server cannot
tested, nor offsets above 4 GiB on wasm32. make the page fetch or allocate what a file's lengths claim. `read()`
of a dataset that would take more than 1 GiB while it is decoded
(stored bytes, the values widened to 64 bits, the result) fails,
naming `readHyperslab`; read windows of large datasets. On wasm32 a
buffer past 2 GiB cannot exist and a failed allocation aborts the
whole module (every open file on the page), which these limits keep
from happening; before 2026-09-27 both did abort it.
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
open. The format code turns file offsets into `usize` to use them
(with a clean error past it), which is 32 bits on wasm32, so nothing
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
tested (with a mock server); files above 200 MB have not been served
for real.
- **Cross-origin servers** must allow CORS for the page's origin and - **Cross-origin servers** must allow CORS for the page's origin and
either expose `Content-Range` (`Access-Control-Expose-Headers`) or either expose `Content-Range` (`Access-Control-Expose-Headers`) or
answer `HEAD` with `Content-Length`. The file is pinned at open by its answer `HEAD` with `Content-Length`. The file is pinned at open by its
`ETag` or `Last-Modified` (and its length): when the page cannot see `ETag` or `Last-Modified` (and its length): when the page cannot see
either header, only a change of length is detected. either header, only a change of length is detected.
- **A server without range support** (it answers `200`) costs a whole - **A server without range support** (it answers `200`) costs a whole
download, up to `maxDownload` (512 MiB), or an error with download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
`fallback: "error"`. with `fallback: "error"`. Every body, this one and each `206`, is read
as it arrives and cut off past its limit (the range asked for, or
`maxDownload`): a server cannot make the page buffer more.
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page - Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
size is not used. No retries: a failed request fails the call (calling size is not used. No retries: a failed request fails the call (calling
again retries it; what was fetched stays cached). again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against - Tested under Node 22 and headless Chromium (Playwright's build) against
a local server; not in Firefox or Safari. a local server, cross-origin included (a page on 127.0.0.1 reading a
file from localhost, with and without exposed headers); not in Firefox
or Safari.
- The native corpus comparison (`tests/lazy.rs` with
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
and which chunk's error is reported depends on the iteration order of
the chunk index (a `HashMap`, seeded per process), so the lazy and
the range-storage reads can name different errors. Both are errors;
not specific to `openUrl` (it predates it).
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are - Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back refused with an error naming the type; attributes of those types come back
as `value: null` with their `dtype`. as `value: null` with their `dtype`.
+30 -13
View File
@@ -54,11 +54,14 @@ r.free();
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block `openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
between calls, default 64 MiB), `fallback` (`"download"`, the default, between calls, default 64 MiB), `maxFetch` (bytes one call may fetch, and
reads the whole file when the server ignores `Range`, up to `maxDownload` so the longest single read, up to 1 GiB, default 512 MiB), `fallback`
bytes, default 512 MiB; `"error"` refuses such a server), `headers` and (`"download"`, the default, reads the whole file when the server ignores
`credentials` (passed to every request), `parallel` (requests in flight, `Range`, up to `maxDownload` bytes, default 512 MiB, at most 1 GiB;
default 6), `fetch` (a `fetch`-compatible function to use). `"error"` refuses such a server), `headers` (a `Headers`, `[name, value]`
pairs or an object) and `credentials` (passed to every request),
`parallel` (requests in flight, 1 to 1024, default 6; when one fails the
others are aborted), `fetch` (a `fetch`-compatible function to use).
How it works: the reader is synchronous and a page cannot block on the How it works: the reader is synchronous and a page cannot block on the
network, so each call runs as a *pass* over the blocks fetched so far. A network, so each call runs as a *pass* over the blocks fetched so far. A
@@ -66,7 +69,9 @@ pass that needs a block not yet fetched is abandoned, the missing blocks
are fetched (in parallel, adjacent blocks in one request), and the pass is are fetched (in parallel, adjacent blocks in one request), and the pass is
run again, until one completes (`docs/design/range-reads.md`, M4). Opening run again, until one completes (`docs/design/range-reads.md`, M4). Opening
costs one request (the first block, which also gives the file's size); costs one request (the first block, which also gives the file's size);
listing a group whose metadata is in blocks already fetched costs none; listing a group whose metadata is in blocks already fetched costs none,
and otherwise a round trip per level of the group's index plus one for
its children's headers, all fetched together;
reading a chunked dataset costs a round trip for its chunk index (a few reading a chunked dataset costs a round trip for its chunk index (a few
for a deep one) and one batch of requests for its chunks. Every answer is for a deep one) and one batch of requests for its chunks. Every answer is
checked — a `206` with exactly the bytes asked for, from the same file checked — a `206` with exactly the bytes asked for, from the same file
@@ -86,9 +91,12 @@ else throws an `Error` naming the type.
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer `Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
`HEAD` with `Content-Length`; if the page cannot see `ETag` or `HEAD` with `Content-Length`; if the page cannot see `ETag` or
`Last-Modified` either, a file replaced on the server is detected only by `Last-Modified` either, a file replaced on the server is detected only by
a change of length. A call keeps what it reads until it finishes, so a a change of length. A call keeps what it reads until it finishes, and
whole read of a large dataset needs its stored bytes in memory; read may fetch at most `maxFetch`; `read()` of a dataset that would take more
windows of large datasets. More in `docs/known-issues.md`. than 1 GiB to decode is refused (read windows of large datasets with
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in
`docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are - Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with refused with an error. Attributes of those types are listed with
`value: null` and their `dtype`. `value: null` and their `dtype`.
@@ -104,7 +112,9 @@ else throws an `Error` naming the type.
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB `fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
`TMPDIR`), serves them with `test/serve.py` (range requests, a request `TMPDIR`), serves them with `test/serve.py` (range requests, a request
counter, and `/norange/...` for a server without range support), then: counter, `/norange/...` for a server without range support,
`/noexpose/...` and `/unexposed/...` for one that does not expose its
headers to CORS), then:
- runs `test/test.mjs` under Node: every dataset (whole and a strided - runs `test/test.mjs` under Node: every dataset (whole and a strided
hyperslab), listing and attribute is compared with what libhdf5 reads hyperslab), listing and attribute is compared with what libhdf5 reads
@@ -114,14 +124,21 @@ counter, and `/norange/...` for a server without range support), then:
`big.h5` (listing it and reading small things of it must take at most 8 `big.h5` (listing it and reading small things of it must take at most 8
requests and under 5% of the file), the download fallback and its requests and under 5% of the file), the download fallback and its
limit, and the errors: HTTP status, a file that changes, a server that limit, and the errors: HTTP status, a file that changes, a server that
sends the wrong bytes or stops honouring `Range`. sends the wrong bytes or stops honouring `Range`. Also the cross-origin
path with no exposed headers (`/noexpose/`: length from `HEAD`), the
size limits on `make_fixture.py`'s `limits.h5`, `hostile_vl.h5` (a heap
collection claiming 2 GiB) and `far.h5` (data at 3 GiB, served by a
mock), bodies longer than asked for, `headers` forms, `parallel`, and
sibling requests aborted after a failure.
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up `CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
to 16 MiB) read by URL with the same file read from bytes; to 16 MiB) read by URL with the same file read from bytes;
- runs `test/browser.sh`: loads the page in headless Chromium with - runs `test/browser.sh`: loads the page in headless Chromium with
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered `?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
tree, types, shapes, attribute and value cells, the request counter, and tree, types, shapes, attribute and value cells, the request counter, and
the error shown for an unsupported type; then a server without range the error shown for an unsupported type; then the file from another
support, and `big.h5` (a small dataset and a window of the large one, origin (localhost), with CORS exposing `Content-Range` and exposing
nothing (`/unexposed/`, where the server must see a `HEAD`), a server
without range support, and `big.h5` (a small dataset and a window of the large one,
with a single-digit percentage of the file fetched). Skipped when no with a single-digit percentage of the file fetched). Skipped when no
Chromium is found (`CHROME` names one; a Playwright download under Chromium is found (`CHROME` names one; a Playwright download under
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker `~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker