docs: openUrl hardening after review (limits, listing passes, CORS tests)
CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB decode limit, the 4 GiB file limit on wasm32, bodies cut off at their length, listing passes, the cross-origin tests, and a pre-existing nondeterministic error choice on cve-2025-2310.h5 that can fail the native corpus comparison), the viewer README (options, how listing costs, tests) and the M4 status in docs/design/range-reads.md. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -50,6 +50,38 @@
|
|||||||
reachable from JavaScript (it was compiled out before), and the promise
|
reachable from JavaScript (it was compiled out before), and the promise
|
||||||
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
|
glue and `remote.js` add JavaScript. Not measured for the docs yet (the
|
||||||
build machine was shared); the viewer README's size table predates M4.
|
build machine was shared); the viewer README's size table predates M4.
|
||||||
|
- **Hardened after review (2026-09-27):**
|
||||||
|
- Sizes a server or a dataset names are errors, never an abort of the
|
||||||
|
wasm module (which took every open file on the page with it): a read
|
||||||
|
past 2 GiB aborted in the lazy cache, reachable by a hostile server
|
||||||
|
claiming a large file and a 2 GiB heap collection, and `read()` of a
|
||||||
|
256 MiB `u8` dataset aborted widening it to 64 bits. New option
|
||||||
|
`maxFetch` (512 MiB, at most 1 GiB): what one call may fetch, and the
|
||||||
|
longest single read, refused before fetching. `read()` refuses a
|
||||||
|
dataset that would take more than 1 GiB to decode, naming
|
||||||
|
`readHyperslab`. A file of 4 GiB or more is refused at open (wasm32
|
||||||
|
reads offsets as 32-bit); `maxDownload` is at most 1 GiB.
|
||||||
|
- Response bodies are read as they arrive and cut off at the length
|
||||||
|
asked for (`maxDownload` for a `200`): a `206` with a gigabyte body
|
||||||
|
was buffered whole before its length was checked.
|
||||||
|
- Listing a group reads every child's header, and every node of each
|
||||||
|
level of the group's index, in one pass: 3000 datasets (h5py, 198 MB)
|
||||||
|
listed in 6 passes and 73 requests at 1 MiB blocks instead of 185
|
||||||
|
passes and 184 serial requests (`libver="latest"`: 9 passes instead of
|
||||||
|
189). In `clawhdf5-format`, the B-tree v1/v2 collectors, the symbol
|
||||||
|
table node loop and the dense-link loop read (without using) the
|
||||||
|
siblings after the first that fails, then return that error: same
|
||||||
|
results and errors, more reads only on failure (free in memory). The
|
||||||
|
lazy cache no longer re-fetches a cached block to merge two requests.
|
||||||
|
- `headers` may be a `Headers` instance or `[name, value]` pairs (a
|
||||||
|
`Headers` was silently dropped); when one range request fails the
|
||||||
|
others in flight are aborted; `parallel` must be an integer from 1 to
|
||||||
|
1024.
|
||||||
|
- Tests: `test/serve.py` serves ranges without exposed `Content-Range`/
|
||||||
|
`ETag` (`/noexpose/`, and `/unexposed/` for a real cross-origin page
|
||||||
|
in Chromium), so the HEAD-length path runs end to end; hostile and
|
||||||
|
oversized files (`make_fixture.py`'s `write_limits`), flooding
|
||||||
|
bodies, aborted siblings.
|
||||||
|
|
||||||
### Range reads, milestone M3: remote files (2026-09-26)
|
### Range reads, milestone M3: remote files (2026-09-26)
|
||||||
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
- **New crate `clawhdf5-remote`.** `open_url("http://host/file.h5")` gives
|
||||||
|
|||||||
@@ -548,7 +548,8 @@ fast path within benchmark noise.
|
|||||||
arithmetic and coalescing (runs of consecutive blocks, a one-block
|
arithmetic and coalescing (runs of consecutive blocks, a one-block
|
||||||
hole filled, at most 8 MiB per request) follow it; the cache is
|
hole filled, at most 8 MiB per request) follow it; the cache is
|
||||||
about 450 lines (with its documentation) in the wasm crate, with no
|
about 450 lines (with its documentation) in the wasm crate, with no
|
||||||
in-flight tracking or HTTP.
|
in-flight tracking or HTTP. (A hole already cached is not filled:
|
||||||
|
it would be fetched again.)
|
||||||
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
|
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
|
||||||
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
|
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
|
||||||
requests at a time, every answer checked (206, `Content-Range` when
|
requests at a time, every answer checked (206, `Content-Range` when
|
||||||
@@ -570,6 +571,17 @@ fast path within benchmark noise.
|
|||||||
attributes and a 10-value window of the 25-million-value dataset
|
attributes and a 10-value window of the 25-million-value dataset
|
||||||
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
|
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
|
||||||
with `open(bytes)`.
|
with `open(bytes)`.
|
||||||
|
- After review (2026-09-27): a call may fetch at most `maxFetch`
|
||||||
|
(512 MiB, at most 1 GiB) and a longer single read is refused before
|
||||||
|
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
|
||||||
|
more is refused at open on wasm32 (offsets become `usize`); every
|
||||||
|
response body is cut off at the length asked for. Parsers that walk
|
||||||
|
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
|
||||||
|
the wasm listing's child headers) read the siblings after the first
|
||||||
|
failure before returning it, so a pass asks for a whole level's
|
||||||
|
missing blocks: listing 3000 datasets went from 185 passes to 6.
|
||||||
|
The 32-bit risk below is covered by a Node test that reads data at
|
||||||
|
3 GiB from a mock server and is refused a 4 GiB file.
|
||||||
|
|
||||||
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
||||||
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
||||||
|
|||||||
+38
-9
@@ -915,26 +915,55 @@ cache, but:
|
|||||||
and is re-run after each wave of misses, so a call costs one round
|
and is re-run after each wave of misses, so a call costs one round
|
||||||
trip per wave, not one for everything: a chunk index is walked a
|
trip per wave, not one for everything: a chunk index is walked a
|
||||||
level (or a node) per round trip, while the chunks of a read are
|
level (or a node) per round trip, while the chunks of a read are
|
||||||
fetched together. Each pass re-parses what the call reads (CPU, not
|
fetched together. Listing a group asks for every child's object
|
||||||
network).
|
header, and every node of a level of the group's index, in one pass
|
||||||
|
(since 2026-09-27; it was one round trip per header block): 3000
|
||||||
|
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
|
||||||
|
`libver="latest"` file (dense links). Each pass re-parses what the
|
||||||
|
call reads (CPU, not network). With headers spread through the file
|
||||||
|
(h5py writes each next to its data) a listing still fetches most of
|
||||||
|
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
|
||||||
- **Memory:** a call keeps every block it reads until it finishes (the
|
- **Memory:** a call keeps every block it reads until it finishes (the
|
||||||
cache budget applies between calls), so reading a large dataset whole
|
cache budget applies between calls). It may fetch at most `maxFetch`
|
||||||
needs its stored bytes in memory next to the result; read windows
|
bytes (512 MiB by default, at most 1 GiB), and a single read longer
|
||||||
(`readHyperslab`) of large datasets. Files above 200 MB have not been
|
than that fails before anything is fetched: a hostile server cannot
|
||||||
tested, nor offsets above 4 GiB on wasm32.
|
make the page fetch or allocate what a file's lengths claim. `read()`
|
||||||
|
of a dataset that would take more than 1 GiB while it is decoded
|
||||||
|
(stored bytes, the values widened to 64 bits, the result) fails,
|
||||||
|
naming `readHyperslab`; read windows of large datasets. On wasm32 a
|
||||||
|
buffer past 2 GiB cannot exist and a failed allocation aborts the
|
||||||
|
whole module (every open file on the page), which these limits keep
|
||||||
|
from happening; before 2026-09-27 both did abort it.
|
||||||
|
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
|
||||||
|
open. The format code turns file offsets into `usize` to use them
|
||||||
|
(with a clean error past it), which is 32 bits on wasm32, so nothing
|
||||||
|
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
|
||||||
|
tested (with a mock server); files above 200 MB have not been served
|
||||||
|
for real.
|
||||||
- **Cross-origin servers** must allow CORS for the page's origin and
|
- **Cross-origin servers** must allow CORS for the page's origin and
|
||||||
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
|
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
|
||||||
answer `HEAD` with `Content-Length`. The file is pinned at open by its
|
answer `HEAD` with `Content-Length`. The file is pinned at open by its
|
||||||
`ETag` or `Last-Modified` (and its length): when the page cannot see
|
`ETag` or `Last-Modified` (and its length): when the page cannot see
|
||||||
either header, only a change of length is detected.
|
either header, only a change of length is detected.
|
||||||
- **A server without range support** (it answers `200`) costs a whole
|
- **A server without range support** (it answers `200`) costs a whole
|
||||||
download, up to `maxDownload` (512 MiB), or an error with
|
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
|
||||||
`fallback: "error"`.
|
with `fallback: "error"`. Every body, this one and each `206`, is read
|
||||||
|
as it arrives and cut off past its limit (the range asked for, or
|
||||||
|
`maxDownload`): a server cannot make the page buffer more.
|
||||||
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
|
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
|
||||||
size is not used. No retries: a failed request fails the call (calling
|
size is not used. No retries: a failed request fails the call (calling
|
||||||
again retries it; what was fetched stays cached).
|
again retries it; what was fetched stays cached).
|
||||||
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
||||||
a local server; not in Firefox or Safari.
|
a local server, cross-origin included (a page on 127.0.0.1 reading a
|
||||||
|
file from localhost, with and without exposed headers); not in Firefox
|
||||||
|
or Safari.
|
||||||
|
- The native corpus comparison (`tests/lazy.rs` with
|
||||||
|
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
|
||||||
|
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
|
||||||
|
and which chunk's error is reported depends on the iteration order of
|
||||||
|
the chunk index (a `HashMap`, seeded per process), so the lazy and
|
||||||
|
the range-storage reads can name different errors. Both are errors;
|
||||||
|
not specific to `openUrl` (it predates it).
|
||||||
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
||||||
refused with an error naming the type; attributes of those types come back
|
refused with an error naming the type; attributes of those types come back
|
||||||
as `value: null` with their `dtype`.
|
as `value: null` with their `dtype`.
|
||||||
|
|||||||
@@ -54,11 +54,14 @@ r.free();
|
|||||||
|
|
||||||
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
|
`openUrl(url, opts)` options, all optional: `blockSize` (bytes per block
|
||||||
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
|
fetched, 512 B to 64 MiB, default 1 MiB), `cacheSize` (bytes of blocks kept
|
||||||
between calls, default 64 MiB), `fallback` (`"download"`, the default,
|
between calls, default 64 MiB), `maxFetch` (bytes one call may fetch, and
|
||||||
reads the whole file when the server ignores `Range`, up to `maxDownload`
|
so the longest single read, up to 1 GiB, default 512 MiB), `fallback`
|
||||||
bytes, default 512 MiB; `"error"` refuses such a server), `headers` and
|
(`"download"`, the default, reads the whole file when the server ignores
|
||||||
`credentials` (passed to every request), `parallel` (requests in flight,
|
`Range`, up to `maxDownload` bytes, default 512 MiB, at most 1 GiB;
|
||||||
default 6), `fetch` (a `fetch`-compatible function to use).
|
`"error"` refuses such a server), `headers` (a `Headers`, `[name, value]`
|
||||||
|
pairs or an object) and `credentials` (passed to every request),
|
||||||
|
`parallel` (requests in flight, 1 to 1024, default 6; when one fails the
|
||||||
|
others are aborted), `fetch` (a `fetch`-compatible function to use).
|
||||||
|
|
||||||
How it works: the reader is synchronous and a page cannot block on the
|
How it works: the reader is synchronous and a page cannot block on the
|
||||||
network, so each call runs as a *pass* over the blocks fetched so far. A
|
network, so each call runs as a *pass* over the blocks fetched so far. A
|
||||||
@@ -66,7 +69,9 @@ pass that needs a block not yet fetched is abandoned, the missing blocks
|
|||||||
are fetched (in parallel, adjacent blocks in one request), and the pass is
|
are fetched (in parallel, adjacent blocks in one request), and the pass is
|
||||||
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
||||||
costs one request (the first block, which also gives the file's size);
|
costs one request (the first block, which also gives the file's size);
|
||||||
listing a group whose metadata is in blocks already fetched costs none;
|
listing a group whose metadata is in blocks already fetched costs none,
|
||||||
|
and otherwise a round trip per level of the group's index plus one for
|
||||||
|
its children's headers, all fetched together;
|
||||||
reading a chunked dataset costs a round trip for its chunk index (a few
|
reading a chunked dataset costs a round trip for its chunk index (a few
|
||||||
for a deep one) and one batch of requests for its chunks. Every answer is
|
for a deep one) and one batch of requests for its chunks. Every answer is
|
||||||
checked — a `206` with exactly the bytes asked for, from the same file
|
checked — a `206` with exactly the bytes asked for, from the same file
|
||||||
@@ -86,9 +91,12 @@ else throws an `Error` naming the type.
|
|||||||
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
|
`Content-Range` (`Access-Control-Expose-Headers: Content-Range`) or answer
|
||||||
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
|
`HEAD` with `Content-Length`; if the page cannot see `ETag` or
|
||||||
`Last-Modified` either, a file replaced on the server is detected only by
|
`Last-Modified` either, a file replaced on the server is detected only by
|
||||||
a change of length. A call keeps what it reads until it finishes, so a
|
a change of length. A call keeps what it reads until it finishes, and
|
||||||
whole read of a large dataset needs its stored bytes in memory; read
|
may fetch at most `maxFetch`; `read()` of a dataset that would take more
|
||||||
windows of large datasets. More in `docs/known-issues.md`.
|
than 1 GiB to decode is refused (read windows of large datasets with
|
||||||
|
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
|
||||||
|
Every response body is cut off past the length asked for. More in
|
||||||
|
`docs/known-issues.md`.
|
||||||
- Compound, reference, opaque and variable-length-sequence datasets are
|
- Compound, reference, opaque and variable-length-sequence datasets are
|
||||||
refused with an error. Attributes of those types are listed with
|
refused with an error. Attributes of those types are listed with
|
||||||
`value: null` and their `dtype`.
|
`value: null` and their `dtype`.
|
||||||
@@ -104,7 +112,9 @@ else throws an `Error` naming the type.
|
|||||||
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
|
`fixture.nc` (netCDF4) with `test/make_fixture.py`, and `big.h5`, a 200 MB
|
||||||
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
|
h5py file (`WASM_BIG_MB` sets its size, 0 leaves it out; it goes under
|
||||||
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
|
`TMPDIR`), serves them with `test/serve.py` (range requests, a request
|
||||||
counter, and `/norange/...` for a server without range support), then:
|
counter, `/norange/...` for a server without range support,
|
||||||
|
`/noexpose/...` and `/unexposed/...` for one that does not expose its
|
||||||
|
headers to CORS), then:
|
||||||
|
|
||||||
- runs `test/test.mjs` under Node: every dataset (whole and a strided
|
- runs `test/test.mjs` under Node: every dataset (whole and a strided
|
||||||
hyperslab), listing and attribute is compared with what libhdf5 reads
|
hyperslab), listing and attribute is compared with what libhdf5 reads
|
||||||
@@ -114,14 +124,21 @@ counter, and `/norange/...` for a server without range support), then:
|
|||||||
`big.h5` (listing it and reading small things of it must take at most 8
|
`big.h5` (listing it and reading small things of it must take at most 8
|
||||||
requests and under 5% of the file), the download fallback and its
|
requests and under 5% of the file), the download fallback and its
|
||||||
limit, and the errors: HTTP status, a file that changes, a server that
|
limit, and the errors: HTTP status, a file that changes, a server that
|
||||||
sends the wrong bytes or stops honouring `Range`.
|
sends the wrong bytes or stops honouring `Range`. Also the cross-origin
|
||||||
|
path with no exposed headers (`/noexpose/`: length from `HEAD`), the
|
||||||
|
size limits on `make_fixture.py`'s `limits.h5`, `hostile_vl.h5` (a heap
|
||||||
|
collection claiming 2 GiB) and `far.h5` (data at 3 GiB, served by a
|
||||||
|
mock), bodies longer than asked for, `headers` forms, `parallel`, and
|
||||||
|
sibling requests aborted after a failure.
|
||||||
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
|
`CLAWHDF5_WASM_CORPUS=DIR` also compares every HDF5 file under `DIR` (up
|
||||||
to 16 MiB) read by URL with the same file read from bytes;
|
to 16 MiB) read by URL with the same file read from bytes;
|
||||||
- runs `test/browser.sh`: loads the page in headless Chromium with
|
- runs `test/browser.sh`: loads the page in headless Chromium with
|
||||||
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
|
`?file=fix/fixture.h5&path=...` for eight objects and checks the rendered
|
||||||
tree, types, shapes, attribute and value cells, the request counter, and
|
tree, types, shapes, attribute and value cells, the request counter, and
|
||||||
the error shown for an unsupported type; then a server without range
|
the error shown for an unsupported type; then the file from another
|
||||||
support, and `big.h5` (a small dataset and a window of the large one,
|
origin (localhost), with CORS exposing `Content-Range` and exposing
|
||||||
|
nothing (`/unexposed/`, where the server must see a `HEAD`), a server
|
||||||
|
without range support, and `big.h5` (a small dataset and a window of the large one,
|
||||||
with a single-digit percentage of the file fetched). Skipped when no
|
with a single-digit percentage of the file fetched). Skipped when no
|
||||||
Chromium is found (`CHROME` names one; a Playwright download under
|
Chromium is found (`CHROME` names one; a Playwright download under
|
||||||
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
|
`~/.cache/ms-playwright` is picked up). Drag-and-drop, the file picker
|
||||||
|
|||||||
Reference in New Issue
Block a user