docs: openUrl hardening after review (limits, listing passes, CORS tests)
CHANGELOG (M4 section), known-issues (wasm limits: maxFetch, the 1 GiB decode limit, the 4 GiB file limit on wasm32, bodies cut off at their length, listing passes, the cross-origin tests, and a pre-existing nondeterministic error choice on cve-2025-2310.h5 that can fail the native corpus comparison), the viewer README (options, how listing costs, tests) and the M4 status in docs/design/range-reads.md. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -548,7 +548,8 @@ fast path within benchmark noise.
|
||||
arithmetic and coalescing (runs of consecutive blocks, a one-block
|
||||
hole filled, at most 8 MiB per request) follow it; the cache is
|
||||
about 450 lines (with its documentation) in the wasm crate, with no
|
||||
in-flight tracking or HTTP.
|
||||
in-flight tracking or HTTP. (A hole already cached is not filled:
|
||||
it would be fetched again.)
|
||||
- **The HTTP is JavaScript** (`crates/clawhdf5-wasm/js/remote.js`,
|
||||
shipped as a wasm-bindgen snippet): `fetch` with `Range`, six
|
||||
requests at a time, every answer checked (206, `Content-Range` when
|
||||
@@ -570,6 +571,17 @@ fast path within benchmark noise.
|
||||
attributes and a 10-value window of the 25-million-value dataset
|
||||
took 5 requests, 6 MiB; with the corpus variable, 622 files agree
|
||||
with `open(bytes)`.
|
||||
- After review (2026-09-27): a call may fetch at most `maxFetch`
|
||||
(512 MiB, at most 1 GiB) and a longer single read is refused before
|
||||
fetching; `read()` refuses a decode over 1 GiB; a file of 4 GiB or
|
||||
more is refused at open on wasm32 (offsets become `usize`); every
|
||||
response body is cut off at the length asked for. Parsers that walk
|
||||
siblings (B-tree v1/v2 collectors, symbol table nodes, dense links,
|
||||
the wasm listing's child headers) read the siblings after the first
|
||||
failure before returning it, so a pass asks for a whole level's
|
||||
missing blocks: listing 3000 datasets went from 185 passes to 6.
|
||||
The 32-bit risk below is covered by a Node test that reads data at
|
||||
3 GiB from a mock server and is refused a 4 GiB file.
|
||||
|
||||
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
||||
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
||||
|
||||
+38
-9
@@ -915,26 +915,55 @@ cache, but:
|
||||
and is re-run after each wave of misses, so a call costs one round
|
||||
trip per wave, not one for everything: a chunk index is walked a
|
||||
level (or a node) per round trip, while the chunks of a read are
|
||||
fetched together. Each pass re-parses what the call reads (CPU, not
|
||||
network).
|
||||
fetched together. Listing a group asks for every child's object
|
||||
header, and every node of a level of the group's index, in one pass
|
||||
(since 2026-09-27; it was one round trip per header block): 3000
|
||||
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
|
||||
`libver="latest"` file (dense links). Each pass re-parses what the
|
||||
call reads (CPU, not network). With headers spread through the file
|
||||
(h5py writes each next to its data) a listing still fetches most of
|
||||
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
|
||||
- **Memory:** a call keeps every block it reads until it finishes (the
|
||||
cache budget applies between calls), so reading a large dataset whole
|
||||
needs its stored bytes in memory next to the result; read windows
|
||||
(`readHyperslab`) of large datasets. Files above 200 MB have not been
|
||||
tested, nor offsets above 4 GiB on wasm32.
|
||||
cache budget applies between calls). It may fetch at most `maxFetch`
|
||||
bytes (512 MiB by default, at most 1 GiB), and a single read longer
|
||||
than that fails before anything is fetched: a hostile server cannot
|
||||
make the page fetch or allocate what a file's lengths claim. `read()`
|
||||
of a dataset that would take more than 1 GiB while it is decoded
|
||||
(stored bytes, the values widened to 64 bits, the result) fails,
|
||||
naming `readHyperslab`; read windows of large datasets. On wasm32 a
|
||||
buffer past 2 GiB cannot exist and a failed allocation aborts the
|
||||
whole module (every open file on the page), which these limits keep
|
||||
from happening; before 2026-09-27 both did abort it.
|
||||
- **File size:** at most 4 GiB - 1 bytes; a larger file is refused at
|
||||
open. The format code turns file offsets into `usize` to use them
|
||||
(with a clean error past it), which is 32 bits on wasm32, so nothing
|
||||
at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are
|
||||
tested (with a mock server); files above 200 MB have not been served
|
||||
for real.
|
||||
- **Cross-origin servers** must allow CORS for the page's origin and
|
||||
either expose `Content-Range` (`Access-Control-Expose-Headers`) or
|
||||
answer `HEAD` with `Content-Length`. The file is pinned at open by its
|
||||
`ETag` or `Last-Modified` (and its length): when the page cannot see
|
||||
either header, only a change of length is detected.
|
||||
- **A server without range support** (it answers `200`) costs a whole
|
||||
download, up to `maxDownload` (512 MiB), or an error with
|
||||
`fallback: "error"`.
|
||||
download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error
|
||||
with `fallback: "error"`. Every body, this one and each `206`, is read
|
||||
as it arrives and cut off past its limit (the range asked for, or
|
||||
`maxDownload`): a server cannot make the page buffer more.
|
||||
- Fixed block size (`blockSize`, 1 MiB by default); a paged file's page
|
||||
size is not used. No retries: a failed request fails the call (calling
|
||||
again retries it; what was fetched stays cached).
|
||||
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
||||
a local server; not in Firefox or Safari.
|
||||
a local server, cross-origin included (a page on 127.0.0.1 reading a
|
||||
file from localhost, with and without exposed headers); not in Firefox
|
||||
or Safari.
|
||||
- The native corpus comparison (`tests/lazy.rs` with
|
||||
`CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file,
|
||||
`cve-2025-2310.h5`: two of its datasets have more than one bad chunk,
|
||||
and which chunk's error is reported depends on the iteration order of
|
||||
the chunk index (a `HashMap`, seeded per process), so the lazy and
|
||||
the range-storage reads can name different errors. Both are errors;
|
||||
not specific to `openUrl` (it predates it).
|
||||
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
||||
refused with an error naming the type; attributes of those types come back
|
||||
as `value: null` with their `dtype`.
|
||||
|
||||
Reference in New Issue
Block a user