wasm: sizes a server or a dataset names are errors, not aborts
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in LazyStorage::assemble (capacity_overflow), taking every open file on the page with it, and a hostile server only had to claim a large length and serve a heap collection of 2 GiB + 4 KiB to get there (after fetching 2 GiB). Reading a large u8 dataset whole aborted the same way when its values were widened to 64 bits. - LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at most 1 GiB): a read longer than it fails at once, before anything is fetched, and an operation whose passes would fetch more than it fails before fetching (Operation::charge). assemble reserves fallibly. - Reader::read refuses a read that would use more than 1 GiB while decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result) with an error naming readHyperslab, before reading. - openUrl refuses a file of 4 GiB or more at open on wasm32: the format code turns offsets into usize, so nothing past 4 GiB can be read there (shown by a new test: data at 3 GiB reads, a 4 GiB file is refused). maxDownload is bounded to 1 GiB. Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8 dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at 3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an error or reads, and that the module survives. Before: RuntimeError: unreachable in Node; the native test read the huge dataset and fetched 2 GiB. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -3,7 +3,9 @@ reads back from them, for the clawhdf5-wasm tests.
|
||||
|
||||
python make_fixture.py OUT_DIR
|
||||
|
||||
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json.
|
||||
writes OUT_DIR/fixture.h5, OUT_DIR/fixture.nc and OUT_DIR/expected.json,
|
||||
and the limit-test files OUT_DIR/limits.h5 and OUT_DIR/hostile_vl.h5 (see
|
||||
write_limits).
|
||||
Both the Rust test (crates/clawhdf5-wasm/tests/h5py_interop.rs, native) and
|
||||
the Node test (test.mjs, the built wasm package) compare against the same
|
||||
expected.json, so the two check the same values.
|
||||
@@ -204,6 +206,72 @@ json.dump({"fixture.h5": describe(h5), "fixture.nc": describe(nc)},
|
||||
open(out / "expected.json", "w"), indent=1, ensure_ascii=False)
|
||||
|
||||
|
||||
HUGE_U8 = 2**28 + 1024
|
||||
# The collection size hostile_vl.h5 claims: past 2 GiB, which a wasm32
|
||||
# buffer cannot hold.
|
||||
HOSTILE_GCOL_SIZE = 2**31 + 4096
|
||||
# The file length a server claims for hostile_vl.h5 (the tests' mock fetch
|
||||
# answers every range with zeros past the real bytes): 3 GiB, within what
|
||||
# wasm32 opens, and room for the collection.
|
||||
HOSTILE_LENGTH = 3 << 30
|
||||
|
||||
|
||||
def write_limits(out):
|
||||
"""Files for the size limits (the tests must get errors, not aborts):
|
||||
|
||||
- limits.h5: /huge_u8, 2^28 + 1024 bytes of u8 in compressed chunks
|
||||
(a small file): read whole it would take over 2 GiB while decoding;
|
||||
its last value is 7.
|
||||
- hostile_vl.h5: a variable-length string dataset /a whose global heap
|
||||
collection claims HOSTILE_GCOL_SIZE bytes, with the superblock's end
|
||||
of file set to HOSTILE_LENGTH (libhdf5 cannot read it; it is only
|
||||
served by a mock that claims that length).
|
||||
- far.h5 and far.json: /x, 16 float64 values, whose contiguous data
|
||||
address is moved FAR_SHIFT bytes on (past 2 GiB, the sign bit of a
|
||||
wasm32 isize) in a file whose end of file is moved as far; the tests'
|
||||
mock serves the data there, to show offsets up to 4 GiB work on
|
||||
wasm32.
|
||||
"""
|
||||
with h5py.File(out / "limits.h5", "w") as f:
|
||||
d = f.create_dataset("huge_u8", shape=(HUGE_U8,), dtype="u1",
|
||||
chunks=(1 << 20,), compression="gzip")
|
||||
d[-1] = 7
|
||||
path = out / "hostile_vl.h5"
|
||||
with h5py.File(path, "w", libver="earliest") as f:
|
||||
f.create_dataset("a", data=["x", "yy"], dtype=h5py.string_dtype())
|
||||
b = bytearray(path.read_bytes())
|
||||
assert b[8] == 0, "a version 0 superblock"
|
||||
b[40:48] = HOSTILE_LENGTH.to_bytes(8, "little") # end of file address
|
||||
at = b.index(b"GCOL")
|
||||
b[at + 8:at + 16] = HOSTILE_GCOL_SIZE.to_bytes(8, "little")
|
||||
path.write_bytes(bytes(b))
|
||||
|
||||
path = out / "far.h5"
|
||||
values = np.arange(16, dtype="<f8") * 1.5
|
||||
with h5py.File(path, "w", libver="earliest") as f:
|
||||
f.create_dataset("x", data=values)
|
||||
data_at = f["x"].id.get_offset()
|
||||
b = bytearray(path.read_bytes())
|
||||
# The layout message: the data's address, then its size.
|
||||
old = data_at.to_bytes(8, "little") + (values.nbytes).to_bytes(8, "little")
|
||||
assert b.count(old) == 1
|
||||
at = b.index(old)
|
||||
b[at:at + 8] = (data_at + FAR_SHIFT).to_bytes(8, "little")
|
||||
length = len(b) + FAR_SHIFT
|
||||
b[40:48] = length.to_bytes(8, "little")
|
||||
path.write_bytes(bytes(b))
|
||||
json.dump({"data_at": data_at, "far_at": data_at + FAR_SHIFT,
|
||||
"nbytes": values.nbytes, "length": length,
|
||||
"values": [float(x) for x in values]},
|
||||
open(out / "far.json", "w"))
|
||||
|
||||
|
||||
# far.h5's data moves this far: past 2 GiB, below 4 GiB.
|
||||
FAR_SHIFT = 3 << 30
|
||||
|
||||
write_limits(out)
|
||||
|
||||
|
||||
def write_big(path, megabytes):
|
||||
"""A large file for the range-request tests (`openUrl`): `/big`, about
|
||||
`megabytes` MB of float64 in 1 MiB chunks, written after a small
|
||||
|
||||
Reference in New Issue
Block a user