wasm: sizes a server or a dataset names are errors, not aborts
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in LazyStorage::assemble (capacity_overflow), taking every open file on the page with it, and a hostile server only had to claim a large length and serve a heap collection of 2 GiB + 4 KiB to get there (after fetching 2 GiB). Reading a large u8 dataset whole aborted the same way when its values were widened to 64 bits. - LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at most 1 GiB): a read longer than it fails at once, before anything is fetched, and an operation whose passes would fetch more than it fails before fetching (Operation::charge). assemble reserves fallibly. - Reader::read refuses a read that would use more than 1 GiB while decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result) with an error naming readHyperslab, before reading. - openUrl refuses a file of 4 GiB or more at open on wasm32: the format code turns offsets into usize, so nothing past 4 GiB can be read there (shown by a new test: data at 3 GiB reads, a 4 GiB file is refused). maxDownload is bounded to 1 GiB. Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8 dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at 3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an error or reads, and that the module survives. Before: RuntimeError: unreachable in Node; the native test read the huge dataset and fetched 2 GiB. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -14,6 +14,15 @@ use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
||||
use clawhdf5_format::storage::Storage;
|
||||
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
||||
|
||||
/// The most memory one read may use while it decodes: the stored bytes,
|
||||
/// the values at 64 bits (integers are widened first) and the values
|
||||
/// returned. A larger read fails with an error naming `readHyperslab`,
|
||||
/// before anything is read: on wasm32 a buffer past 2 GiB cannot be
|
||||
/// allocated at all, and failing to allocate aborts the module (every open
|
||||
/// file on the page with it). 1 GiB leaves room in wasm32's 4 GiB for the
|
||||
/// file's cached blocks and the JavaScript copy of the result.
|
||||
pub const MAX_READ_BYTES: u64 = 1 << 30;
|
||||
|
||||
/// Errors are reported to JavaScript as messages.
|
||||
pub type Result<T> = std::result::Result<T, String>;
|
||||
|
||||
@@ -241,13 +250,27 @@ impl Reader {
|
||||
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
|
||||
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
|
||||
}
|
||||
let raw = ds.read_selection(&selection).map_err(err)?;
|
||||
let data = self.decode(&raw, &dt)?;
|
||||
out_shape.extend(element_shape(&dt));
|
||||
let expected = out_shape
|
||||
.iter()
|
||||
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
|
||||
.ok_or("selection size overflows")?;
|
||||
let cost = expected.saturating_mul(bytes_per_value(&dt));
|
||||
if cost > MAX_READ_BYTES {
|
||||
return Err(format!(
|
||||
"reading {path}{} would take about {} MiB of memory, more than the {} MiB \
|
||||
one read may use; read it in parts (readHyperslab)",
|
||||
if slab.is_some() {
|
||||
" (this selection)"
|
||||
} else {
|
||||
" whole"
|
||||
},
|
||||
cost >> 20,
|
||||
MAX_READ_BYTES >> 20
|
||||
));
|
||||
}
|
||||
let raw = ds.read_selection(&selection).map_err(err)?;
|
||||
let data = self.decode(&raw, &dt)?;
|
||||
if data.len() as u64 != expected {
|
||||
return Err(format!(
|
||||
"read {} values for shape {out_shape:?} ({expected} expected)",
|
||||
@@ -316,6 +339,27 @@ impl Reader {
|
||||
}
|
||||
}
|
||||
|
||||
/// Memory one value of type `dt` takes while [`Reader::read`] decodes it
|
||||
/// (an array type's elements count as values): its stored bytes, plus what
|
||||
/// [`Reader::decode`] builds from them. A string counts its `String` (24
|
||||
/// bytes on 64-bit targets, less on wasm32) and, for a fixed-length one,
|
||||
/// its text; a variable-length string's text lives in the heap and is
|
||||
/// bounded by the storage's own read limit.
|
||||
fn bytes_per_value(dt: &Datatype) -> u64 {
|
||||
let base = array_base(dt);
|
||||
let stored = u64::from(base.type_size());
|
||||
stored
|
||||
+ match base {
|
||||
Datatype::FloatingPoint { size, .. } if *size <= 4 => 4,
|
||||
Datatype::FloatingPoint { .. } => 8,
|
||||
// Widened to 64 bits, then narrowed to a new vector.
|
||||
Datatype::FixedPoint { .. } => 8 + stored,
|
||||
Datatype::String { .. } => 24 + stored,
|
||||
Datatype::VariableLength { .. } | Datatype::Enumeration { .. } => 24,
|
||||
_ => 0,
|
||||
}
|
||||
}
|
||||
|
||||
/// Narrow integers read at 64 bits to the dataset's own width. The source is
|
||||
/// that width, so this cannot fail on correct input; it is checked anyway.
|
||||
fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T>> {
|
||||
|
||||
Reference in New Issue
Block a user