- opts.headers is read as fetch reads it (a Headers, [name, value] pairs
or a plain object); it was spread as an object, which silently dropped
a Headers instance (a common way to pass Authorization). A caller's
Range is not sent.
- When one range request of a batch fails, the others in flight are
aborted (one AbortController per batch, its signal passed to fetch)
and no new ones start; the first failure is the error. The workers
used to go on issuing requests nobody waited for.
- parallel must be an integer from 1 to 1024 (openUrl) or a positive
integer (fetchRanges): a non-number gave NaN workers, so none ran and
fetchRanges returned nothing.
Tests (test.mjs): headers as object, Headers and pairs reach the fetch;
bad parallel values are option errors; fetchRanges with a 500 on the
third of 20 ranges at parallel 3 starts 3 requests and aborts the 2 in
flight. Before: the Headers case sent no header, 17 requests started
after the failure with none aborted, and parallel "x" returned.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.
- core::Reader::list reads every child's header before returning the
first error (the same error, in listing order, Group::groups/datasets
return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
and v2 collectors, the symbol table node loop and the dense-link loop
go on reading (not using) the remaining siblings, then return that
first error: results and errors are unchanged, only failing
traversals read more, and in memory that is free (storage::touch).
A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
(it was fetched again: 215 MB fetched from a 198 MB file).
Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
libver earliest, 64 KiB: 537/536 -> 8/531 (6 in flight)
libver latest, 1 MiB: 189/188 -> 9/98
libver latest, 64 KiB: 453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in
LazyStorage::assemble (capacity_overflow), taking every open file on the
page with it, and a hostile server only had to claim a large length and
serve a heap collection of 2 GiB + 4 KiB to get there (after fetching
2 GiB). Reading a large u8 dataset whole aborted the same way when its
values were widened to 64 bits.
- LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at
most 1 GiB): a read longer than it fails at once, before anything is
fetched, and an operation whose passes would fetch more than it fails
before fetching (Operation::charge). assemble reserves fallibly.
- Reader::read refuses a read that would use more than 1 GiB while
decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result)
with an error naming readHyperslab, before reading.
- openUrl refuses a file of 4 GiB or more at open on wasm32: the format
code turns offsets into usize, so nothing past 4 GiB can be read there
(shown by a new test: data at 3 GiB reads, a 4 GiB file is refused).
maxDownload is bounded to 1 GiB.
Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8
dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at
3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an
error or reads, and that the module survives. Before: RuntimeError:
unreachable in Node; the native test read the huge dataset and fetched
2 GiB.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
openUrl(url, opts) returns a RemoteFile with the methods of H5File
(kind, list, info, attrs, attrErrors, read, readHyperslab), each a
promise, and stats(). It runs every call through the restartable
LazyStorage: a pass that misses reports the byte ranges, js/remote.js
fetches them with fetch() and Range headers (six at a time), and the
pass is re-run. This keeps the main thread free without a Worker or
synchronous XHR (h5wasm's lazy files need both), as the design doc
recommends; the cost is re-running a pass per wave of misses.
Every answer is checked: a 206 with exactly the bytes asked for, and
the same ETag/Last-Modified and length as at open, else an error (never
data). A server that ignores Range (200) is downloaded whole, up to
maxDownload (512 MiB), unless fallback: "error". Options: blockSize,
cacheSize, headers, credentials, parallel, fetch.
test/serve.py is a range-capable static server with request counting
(and /norange/ for a server without range support). test.mjs repeats
every fixture check on files opened by URL (1 MiB and 512 B blocks),
checks the request budget on a 200 MB h5py file (list, three small
reads and a window of the big dataset: 5 requests, 6 MiB), the
download fallback, and HTTP errors, changed files and wrong answers;
with CLAWHDF5_WASM_CORPUS every corpus file is compared with open(bytes).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
LazyStorage holds the blocks of a remote file fetched so far. An
operation runs as passes over it: a read that misses records the
missing blocks and fails, the pass's result is dropped whatever it is
(a parser may have caught the error and carried on), and the caller
fetches the reported ranges and re-runs the pass. No block is evicted
while an operation is in flight, so every pass that does not finish
asks for at least one new block and the operation ends. Blocks sit in
an LRU with a byte budget, trimmed between operations, bulk (raw data)
blocks first.
Reader::open_storage opens a file through any Storage, and variable-
length strings resolve through the file's storage instead of
File::as_bytes, which panics for a file not in memory.
tests/lazy.rs compares, file by file, what the viewer can show (kinds,
listings, attributes, info, whole reads and hyperslabs) read lazily
with the facade's range-storage path and the in-memory reader: files
built here at 512 B to 1 MiB blocks, the h5py/netCDF4 fixture, and with
CLAWHDF5_WASM_CORPUS the conformance corpus (656 files agree). A
listing plus a small read of a 48 MB file fetches 3 ranges.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
h5rs (dump, ls, diff, check --data) kept its own lenient VL decoder:
a heap object longer than its element was cut to the element's length
(libhdf5 and h5py refuse it), a null string printed "" where h5dump
prints NULL, the stored element size was trusted, and every heap
collection was kept as an owned copy for the whole run. It now resolves
each element with VlResolver::element / string_element (new: one element
in place, borrowing from the file), and refuses a VL type whose stored
element size is not 4 + offset size + 4, as File does. H5::heap_object
and its cache are gone. h5diff compares a null VL string equal to an
empty one; so does h5rs diff.
clawhdf5-wasm already resolved VL strings with read_vl_strings; it now
uses VlResolver and checks the stored element size before reading, as
File::read_string does.
Tests (h5py writes the files, patched for "a\0b", a null element and
mis-sized heap objects, with 8- and 4-byte offsets):
- h5rs_interop dump_prints_vl_data_like_h5dump: byte-identical to h5dump;
- dump_json_vl_values_match_h5py: h5py's values, errors where h5py fails;
- check_data_flags_mis_sized_vl_heap_objects;
- clawhdf5-wasm tests/vl_strings.rs: wasm, File and h5py agree.
All four fail before. check --data over the 150 cve_hdf5 CVE and fuzzer
files now passes 15 (h5dump rejects 8 of them), was 16 and 9: the
stored-size check flags cve-2024-32608. h5rs-check-ok-files.sh --data:
0 of 422 flagged; h5rs-fuzz.sh: clean on 180 files.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
open(bytes) -> H5File with kind/list/info/attrs/attrErrors/read/
readHyperslab. Numeric data comes back in the typed array of the
stored width (Int16Array for i16, BigInt64Array for i64, Float32Array
for f32/f16, ...), strings and enum names as string arrays, array
datatypes flattened with their dims appended to the shape. Compound,
reference, opaque and VL-sequence datasets are refused with an error
naming the type; nothing is returned as reinterpreted bytes.
The logic is in a plain-Rust core module, tested natively: unit tests,
and h5py_interop, which compares every dataset, hyperslab, listing and
attribute of an h5py- and a netCDF4-written file with what libhdf5
reads back (generator shared with the Node test of the built package).
No mmap, no threads; lz4 is on, zstd/szip (C) are not. A
wasm-release profile (opt-level s, LTO) serves the browser build.
ci-test.sh lints the crate for wasm32 and checks it builds no C.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>