Commit Graph
5 Commits
Author SHA1 Message Date
osobhandClaude Opus 5.5 761bdbf24f format: look a name up in a v1 group down its B-tree, as libhdf5 does
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).

`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.

Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
  earliest, 1 MiB:  7/74/193.6 MB -> 6/5/5.2 MB
  earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
  latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
  -> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
  commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:47:47 -05:00
osobhandClaude Opus 5.5 6f5d14fd62 format: group walks go on past a failed node and hint what they read next
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.

- The v1 and v2 B-tree collectors descend into every child of a node
  after one fails (they only read the siblings before, so a sibling's
  subtree came a pass later), then return the first error: results and
  errors unchanged. The v2 walk stops once its record budget is spent,
  so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
  a group B-tree node's and a symbol table node's body (read once their
  header gives a length, a round trip later when the body is in the
  next block), an object header's first chunk and its continuation
  chunks, the symbol table nodes a B-tree leaf names, a dense group's
  name index header and the heap's root block (both read right after
  the heap header). A listing also hints every child's object header as
  its entry is read, even after a failure, and every direct block of a
  dense group's heap (reading the indirect blocks, at most 4096 entries
  and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
  n entries end) is one helper used by the object reads and the hints.

Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
  earliest, 1 MiB:  6/73/192.5 MB -> 4/68/192.5 MB
  earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
  latest,   1 MiB:  9/98/196.5 MB -> 5/86/196.5 MB
  latest,   64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 22:47:44 -05:00
osobhandClaude Opus 5.5 e553153e48 wasm, format: listing a group asks for all its missing blocks per pass
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.

- core::Reader::list reads every child's header before returning the
  first error (the same error, in listing order, Group::groups/datasets
  return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
  and v2 collectors, the symbol table node loop and the dense-link loop
  go on reading (not using) the remaining siblings, then return that
  first error: results and errors are unchanged, only failing
  traversals read more, and in memory that is free (storage::touch).
  A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
  (it was fetched again: 215 MB fetched from a 198 MB file).

Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
  libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
  libver earliest, 64 KiB:       537/536 -> 8/531 (6 in flight)
  libver latest,   1 MiB:        189/188 -> 9/98
  libver latest,   64 KiB:       453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:48:49 -05:00
osobhandClaude Opus 5.5 dbafa952ac wasm: sizes a server or a dataset names are errors, not aborts
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in
LazyStorage::assemble (capacity_overflow), taking every open file on the
page with it, and a hostile server only had to claim a large length and
serve a heap collection of 2 GiB + 4 KiB to get there (after fetching
2 GiB). Reading a large u8 dataset whole aborted the same way when its
values were widened to 64 bits.

- LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at
  most 1 GiB): a read longer than it fails at once, before anything is
  fetched, and an operation whose passes would fetch more than it fails
  before fetching (Operation::charge). assemble reserves fallibly.
- Reader::read refuses a read that would use more than 1 GiB while
  decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result)
  with an error naming readHyperslab, before reading.
- openUrl refuses a file of 4 GiB or more at open on wasm32: the format
  code turns offsets into usize, so nothing past 4 GiB can be read there
  (shown by a new test: data at 3 GiB reads, a 4 GiB file is refused).
  maxDownload is bounded to 1 GiB.

Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8
dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at
3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an
error or reads, and that the module survives. Before: RuntimeError:
unreachable in Node; the native test read the huge dataset and fetched
2 GiB.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 07:36:21 -05:00
osobhandClaude Opus 5.5 076feb089a wasm: restartable NeedBytes storage for lazy reads (range-read M4, core)
LazyStorage holds the blocks of a remote file fetched so far. An
operation runs as passes over it: a read that misses records the
missing blocks and fails, the pass's result is dropped whatever it is
(a parser may have caught the error and carried on), and the caller
fetches the reported ranges and re-runs the pass. No block is evicted
while an operation is in flight, so every pass that does not finish
asks for at least one new block and the operation ends. Blocks sit in
an LRU with a byte budget, trimmed between operations, bulk (raw data)
blocks first.

Reader::open_storage opens a file through any Storage, and variable-
length strings resolve through the file's storage instead of
File::as_bytes, which panics for a file not in memory.

tests/lazy.rs compares, file by file, what the viewer can show (kinds,
listings, attributes, info, whole reads and hyperslabs) read lazily
with the facade's range-storage path and the in-memory reader: files
built here at 512 B to 1 MiB blocks, the h5py/netCDF4 fixture, and with
CLAWHDF5_WASM_CORPUS the conformance corpus (656 files agree). A
listing plus a small read of a 48 MB file fetches 3 ranges.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 06:35:21 -05:00