cab952bb007b1435c9cb034a92d729f33a7ed05a
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
761bdbf24f |
format: look a name up in a v1 group down its B-tree, as libhdf5 does
Opening one dataset of a v1 (symbol table) group read every symbol table node and every name of the group to find it: over openUrl, 74 requests and 193 MB to read one 64 KiB dataset of the reviewer's 3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests and 34 MB at 64 KiB. Locally it made a lookup O(entries). `group_v1::find_v1_entry` looks the name up as libhdf5's `H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node with `H5G__node_cmp3` (left key < name <= right key, keys being names in the local heap, compared bytewise like strcmp), then the one symbol table node, after the heap's free list is checked as the listing does. The heap's data segment (up to 1 MiB) is hinted, since the keys are read one after another. Path resolution uses it for a v1 group; only when it does not find a hard link of that name (a soft link, or a B-tree out of name order, damaged or hand-made, where libhdf5 would report the name missing) does it read every entry as before. A storage error is returned as is (over the lazy reader, a miss: reading every entry would not get further). The result differs from before only in a group holding two entries of one name, where the B-tree's is now the one found, as in libhdf5. Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read one 64 KiB dataset of the reviewer-like file, passes/requests/bytes, before -> after (open included): earliest, 1 MiB: 7/74/193.6 MB -> 6/5/5.2 MB earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB -> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous commit's hints: the name index header with the heap header) New tests, failing before: every child of v1_groups_400.h5 resolves to its listed address reading under 1/8 of the listing's bytes, and missing names are not found; a name moved out of B-tree order is still found (by the fallback); reading one of 2000 datasets lazily at 512-byte blocks takes at most 7 passes and 6 requests (earliest; 529 requests, 333 kB before) and 8 passes, 9 requests (latest). Conformance 600 of 697 (baseline 600), no file's class or detail changed against a run of main. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
6f5d14fd62 |
format: group walks go on past a failed node and hint what they read next
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.
- The v1 and v2 B-tree collectors descend into every child of a node
after one fails (they only read the siblings before, so a sibling's
subtree came a pass later), then return the first error: results and
errors unchanged. The v2 walk stops once its record budget is spent,
so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
a group B-tree node's and a symbol table node's body (read once their
header gives a length, a round trip later when the body is in the
next block), an object header's first chunk and its continuation
chunks, the symbol table nodes a B-tree leaf names, a dense group's
name index header and the heap's root block (both read right after
the heap header). A listing also hints every child's object header as
its entry is read, even after a failure, and every direct block of a
dense group's heap (reading the indirect blocks, at most 4096 entries
and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
n entries end) is one helper used by the object reads and the hints.
Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
earliest, 1 MiB: 6/73/192.5 MB -> 4/68/192.5 MB
earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
latest, 1 MiB: 9/98/196.5 MB -> 5/86/196.5 MB
latest, 64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
||
|
|
e553153e48 |
wasm, format: listing a group asks for all its missing blocks per pass
Listing a group read every child's object header and stopped at the
first that was not fetched yet, and so did the traversals of the group's
index (v1 B-tree and symbol table nodes, the local heap's names, v2
B-tree nodes and fractal heap objects). Over openUrl's restartable
reader each block cost its own pass and round trip: 184 serial requests
to list 3000 datasets at 1 MiB blocks, 536 at 64 KiB.
- core::Reader::list reads every child's header before returning the
first error (the same error, in listing order, Group::groups/datasets
return), classifying them as those do.
- clawhdf5-format: after the first sibling that fails, the B-tree v1
and v2 collectors, the symbol table node loop and the dense-link loop
go on reading (not using) the remaining siblings, then return that
first error: results and errors are unchanged, only failing
traversals read more, and in memory that is free (storage::touch).
A v1 group's local heap segment (names) is read at once, up to 1 MiB.
- LazyStorage no longer fills a one-block hole that is already cached
(it was fetched again: 215 MB fetched from a 198 MB file).
Measured with tests/lazy.rs listing_cost_of_a_given_file on the
reviewer's file (h5py, 3000 datasets of 64 KiB, 198 MB), list('/'):
libver earliest, 1 MiB blocks: 185 passes/184 requests -> 6/73
libver earliest, 64 KiB: 537/536 -> 8/531 (6 in flight)
libver latest, 1 MiB: 189/188 -> 9/98
libver latest, 64 KiB: 453/452 -> 11/452
Bytes fetched are unchanged (the headers are spread through the file).
New test listing_a_large_group_takes_a_few_passes (512-byte blocks):
FileBuilder 600 children 102 -> 5 passes; h5py earliest/latest 2000
children 8 and 11 passes. Conformance 600 of 697 (baseline 600);
check-32bit-casts, check-nostd and h5rs-fuzz over the CVE corpus clean.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
||
|
|
dbafa952ac |
wasm: sizes a server or a dataset names are errors, not aborts
A read longer than isize::MAX (2 GiB on wasm32) aborted the module in LazyStorage::assemble (capacity_overflow), taking every open file on the page with it, and a hostile server only had to claim a large length and serve a heap collection of 2 GiB + 4 KiB to get there (after fetching 2 GiB). Reading a large u8 dataset whole aborted the same way when its values were widened to 64 bits. - LazyConfig::max_fetch (openUrl option maxFetch, default 512 MiB, at most 1 GiB): a read longer than it fails at once, before anything is fetched, and an operation whose passes would fetch more than it fails before fetching (Operation::charge). assemble reserves fallibly. - Reader::read refuses a read that would use more than 1 GiB while decoding (core::MAX_READ_BYTES: stored bytes + 64-bit values + result) with an error naming readHyperslab, before reading. - openUrl refuses a file of 4 GiB or more at open on wasm32: the format code turns offsets into usize, so nothing past 4 GiB can be read there (shown by a new test: data at 3 GiB reads, a 4 GiB file is refused). maxDownload is bounded to 1 GiB. Tests: make_fixture.py writes limits.h5 (a sparse 2^28 + 1024 byte u8 dataset), hostile_vl.h5 (the reviewer's collection) and far.h5 (data at 3 GiB); test.mjs (wasm32) and tests/lazy.rs (native) check each is an error or reads, and that the module survives. Before: RuntimeError: unreachable in Node; the native test read the huge dataset and fetched 2 GiB. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |
||
|
|
076feb089a |
wasm: restartable NeedBytes storage for lazy reads (range-read M4, core)
LazyStorage holds the blocks of a remote file fetched so far. An operation runs as passes over it: a read that misses records the missing blocks and fails, the pass's result is dropped whatever it is (a parser may have caught the error and carried on), and the caller fetches the reported ranges and re-runs the pass. No block is evicted while an operation is in flight, so every pass that does not finish asks for at least one new block and the operation ends. Blocks sit in an LRU with a byte budget, trimmed between operations, bulk (raw data) blocks first. Reader::open_storage opens a file through any Storage, and variable- length strings resolve through the file's storage instead of File::as_bytes, which panics for a file not in memory. tests/lazy.rs compares, file by file, what the viewer can show (kinds, listings, attributes, info, whole reads and hyperslabs) read lazily with the facade's range-storage path and the in-memory reader: files built here at 512 B to 1 MiB blocks, the h5py/netCDF4 fixture, and with CLAWHDF5_WASM_CORPUS the conformance corpus (656 files agree). A listing plus a small read of a 48 MB file fetches 3 ranges. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> |