Regenerated on tank: ok 600 -> 602 (attr_datatypes.hdf5 and
tcomplex_be.h5, compared against h5py's big-endian VL values corrected),
mismatch 2 -> 0; cve-2025-2308.h5 and cve-2025-44904.h5 are ref-bug
(h5py's values varied across six heaps in this run);
bad_nbit_parms_walk.h5 read the same in all six this run, so it stays
our-error, as the classification rule requires. Baseline raised.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased): the v1 B-tree lookup, `Storage::hint`, the walks
that go on past a missing node, and the counts before and after on an
h5py file like the reviewer's (3000 datasets, 198 MB, earliest and
latest libver, 1 MiB and 64 KiB blocks), the corpus read lazily at
512 B and 64 KiB blocks, and the Node/Chromium suite.
known-issues (browser limits, round trips): the new counts, why the
passes cannot go lower (the chain of addresses), why merging nearby
requests does not help such a file, and that a second listing refetches
at 1 MiB blocks when the file's metadata blocks exceed `cacheSize`.
range-reads.md M4 status and the viewer README follow.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).
`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.
Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
earliest, 1 MiB: 7/74/193.6 MB -> 6/5/5.2 MB
earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
-> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.
- The v1 and v2 B-tree collectors descend into every child of a node
after one fails (they only read the siblings before, so a sibling's
subtree came a pass later), then return the first error: results and
errors unchanged. The v2 walk stops once its record budget is spent,
so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
a group B-tree node's and a symbol table node's body (read once their
header gives a length, a round trip later when the body is in the
next block), an object header's first chunk and its continuation
chunks, the symbol table nodes a B-tree leaf names, a dense group's
name index header and the heap's root block (both read right after
the heap header). A listing also hints every child's object header as
its entry is read, even after a failure, and every direct block of a
dense group's heap (reading the indirect blocks, at most 4096 entries
and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
n entries end) is one helper used by the object reads and the hints.
Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
earliest, 1 MiB: 6/73/192.5 MB -> 4/68/192.5 MB
earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
latest, 1 MiB: 9/98/196.5 MB -> 5/86/196.5 MB
latest, 64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A parser often learns where the next structures are (a node's children,
a structure's body once its prefix gives its length) before it reads
them one at a time. Over openUrl's restartable reader a structure only
reached after a miss costs a pass, and a round trip, of its own.
- clawhdf5-format: `Storage::hint(offset, len)`, "about to be read by
this operation". Default: nothing (every backend that reads when
asked); `&T`, `Box`, `Arc` and the facade's `FileData` forward it
(shifted past a user block, clamped to the file).
- clawhdf5-wasm `LazyStorage` records hinted blocks it lacks. A pass
that misses nothing ignores them (a hint never adds a round trip); a
pass that misses also asks for them, in file order, while the pass
stays within what is left of the operation's `maxFetch` budget (a
hint never makes a call fail). At most 1 MiB (or a block) per hint
and 65536 blocks per pass are recorded, whatever a file makes a
parser hint. `Operation::attempt` follows hints; the plain
`LazyStorage::attempt` does not. `run_blocking` and the browser's
driver use the former. `LazyStats::hinted_blocks` counts them.
No parser hints yet: results, passes and requests are unchanged.
Tests: hinted blocks come with a miss and never alone (cached ones
skipped, a plain attempt ignores them), and stay within the fetch
budget, a hint past the end or longer than the file harmless.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A ref-bug file's refused objects were still listed as our-error root
causes. The summary line now names every non-ok class and its count, and
an empty root-cause table says None.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
4313917 kept parse_v1_messages out of line (#[inline(never)]) so the
generic parser would not carry the loop; since ef428d7 the slice path is
compiled once in this crate, and the call itself was the remaining cost
of the chunk queue: A/B builds of object_header_parse_x401 with only this
attribute changed put #[inline(never)] and no attribute at 24.5-24.9 us
and #[inline] at 23.6-24.0 us, with 8f59b2e at 23.6-24.1 us. Lazily
creating the chunk list only when a continuation is found (tried too)
measured no faster and was not kept.
Same code otherwise: every chunk-queue check (65,536 chunks, cycles, file
size budget, one chunk buffer at a time, libhdf5 order, overlap allowed)
is unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us, on a heuristic (any big-endian VL mismatch) and
a fixed list.
- ref.py checks that the installed h5py returns big-endian VL elements
with the file's bytes under a little-endian dtype (writing and reading
a vlen('>f4') in memory) and, if so, relabels them with the file's byte
order before hashing, marking the object `ref_fix`. The values are now
compared: attr_datatypes.hdf5 /@vlen_uint64 and tcomplex_be.h5
/VariableLengthDatasetFloatComplex are identical to ours (h5dump 1.14.6
prints the same (1, 2), (3, 4, 5), (42)).
- ref_bugs.py re-reads each object h5py reads only through a libhdf5 bug
in six processes with different heaps (import order, MALLOC_PERTURB_).
Values the file determines are the same every time; these three change
(6, 6 and 3 distinct results), so they are over-read memory, not data
clawhdf5 could match. compare.py classifies a file `ref-bug` only when
every difference is such an object confirmed in the same run.
- report.py: the ref-bug class, the evidence table, the corrected
objects; test_ref.py covers both (run in the nightly job).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>