Three items: fewer round trips when the browser reader opens a remote file, ObjectHeader::parse back to its pre-#18 speed, and the last conformance mismatches either fixed or shown with evidence not to be ours. Conformance: 600 → 602 of 697, 0 mismatches.
Browser: fewer round trips for remote files (range-read M4 follow-up)
Storage::hint(offset, len): new in clawhdf5-format, and a no-op by default. It lets a parser name the blocks it will read next. The wasm lazy reader fetches hinted blocks only alongside a pass's real misses and within the call's maxFetch, so a hint never adds a round trip, never fails a call, and never changes a result. Each hint is capped at 1 MiB, with a limit per pass.
Hints are added for:
group B-tree and symbol-table node bodies;
object headers' first and continuation chunks;
every child's header as soon as a listing reads its entry;
a dense group's name index and heap blocks.
Group walkers keep going past a missing block, so one pass finds every child. They still return the first error, so results and errors are unchanged.
Lookups in v1 groups now binary-search the group's B-tree, as libhdf5's H5G__stab_lookup does, instead of reading every entry. If the name isn't found that way, they fall back to the full read.
Measured on a 198 MB h5py file with 3000 datasets. These are request counts, so machine load doesn't affect them.
before (passes / requests / bytes)
after
open + read one dataset, 64 KiB blocks, earliest
9 / 515 / 34.1 MB
8 / 7 / 524 KB
open + read one dataset, 1 MiB blocks, earliest
7 / 74 / 193.6 MB
6 / 5 / 5.2 MB
list('/'), 1 MiB blocks, latest
9 / 98 / 196.5 MB
5 / 86 / 196.5 MB
list('/'), 64 KiB blocks, latest
11 / 452 / 29.6 MB
6 / 454 / 30.5 MB
Listing still reads most of this file, because h5py spreads the child headers through it. Classifying children without reading their headers could mislabel damaged files, so that wasn't done; it's documented in known-issues. Test budgets are tightened to the new numbers.
Self-check:
the corpus comparison at 512 B, 4 KiB and 64 KiB blocks matches the in-memory and range-storage readers on 656 files, errors included;
the h5rs fuzz sweep with mutations is clean;
it caught one bug, where a shared-subtree work bound kept walking after its budget ran out, which is fixed.
ObjectHeader::parse back to pre-#18 speed
The remaining cost was a single #[inline(never)] on the version-1 message loop from 4313917; switching it to #[inline] removes it. All safety properties are unchanged: bounded continuation chains, cycle refusal, the file-size budget, one chunk buffer at a time, libhdf5 message order, and overlapping chunks.
Idle A/B (tank, every round started at load 1.0–1.6 with no rustc running; separate binaries alternating, taskset -c 5, 3 rounds; median):
main 425585e
this branch
object_header_parse_x401
24.61 µs
23.33 µs (−5.2%)
snod_parse_all
1.865 µs
1.849 µs
btree_v1_walk
351 ns
343 ns
facade_list_400_groups
8.17 ms
8.16 ms
The agent's own A/B against 8f59b2e, from before #18, puts it 1.0–2.6% faster. The known-issues entry is closed.
The last non-ok conformance files
2 mismatches, now ok.attr_datatypes.hdf5 and tcomplex_be.h5 are h5py's big-endian VL bug: h5py returns the file's bytes under a little-endian dtype. h5dump prints what we read. ref.py checks at runtime that the installed h5py still has the bug, and if so compares against the corrected values.
3 files where libhdf5 reads past its buffers. New class ref-bug, with the evidence in the report:
cve-2025-2308: scale-offset codes need 17 bytes but the chunk holds 5.
cve-2025-44904: unfiltered chunks are stored short.
bad_nbit_parms_walk: the N-Bit parameter list is one value short. libhdf5's own test requires this read to fail.
conformance/ref_bugs.py re-reads each object in 6 processes with different heaps on every run. A file counts as ref-bug only while h5py's values actually vary in that run; otherwise it goes back to our-error. In this run bad_nbit_parms_walk.h5 read the same all 6 times, so it's counted as our error. The two CVE files are ref-bug.
Verification (tank, 2026-09-27)
scripts/ci-test.sh: 27/27 pass.
Browser suite under Node and headless Chromium: 256 package checks, 1240 openUrl checks, all browser checks.
Conformance: 602 of 697, 0 panics, hangs or crashes. The gate passes and the baseline is raised.
h5rs check --data: 0 of 437 fully-read files flagged.
Three items: fewer round trips when the browser reader opens a remote file, `ObjectHeader::parse` back to its pre-#18 speed, and the last conformance mismatches either fixed or shown with evidence not to be ours. **Conformance: 600 → 602 of 697, 0 mismatches.**
## Browser: fewer round trips for remote files (range-read M4 follow-up)
- **`Storage::hint(offset, len)`:** new in clawhdf5-format, and a no-op by default. It lets a parser name the blocks it will read next. The wasm lazy reader fetches hinted blocks only alongside a pass's real misses and within the call's `maxFetch`, so a hint never adds a round trip, never fails a call, and never changes a result. Each hint is capped at 1 MiB, with a limit per pass.
- **Hints are added for:**
- group B-tree and symbol-table node bodies;
- object headers' first and continuation chunks;
- every child's header as soon as a listing reads its entry;
- a dense group's name index and heap blocks.
- **Group walkers keep going past a missing block,** so one pass finds every child. They still return the first error, so results and errors are unchanged.
- **Lookups in v1 groups** now binary-search the group's B-tree, as libhdf5's `H5G__stab_lookup` does, instead of reading every entry. If the name isn't found that way, they fall back to the full read.
- **Measured** on a 198 MB h5py file with 3000 datasets. These are request counts, so machine load doesn't affect them.
| | before (passes / requests / bytes) | after |
|---|---|---|
| open + read one dataset, 64 KiB blocks, earliest | 9 / 515 / 34.1 MB | **8 / 7 / 524 KB** |
| open + read one dataset, 1 MiB blocks, earliest | 7 / 74 / 193.6 MB | **6 / 5 / 5.2 MB** |
| `list('/')`, 1 MiB blocks, latest | 9 / 98 / 196.5 MB | 5 / 86 / 196.5 MB |
| `list('/')`, 64 KiB blocks, latest | 11 / 452 / 29.6 MB | 6 / 454 / 30.5 MB |
Listing still reads most of this file, because h5py spreads the child headers through it. Classifying children without reading their headers could mislabel damaged files, so that wasn't done; it's documented in known-issues. Test budgets are tightened to the new numbers.
- **Self-check:**
- the corpus comparison at 512 B, 4 KiB and 64 KiB blocks matches the in-memory and range-storage readers on 656 files, errors included;
- the `h5rs` fuzz sweep with mutations is clean;
- it caught one bug, where a shared-subtree work bound kept walking after its budget ran out, which is fixed.
## `ObjectHeader::parse` back to pre-#18 speed
The remaining cost was a single `#[inline(never)]` on the version-1 message loop from `4313917`; switching it to `#[inline]` removes it. All safety properties are unchanged: bounded continuation chains, cycle refusal, the file-size budget, one chunk buffer at a time, libhdf5 message order, and overlapping chunks.
Idle A/B (tank, every round started at load 1.0–1.6 with no rustc running; separate binaries alternating, `taskset -c 5`, 3 rounds; median):
| | main `425585e` | this branch |
|---|---:|---:|
| `object_header_parse_x401` | 24.61 µs | **23.33 µs (−5.2%)** |
| `snod_parse_all` | 1.865 µs | 1.849 µs |
| `btree_v1_walk` | 351 ns | 343 ns |
| `facade_list_400_groups` | 8.17 ms | 8.16 ms |
The agent's own A/B against `8f59b2e`, from before #18, puts it 1.0–2.6% faster. The known-issues entry is closed.
## The last non-ok conformance files
- **2 mismatches, now ok.** `attr_datatypes.hdf5` and `tcomplex_be.h5` are h5py's big-endian VL bug: h5py returns the file's bytes under a little-endian dtype. `h5dump` prints what we read. `ref.py` checks at runtime that the installed h5py still has the bug, and if so compares against the corrected values.
- **3 files where libhdf5 reads past its buffers.** New class `ref-bug`, with the evidence in the report:
- `cve-2025-2308`: scale-offset codes need 17 bytes but the chunk holds 5.
- `cve-2025-44904`: unfiltered chunks are stored short.
- `bad_nbit_parms_walk`: the N-Bit parameter list is one value short. libhdf5's own test requires this read to fail.
- `conformance/ref_bugs.py` re-reads each object in 6 processes with different heaps on every run. A file counts as `ref-bug` only while h5py's values actually vary in that run; otherwise it goes back to `our-error`. In this run `bad_nbit_parms_walk.h5` read the same all 6 times, so it's counted as our error. The two CVE files are `ref-bug`.
## Verification (tank, 2026-09-27)
- `scripts/ci-test.sh`: **27/27 pass**.
- Browser suite under Node and headless Chromium: 256 package checks, 1240 `openUrl` checks, all browser checks.
- Conformance: 602 of 697, 0 panics, hangs or crashes. The gate passes and the baseline is raised.
- `h5rs check --data`: 0 of 437 fully-read files flagged.
- ClawBrainHub: builds and 204 of 204 tests pass.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us, on a heuristic (any big-endian VL mismatch) and
a fixed list.
- ref.py checks that the installed h5py returns big-endian VL elements
with the file's bytes under a little-endian dtype (writing and reading
a vlen('>f4') in memory) and, if so, relabels them with the file's byte
order before hashing, marking the object `ref_fix`. The values are now
compared: attr_datatypes.hdf5 /@vlen_uint64 and tcomplex_be.h5
/VariableLengthDatasetFloatComplex are identical to ours (h5dump 1.14.6
prints the same (1, 2), (3, 4, 5), (42)).
- ref_bugs.py re-reads each object h5py reads only through a libhdf5 bug
in six processes with different heaps (import order, MALLOC_PERTURB_).
Values the file determines are the same every time; these three change
(6, 6 and 3 distinct results), so they are over-read memory, not data
clawhdf5 could match. compare.py classifies a file `ref-bug` only when
every difference is such an object confirmed in the same run.
- report.py: the ref-bug class, the evidence table, the corrected
objects; test_ref.py covers both (run in the nightly job).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
4313917 kept parse_v1_messages out of line (#[inline(never)]) so the
generic parser would not carry the loop; since ef428d7 the slice path is
compiled once in this crate, and the call itself was the remaining cost
of the chunk queue: A/B builds of object_header_parse_x401 with only this
attribute changed put #[inline(never)] and no attribute at 24.5-24.9 us
and #[inline] at 23.6-24.0 us, with 8f59b2e at 23.6-24.1 us. Lazily
creating the chunk list only when a continuation is found (tried too)
measured no faster and was not kept.
Same code otherwise: every chunk-queue check (65,536 chunks, cycles, file
size budget, one chunk buffer at a time, libhdf5 order, overlap allowed)
is unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A ref-bug file's refused objects were still listed as our-error root
causes. The summary line now names every non-ok class and its count, and
an empty root-cause table says None.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A parser often learns where the next structures are (a node's children,
a structure's body once its prefix gives its length) before it reads
them one at a time. Over openUrl's restartable reader a structure only
reached after a miss costs a pass, and a round trip, of its own.
- clawhdf5-format: `Storage::hint(offset, len)`, "about to be read by
this operation". Default: nothing (every backend that reads when
asked); `&T`, `Box`, `Arc` and the facade's `FileData` forward it
(shifted past a user block, clamped to the file).
- clawhdf5-wasm `LazyStorage` records hinted blocks it lacks. A pass
that misses nothing ignores them (a hint never adds a round trip); a
pass that misses also asks for them, in file order, while the pass
stays within what is left of the operation's `maxFetch` budget (a
hint never makes a call fail). At most 1 MiB (or a block) per hint
and 65536 blocks per pass are recorded, whatever a file makes a
parser hint. `Operation::attempt` follows hints; the plain
`LazyStorage::attempt` does not. `run_blocking` and the browser's
driver use the former. `LazyStats::hinted_blocks` counts them.
No parser hints yet: results, passes and requests are unchanged.
Tests: hinted blocks come with a miss and never alone (cached ones
skipped, a plain attempt ignores them), and stay within the fetch
budget, a hint past the end or longer than the file harmless.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Listing a large group over openUrl still took 6-11 passes (network round
trips) for the reviewer's 3000-dataset h5py file: each pass only found
the structures the walk reached before its first miss.
- The v1 and v2 B-tree collectors descend into every child of a node
after one fails (they only read the siblings before, so a sibling's
subtree came a pass later), then return the first error: results and
errors unchanged. The v2 walk stops once its record budget is spent,
so a shared-subtree tree still cannot multiply the work.
- Hints (`Storage::hint`, a no-op for every backend but the lazy one):
a group B-tree node's and a symbol table node's body (read once their
header gives a length, a round trip later when the body is in the
next block), an object header's first chunk and its continuation
chunks, the symbol table nodes a B-tree leaf names, a dense group's
name index header and the heap's root block (both read right after
the heap header). A listing also hints every child's object header as
its entry is read, even after a failure, and every direct block of a
dense group's heap (reading the indirect blocks, at most 4096 entries
and 4 levels deep); a lookup does not.
- The fractal heap's indirect-block layout (entry sizes, where the first
n entries end) is one helper used by the object reads and the hints.
Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file
like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'),
passes/requests/bytes, before -> after:
earliest, 1 MiB: 6/73/192.5 MB -> 4/68/192.5 MB
earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB
latest, 1 MiB: 9/98/196.5 MB -> 5/86/196.5 MB
latest, 64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB
listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets
tightened to the new counts: FileBuilder 600 children 5 -> 4 passes,
h5py 2000 children earliest 8 -> 5, latest 11 -> 6.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).
`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.
Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
earliest, 1 MiB: 7/74/193.6 MB -> 6/5/5.2 MB
earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
-> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
CHANGELOG (Unreleased): the v1 B-tree lookup, `Storage::hint`, the walks
that go on past a missing node, and the counts before and after on an
h5py file like the reviewer's (3000 datasets, 198 MB, earliest and
latest libver, 1 MiB and 64 KiB blocks), the corpus read lazily at
512 B and 64 KiB blocks, and the Node/Chromium suite.
known-issues (browser limits, round trips): the new counts, why the
passes cannot go lower (the chain of addresses), why merging nearby
requests does not help such a file, and that a second listing refetches
at 1 MiB blocks when the file's metadata blocks exceed `cacheSize`.
range-reads.md M4 status and the viewer README follow.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Regenerated on tank: ok 600 -> 602 (attr_datatypes.hdf5 and
tcomplex_be.h5, compared against h5py's big-endian VL values corrected),
mismatch 2 -> 0; cve-2025-2308.h5 and cve-2025-44904.h5 are ref-bug
(h5py's values varied across six heaps in this run);
bad_nbit_parms_walk.h5 read the same in all six this run, so it stays
our-error, as the classification rule requires. Baseline raised.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
osobh
merged commit 9b5803f587 into main2026-09-28 11:57:04 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Three items: fewer round trips when the browser reader opens a remote file,
ObjectHeader::parseback to its pre-#18 speed, and the last conformance mismatches either fixed or shown with evidence not to be ours. Conformance: 600 → 602 of 697, 0 mismatches.Browser: fewer round trips for remote files (range-read M4 follow-up)
Storage::hint(offset, len): new in clawhdf5-format, and a no-op by default. It lets a parser name the blocks it will read next. The wasm lazy reader fetches hinted blocks only alongside a pass's real misses and within the call'smaxFetch, so a hint never adds a round trip, never fails a call, and never changes a result. Each hint is capped at 1 MiB, with a limit per pass.H5G__stab_lookupdoes, instead of reading every entry. If the name isn't found that way, they fall back to the full read.list('/'), 1 MiB blocks, latestlist('/'), 64 KiB blocks, latestListing still reads most of this file, because h5py spreads the child headers through it. Classifying children without reading their headers could mislabel damaged files, so that wasn't done; it's documented in known-issues. Test budgets are tightened to the new numbers.
h5rsfuzz sweep with mutations is clean;ObjectHeader::parseback to pre-#18 speedThe remaining cost was a single
#[inline(never)]on the version-1 message loop from4313917; switching it to#[inline]removes it. All safety properties are unchanged: bounded continuation chains, cycle refusal, the file-size budget, one chunk buffer at a time, libhdf5 message order, and overlapping chunks.Idle A/B (tank, every round started at load 1.0–1.6 with no rustc running; separate binaries alternating,
taskset -c 5, 3 rounds; median):425585eobject_header_parse_x401snod_parse_allbtree_v1_walkfacade_list_400_groupsThe agent's own A/B against
8f59b2e, from before #18, puts it 1.0–2.6% faster. The known-issues entry is closed.The last non-ok conformance files
attr_datatypes.hdf5andtcomplex_be.h5are h5py's big-endian VL bug: h5py returns the file's bytes under a little-endian dtype.h5dumpprints what we read.ref.pychecks at runtime that the installed h5py still has the bug, and if so compares against the corrected values.ref-bug, with the evidence in the report:cve-2025-2308: scale-offset codes need 17 bytes but the chunk holds 5.cve-2025-44904: unfiltered chunks are stored short.bad_nbit_parms_walk: the N-Bit parameter list is one value short. libhdf5's own test requires this read to fail.conformance/ref_bugs.pyre-reads each object in 6 processes with different heaps on every run. A file counts asref-bugonly while h5py's values actually vary in that run; otherwise it goes back toour-error. In this runbad_nbit_parms_walk.h5read the same all 6 times, so it's counted as our error. The two CVE files areref-bug.Verification (tank, 2026-09-27)
scripts/ci-test.sh: 27/27 pass.openUrlchecks, all browser checks.h5rs check --data: 0 of 437 fully-read files flagged.🤖 Generated with Claude Code
The last 3 our-errors and 2 mismatches were documented as not ours but still counted against us, on a heuristic (any big-endian VL mismatch) and a fixed list. - ref.py checks that the installed h5py returns big-endian VL elements with the file's bytes under a little-endian dtype (writing and reading a vlen('>f4') in memory) and, if so, relabels them with the file's byte order before hashing, marking the object `ref_fix`. The values are now compared: attr_datatypes.hdf5 /@vlen_uint64 and tcomplex_be.h5 /VariableLengthDatasetFloatComplex are identical to ours (h5dump 1.14.6 prints the same (1, 2), (3, 4, 5), (42)). - ref_bugs.py re-reads each object h5py reads only through a libhdf5 bug in six processes with different heaps (import order, MALLOC_PERTURB_). Values the file determines are the same every time; these three change (6, 6 and 3 distinct results), so they are over-read memory, not data clawhdf5 could match. compare.py classifies a file `ref-bug` only when every difference is such an object confirmed in the same run. - report.py: the ref-bug class, the evidence table, the corrected objects; test_ref.py covers both (run in the nightly job). Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>Listing a large group over openUrl still took 6-11 passes (network round trips) for the reviewer's 3000-dataset h5py file: each pass only found the structures the walk reached before its first miss. - The v1 and v2 B-tree collectors descend into every child of a node after one fails (they only read the siblings before, so a sibling's subtree came a pass later), then return the first error: results and errors unchanged. The v2 walk stops once its record budget is spent, so a shared-subtree tree still cannot multiply the work. - Hints (`Storage::hint`, a no-op for every backend but the lazy one): a group B-tree node's and a symbol table node's body (read once their header gives a length, a round trip later when the body is in the next block), an object header's first chunk and its continuation chunks, the symbol table nodes a B-tree leaf names, a dense group's name index header and the heap's root block (both read right after the heap header). A listing also hints every child's object header as its entry is read, even after a failure, and every direct block of a dense group's heap (reading the indirect blocks, at most 4096 entries and 4 levels deep); a lookup does not. - The fractal heap's indirect-block layout (entry sizes, where the first n entries end) is one helper used by the object reads and the hints. Measured with tests/lazy.rs listing_cost_of_a_given_file on an h5py file like the reviewer's (3000 datasets of 64 KiB, 198 MB), list('/'), passes/requests/bytes, before -> after: earliest, 1 MiB: 6/73/192.5 MB -> 4/68/192.5 MB earliest, 64 KiB: 8/531/35.2 MB -> 5/530/35.3 MB latest, 1 MiB: 9/98/196.5 MB -> 5/86/196.5 MB latest, 64 KiB: 11/452/29.6 MB -> 6/454/30.5 MB listing_a_large_group_takes_a_few_passes (512-byte blocks), budgets tightened to the new counts: FileBuilder 600 children 5 -> 4 passes, h5py 2000 children earliest 8 -> 5, latest 11 -> 6. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>8f59b2e) 5b45c60c9c