docs: fewer round trips for remote files in the browser, counted

CHANGELOG (Unreleased): the v1 B-tree lookup, `Storage::hint`, the walks
that go on past a missing node, and the counts before and after on an
h5py file like the reviewer's (3000 datasets, 198 MB, earliest and
latest libver, 1 MiB and 64 KiB blocks), the corpus read lazily at
512 B and 64 KiB blocks, and the Node/Chromium suite.
known-issues (browser limits, round trips): the new counts, why the
passes cannot go lower (the chain of addresses), why merging nearby
requests does not help such a file, and that a second listing refetches
at 1 MiB blocks when the file's metadata blocks exceed `cacheSize`.
range-reads.md M4 status and the viewer README follow.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 22:48:47 -05:00
co-authored by Claude Opus 5.5
parent 761bdbf24f
commit 2e5b059530
4 changed files with 92 additions and 6 deletions
+55
View File
@@ -2,6 +2,61 @@
## Unreleased
### Remote files in the browser: fewer round trips to list a group or open a dataset (2026-09-27)
- **Opening one dataset of a v1 (symbol table) group no longer reads the
whole group.** A name is looked up down the group's B-tree, as
libhdf5's `H5G__stab_lookup` does (binary search on the node keys, names
in the local heap compared bytewise, then one symbol table node); only
when that finds no hard link of that name (a soft link, or a B-tree out
of name order, where libhdf5 would report it missing) is every entry
read, as before. Local files benefit too (a lookup read O(entries)).
In a group holding two entries of one name, the B-tree's is now the one
found, as in libhdf5.
- **`Storage::hint(offset, len)`** (clawhdf5-format): a parser says what
it reads next — a B-tree node's or symbol table node's body, an object
header's first chunk and continuation chunks, the symbol table nodes a
B-tree leaf names, a dense group's name index header and heap blocks,
and in a listing every child's object header. Every backend ignores it
but the browser's restartable reader (`clawhdf5_wasm::lazy`), which
fetches the hinted blocks it lacks together with the blocks a pass
missed, within the call's `maxFetch` budget; a pass that misses
nothing ignores them, so a hint never adds a round trip, and results
never depend on hints.
- The v1 and v2 B-tree walks of a listing descend into every child after
one fails (before, the siblings were only read, so their subtrees came
a pass later), then return the first error: same results and errors.
- Counted on tank, 2026-09-27, with `CLAWHDF5_WASM_LIST_FILE=<file>
CLAWHDF5_WASM_READ=/d1500 cargo test --release -p clawhdf5-wasm --test
lazy listing_cost_of_a_given_file -- --nocapture` on an h5py file like
the reviewer's (3000 datasets of 16384 `f32`, 198 MB, h5py 3.16 /
HDF5 2.0), passes / requests / bytes, before -> after:
| file, block size | `list('/')` | open + read one dataset |
|---|---|---|
| earliest, 1 MiB | 6 / 73 / 192.5 MB -> 4 / 68 / 192.5 MB | 7 / 74 / 193.6 MB -> 6 / 5 / 5.2 MB |
| earliest, 64 KiB | 8 / 531 / 35.2 MB -> 5 / 530 / 35.3 MB | 9 / 515 / 34.1 MB -> 8 / 7 / 0.52 MB |
| latest, 1 MiB | 9 / 98 / 196.5 MB -> 5 / 86 / 196.5 MB | 8 / 7 / 6.7 MB -> 7 / 7 / 6.7 MB |
| latest, 64 KiB | 11 / 452 / 29.6 MB -> 6 / 454 / 30.5 MB | 9 / 8 / 0.58 MB -> 8 / 8 / 0.58 MB |
The listing's passes now follow the depth of the group's index (the
chain index levels -> symbol table nodes or heap objects -> child
headers); its bytes are the child headers, which h5py spreads through
the file (at 1 MiB blocks most of it). The whole corpus read lazily
(`CLAWHDF5_WASM_CORPUS`, 656 files, every object listed, described and
read): 27 513 -> 27 496 passes, 3 001 -> 2 962 requests and 194.8 ->
195.3 MB at 64 KiB blocks; 41 342 -> 37 517 passes, 32 578 -> 29 511
requests, 65.4 -> 65.7 MB at 512 B. The Node and Chromium suite
(`examples/wasm-viewer/test/run.sh`) passes unchanged (the 200 MB file
still takes 5 requests, 6 MiB); its corpus comparison fetched 33.54 ->
33.61 MB.
- Tests: the listing budgets (`listing_a_large_group_takes_a_few_passes`,
512-byte blocks) are tightened to the new counts (FileBuilder, 600
children: 5 -> 4 passes; h5py, 2000 children: 8 -> 5 and 11 -> 6); new
`reading_one_dataset_of_a_large_group_fetches_a_few_blocks` (h5py
earliest: 529 requests, 333 kB -> at most 6 requests, 27 kB), v1
lookups against the listing (and with a name moved out of B-tree
order), hints riding only on misses and within the fetch budget.
### Deterministic errors on damaged chunked datasets (2026-09-27)
- A read through the file's chunk cache listed a damaged dataset's chunks in
hash-map order, seeded per `File`, so two opens of the same file could