Remote files open in a few requests; ObjectHeader::parse back to speed; last conformance mismatches resolved (602/697) #21
@@ -2,6 +2,61 @@
|
|||||||
|
|
||||||
## Unreleased
|
## Unreleased
|
||||||
|
|
||||||
|
### Remote files in the browser: fewer round trips to list a group or open a dataset (2026-09-27)
|
||||||
|
- **Opening one dataset of a v1 (symbol table) group no longer reads the
|
||||||
|
whole group.** A name is looked up down the group's B-tree, as
|
||||||
|
libhdf5's `H5G__stab_lookup` does (binary search on the node keys, names
|
||||||
|
in the local heap compared bytewise, then one symbol table node); only
|
||||||
|
when that finds no hard link of that name (a soft link, or a B-tree out
|
||||||
|
of name order, where libhdf5 would report it missing) is every entry
|
||||||
|
read, as before. Local files benefit too (a lookup read O(entries)).
|
||||||
|
In a group holding two entries of one name, the B-tree's is now the one
|
||||||
|
found, as in libhdf5.
|
||||||
|
- **`Storage::hint(offset, len)`** (clawhdf5-format): a parser says what
|
||||||
|
it reads next — a B-tree node's or symbol table node's body, an object
|
||||||
|
header's first chunk and continuation chunks, the symbol table nodes a
|
||||||
|
B-tree leaf names, a dense group's name index header and heap blocks,
|
||||||
|
and in a listing every child's object header. Every backend ignores it
|
||||||
|
but the browser's restartable reader (`clawhdf5_wasm::lazy`), which
|
||||||
|
fetches the hinted blocks it lacks together with the blocks a pass
|
||||||
|
missed, within the call's `maxFetch` budget; a pass that misses
|
||||||
|
nothing ignores them, so a hint never adds a round trip, and results
|
||||||
|
never depend on hints.
|
||||||
|
- The v1 and v2 B-tree walks of a listing descend into every child after
|
||||||
|
one fails (before, the siblings were only read, so their subtrees came
|
||||||
|
a pass later), then return the first error: same results and errors.
|
||||||
|
- Counted on tank, 2026-09-27, with `CLAWHDF5_WASM_LIST_FILE=<file>
|
||||||
|
CLAWHDF5_WASM_READ=/d1500 cargo test --release -p clawhdf5-wasm --test
|
||||||
|
lazy listing_cost_of_a_given_file -- --nocapture` on an h5py file like
|
||||||
|
the reviewer's (3000 datasets of 16384 `f32`, 198 MB, h5py 3.16 /
|
||||||
|
HDF5 2.0), passes / requests / bytes, before -> after:
|
||||||
|
|
||||||
|
| file, block size | `list('/')` | open + read one dataset |
|
||||||
|
|---|---|---|
|
||||||
|
| earliest, 1 MiB | 6 / 73 / 192.5 MB -> 4 / 68 / 192.5 MB | 7 / 74 / 193.6 MB -> 6 / 5 / 5.2 MB |
|
||||||
|
| earliest, 64 KiB | 8 / 531 / 35.2 MB -> 5 / 530 / 35.3 MB | 9 / 515 / 34.1 MB -> 8 / 7 / 0.52 MB |
|
||||||
|
| latest, 1 MiB | 9 / 98 / 196.5 MB -> 5 / 86 / 196.5 MB | 8 / 7 / 6.7 MB -> 7 / 7 / 6.7 MB |
|
||||||
|
| latest, 64 KiB | 11 / 452 / 29.6 MB -> 6 / 454 / 30.5 MB | 9 / 8 / 0.58 MB -> 8 / 8 / 0.58 MB |
|
||||||
|
|
||||||
|
The listing's passes now follow the depth of the group's index (the
|
||||||
|
chain index levels -> symbol table nodes or heap objects -> child
|
||||||
|
headers); its bytes are the child headers, which h5py spreads through
|
||||||
|
the file (at 1 MiB blocks most of it). The whole corpus read lazily
|
||||||
|
(`CLAWHDF5_WASM_CORPUS`, 656 files, every object listed, described and
|
||||||
|
read): 27 513 -> 27 496 passes, 3 001 -> 2 962 requests and 194.8 ->
|
||||||
|
195.3 MB at 64 KiB blocks; 41 342 -> 37 517 passes, 32 578 -> 29 511
|
||||||
|
requests, 65.4 -> 65.7 MB at 512 B. The Node and Chromium suite
|
||||||
|
(`examples/wasm-viewer/test/run.sh`) passes unchanged (the 200 MB file
|
||||||
|
still takes 5 requests, 6 MiB); its corpus comparison fetched 33.54 ->
|
||||||
|
33.61 MB.
|
||||||
|
- Tests: the listing budgets (`listing_a_large_group_takes_a_few_passes`,
|
||||||
|
512-byte blocks) are tightened to the new counts (FileBuilder, 600
|
||||||
|
children: 5 -> 4 passes; h5py, 2000 children: 8 -> 5 and 11 -> 6); new
|
||||||
|
`reading_one_dataset_of_a_large_group_fetches_a_few_blocks` (h5py
|
||||||
|
earliest: 529 requests, 333 kB -> at most 6 requests, 27 kB), v1
|
||||||
|
lookups against the listing (and with a name moved out of B-tree
|
||||||
|
order), hints riding only on misses and within the fetch budget.
|
||||||
|
|
||||||
### Deterministic errors on damaged chunked datasets (2026-09-27)
|
### Deterministic errors on damaged chunked datasets (2026-09-27)
|
||||||
- A read through the file's chunk cache listed a damaged dataset's chunks in
|
- A read through the file's chunk cache listed a damaged dataset's chunks in
|
||||||
hash-map order, seeded per `File`, so two opens of the same file could
|
hash-map order, seeded per `File`, so two opens of the same file could
|
||||||
|
|||||||
@@ -607,6 +607,16 @@ fast path within benchmark noise.
|
|||||||
missing blocks: listing 3000 datasets went from 185 passes to 6.
|
missing blocks: listing 3000 datasets went from 185 passes to 6.
|
||||||
The 32-bit risk below is covered by a Node test that reads data at
|
The 32-bit risk below is covered by a Node test that reads data at
|
||||||
3 GiB from a mock server and is refused a 4 GiB file.
|
3 GiB from a mock server and is refused a 4 GiB file.
|
||||||
|
- Fewer round trips (2026-09-27, later): the walks descend into every
|
||||||
|
child after a failure (not only read the siblings), and parsers
|
||||||
|
call `Storage::hint` for what they read next (node bodies, object
|
||||||
|
header chunks, a dense group's heap blocks, a listing's child
|
||||||
|
headers); `LazyStorage` fetches hinted blocks only along with a
|
||||||
|
pass's real misses and within `maxFetch`, so hints never add a
|
||||||
|
round trip nor change a result. A v1 group's name is looked up down
|
||||||
|
its B-tree (`H5G__stab_lookup`), not by listing it. 3000 datasets
|
||||||
|
list in 4 passes (earliest) and 5 (latest) at 1 MiB blocks, and
|
||||||
|
opening one of them costs 5 requests, not 74 (CHANGELOG).
|
||||||
|
|
||||||
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
||||||
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
||||||
|
|||||||
+23
-4
@@ -1051,10 +1051,29 @@ cache, but:
|
|||||||
header, and every node of a level of the group's index, in one pass
|
header, and every node of a level of the group's index, in one pass
|
||||||
(since 2026-09-27; it was one round trip per header block): 3000
|
(since 2026-09-27; it was one round trip per header block): 3000
|
||||||
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
|
datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a
|
||||||
`libver="latest"` file (dense links). Each pass re-parses what the
|
`libver="latest"` file (dense links). **Since 2026-09-27 (later):**
|
||||||
call reads (CPU, not network). With headers spread through the file
|
4 and 5 passes (5 and 6 at 64 KiB, from 8 and 11): the index walks go
|
||||||
(h5py writes each next to its data) a listing still fetches most of
|
on past a missing node, and parsers hint what they read next
|
||||||
the file at 1 MiB blocks; a smaller `blockSize` fetches less.
|
(`Storage::hint`: node bodies, the heap's blocks, each child's
|
||||||
|
header), which the lazy reader fetches with a pass's misses. That is
|
||||||
|
the depth of the chain (index levels, then symbol table nodes or
|
||||||
|
heap objects, then headers) plus the pass that finishes; it cannot
|
||||||
|
go lower without reading structures before their addresses are
|
||||||
|
known. Opening one dataset of a v1 group looks its name up down the
|
||||||
|
group's B-tree (it read every entry: 74 requests, 193 MB at 1 MiB
|
||||||
|
blocks for one 64 KiB dataset of the 3000; now 5 requests, 5 MB).
|
||||||
|
Each pass re-parses what the call reads (CPU, not network). With
|
||||||
|
headers spread through the file (h5py writes each next to its data)
|
||||||
|
a listing still fetches most of the file at 1 MiB blocks (192 of
|
||||||
|
198 MB; 35 MB in 530 requests at 64 KiB); a smaller `blockSize`
|
||||||
|
fetches less. Merging nearby requests does not help such a file: the
|
||||||
|
blocks a listing needs are five or six apart at 64 KiB, so fewer requests would
|
||||||
|
mean fetching most of the file. Listing it a second time is free at
|
||||||
|
64 KiB blocks, but at 1 MiB its metadata blocks (192 MB) exceed the
|
||||||
|
64 MiB `cacheSize`, so they are fetched again (the earliest file: 4
|
||||||
|
passes, 50 requests); a larger `cacheSize` keeps them. A file's paged
|
||||||
|
aggregation (metadata in pages) is not used to fetch its metadata in
|
||||||
|
one request.
|
||||||
- **Memory:** a call keeps every block it reads until it finishes (the
|
- **Memory:** a call keeps every block it reads until it finishes (the
|
||||||
cache budget applies between calls). It may fetch at most `maxFetch`
|
cache budget applies between calls). It may fetch at most `maxFetch`
|
||||||
bytes (512 MiB by default, at most 1 GiB), and a single read longer
|
bytes (512 MiB by default, at most 1 GiB), and a single read longer
|
||||||
|
|||||||
@@ -70,8 +70,10 @@ are fetched (in parallel, adjacent blocks in one request), and the pass is
|
|||||||
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
run again, until one completes (`docs/design/range-reads.md`, M4). Opening
|
||||||
costs one request (the first block, which also gives the file's size);
|
costs one request (the first block, which also gives the file's size);
|
||||||
listing a group whose metadata is in blocks already fetched costs none,
|
listing a group whose metadata is in blocks already fetched costs none,
|
||||||
and otherwise a round trip per level of the group's index plus one for
|
and otherwise about a round trip per level of the group's index plus one
|
||||||
its children's headers, all fetched together;
|
for its children's headers, all fetched together (the reader fetches
|
||||||
|
what it knows it reads next along with what a pass missed); opening one
|
||||||
|
object looks its name up in the group's index, not the whole group;
|
||||||
reading a chunked dataset costs a round trip for its chunk index (a few
|
reading a chunked dataset costs a round trip for its chunk index (a few
|
||||||
for a deep one) and one batch of requests for its chunks. Every answer is
|
for a deep one) and one batch of requests for its chunks. Every answer is
|
||||||
checked — a `206` with exactly the bytes asked for, from the same file
|
checked — a `206` with exactly the bytes asked for, from the same file
|
||||||
|
|||||||
Reference in New Issue
Block a user