format: look a name up in a v1 group down its B-tree, as libhdf5 does

Opening one dataset of a v1 (symbol table) group read every symbol
table node and every name of the group to find it: over openUrl, 74
requests and 193 MB to read one 64 KiB dataset of the reviewer's
3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests
and 34 MB at 64 KiB. Locally it made a lookup O(entries).

`group_v1::find_v1_entry` looks the name up as libhdf5's
`H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node
with `H5G__node_cmp3` (left key < name <= right key, keys being names in
the local heap, compared bytewise like strcmp), then the one symbol
table node, after the heap's free list is checked as the listing does.
The heap's data segment (up to 1 MiB) is hinted, since the keys are read
one after another. Path resolution uses it for a v1 group; only when it
does not find a hard link of that name (a soft link, or a B-tree out of
name order, damaged or hand-made, where libhdf5 would report the name
missing) does it read every entry as before. A storage error is returned
as is (over the lazy reader, a miss: reading every entry would not get
further). The result differs from before only in a group holding two
entries of one name, where the B-tree's is now the one found, as in
libhdf5.

Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read
one 64 KiB dataset of the reviewer-like file, passes/requests/bytes,
before -> after (open included):
  earliest, 1 MiB:  7/74/193.6 MB -> 6/5/5.2 MB
  earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB
  latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB
  -> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous
  commit's hints: the name index header with the heap header)
New tests, failing before: every child of v1_groups_400.h5 resolves
to its listed address reading under 1/8 of the listing's bytes, and
missing names are not found; a name moved out of B-tree order is still
found (by the fallback); reading one of 2000 datasets lazily at 512-byte
blocks takes at most 7 passes and 6 requests (earliest; 529 requests,
333 kB before) and 8 passes, 9 requests (latest).
Conformance 600 of 697 (baseline 600), no file's class or detail
changed against a run of main.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 22:47:47 -05:00
co-authored by Claude Opus 5.5
parent 6f5d14fd62
commit 761bdbf24f
4 changed files with 265 additions and 10 deletions
+73 -8
View File
@@ -475,11 +475,17 @@ fn corpus_files_read_the_same_lazily() {
);
}
/// Passes and requests `list(path)` takes on a file opened lazily at
/// `block`-byte blocks (the open not counted), checking the listing against
/// the in-memory one.
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
/// What one call cost on a file opened lazily (the open not counted).
#[derive(Debug, Clone, Copy)]
struct Cost {
passes: u64,
requests: u64,
bytes: u64,
}
/// Open `data` lazily at `block`-byte blocks, run `op`, and return its
/// result and what it cost (the open not counted).
fn cost_of<T>(data: &[u8], block: u64, op: impl Fn(&Reader) -> T) -> (T, Cost) {
let lazy = Lazy::open(
data.to_vec(),
LazyConfig {
@@ -489,14 +495,41 @@ fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
)
.unwrap();
let before = lazy.storage.stats();
assert_eq!(lazy.call(|r| r.list(path)).unwrap(), want);
let out = lazy.call(op);
let after = lazy.storage.stats();
(
after.passes - before.passes,
after.requests - before.requests,
out,
Cost {
passes: after.passes - before.passes,
requests: after.requests - before.requests,
bytes: after.bytes_fetched - before.bytes_fetched,
},
)
}
/// Passes and requests `list(path)` takes on a file opened lazily at
/// `block`-byte blocks (the open not counted), checking the listing against
/// the in-memory one.
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
let (got, cost) = cost_of(data, block, |r| r.list(path));
assert_eq!(got.unwrap(), want);
(cost.passes, cost.requests)
}
/// What reading the dataset at `path` whole takes on a file opened lazily
/// at `block`-byte blocks (the open not counted), checking the values
/// against the in-memory read.
fn read_cost(data: &[u8], path: &str, block: u64) -> Cost {
let want = Reader::open(data.to_vec())
.unwrap()
.read(path, None)
.unwrap();
let (got, cost) = cost_of(data, block, |r| r.read(path, None));
assert_eq!(got.unwrap(), want, "{path}");
cost
}
/// An h5py file of `n` datasets of 256 `f32` each (`d0` ... ) in the root
/// group, written with `libver`.
fn h5py_many(dir: &Path, libver: &str, n: usize) -> Vec<u8> {
@@ -561,6 +594,38 @@ fn listing_a_large_group_takes_a_few_passes() {
}
}
/// Opening one dataset of a large group looks its name up, not the whole
/// group: in a v1 (symbol table) group down its B-tree as libhdf5 does, in
/// a dense group down its name index. Reading one small dataset of 2000
/// fetches a few blocks, where it used to read every symbol table node and
/// name of a v1 group (529 requests, 333 kB at 512-byte blocks, before
/// 2026-09-27).
#[test]
fn reading_one_dataset_of_a_large_group_fetches_a_few_blocks() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
python()
);
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
return;
}
let dir = tempfile::tempdir().unwrap();
for (libver, most_passes, most_requests) in [("earliest", 7, 6), ("latest", 8, 9)] {
let data = h5py_many(dir.path(), libver, 2000);
for path in ["/d0", "/d1234", "/d1999"] {
let cost = read_cost(&data, path, 512);
eprintln!("h5py libver={libver}, 2000 children, read {path}: {cost:?}");
assert!(
cost.passes <= most_passes && cost.requests <= most_requests,
"{libver} {path}: {cost:?}"
);
assert!(cost.bytes <= 64 * 1024, "{libver} {path}: {cost:?}");
}
}
}
/// `CLAWHDF5_WASM_LIST_FILE=file.h5`: what listing the root group of that
/// file costs lazily, and opening it and reading one dataset whole
/// (`CLAWHDF5_WASM_READ`, by default the middle dataset of the listing),