format: look a name up in a v1 group down its B-tree, as libhdf5 does
Opening one dataset of a v1 (symbol table) group read every symbol table node and every name of the group to find it: over openUrl, 74 requests and 193 MB to read one 64 KiB dataset of the reviewer's 3000-dataset h5py file (libver earliest) at 1 MiB blocks, 515 requests and 34 MB at 64 KiB. Locally it made a lookup O(entries). `group_v1::find_v1_entry` looks the name up as libhdf5's `H5G__stab_lookup` does: `H5B_find`'s binary search at each B-tree node with `H5G__node_cmp3` (left key < name <= right key, keys being names in the local heap, compared bytewise like strcmp), then the one symbol table node, after the heap's free list is checked as the listing does. The heap's data segment (up to 1 MiB) is hinted, since the keys are read one after another. Path resolution uses it for a v1 group; only when it does not find a hard link of that name (a soft link, or a B-tree out of name order, damaged or hand-made, where libhdf5 would report the name missing) does it read every entry as before. A storage error is returned as is (over the lazy reader, a miss: reading every entry would not get further). The result differs from before only in a group holding two entries of one name, where the B-tree's is now the one found, as in libhdf5. Measured with tests/lazy.rs listing_cost_of_a_given_file, open + read one 64 KiB dataset of the reviewer-like file, passes/requests/bytes, before -> after (open included): earliest, 1 MiB: 7/74/193.6 MB -> 6/5/5.2 MB earliest, 64 KiB: 9/515/34.1 MB -> 8/7/524 KB latest (dense groups, already a name-index lookup): 1 MiB 8/7/6.7 MB -> 7/7/6.7 MB, 64 KiB 9/8/581 KB -> 8/8/581 KB (the previous commit's hints: the name index header with the heap header) New tests, failing before: every child of v1_groups_400.h5 resolves to its listed address reading under 1/8 of the listing's bytes, and missing names are not found; a name moved out of B-tree order is still found (by the fallback); reading one of 2000 datasets lazily at 512-byte blocks takes at most 7 passes and 6 requests (earliest; 529 requests, 333 kB before) and 8 passes, 9 requests (latest). Conformance 600 of 697 (baseline 600), no file's class or detail changed against a run of main. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -475,11 +475,17 @@ fn corpus_files_read_the_same_lazily() {
|
||||
);
|
||||
}
|
||||
|
||||
/// Passes and requests `list(path)` takes on a file opened lazily at
|
||||
/// `block`-byte blocks (the open not counted), checking the listing against
|
||||
/// the in-memory one.
|
||||
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
|
||||
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
|
||||
/// What one call cost on a file opened lazily (the open not counted).
|
||||
#[derive(Debug, Clone, Copy)]
|
||||
struct Cost {
|
||||
passes: u64,
|
||||
requests: u64,
|
||||
bytes: u64,
|
||||
}
|
||||
|
||||
/// Open `data` lazily at `block`-byte blocks, run `op`, and return its
|
||||
/// result and what it cost (the open not counted).
|
||||
fn cost_of<T>(data: &[u8], block: u64, op: impl Fn(&Reader) -> T) -> (T, Cost) {
|
||||
let lazy = Lazy::open(
|
||||
data.to_vec(),
|
||||
LazyConfig {
|
||||
@@ -489,14 +495,41 @@ fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
|
||||
)
|
||||
.unwrap();
|
||||
let before = lazy.storage.stats();
|
||||
assert_eq!(lazy.call(|r| r.list(path)).unwrap(), want);
|
||||
let out = lazy.call(op);
|
||||
let after = lazy.storage.stats();
|
||||
(
|
||||
after.passes - before.passes,
|
||||
after.requests - before.requests,
|
||||
out,
|
||||
Cost {
|
||||
passes: after.passes - before.passes,
|
||||
requests: after.requests - before.requests,
|
||||
bytes: after.bytes_fetched - before.bytes_fetched,
|
||||
},
|
||||
)
|
||||
}
|
||||
|
||||
/// Passes and requests `list(path)` takes on a file opened lazily at
|
||||
/// `block`-byte blocks (the open not counted), checking the listing against
|
||||
/// the in-memory one.
|
||||
fn listing_cost(data: &[u8], path: &str, block: u64) -> (u64, u64) {
|
||||
let want = Reader::open(data.to_vec()).unwrap().list(path).unwrap();
|
||||
let (got, cost) = cost_of(data, block, |r| r.list(path));
|
||||
assert_eq!(got.unwrap(), want);
|
||||
(cost.passes, cost.requests)
|
||||
}
|
||||
|
||||
/// What reading the dataset at `path` whole takes on a file opened lazily
|
||||
/// at `block`-byte blocks (the open not counted), checking the values
|
||||
/// against the in-memory read.
|
||||
fn read_cost(data: &[u8], path: &str, block: u64) -> Cost {
|
||||
let want = Reader::open(data.to_vec())
|
||||
.unwrap()
|
||||
.read(path, None)
|
||||
.unwrap();
|
||||
let (got, cost) = cost_of(data, block, |r| r.read(path, None));
|
||||
assert_eq!(got.unwrap(), want, "{path}");
|
||||
cost
|
||||
}
|
||||
|
||||
/// An h5py file of `n` datasets of 256 `f32` each (`d0` ... ) in the root
|
||||
/// group, written with `libver`.
|
||||
fn h5py_many(dir: &Path, libver: &str, n: usize) -> Vec<u8> {
|
||||
@@ -561,6 +594,38 @@ fn listing_a_large_group_takes_a_few_passes() {
|
||||
}
|
||||
}
|
||||
|
||||
/// Opening one dataset of a large group looks its name up, not the whole
|
||||
/// group: in a v1 (symbol table) group down its B-tree as libhdf5 does, in
|
||||
/// a dense group down its name index. Reading one small dataset of 2000
|
||||
/// fetches a few blocks, where it used to read every symbol table node and
|
||||
/// name of a v1 group (529 requests, 333 kB at 512-byte blocks, before
|
||||
/// 2026-09-27).
|
||||
#[test]
|
||||
fn reading_one_dataset_of_a_large_group_fetches_a_few_blocks() {
|
||||
if !python_available() {
|
||||
assert!(
|
||||
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||
"CLAWHDF5_REQUIRE_INTEROP=1 but {} lacks h5py/netCDF4/numpy",
|
||||
python()
|
||||
);
|
||||
eprintln!("skipping: {} lacks h5py/netCDF4/numpy", python());
|
||||
return;
|
||||
}
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
for (libver, most_passes, most_requests) in [("earliest", 7, 6), ("latest", 8, 9)] {
|
||||
let data = h5py_many(dir.path(), libver, 2000);
|
||||
for path in ["/d0", "/d1234", "/d1999"] {
|
||||
let cost = read_cost(&data, path, 512);
|
||||
eprintln!("h5py libver={libver}, 2000 children, read {path}: {cost:?}");
|
||||
assert!(
|
||||
cost.passes <= most_passes && cost.requests <= most_requests,
|
||||
"{libver} {path}: {cost:?}"
|
||||
);
|
||||
assert!(cost.bytes <= 64 * 1024, "{libver} {path}: {cost:?}");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// `CLAWHDF5_WASM_LIST_FILE=file.h5`: what listing the root group of that
|
||||
/// file costs lazily, and opening it and reading one dataset whole
|
||||
/// (`CLAWHDF5_WASM_READ`, by default the middle dataset of the listing),
|
||||
|
||||
Reference in New Issue
Block a user