docs: local metadata and data reads after range-read M2/M3, rechecked

8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.

The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 15:35:59 -05:00
co-authored by Claude Opus 5.5
parent 7a8fae0357
commit d239292655
3 changed files with 119 additions and 0 deletions
+17
View File
@@ -477,6 +477,19 @@ Design: `docs/design/swmr.md`.
`main` and it calls only format-crate code; its earlier +5–10% (323–355
vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured
357–374 ns against `main`'s 355–371 in the same rounds).
**Rechecked 2026-09-27** (still provisional: two orphaned h5py test
processes kept tank's 1-minute load at 2.1–2.6, the load < 2 gate was
not met in 2 hours, load 2.3–3.3 during the runs; `8f59b2e` vs `main`
`7a8fae0` as separate binaries, pinned to one CPU, 6 alternating rounds,
median (range) over rounds; `BENCHMARKS.md`, "Local metadata and data
reads after range-read M2/M3"): facade listing 8.26 (8.08–8.45) → 8.20
(8.14–8.36) ms, −0.8%, so the listing regression is gone; symbol-table
nodes −0.4%, B-tree walk +0.5%; `ObjectHeader::parse` ×401 24.13
(23.97–25.56) → 25.11 (24.73–25.21) µs, **+4.1%, still open**
(`docs/known-issues.md`). Data reads (`concurrent_read --decode-threads
1`, 3 alternating rounds): full reads of deflate data 7–23% faster,
contiguous full reads and deflate hyperslabs within ±2%, one-thread
contiguous hyperslabs −5.6% (open, same entry).
- Tests (2026-09-26, tank):
- `clawhdf5-format/tests/storage_equivalence.rs` now also reads every
dataset — whole, fill-aware, through a chunk cache (twice) and the
@@ -631,6 +644,10 @@ Design: `docs/design/swmr.md`.
instead of being inlined into the generic parser. Provisional (tank load
3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs
24.85–25.15 µs on `main`.
Rechecked 2026-09-27 with 6 alternating rounds (load 2.3–3.3, still
provisional): 24.73–25.21 µs against 23.97–24.29 µs (one outlier 25.56)
for `8f59b2e`, median +4.1%; a residual cost remains (see the M2 speed
note and `docs/known-issues.md`).
- Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py
`earliest`/`v110`/`latest` and clawhdf5-written files; structure
comparisons with libhdf5 for version-2 B-trees, shrink on every index,