docs: local read A/B re-run on an idle machine
CI / test-arm64 (pull_request) Successful in 1m25s
CI / test (pull_request) Successful in 15m57s

The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 19:43:07 -05:00
co-authored by Claude Opus 5.5
parent b39a705e77
commit 7179006aee
3 changed files with 79 additions and 106 deletions
+16 -17
View File
@@ -477,19 +477,19 @@ Design: `docs/design/swmr.md`.
`main` and it calls only format-crate code; its earlier +5–10% (323–355
vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured
357–374 ns against `main`'s 355–371 in the same rounds).
**Rechecked 2026-09-27** (still provisional: two orphaned h5py test
processes kept tank's 1-minute load at 2.1–2.6, the load < 2 gate was
not met in 2 hours, load 2.3–3.3 during the runs; `8f59b2e` vs `main`
`7a8fae0` as separate binaries, pinned to one CPU, 6 alternating rounds,
median (range) over rounds; `BENCHMARKS.md`, "Local metadata and data
reads after range-read M2/M3"): facade listing 8.26 (8.08–8.45) → 8.20
(8.14–8.36) ms, −0.8%, so the listing regression is gone; symbol-table
nodes −0.4%, B-tree walk +0.5%; `ObjectHeader::parse` ×401 24.13
(23.97–25.56) → 25.11 (24.73–25.21) µs, **+4.1%, still open**
(`docs/known-issues.md`). Data reads (`concurrent_read --decode-threads
1`, 3 alternating rounds): full reads of deflate data 7–23% faster,
contiguous full reads and deflate hyperslabs within ±2%, one-thread
contiguous hyperslabs −5.6% (open, same entry).
**Rechecked 2026-09-27 on an idle tank** (load below 2 at every round;
`8f59b2e` vs `main` `7a8fae0` as separate binaries, pinned to one CPU, 3
alternating rounds, median (range); `BENCHMARKS.md`, "Local metadata and
data reads after range-read M2/M3"): facade listing 8.24 (8.04–8.28) →
8.10 (8.04–8.13) ms, **−1.7%**, so the listing regression is gone;
symbol-table nodes −0.5%, B-tree walk +1.1%; `ObjectHeader::parse` ×401
23.86 (23.79–23.96) → 24.86 (24.69–24.98) µs, **+4.2%, open**
(`docs/known-issues.md`; about 2.5 ns per header, not visible in the
listing). Data reads (`concurrent_read --decode-threads 1`): full reads
of deflate data +1.7% (1 thread) to +36% (16 threads), contiguous reads
and deflate hyperslabs within ±2%. (A run earlier that day under load
2.3–3.3 put single-thread contiguous hyperslabs at −5.6%; the idle rerun
does not reproduce it.)
- Tests (2026-09-26, tank):
- `clawhdf5-format/tests/storage_equivalence.rs` now also reads every
dataset — whole, fill-aware, through a chunk cache (twice) and the
@@ -644,10 +644,9 @@ Design: `docs/design/swmr.md`.
instead of being inlined into the generic parser. Provisional (tank load
3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs
24.85–25.15 µs on `main`.
Rechecked 2026-09-27 with 6 alternating rounds (load 2.3–3.3, still
provisional): 24.73–25.21 µs against 23.97–24.29 µs (one outlier 25.56)
for `8f59b2e`, median +4.1%; a residual cost remains (see the M2 speed
note and `docs/known-issues.md`).
Rechecked 2026-09-27 on an idle tank (3 alternating rounds, load below
2): 24.69–24.98 µs against 23.79–23.96 µs for `8f59b2e`, median +4.2%; a
residual cost remains (see the M2 speed note and `docs/known-issues.md`).
- Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py
`earliest`/`v110`/`latest` and clawhdf5-written files; structure
comparisons with libhdf5 for version-2 B-trees, shrink on every index,