docs: local read A/B re-run on an idle machine
CI / test-arm64 (pull_request) Successful in 1m25s
CI / test (pull_request) Successful in 15m57s

The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 19:43:07 -05:00
co-authored by Claude Opus 5.5
parent b39a705e77
commit 7179006aee
3 changed files with 79 additions and 106 deletions
+16 -20
View File
@@ -7,27 +7,23 @@ deleting it.
---
## Local-read slowdowns left after range-read M2/M3 (measured 2026-09-27)
## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27)
**Status:** open (speed only; values are correct). Measured on tank,
`main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`,
alternating separate binaries (`BENCHMARKS.md`, "Local metadata and data
reads after range-read M2/M3"). Provisional: the machine never went idle
(1-minute load 2.3–3.3 for the metadata runs, two orphaned h5py processes
using a core each), so re-run both on an idle tank before acting.
- `ObjectHeader::parse` of 401 small headers (`local_metadata_bench`,
`object_header_parse_x401`): 24.13 → 25.11 µs median over 6 rounds
(+4.1%, about 2.4 ns per header; ranges 23.97–25.56 vs 24.73–25.21, five
of six base rounds at or below 24.29). The header parser has read its
continuation chunks from a queue since `a69c5be`; `4313917` removed its
per-header allocations but not all of the cost. Listing the 400-group
file through the facade, which parses those headers, is not slower
(−0.8%).
- One thread reading 256 x 256 hyperslabs of a contiguous dataset
(`concurrent_read`, `contiguous same`, 1 thread): 22682 → 21421 MB/s
median over 3 rounds (−5.6%; 22658–23411 vs 20358–21908), about 0.6 µs
more per 256 KiB slab. The multi-thread rows of that mode, which are
cache-bound and noisy, did not slow down; full reads did not either.
**Status:** open (speed only; values are correct). Measured on an idle tank
(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`)
against `main` `7a8fae0`, separate binaries alternating, 3 rounds
(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"):
`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges
23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read
continuation chunks from a bounded queue since `a69c5be` (bounding what a
crafted header can make it read); `4313917` removed its per-header
allocations but not all of the cost. Listing the 400-group file through the
facade, which parses the same headers, is 1.7% faster, so no user-visible
path is slower.
An earlier run the same day under load also listed single-thread contiguous
hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with
overlapping ranges (noise), so that item is withdrawn.
## Files a SWMR writer had open could not be read past a stale end of file