docs: local metadata and data reads after range-read M2/M3, rechecked

8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.

The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 15:35:59 -05:00
co-authored by Claude Opus 5.5
parent 7a8fae0357
commit d239292655
3 changed files with 119 additions and 0 deletions
+22
View File
@@ -7,6 +7,28 @@ deleting it.
---
## Local-read slowdowns left after range-read M2/M3 (measured 2026-09-27)
**Status:** open (speed only; values are correct). Measured on tank,
`main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`,
alternating separate binaries (`BENCHMARKS.md`, "Local metadata and data
reads after range-read M2/M3"). Provisional: the machine never went idle
(1-minute load 2.3–3.3 for the metadata runs, two orphaned h5py processes
using a core each), so re-run both on an idle tank before acting.
- `ObjectHeader::parse` of 401 small headers (`local_metadata_bench`,
`object_header_parse_x401`): 24.13 → 25.11 µs median over 6 rounds
(+4.1%, about 2.4 ns per header; ranges 23.97–25.56 vs 24.73–25.21, five
of six base rounds at or below 24.29). The header parser has read its
continuation chunks from a queue since `a69c5be`; `4313917` removed its
per-header allocations but not all of the cost. Listing the 400-group
file through the facade, which parses those headers, is not slower
(−0.8%).
- One thread reading 256 x 256 hyperslabs of a contiguous dataset
(`concurrent_read`, `contiguous same`, 1 thread): 22682 → 21421 MB/s
median over 3 rounds (−5.6%; 22658–23411 vs 20358–21908), about 0.6 µs
more per 256 KiB slab. The multi-thread rows of that mode, which are
cache-bound and noisy, did not slow down; full reads did not either.
## Files a SWMR writer had open could not be read past a stale end of file
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any