docs: local read A/B re-run on an idle machine
The first run of the day could not get an idle tank: two orphaned h5py SWMR reader processes from earlier interop tests (since stopped) kept a core each busy. Re-run with the load below 2 at every round: ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5 ns per header, not visible in the facade listing, which is -1.7%); full deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single- thread contiguous hyperslab result from the loaded run is noise (-1.7%, overlapping ranges) and is withdrawn from known-issues. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
+16
-20
@@ -7,27 +7,23 @@ deleting it.
|
||||
|
||||
---
|
||||
|
||||
## Local-read slowdowns left after range-read M2/M3 (measured 2026-09-27)
|
||||
## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27)
|
||||
|
||||
**Status:** open (speed only; values are correct). Measured on tank,
|
||||
`main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`,
|
||||
alternating separate binaries (`BENCHMARKS.md`, "Local metadata and data
|
||||
reads after range-read M2/M3"). Provisional: the machine never went idle
|
||||
(1-minute load 2.3–3.3 for the metadata runs, two orphaned h5py processes
|
||||
using a core each), so re-run both on an idle tank before acting.
|
||||
- `ObjectHeader::parse` of 401 small headers (`local_metadata_bench`,
|
||||
`object_header_parse_x401`): 24.13 → 25.11 µs median over 6 rounds
|
||||
(+4.1%, about 2.4 ns per header; ranges 23.97–25.56 vs 24.73–25.21, five
|
||||
of six base rounds at or below 24.29). The header parser has read its
|
||||
continuation chunks from a queue since `a69c5be`; `4313917` removed its
|
||||
per-header allocations but not all of the cost. Listing the 400-group
|
||||
file through the facade, which parses those headers, is not slower
|
||||
(−0.8%).
|
||||
- One thread reading 256 x 256 hyperslabs of a contiguous dataset
|
||||
(`concurrent_read`, `contiguous same`, 1 thread): 22682 → 21421 MB/s
|
||||
median over 3 rounds (−5.6%; 22658–23411 vs 20358–21908), about 0.6 µs
|
||||
more per 256 KiB slab. The multi-thread rows of that mode, which are
|
||||
cache-bound and noisy, did not slow down; full reads did not either.
|
||||
**Status:** open (speed only; values are correct). Measured on an idle tank
|
||||
(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`)
|
||||
against `main` `7a8fae0`, separate binaries alternating, 3 rounds
|
||||
(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"):
|
||||
`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges
|
||||
23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read
|
||||
continuation chunks from a bounded queue since `a69c5be` (bounding what a
|
||||
crafted header can make it read); `4313917` removed its per-header
|
||||
allocations but not all of the cost. Listing the 400-group file through the
|
||||
facade, which parses the same headers, is 1.7% faster, so no user-visible
|
||||
path is slower.
|
||||
|
||||
An earlier run the same day under load also listed single-thread contiguous
|
||||
hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with
|
||||
overlapping ranges (noise), so that item is withdrawn.
|
||||
|
||||
## Files a SWMR writer had open could not be read past a stale end of file
|
||||
|
||||
|
||||
Reference in New Issue
Block a user