From d2392926551f561efc8f0d99036ce0209de83870 Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 15:35:59 -0500 Subject: [PATCH] docs: local metadata and data reads after range-read M2/M3, rechecked 8f59b2e (main before PR #18) against 7a8fae0, separate binaries, alternating rounds on tank (6 Criterion rounds of local_metadata_bench, 3 of concurrent_read --decode-threads 1). Still provisional: two orphaned h5py test processes held the load at 2.1-2.6 and the load < 2 gate was not met in 2 hours. The facade listing regression is gone (-0.8%). ObjectHeader::parse x401 is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as open in known-issues. Full deflate reads are 7-23% faster. Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 80 ++++++++++++++++++++++++++++++++++++++++++++ CHANGELOG.md | 17 ++++++++++ docs/known-issues.md | 22 ++++++++++++ 3 files changed, 119 insertions(+) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 62beeb6..7937d0e 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -482,6 +482,86 @@ The rows and columns of the uncompressed layouts are within 20% (chunked column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not explain the slower windows. +## Local file speed after range reads + +### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) + +Range-read M2/M3 (PR #18) moved the facade onto the `Storage` abstraction. +This compares `main` just before it, `8f59b2e`, with `main` at `7a8fae0` +(PR #18 and PR #19 merged, including the slice-entry-point fix `ef428d7` +and the object-header fix `4313917`), each built in its own worktree and +target directory and run as separate binaries, alternating base, +candidate, base, candidate. + +**Provisional: tank was not idle.** Two orphaned h5py processes (leftover +SWMR test readers from earlier test runs, about 100% of a core each) held +the 1-minute load average at 2.1-2.6 throughout; the load < 2 gate was +polled every 20 s for 2 hours and never met, so the runs went ahead +anyway. Load average (1-minute) during the runs: Criterion rounds 1-3 +2.29-3.31, rounds 4-6 2.41-3.12 (each started below 2.5); +`concurrent_read` 3.04-6.49 (that figure includes its own 16 threads). +Machine: AMD Ryzen 7 7800X3D (8 cores, 16 threads). + +Metadata, Criterion `crates/clawhdf5/benches/local_metadata_bench.rs` over +`clawhdf5-format/tests/fixtures/v1_groups_400.h5`, pinned to one CPU, 6 +alternating rounds each (median of each round's Criterion estimate, then +median and range over the rounds): + +```text +cargo bench -p clawhdf5 --bench local_metadata_bench --no-run +cd crates/clawhdf5 && taskset -c 5 --bench --noplot \ + --warm-up-time 3 --measurement-time 10 +``` + +| function | `8f59b2e` median (range) | `7a8fae0` median (range) | change | +|---|---:|---:|---:| +| `object_header_parse_x401` | 24.13 us (23.97-25.56) | 25.11 us (24.73-25.21) | +4.1% | +| `snod_parse_all` | 1.866 us (1.855-1.888) | 1.858 us (1.846-1.884) | -0.4% | +| `btree_v1_walk` | 344 ns (337-355) | 346 ns (340-355) | +0.5% | +| `facade_list_400_groups` | 8.26 ms (8.08-8.45) | 8.20 ms (8.14-8.36) | -0.8% | + +The facade listing's +7-10% seen during PR #18 is gone (-0.8%, within the +ranges). `ObjectHeader::parse` of 401 small headers is still slower: five +of six base rounds measured 23.97-24.29 us (one outlier 25.56), every +candidate round 24.73-25.21 us, so +4% (about 2.4 ns per header) looks +real rather than noise; it does not show in the listing that parses those +headers. Listed in `docs/known-issues.md`. + +Data reads, `concurrent_read` (the "Concurrent reads" workload below: 64 +datasets of 64 MiB, `--decode-threads 1`, 3 repetitions per point, median +MB/s reported by the harness), 3 alternating rounds each; median and range +over the rounds: + +```text +cargo build --release -p clawhdf5-bench --bin concurrent_read +concurrent_read --dir ~/.cache/concurrent-read --decode-threads 1 --reps 3 +``` + +| layout | mode | threads | `8f59b2e` MB/s (range) | `7a8fae0` MB/s (range) | change | +|---|---|---:|---:|---:|---:| +| deflate | distinct | 1 | 815 (810-816) | 875 (874-877) | +7.4% | +| deflate | distinct | 4 | 2964 (2959-2969) | 3302 (3301-3309) | +11.4% | +| deflate | distinct | 8 | 4401 (4379-4414) | 5074 (5072-5080) | +15.3% | +| deflate | distinct | 16 | 5568 (5500-5598) | 6862 (6781-6962) | +23.2% | +| deflate | same | 1 | 228 (228-230) | 230 (230-231) | +1.1% | +| deflate | same | 4 | 896 (893-900) | 893 (890-896) | -0.3% | +| deflate | same | 16 | 1837 (1829-1839) | 1844 (1805-1865) | +0.4% | +| contiguous | distinct | 1 | 10360 (9636-10625) | 10143 (10063-10402) | -2.1% | +| contiguous | distinct | 4 | 13658 (13543-13750) | 13704 (13448-13770) | +0.3% | +| contiguous | distinct | 16 | 12311 (12256-12354) | 12307 (12273-12337) | -0.0% | +| contiguous | same | 1 | 22682 (22658-23411) | 21421 (20358-21908) | -5.6% | +| contiguous | same | 4 | 97424 (96473-100871) | 99417 (98952-105308) | +2.0% | +| contiguous | same | 16 | 137754 (126322-156061) | 152767 (136109-154669) | +10.9% | + +(2- and 8-thread rows are in line with their neighbours.) Full reads of +chunked deflate data are 7-23% *faster* at `7a8fae0`, with tight ranges; +which commit did it was not bisected. Contiguous full reads and deflate +hyperslabs are unchanged. One row is slower: one thread reading 256 x 256 +hyperslabs of a contiguous dataset, -5.6% with no overlap between the +rounds (about 0.6 us more per 256 KiB slab); the multi-thread `contiguous +same` rows, which are cache-bound and noisy, went the other way. Listed in +`docs/known-issues.md` as open, pending an idle re-run. + ## Concurrent reads ### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`) diff --git a/CHANGELOG.md b/CHANGELOG.md index f7d44f1..2743744 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -477,6 +477,19 @@ Design: `docs/design/swmr.md`. `main` and it calls only format-crate code; its earlier +5–10% (323–355 vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured 357–374 ns against `main`'s 355–371 in the same rounds). + **Rechecked 2026-09-27** (still provisional: two orphaned h5py test + processes kept tank's 1-minute load at 2.1–2.6, the load < 2 gate was + not met in 2 hours, load 2.3–3.3 during the runs; `8f59b2e` vs `main` + `7a8fae0` as separate binaries, pinned to one CPU, 6 alternating rounds, + median (range) over rounds; `BENCHMARKS.md`, "Local metadata and data + reads after range-read M2/M3"): facade listing 8.26 (8.08–8.45) → 8.20 + (8.14–8.36) ms, −0.8%, so the listing regression is gone; symbol-table + nodes −0.4%, B-tree walk +0.5%; `ObjectHeader::parse` ×401 24.13 + (23.97–25.56) → 25.11 (24.73–25.21) µs, **+4.1%, still open** + (`docs/known-issues.md`). Data reads (`concurrent_read --decode-threads + 1`, 3 alternating rounds): full reads of deflate data 7–23% faster, + contiguous full reads and deflate hyperslabs within ±2%, one-thread + contiguous hyperslabs −5.6% (open, same entry). - Tests (2026-09-26, tank): - `clawhdf5-format/tests/storage_equivalence.rs` now also reads every dataset — whole, fill-aware, through a chunk cache (twice) and the @@ -631,6 +644,10 @@ Design: `docs/design/swmr.md`. instead of being inlined into the generic parser. Provisional (tank load 3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs 24.85–25.15 µs on `main`. + Rechecked 2026-09-27 with 6 alternating rounds (load 2.3–3.3, still + provisional): 24.73–25.21 µs against 23.97–24.29 µs (one outlier 25.56) + for `8f59b2e`, median +4.1%; a residual cost remains (see the M2 speed + note and `docs/known-issues.md`). - Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py `earliest`/`v110`/`latest` and clawhdf5-written files; structure comparisons with libhdf5 for version-2 B-trees, shrink on every index, diff --git a/docs/known-issues.md b/docs/known-issues.md index 7eea8e7..613a729 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -7,6 +7,28 @@ deleting it. --- +## Local-read slowdowns left after range-read M2/M3 (measured 2026-09-27) + +**Status:** open (speed only; values are correct). Measured on tank, +`main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`, +alternating separate binaries (`BENCHMARKS.md`, "Local metadata and data +reads after range-read M2/M3"). Provisional: the machine never went idle +(1-minute load 2.3–3.3 for the metadata runs, two orphaned h5py processes +using a core each), so re-run both on an idle tank before acting. +- `ObjectHeader::parse` of 401 small headers (`local_metadata_bench`, + `object_header_parse_x401`): 24.13 → 25.11 µs median over 6 rounds + (+4.1%, about 2.4 ns per header; ranges 23.97–25.56 vs 24.73–25.21, five + of six base rounds at or below 24.29). The header parser has read its + continuation chunks from a queue since `a69c5be`; `4313917` removed its + per-header allocations but not all of the cost. Listing the 400-group + file through the facade, which parses those headers, is not slower + (−0.8%). +- One thread reading 256 x 256 hyperslabs of a contiguous dataset + (`concurrent_read`, `contiguous same`, 1 thread): 22682 → 21421 MB/s + median over 3 rounds (−5.6%; 22658–23411 vs 20358–21908), about 0.6 µs + more per 256 KiB slab. The multi-thread rows of that mode, which are + cache-bound and noisy, did not slow down; full reads did not either. + ## Files a SWMR writer had open could not be read past a stale end of file **Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any