docs: local metadata and data reads after range-read M2/M3, rechecked
8f59b2e(main before PR #18) against7a8fae0, separate binaries, alternating rounds on tank (6 Criterion rounds of local_metadata_bench, 3 of concurrent_read --decode-threads 1). Still provisional: two orphaned h5py test processes held the load at 2.1-2.6 and the load < 2 gate was not met in 2 hours. The facade listing regression is gone (-0.8%). ObjectHeader::parse x401 is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as open in known-issues. Full deflate reads are 7-23% faster. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
@@ -482,6 +482,86 @@ The rows and columns of the uncompressed layouts are within 20% (chunked
|
|||||||
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
|
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
|
||||||
explain the slower windows.
|
explain the slower windows.
|
||||||
|
|
||||||
|
## Local file speed after range reads
|
||||||
|
|
||||||
|
### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank)
|
||||||
|
|
||||||
|
Range-read M2/M3 (PR #18) moved the facade onto the `Storage` abstraction.
|
||||||
|
This compares `main` just before it, `8f59b2e`, with `main` at `7a8fae0`
|
||||||
|
(PR #18 and PR #19 merged, including the slice-entry-point fix `ef428d7`
|
||||||
|
and the object-header fix `4313917`), each built in its own worktree and
|
||||||
|
target directory and run as separate binaries, alternating base,
|
||||||
|
candidate, base, candidate.
|
||||||
|
|
||||||
|
**Provisional: tank was not idle.** Two orphaned h5py processes (leftover
|
||||||
|
SWMR test readers from earlier test runs, about 100% of a core each) held
|
||||||
|
the 1-minute load average at 2.1-2.6 throughout; the load < 2 gate was
|
||||||
|
polled every 20 s for 2 hours and never met, so the runs went ahead
|
||||||
|
anyway. Load average (1-minute) during the runs: Criterion rounds 1-3
|
||||||
|
2.29-3.31, rounds 4-6 2.41-3.12 (each started below 2.5);
|
||||||
|
`concurrent_read` 3.04-6.49 (that figure includes its own 16 threads).
|
||||||
|
Machine: AMD Ryzen 7 7800X3D (8 cores, 16 threads).
|
||||||
|
|
||||||
|
Metadata, Criterion `crates/clawhdf5/benches/local_metadata_bench.rs` over
|
||||||
|
`clawhdf5-format/tests/fixtures/v1_groups_400.h5`, pinned to one CPU, 6
|
||||||
|
alternating rounds each (median of each round's Criterion estimate, then
|
||||||
|
median and range over the rounds):
|
||||||
|
|
||||||
|
```text
|
||||||
|
cargo bench -p clawhdf5 --bench local_metadata_bench --no-run
|
||||||
|
cd crates/clawhdf5 && taskset -c 5 <bench binary> --bench --noplot \
|
||||||
|
--warm-up-time 3 --measurement-time 10
|
||||||
|
```
|
||||||
|
|
||||||
|
| function | `8f59b2e` median (range) | `7a8fae0` median (range) | change |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| `object_header_parse_x401` | 24.13 us (23.97-25.56) | 25.11 us (24.73-25.21) | +4.1% |
|
||||||
|
| `snod_parse_all` | 1.866 us (1.855-1.888) | 1.858 us (1.846-1.884) | -0.4% |
|
||||||
|
| `btree_v1_walk` | 344 ns (337-355) | 346 ns (340-355) | +0.5% |
|
||||||
|
| `facade_list_400_groups` | 8.26 ms (8.08-8.45) | 8.20 ms (8.14-8.36) | -0.8% |
|
||||||
|
|
||||||
|
The facade listing's +7-10% seen during PR #18 is gone (-0.8%, within the
|
||||||
|
ranges). `ObjectHeader::parse` of 401 small headers is still slower: five
|
||||||
|
of six base rounds measured 23.97-24.29 us (one outlier 25.56), every
|
||||||
|
candidate round 24.73-25.21 us, so +4% (about 2.4 ns per header) looks
|
||||||
|
real rather than noise; it does not show in the listing that parses those
|
||||||
|
headers. Listed in `docs/known-issues.md`.
|
||||||
|
|
||||||
|
Data reads, `concurrent_read` (the "Concurrent reads" workload below: 64
|
||||||
|
datasets of 64 MiB, `--decode-threads 1`, 3 repetitions per point, median
|
||||||
|
MB/s reported by the harness), 3 alternating rounds each; median and range
|
||||||
|
over the rounds:
|
||||||
|
|
||||||
|
```text
|
||||||
|
cargo build --release -p clawhdf5-bench --bin concurrent_read
|
||||||
|
concurrent_read --dir ~/.cache/concurrent-read --decode-threads 1 --reps 3
|
||||||
|
```
|
||||||
|
|
||||||
|
| layout | mode | threads | `8f59b2e` MB/s (range) | `7a8fae0` MB/s (range) | change |
|
||||||
|
|---|---|---:|---:|---:|---:|
|
||||||
|
| deflate | distinct | 1 | 815 (810-816) | 875 (874-877) | +7.4% |
|
||||||
|
| deflate | distinct | 4 | 2964 (2959-2969) | 3302 (3301-3309) | +11.4% |
|
||||||
|
| deflate | distinct | 8 | 4401 (4379-4414) | 5074 (5072-5080) | +15.3% |
|
||||||
|
| deflate | distinct | 16 | 5568 (5500-5598) | 6862 (6781-6962) | +23.2% |
|
||||||
|
| deflate | same | 1 | 228 (228-230) | 230 (230-231) | +1.1% |
|
||||||
|
| deflate | same | 4 | 896 (893-900) | 893 (890-896) | -0.3% |
|
||||||
|
| deflate | same | 16 | 1837 (1829-1839) | 1844 (1805-1865) | +0.4% |
|
||||||
|
| contiguous | distinct | 1 | 10360 (9636-10625) | 10143 (10063-10402) | -2.1% |
|
||||||
|
| contiguous | distinct | 4 | 13658 (13543-13750) | 13704 (13448-13770) | +0.3% |
|
||||||
|
| contiguous | distinct | 16 | 12311 (12256-12354) | 12307 (12273-12337) | -0.0% |
|
||||||
|
| contiguous | same | 1 | 22682 (22658-23411) | 21421 (20358-21908) | -5.6% |
|
||||||
|
| contiguous | same | 4 | 97424 (96473-100871) | 99417 (98952-105308) | +2.0% |
|
||||||
|
| contiguous | same | 16 | 137754 (126322-156061) | 152767 (136109-154669) | +10.9% |
|
||||||
|
|
||||||
|
(2- and 8-thread rows are in line with their neighbours.) Full reads of
|
||||||
|
chunked deflate data are 7-23% *faster* at `7a8fae0`, with tight ranges;
|
||||||
|
which commit did it was not bisected. Contiguous full reads and deflate
|
||||||
|
hyperslabs are unchanged. One row is slower: one thread reading 256 x 256
|
||||||
|
hyperslabs of a contiguous dataset, -5.6% with no overlap between the
|
||||||
|
rounds (about 0.6 us more per 256 KiB slab); the multi-thread `contiguous
|
||||||
|
same` rows, which are cache-bound and noisy, went the other way. Listed in
|
||||||
|
`docs/known-issues.md` as open, pending an idle re-run.
|
||||||
|
|
||||||
## Concurrent reads
|
## Concurrent reads
|
||||||
|
|
||||||
### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`)
|
### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`)
|
||||||
|
|||||||
@@ -477,6 +477,19 @@ Design: `docs/design/swmr.md`.
|
|||||||
`main` and it calls only format-crate code; its earlier +5–10% (323–355
|
`main` and it calls only format-crate code; its earlier +5–10% (323–355
|
||||||
vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured
|
vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured
|
||||||
357–374 ns against `main`'s 355–371 in the same rounds).
|
357–374 ns against `main`'s 355–371 in the same rounds).
|
||||||
|
**Rechecked 2026-09-27** (still provisional: two orphaned h5py test
|
||||||
|
processes kept tank's 1-minute load at 2.1–2.6, the load < 2 gate was
|
||||||
|
not met in 2 hours, load 2.3–3.3 during the runs; `8f59b2e` vs `main`
|
||||||
|
`7a8fae0` as separate binaries, pinned to one CPU, 6 alternating rounds,
|
||||||
|
median (range) over rounds; `BENCHMARKS.md`, "Local metadata and data
|
||||||
|
reads after range-read M2/M3"): facade listing 8.26 (8.08–8.45) → 8.20
|
||||||
|
(8.14–8.36) ms, −0.8%, so the listing regression is gone; symbol-table
|
||||||
|
nodes −0.4%, B-tree walk +0.5%; `ObjectHeader::parse` ×401 24.13
|
||||||
|
(23.97–25.56) → 25.11 (24.73–25.21) µs, **+4.1%, still open**
|
||||||
|
(`docs/known-issues.md`). Data reads (`concurrent_read --decode-threads
|
||||||
|
1`, 3 alternating rounds): full reads of deflate data 7–23% faster,
|
||||||
|
contiguous full reads and deflate hyperslabs within ±2%, one-thread
|
||||||
|
contiguous hyperslabs −5.6% (open, same entry).
|
||||||
- Tests (2026-09-26, tank):
|
- Tests (2026-09-26, tank):
|
||||||
- `clawhdf5-format/tests/storage_equivalence.rs` now also reads every
|
- `clawhdf5-format/tests/storage_equivalence.rs` now also reads every
|
||||||
dataset — whole, fill-aware, through a chunk cache (twice) and the
|
dataset — whole, fill-aware, through a chunk cache (twice) and the
|
||||||
@@ -631,6 +644,10 @@ Design: `docs/design/swmr.md`.
|
|||||||
instead of being inlined into the generic parser. Provisional (tank load
|
instead of being inlined into the generic parser. Provisional (tank load
|
||||||
3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs
|
3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs
|
||||||
24.85–25.15 µs on `main`.
|
24.85–25.15 µs on `main`.
|
||||||
|
Rechecked 2026-09-27 with 6 alternating rounds (load 2.3–3.3, still
|
||||||
|
provisional): 24.73–25.21 µs against 23.97–24.29 µs (one outlier 25.56)
|
||||||
|
for `8f59b2e`, median +4.1%; a residual cost remains (see the M2 speed
|
||||||
|
note and `docs/known-issues.md`).
|
||||||
- Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py
|
- Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py
|
||||||
`earliest`/`v110`/`latest` and clawhdf5-written files; structure
|
`earliest`/`v110`/`latest` and clawhdf5-written files; structure
|
||||||
comparisons with libhdf5 for version-2 B-trees, shrink on every index,
|
comparisons with libhdf5 for version-2 B-trees, shrink on every index,
|
||||||
|
|||||||
@@ -7,6 +7,28 @@ deleting it.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Local-read slowdowns left after range-read M2/M3 (measured 2026-09-27)
|
||||||
|
|
||||||
|
**Status:** open (speed only; values are correct). Measured on tank,
|
||||||
|
`main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`,
|
||||||
|
alternating separate binaries (`BENCHMARKS.md`, "Local metadata and data
|
||||||
|
reads after range-read M2/M3"). Provisional: the machine never went idle
|
||||||
|
(1-minute load 2.3–3.3 for the metadata runs, two orphaned h5py processes
|
||||||
|
using a core each), so re-run both on an idle tank before acting.
|
||||||
|
- `ObjectHeader::parse` of 401 small headers (`local_metadata_bench`,
|
||||||
|
`object_header_parse_x401`): 24.13 → 25.11 µs median over 6 rounds
|
||||||
|
(+4.1%, about 2.4 ns per header; ranges 23.97–25.56 vs 24.73–25.21, five
|
||||||
|
of six base rounds at or below 24.29). The header parser has read its
|
||||||
|
continuation chunks from a queue since `a69c5be`; `4313917` removed its
|
||||||
|
per-header allocations but not all of the cost. Listing the 400-group
|
||||||
|
file through the facade, which parses those headers, is not slower
|
||||||
|
(−0.8%).
|
||||||
|
- One thread reading 256 x 256 hyperslabs of a contiguous dataset
|
||||||
|
(`concurrent_read`, `contiguous same`, 1 thread): 22682 → 21421 MB/s
|
||||||
|
median over 3 rounds (−5.6%; 22658–23411 vs 20358–21908), about 0.6 µs
|
||||||
|
more per 256 KiB slab. The multi-thread rows of that mode, which are
|
||||||
|
cache-bound and noisy, did not slow down; full reads did not either.
|
||||||
|
|
||||||
## Files a SWMR writer had open could not be read past a stale end of file
|
## Files a SWMR writer had open could not be read past a stale end of file
|
||||||
|
|
||||||
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any
|
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any
|
||||||
|
|||||||
Reference in New Issue
Block a user