From d2392926551f561efc8f0d99036ce0209de83870 Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 15:35:59 -0500 Subject: [PATCH 1/3] docs: local metadata and data reads after range-read M2/M3, rechecked 8f59b2e (main before PR #18) against 7a8fae0, separate binaries, alternating rounds on tank (6 Criterion rounds of local_metadata_bench, 3 of concurrent_read --decode-threads 1). Still provisional: two orphaned h5py test processes held the load at 2.1-2.6 and the load < 2 gate was not met in 2 hours. The facade listing regression is gone (-0.8%). ObjectHeader::parse x401 is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as open in known-issues. Full deflate reads are 7-23% faster. Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 80 ++++++++++++++++++++++++++++++++++++++++++++ CHANGELOG.md | 17 ++++++++++ docs/known-issues.md | 22 ++++++++++++ 3 files changed, 119 insertions(+) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 62beeb6..7937d0e 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -482,6 +482,86 @@ The rows and columns of the uncompressed layouts are within 20% (chunked column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not explain the slower windows. +## Local file speed after range reads + +### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) + +Range-read M2/M3 (PR #18) moved the facade onto the `Storage` abstraction. +This compares `main` just before it, `8f59b2e`, with `main` at `7a8fae0` +(PR #18 and PR #19 merged, including the slice-entry-point fix `ef428d7` +and the object-header fix `4313917`), each built in its own worktree and +target directory and run as separate binaries, alternating base, +candidate, base, candidate. + +**Provisional: tank was not idle.** Two orphaned h5py processes (leftover +SWMR test readers from earlier test runs, about 100% of a core each) held +the 1-minute load average at 2.1-2.6 throughout; the load < 2 gate was +polled every 20 s for 2 hours and never met, so the runs went ahead +anyway. Load average (1-minute) during the runs: Criterion rounds 1-3 +2.29-3.31, rounds 4-6 2.41-3.12 (each started below 2.5); +`concurrent_read` 3.04-6.49 (that figure includes its own 16 threads). +Machine: AMD Ryzen 7 7800X3D (8 cores, 16 threads). + +Metadata, Criterion `crates/clawhdf5/benches/local_metadata_bench.rs` over +`clawhdf5-format/tests/fixtures/v1_groups_400.h5`, pinned to one CPU, 6 +alternating rounds each (median of each round's Criterion estimate, then +median and range over the rounds): + +```text +cargo bench -p clawhdf5 --bench local_metadata_bench --no-run +cd crates/clawhdf5 && taskset -c 5 --bench --noplot \ + --warm-up-time 3 --measurement-time 10 +``` + +| function | `8f59b2e` median (range) | `7a8fae0` median (range) | change | +|---|---:|---:|---:| +| `object_header_parse_x401` | 24.13 us (23.97-25.56) | 25.11 us (24.73-25.21) | +4.1% | +| `snod_parse_all` | 1.866 us (1.855-1.888) | 1.858 us (1.846-1.884) | -0.4% | +| `btree_v1_walk` | 344 ns (337-355) | 346 ns (340-355) | +0.5% | +| `facade_list_400_groups` | 8.26 ms (8.08-8.45) | 8.20 ms (8.14-8.36) | -0.8% | + +The facade listing's +7-10% seen during PR #18 is gone (-0.8%, within the +ranges). `ObjectHeader::parse` of 401 small headers is still slower: five +of six base rounds measured 23.97-24.29 us (one outlier 25.56), every +candidate round 24.73-25.21 us, so +4% (about 2.4 ns per header) looks +real rather than noise; it does not show in the listing that parses those +headers. Listed in `docs/known-issues.md`. + +Data reads, `concurrent_read` (the "Concurrent reads" workload below: 64 +datasets of 64 MiB, `--decode-threads 1`, 3 repetitions per point, median +MB/s reported by the harness), 3 alternating rounds each; median and range +over the rounds: + +```text +cargo build --release -p clawhdf5-bench --bin concurrent_read +concurrent_read --dir ~/.cache/concurrent-read --decode-threads 1 --reps 3 +``` + +| layout | mode | threads | `8f59b2e` MB/s (range) | `7a8fae0` MB/s (range) | change | +|---|---|---:|---:|---:|---:| +| deflate | distinct | 1 | 815 (810-816) | 875 (874-877) | +7.4% | +| deflate | distinct | 4 | 2964 (2959-2969) | 3302 (3301-3309) | +11.4% | +| deflate | distinct | 8 | 4401 (4379-4414) | 5074 (5072-5080) | +15.3% | +| deflate | distinct | 16 | 5568 (5500-5598) | 6862 (6781-6962) | +23.2% | +| deflate | same | 1 | 228 (228-230) | 230 (230-231) | +1.1% | +| deflate | same | 4 | 896 (893-900) | 893 (890-896) | -0.3% | +| deflate | same | 16 | 1837 (1829-1839) | 1844 (1805-1865) | +0.4% | +| contiguous | distinct | 1 | 10360 (9636-10625) | 10143 (10063-10402) | -2.1% | +| contiguous | distinct | 4 | 13658 (13543-13750) | 13704 (13448-13770) | +0.3% | +| contiguous | distinct | 16 | 12311 (12256-12354) | 12307 (12273-12337) | -0.0% | +| contiguous | same | 1 | 22682 (22658-23411) | 21421 (20358-21908) | -5.6% | +| contiguous | same | 4 | 97424 (96473-100871) | 99417 (98952-105308) | +2.0% | +| contiguous | same | 16 | 137754 (126322-156061) | 152767 (136109-154669) | +10.9% | + +(2- and 8-thread rows are in line with their neighbours.) Full reads of +chunked deflate data are 7-23% *faster* at `7a8fae0`, with tight ranges; +which commit did it was not bisected. Contiguous full reads and deflate +hyperslabs are unchanged. One row is slower: one thread reading 256 x 256 +hyperslabs of a contiguous dataset, -5.6% with no overlap between the +rounds (about 0.6 us more per 256 KiB slab); the multi-thread `contiguous +same` rows, which are cache-bound and noisy, went the other way. Listed in +`docs/known-issues.md` as open, pending an idle re-run. + ## Concurrent reads ### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`) diff --git a/CHANGELOG.md b/CHANGELOG.md index f7d44f1..2743744 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -477,6 +477,19 @@ Design: `docs/design/swmr.md`. `main` and it calls only format-crate code; its earlier +5–10% (323–355 vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured 357–374 ns against `main`'s 355–371 in the same rounds). + **Rechecked 2026-09-27** (still provisional: two orphaned h5py test + processes kept tank's 1-minute load at 2.1–2.6, the load < 2 gate was + not met in 2 hours, load 2.3–3.3 during the runs; `8f59b2e` vs `main` + `7a8fae0` as separate binaries, pinned to one CPU, 6 alternating rounds, + median (range) over rounds; `BENCHMARKS.md`, "Local metadata and data + reads after range-read M2/M3"): facade listing 8.26 (8.08–8.45) → 8.20 + (8.14–8.36) ms, −0.8%, so the listing regression is gone; symbol-table + nodes −0.4%, B-tree walk +0.5%; `ObjectHeader::parse` ×401 24.13 + (23.97–25.56) → 25.11 (24.73–25.21) µs, **+4.1%, still open** + (`docs/known-issues.md`). Data reads (`concurrent_read --decode-threads + 1`, 3 alternating rounds): full reads of deflate data 7–23% faster, + contiguous full reads and deflate hyperslabs within ±2%, one-thread + contiguous hyperslabs −5.6% (open, same entry). - Tests (2026-09-26, tank): - `clawhdf5-format/tests/storage_equivalence.rs` now also reads every dataset — whole, fill-aware, through a chunk cache (twice) and the @@ -631,6 +644,10 @@ Design: `docs/design/swmr.md`. instead of being inlined into the generic parser. Provisional (tank load 3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs 24.85–25.15 µs on `main`. + Rechecked 2026-09-27 with 6 alternating rounds (load 2.3–3.3, still + provisional): 24.73–25.21 µs against 23.97–24.29 µs (one outlier 25.56) + for `8f59b2e`, median +4.1%; a residual cost remains (see the M2 speed + note and `docs/known-issues.md`). - Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py `earliest`/`v110`/`latest` and clawhdf5-written files; structure comparisons with libhdf5 for version-2 B-trees, shrink on every index, diff --git a/docs/known-issues.md b/docs/known-issues.md index 7eea8e7..613a729 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -7,6 +7,28 @@ deleting it. --- +## Local-read slowdowns left after range-read M2/M3 (measured 2026-09-27) + +**Status:** open (speed only; values are correct). Measured on tank, +`main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`, +alternating separate binaries (`BENCHMARKS.md`, "Local metadata and data +reads after range-read M2/M3"). Provisional: the machine never went idle +(1-minute load 2.3–3.3 for the metadata runs, two orphaned h5py processes +using a core each), so re-run both on an idle tank before acting. +- `ObjectHeader::parse` of 401 small headers (`local_metadata_bench`, + `object_header_parse_x401`): 24.13 → 25.11 µs median over 6 rounds + (+4.1%, about 2.4 ns per header; ranges 23.97–25.56 vs 24.73–25.21, five + of six base rounds at or below 24.29). The header parser has read its + continuation chunks from a queue since `a69c5be`; `4313917` removed its + per-header allocations but not all of the cost. Listing the 400-group + file through the facade, which parses those headers, is not slower + (−0.8%). +- One thread reading 256 x 256 hyperslabs of a contiguous dataset + (`concurrent_read`, `contiguous same`, 1 thread): 22682 → 21421 MB/s + median over 3 rounds (−5.6%; 22658–23411 vs 20358–21908), about 0.6 µs + more per 256 KiB slab. The multi-thread rows of that mode, which are + cache-bound and noisy, did not slow down; full reads did not either. + ## Files a SWMR writer had open could not be read past a stale end of file **Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any From b39a705e77651da24815b63258f610710275046c Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 19:25:36 -0500 Subject: [PATCH 2/3] docs: LongMemEval with real embeddings re-run on 2026-09-27 Oracle and full longmemeval_s haystack, f32 and --float16, plus --sweep (both corpora) and --rerank-sweep, on tank at 7a8fae0 (MiniLM on the RTX 5060 Ti). Not idle: two orphaned h5py test processes kept the 1-minute load at 2.1-2.7 (up to 8.4 during runs), so no latency figure was updated. The headline (hybrid 0.4/0.6 turn Hit@5 81.4%) and the float16 claim reproduce. Oracle hybrid is 86.8%, not 85.2%: the old figure was at the 0.7/0.3 default of the time (today's 0.7/0.3 gives 85.2%). The ablation's 0.7/0.3 row, vector-only MRR and seven sweep rows move in the last digit or by one or two questions, most likely from the index tie-break of 3ed0489. RRF turn MRR 0.5967 -> 0.5969 is not explained. Recency counts vary by one question between identical runs, so the float16 "flips" note is corrected. ROADMAP 8.2 no longer repeats the retracted session-level and MemX claims. Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 155 ++++++++++++++++++++++++++++++++++++++++++-------- ROADMAP.md | 4 +- 2 files changed, 132 insertions(+), 27 deletions(-) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 7937d0e..b5b3b07 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -30,12 +30,15 @@ target: Criterion stretched it where 5 s could not hold the samples it needed > and in memory), Consolidation Efficiency, Ephemeral Tier, Multi-modal Search, > the Search and Read harnesses, and the "h5bench-Equivalent I/O Benchmarks" > and "Independent Validation: tank" sections. What does not yet meet that bar: -> the LongMemEval rows that need real embeddings (not re-run here, except the -> dated float16 comparison), the Consolidation Efficiency 100K cycle row and +> the Consolidation Efficiency 100K cycle row and > memory-reduction part (the 2026-09-24 run was stopped before it produced > them), the int8 side of "Quantising the index copy" (not re-run), and the > i7-12650H and macOS M3 Max rows under Cross-Platform Notes. That is a > known, tracked documentation gap, not a claim that those numbers are wrong. +> The LongMemEval rows that need real embeddings (vector-only, hybrid, RRF, +> stemmed hybrid, re-ranking, the weight sweep, the oracle variant) were +> re-run on 2026-09-27 on tank; see "Re-run with real embeddings" under +> LongMemEval Results. > > **Correctness note (2026-08-06).** Being dated and reproducible is necessary but > not sufficient — a number can be perfectly reproducible and still measure the @@ -325,6 +328,19 @@ which of two gold sessions ranks first, out of ~320. Those flips show the half-precision path was in effect; they do not change a single hit. The f32 run reproduces the published hybrid numbers exactly. +**Re-checked 2026-09-27** (tank, commit 7a8fae0, the same pair of runs on +`longmemeval_s_cleaned.json`, which is the same file; the machine was not +idle, which does not affect recall): the result is the same. The table above +reproduced exactly. Every Hit@k and MRR of the eight modes matched between f32 +and float16 at both levels, with two exceptions: RRF's session MRR (0.9253 vs +0.9254) and two per-type session MRRs in the fourth decimal. Three modes +differed by one question in the recency count. That re-run also corrects the +sentence above: two f32 runs on the same day differed by one question in +recency as well, so those flips are run-to-run variation and do not show that +the half-precision path was in effect. `--float16` is what shows that: the +harness prints "Stores use MemoryConfig::float16" and `MemoryConfig::float16` +is set on every store. + ### Opening a store (`read_from_disk`) `HDF5Memory::open` memory-mapped the file, copied the whole mapping into a @@ -1490,7 +1506,67 @@ turn-level row plus its session Hit@1 in the tokenizer table. The run also produced figures this document does not publish (stemmed session Hit@5, Hit@10 and MRR, and per-type Hit@5/Hit@10/MRR for both modes), so there was nothing to compare them with. Rows that need real embeddings (vector-only, -hybrid, RRF, re-ranking, the weight sweep) were not re-run. +hybrid, RRF, re-ranking, the weight sweep) were not re-run then; they were on +2026-09-27 (next section). + +### Re-run with real embeddings (2026-09-27, tank) + +Every recall row in this section that needs real embeddings was measured +again on 2026-09-27 on tank (AMD Ryzen 7 7800X3D, 16 threads; MiniLM +embeddings on an RTX 5060 Ti, retrieval on the CPU), with the search code of +commit 7a8fae0. Six runs: + +```bash +cargo build --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings-cuda +B=target/release/longmemeval_bench W=weights/all-minilm-l6-v2 +$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --float16 +$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W --sweep +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --sweep +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --rerank-sweep +``` + +(`longmemeval_s.json`, used by the commands elsewhere in this section, is the +same file as `longmemeval_s_cleaned.json`.) + +**The machine was not idle.** The 1-minute load never fell below 2 in a +2-hour wait, because two stray test processes were each holding a core; it +was 2.1–2.7 when each run started and up to 8.4 while the runs were going. +Recall does not depend on load. Latency does, so no latency figure in this +file was updated from these runs. + +**What reproduced exactly:** every published full-haystack Hit@1/5/10 and MRR +at turn and session level for BM25, vector-only, hybrid 0.4/0.6, both stemmed +modes, and every re-ranking row; RRF except as below; the float16 table; 4 of +the 11 weight-sweep rows (0.0, 0.1, 0.3 and 0.9); and the oracle vector-only +figure. The headline, +hybrid 0.4/0.6 turn Hit@5 **81.4%**, is unchanged. + +**What changed.** Old values are kept here; the tables below show the new ones. + +| Figure | Published | 2026-09-27 | Why | +|---|---:|---:|---| +| Oracle, hybrid turn Hit@5 | 85.2% (2026-08-07, c913cd1) | **86.8%** | Weights. 85.2% was measured at 0.7/0.3, the default then. The default has been 0.4/0.6 since 29baabb. Today's oracle sweep gives 85.2% at 0.7/0.3 and 86.8% at 0.4/0.6. | +| Oracle, BM25-only turn Hit@5 with embeddings | 84.2% (2026-08-07) | 84.4% | Now equal to the zero-embedding figure, as it already was on the full haystack. At c913cd1, tied candidates were ordered by `HashMap` iteration. Since 3ed0489 (2026-09-19) they are broken by index. 3ed0489 is the likely cause; it was not bisected. | +| Ablation / sweep, hybrid 0.7/0.3, turn | 44.4% / 79.2% / 86.0% / 0.5868 | 44.2% / 79.2% / 85.8% / 0.5856 | The same tie-breaking change. 29baabb's re-run on 2026-09-19, made after 3ed0489 that same day, already had 44.2% / 85.8% / 0.5856. | +| Ablation, hybrid 0.7/0.3, session | 88.2% / 95.8% / 97.8% / 0.9158 | 88.0% / 95.8% / 97.6% / 0.9146 | Same. | +| Ablation / sweep, vector-only turn MRR | 0.5027 | 0.5031 | Same. The fusion and float16 tables already had 0.5031. | +| Sweep 0.2/0.8, turn | 53.6% / 78.2% / 0.6440 | 53.4% / 78.0% / 0.6429 | Same (the sweep was measured 2026-08-07, 1537a94). | +| Sweep 0.4/0.6, 0.5/0.5, 0.8/0.2 turn MRR | 0.6429 / 0.6234 / 0.5571 | 0.6430 / 0.6232 / 0.5574 | Same. | +| Sweep 0.6/0.4 | Hit@5 79.8%, MRR 0.6069, session Hit@5 96.6% | 79.6%, 0.6067, 96.4% | Same. | +| RRF turn MRR | 0.5967 (2026-09-19, aa92fef) | 0.5969 | Not explained. It may be a later search-path change, such as the int8 index becoming the default in 8b85d93. It may also be the run-to-run variation described next. It was not bisected. | + +**Not exactly deterministic.** Between two f32 runs today, the recency count +(the share of `knowledge-update` questions where the newest gold session +ranks first) differed by one question in two modes: hybrid 0.4/0.6 was +144/320 in one run and 145/320 in the other, and re-rank with a 1-day +half-life was 165/319 and 166/319. Each question gets a fresh store, so this +is not state carried between modes. The cause was not found; a candidate is +that the GPU embeddings are not bit-for-bit identical from run to run. Hit@k +and MRR agreed in every run that repeated a mode (hybrid 0.4/0.6 was measured +four times: three f32 runs and one float16 run), but a last-digit change in an +MRR, or a one-question change in recency, is within this variation. ### Full haystack — `longmemeval_s`, n=500 (the number to cite) @@ -1520,13 +1596,16 @@ the question. Real 384-d `all-MiniLM-L6-v2` embeddings, 190,015 unique texts encoded once on an RTX 5060 Ti (~13 min; the same work on the 8-core CPU was still unfinished after -30 minutes, so the GPU path is not a convenience here). Turn-level: +30 minutes, so the GPU path is not a convenience here). First measured +2026-08-07 (c913cd1); the values below are the 2026-09-27 re-run, which moved +the vector-only MRR and the `0.7/0.3` row (see the re-run section above). +Turn-level: | Mode | Hit@1 | Hit@5 | Hit@10 | MRR | |------|-------|-------|--------|-----| | BM25 only (`0.0`/`1.0`) | **53.8%** | 75.0% | 81.6% | **0.6320** | -| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5027 | -| Hybrid (`0.7`/`0.3`) | 44.4% | **79.2%** | **86.0%** | 0.5868 | +| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5031 | +| Hybrid (`0.7`/`0.3`) | 44.2% | **79.2%** | **85.8%** | 0.5856 | Session-level: @@ -1534,7 +1613,7 @@ Session-level: |------|-------|-------|--------|-----| | BM25 only | 86.2% | 93.6% | 96.6% | 0.8948 | | Vector only | 85.4% | 94.2% | 96.6% | 0.8901 | -| Hybrid | **88.2%** | **95.8%** | **97.8%** | **0.9158** | +| Hybrid | **88.0%** | **95.8%** | **97.6%** | **0.9146** | ### Fusion method — weighted vs. RRF, full haystack, n=500 @@ -1548,7 +1627,7 @@ takes a `Fusion`, and both run over the same HNSW + BM25 candidates: | BM25 only | **53.8%** | 75.0% | 81.6% | 0.6320 | 86.2% | 0.8948 | | Vector only | 36.0% | 71.8% | 81.6% | 0.5031 | 85.4% | 0.8901 | | **Weighted 0.4 / 0.6** | 51.6% | **81.4%** | **87.8%** | **0.6430** | **91.0%** | **0.9347** | -| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5967 | 89.6% | 0.9253 | +| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5969 | 89.6% | 0.9253 | **RRF loses to the tuned weighted sum** — 6.6pp of turn Hit@1 and 0.046 of MRR — and lands almost exactly where the old `0.7/0.3` weighting did (44.2% / @@ -1613,6 +1692,12 @@ one (see `newest_gold_first`); ~45% is chance. | + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% | | + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% | +Re-run on 2026-09-27 (`--rerank-sweep`, tank, commit 7a8fae0): every Hit@k and +MRR above reproduced exactly. The recency column came out 45.3%, 87.5%, +52.0%, 52.2%, 51.7% and 50.5%. Each of those is within one question of the +value in the table, which is the run-to-run variation described under +"Re-run with real embeddings" above, so the table was left as it was. + **The pre-fix row is the finding.** Ordering candidates by recency alone costs 40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in the pool rather than the ones that answer the question. It does ace the recency @@ -1626,9 +1711,9 @@ cannot reach the 87.5% the degenerate ordering gets. Those two rows are the ends of a trade-off, and the default sits deliberately near the relevance end. **Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency -moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because -the temporal term is capped by its weight (0.3) while relevance differences -between candidates are larger. The 24-hour default is kept; there is no +moves 1.4pp (1.7pp in the 2026-09-27 re-run) and MRR 0.003 — inside the +noise of a 500-question run — because the temporal term is capped by its +weight (0.3) while relevance differences between candidates are larger. The 24-hour default is kept; there is no measured reason to change it, and a corpus-matched value is not the lever it looks like. @@ -1636,24 +1721,28 @@ looks like. `0.7/0.3` was a documented default, never a searched one. Sweeping `vector_weight` from 0.0 to 1.0 (`--sweep`, reusing the one-time embedding -table) shows it is not merely suboptimal but **strictly dominated**: +table) shows it is not merely suboptimal but **strictly dominated**. First +measured 2026-08-07 (1537a94); the values below are the 2026-09-27 re-run, +which changed the 0.2, 0.4, 0.5, 0.6, 0.7, 0.8 and 1.0 rows in the last digit +or by one or two questions (see the re-run section above): | vector / keyword | Hit@1 | Hit@5 | Hit@10 | MRR | session Hit@5 | |---|---|---|---|---|---| | 0.0 / 1.0 (BM25) | **53.8%** | 75.0% | 81.6% | 0.6320 | 93.6% | | 0.1 / 0.9 | 53.2% | 77.4% | 83.8% | 0.6374 | 95.0% | -| 0.2 / 0.8 | 53.6% | 78.2% | 85.6% | 0.6440 | 95.4% | +| 0.2 / 0.8 | 53.4% | 78.0% | 85.6% | 0.6429 | 95.4% | | 0.3 / 0.7 | 53.2% | 78.8% | 87.2% | **0.6463** | 96.0% | -| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6429 | 96.8% | -| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6234 | **97.4%** | -| 0.6 / 0.4 | 46.6% | 79.8% | 87.4% | 0.6069 | 96.6% | -| 0.7 / 0.3 *(old default)* | 44.4% | 79.2% | 86.0% | 0.5868 | 95.8% | -| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5571 | 95.2% | +| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6430 | 96.8% | +| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6232 | **97.4%** | +| 0.6 / 0.4 | 46.6% | 79.6% | 87.4% | 0.6067 | 96.4% | +| 0.7 / 0.3 *(old default)* | 44.2% | 79.2% | 85.8% | 0.5856 | 95.8% | +| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5574 | 95.2% | | 0.9 / 0.1 | 37.8% | 73.4% | 84.6% | 0.5289 | 94.2% | -| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5027 | 94.2% | +| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5031 | 94.2% | **`0.4/0.6` beats `0.7/0.3` on every metric at both granularities** — Hit@1 -+7.2pp, Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. There is no trade being made; the ++7.4pp, Hit@5 +2.2, Hit@10 +2.0, MRR +0.057 (2026-09-27 figures; +7.2pp, ++2.2, +1.8 and +0.056 as measured on 2026-08-07). There is no trade being made; the old default was simply on the wrong side of the peak. **`0.4/0.6` is the recommended setting**, with `0.3/0.7` preferable if rank-1 precision matters most (it takes the best MRR in the sweep and gives up only 0.6pp of Hit@1 @@ -1714,11 +1803,27 @@ price of the harder corpus, and is the reason oracle-only numbers should not be presented as LongMemEval results. Session-level figures on this variant are degenerate — see below. -With real embeddings the same oracle corpus gives BM25-only 84.2% / vector-only -80.4% / hybrid **85.2%** Hit@5 turn-level — hybrid ahead at Hit@5 and Hit@10 and -behind at Hit@1, matching the full-haystack pattern above. (BM25-only reads 84.2% -here against 84.4% with zero embedding vectors: one question of 500 changes rank, -with MRR identical at 0.6597. On the full haystack the two agree exactly.) +With real embeddings, the same oracle corpus gives these turn-level figures +(2026-09-27, tank, commit 7a8fae0; command and load in "Re-run with real +embeddings" above): + +| Mode | Hit@1 | Hit@5 | Hit@10 | MRR | +|---|---:|---:|---:|---:| +| BM25 only | **52.6%** | 84.4% | 90.4% | 0.6597 | +| Vector only | 39.6% | 80.4% | 91.0% | 0.5605 | +| Hybrid 0.4 / 0.6 (default) | 52.4% | **86.8%** | **92.4%** | **0.6678** | +| Hybrid 0.7 / 0.3 (old default) | 48.4% | 85.2% | 92.2% | 0.6382 | + +Hybrid 0.4/0.6 leads at Hit@5, Hit@10 and MRR and is 0.2pp (one question) +behind BM25 at Hit@1, as on the full haystack. + +Until 2026-09-27 this paragraph gave BM25-only 84.2%, vector-only 80.4% and +hybrid **85.2%** Hit@5. Those were measured on 2026-08-07 (c913cd1), when the +hybrid default was 0.7/0.3; today's 0.7/0.3 row reproduces the 85.2%. The +0.4/0.6 default (29baabb) is what moves hybrid to 86.8%. The earlier BM25-only +84.2% with embeddings, one question below the zero-embedding 84.4%, predates +the index tie-break of 3ed0489; the two now agree, as they always did on the +full haystack. ### Retracted: session-level recall and the MemX comparison diff --git a/ROADMAP.md b/ROADMAP.md index 914b02b..319a27e 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -135,7 +135,7 @@ **Crates:** `clawhdf5-agent`, `clawhdf5-bench` - [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547 -- [x] **8.2** LongMemEval benchmark — 500 queries, session Hit@1=100%, turn Hit@5=84.4% (beats MemX 51.6%), MRR=0.660 +- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results) - [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal - [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion - [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss @@ -168,7 +168,7 @@ Verified against current repo state on 2026-08-05 (see also `docs/superpowers/pl ### Recently closed out (2026-08-05, Tier 3–4 hardening pass) -- [x] Academic benchmark cross-validation — LongMemEval reproduced against MemX on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% vs MemX's 51.6%; recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05) +- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05) - [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer - [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories - [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open From 7179006aee8267414029428a5e520afaae54b4d5 Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 19:43:07 -0500 Subject: [PATCH 3/3] docs: local read A/B re-run on an idle machine The first run of the day could not get an idle tank: two orphaned h5py SWMR reader processes from earlier interop tests (since stopped) kept a core each busy. Re-run with the load below 2 at every round: ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5 ns per header, not visible in the facade listing, which is -1.7%); full deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single- thread contiguous hyperslab result from the loaded run is noise (-1.7%, overlapping ranges) and is withdrawn from known-issues. Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 116 ++++++++++++++++++------------------------- CHANGELOG.md | 33 ++++++------ docs/known-issues.md | 36 ++++++-------- 3 files changed, 79 insertions(+), 106 deletions(-) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index b5b3b07..f6e8a3c 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -502,83 +502,61 @@ explain the slower windows. ### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) -Range-read M2/M3 (PR #18) moved the facade onto the `Storage` abstraction. -This compares `main` just before it, `8f59b2e`, with `main` at `7a8fae0` -(PR #18 and PR #19 merged, including the slice-entry-point fix `ef428d7` -and the object-header fix `4313917`), each built in its own worktree and -target directory and run as separate binaries, alternating base, -candidate, base, candidate. +`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main` +`7a8fae0` (PRs #18 and #19), each built in its own worktree and run as +separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen +7 7800X3D, 16 threads). **Idle:** every round started with the 1-minute load +average below 2 (1.05–1.98; `target/ab-results2/load.log`). Criterion: +`taskset -c 5 local_metadata_bench --bench --warm-up-time 3 +--measurement-time 10`, 3 rounds each. Reads: `concurrent_read --dir +~/.cache/concurrent-read --decode-threads 1 --reps 3`, 3 rounds each. +Median (range) over the rounds. -**Provisional: tank was not idle.** Two orphaned h5py processes (leftover -SWMR test readers from earlier test runs, about 100% of a core each) held -the 1-minute load average at 2.1-2.6 throughout; the load < 2 gate was -polled every 20 s for 2 hours and never met, so the runs went ahead -anyway. Load average (1-minute) during the runs: Criterion rounds 1-3 -2.29-3.31, rounds 4-6 2.41-3.12 (each started below 2.5); -`concurrent_read` 3.04-6.49 (that figure includes its own 16 threads). -Machine: AMD Ryzen 7 7800X3D (8 cores, 16 threads). +`local_metadata_bench` (the 400-group v1 fixture): -Metadata, Criterion `crates/clawhdf5/benches/local_metadata_bench.rs` over -`clawhdf5-format/tests/fixtures/v1_groups_400.h5`, pinned to one CPU, 6 -alternating rounds each (median of each round's Criterion estimate, then -median and range over the rounds): - -```text -cargo bench -p clawhdf5 --bench local_metadata_bench --no-run -cd crates/clawhdf5 && taskset -c 5 --bench --noplot \ - --warm-up-time 3 --measurement-time 10 -``` - -| function | `8f59b2e` median (range) | `7a8fae0` median (range) | change | +| function | 8f59b2e | 7a8fae0 | change | |---|---:|---:|---:| -| `object_header_parse_x401` | 24.13 us (23.97-25.56) | 25.11 us (24.73-25.21) | +4.1% | -| `snod_parse_all` | 1.866 us (1.855-1.888) | 1.858 us (1.846-1.884) | -0.4% | -| `btree_v1_walk` | 344 ns (337-355) | 346 ns (340-355) | +0.5% | -| `facade_list_400_groups` | 8.26 ms (8.08-8.45) | 8.20 ms (8.14-8.36) | -0.8% | +| `object_header_parse_x401` | 23.86 µs (23.79–23.96) | 24.86 µs (24.69–24.98) | **+4.2%** | +| `snod_parse_all` | 1.842 µs (1.835–1.847) | 1.833 µs (1.833–1.854) | −0.5% | +| `btree_v1_walk` | 344 ns (338–346) | 348 ns (337–356) | +1.1% | +| `facade_list_400_groups` | 8.24 ms (8.04–8.28) | 8.10 ms (8.04–8.13) | −1.7% | -The facade listing's +7-10% seen during PR #18 is gone (-0.8%, within the -ranges). `ObjectHeader::parse` of 401 small headers is still slower: five -of six base rounds measured 23.97-24.29 us (one outlier 25.56), every -candidate round 24.73-25.21 us, so +4% (about 2.4 ns per header) looks -real rather than noise; it does not show in the listing that parses those -headers. Listed in `docs/known-issues.md`. +`concurrent_read`, MB/s (64 datasets of 64 MiB `f32`; deflate chunks +256 x 256, level 4): -Data reads, `concurrent_read` (the "Concurrent reads" workload below: 64 -datasets of 64 MiB, `--decode-threads 1`, 3 repetitions per point, median -MB/s reported by the harness), 3 alternating rounds each; median and range -over the rounds: - -```text -cargo build --release -p clawhdf5-bench --bin concurrent_read -concurrent_read --dir ~/.cache/concurrent-read --decode-threads 1 --reps 3 -``` - -| layout | mode | threads | `8f59b2e` MB/s (range) | `7a8fae0` MB/s (range) | change | +| layout | mode | threads | 8f59b2e | 7a8fae0 | change | |---|---|---:|---:|---:|---:| -| deflate | distinct | 1 | 815 (810-816) | 875 (874-877) | +7.4% | -| deflate | distinct | 4 | 2964 (2959-2969) | 3302 (3301-3309) | +11.4% | -| deflate | distinct | 8 | 4401 (4379-4414) | 5074 (5072-5080) | +15.3% | -| deflate | distinct | 16 | 5568 (5500-5598) | 6862 (6781-6962) | +23.2% | -| deflate | same | 1 | 228 (228-230) | 230 (230-231) | +1.1% | -| deflate | same | 4 | 896 (893-900) | 893 (890-896) | -0.3% | -| deflate | same | 16 | 1837 (1829-1839) | 1844 (1805-1865) | +0.4% | -| contiguous | distinct | 1 | 10360 (9636-10625) | 10143 (10063-10402) | -2.1% | -| contiguous | distinct | 4 | 13658 (13543-13750) | 13704 (13448-13770) | +0.3% | -| contiguous | distinct | 16 | 12311 (12256-12354) | 12307 (12273-12337) | -0.0% | -| contiguous | same | 1 | 22682 (22658-23411) | 21421 (20358-21908) | -5.6% | -| contiguous | same | 4 | 97424 (96473-100871) | 99417 (98952-105308) | +2.0% | -| contiguous | same | 16 | 137754 (126322-156061) | 152767 (136109-154669) | +10.9% | +| deflate | distinct | 1 | 891 (880–892) | 907 (906–908) | +1.7% | +| deflate | distinct | 2 | 1692 (1680–1699) | 1777 (1767–1779) | +5.0% | +| deflate | distinct | 4 | 3102 (3099–3102) | 3348 (3347–3350) | +8.0% | +| deflate | distinct | 8 | 5306 (5168–5345) | 6240 (6226–6248) | +17.6% | +| deflate | distinct | 16 | 6258 (6109–6442) | 8525 (8513–8561) | **+36.2%** | +| deflate | same | 1 | 234 (234–235) | 233 (233–234) | −0.5% | +| deflate | same | 16 | 2443 (2038–2444) | 2468 (2424–2473) | +1.0% | +| contiguous | distinct | 1 | 13477 (12949–13874) | 13302 (13287–13578) | −1.3% | +| contiguous | distinct | 16 | 12501 (12484–12530) | 12566 (12449–12567) | +0.5% | +| contiguous | same | 1 | 29866 (28832–30207) | 29364 (28838–29780) | −1.7% | +| contiguous | same | 16 | 163415 (161097–229146) | 233280 (159377–238440) | (noise) | -(2- and 8-thread rows are in line with their neighbours.) Full reads of -chunked deflate data are 7-23% *faster* at `7a8fae0`, with tight ranges; -which commit did it was not bisected. Contiguous full reads and deflate -hyperslabs are unchanged. One row is slower: one thread reading 256 x 256 -hyperslabs of a contiguous dataset, -5.6% with no overlap between the -rounds (about 0.6 us more per 256 KiB slab); the multi-thread `contiguous -same` rows, which are cache-bound and noisy, went the other way. Listed in -`docs/known-issues.md` as open, pending an idle re-run. +What this shows: +- **Local metadata reads are at parity or faster.** Listing the 400-group + file through the facade is 1.7% faster than before M2/M3; the +7–10% + listing regression found while merging #18 is gone. +- **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header; + the base and candidate ranges do not overlap). It is the cost of reading + continuation chunks from a bounded queue (the fix for unbounded reads on + crafted headers) and does not show in the listing. Kept open in + `docs/known-issues.md`. +- **Full reads of deflate data got faster** after #18 (in-place chunk + decoding into the typed output and per-thread scratch buffers): +1.7% on + one thread, +36% at 16. +- Single-thread contiguous hyperslabs are within noise (−1.7%, overlapping + ranges). The multi-thread `contiguous same` rows read one 64 MiB dataset + out of the CPU caches and swing widely between rounds of the same build. -## Concurrent reads +An earlier run the same day at load 2.3–3.3 (two orphaned h5py processes, +since stopped, each using a core) reported that single-thread contiguous +hyperslab row as −5.6%; the idle rerun above does not reproduce it. ### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2743744..306dd13 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -477,19 +477,19 @@ Design: `docs/design/swmr.md`. `main` and it calls only format-crate code; its earlier +5–10% (323–355 vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured 357–374 ns against `main`'s 355–371 in the same rounds). - **Rechecked 2026-09-27** (still provisional: two orphaned h5py test - processes kept tank's 1-minute load at 2.1–2.6, the load < 2 gate was - not met in 2 hours, load 2.3–3.3 during the runs; `8f59b2e` vs `main` - `7a8fae0` as separate binaries, pinned to one CPU, 6 alternating rounds, - median (range) over rounds; `BENCHMARKS.md`, "Local metadata and data - reads after range-read M2/M3"): facade listing 8.26 (8.08–8.45) → 8.20 - (8.14–8.36) ms, −0.8%, so the listing regression is gone; symbol-table - nodes −0.4%, B-tree walk +0.5%; `ObjectHeader::parse` ×401 24.13 - (23.97–25.56) → 25.11 (24.73–25.21) µs, **+4.1%, still open** - (`docs/known-issues.md`). Data reads (`concurrent_read --decode-threads - 1`, 3 alternating rounds): full reads of deflate data 7–23% faster, - contiguous full reads and deflate hyperslabs within ±2%, one-thread - contiguous hyperslabs −5.6% (open, same entry). + **Rechecked 2026-09-27 on an idle tank** (load below 2 at every round; + `8f59b2e` vs `main` `7a8fae0` as separate binaries, pinned to one CPU, 3 + alternating rounds, median (range); `BENCHMARKS.md`, "Local metadata and + data reads after range-read M2/M3"): facade listing 8.24 (8.04–8.28) → + 8.10 (8.04–8.13) ms, **−1.7%**, so the listing regression is gone; + symbol-table nodes −0.5%, B-tree walk +1.1%; `ObjectHeader::parse` ×401 + 23.86 (23.79–23.96) → 24.86 (24.69–24.98) µs, **+4.2%, open** + (`docs/known-issues.md`; about 2.5 ns per header, not visible in the + listing). Data reads (`concurrent_read --decode-threads 1`): full reads + of deflate data +1.7% (1 thread) to +36% (16 threads), contiguous reads + and deflate hyperslabs within ±2%. (A run earlier that day under load + 2.3–3.3 put single-thread contiguous hyperslabs at −5.6%; the idle rerun + does not reproduce it.) - Tests (2026-09-26, tank): - `clawhdf5-format/tests/storage_equivalence.rs` now also reads every dataset — whole, fill-aware, through a chunk cache (twice) and the @@ -644,10 +644,9 @@ Design: `docs/design/swmr.md`. instead of being inlined into the generic parser. Provisional (tank load 3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs 24.85–25.15 µs on `main`. - Rechecked 2026-09-27 with 6 alternating rounds (load 2.3–3.3, still - provisional): 24.73–25.21 µs against 23.97–24.29 µs (one outlier 25.56) - for `8f59b2e`, median +4.1%; a residual cost remains (see the M2 speed - note and `docs/known-issues.md`). + Rechecked 2026-09-27 on an idle tank (3 alternating rounds, load below + 2): 24.69–24.98 µs against 23.79–23.96 µs for `8f59b2e`, median +4.2%; a + residual cost remains (see the M2 speed note and `docs/known-issues.md`). - Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py `earliest`/`v110`/`latest` and clawhdf5-written files; structure comparisons with libhdf5 for version-2 B-trees, shrink on every index, diff --git a/docs/known-issues.md b/docs/known-issues.md index 613a729..98e1fac 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -7,27 +7,23 @@ deleting it. --- -## Local-read slowdowns left after range-read M2/M3 (measured 2026-09-27) +## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27) -**Status:** open (speed only; values are correct). Measured on tank, -`main` before range-read M2/M3 (`8f59b2e`) against `main` `7a8fae0`, -alternating separate binaries (`BENCHMARKS.md`, "Local metadata and data -reads after range-read M2/M3"). Provisional: the machine never went idle -(1-minute load 2.3–3.3 for the metadata runs, two orphaned h5py processes -using a core each), so re-run both on an idle tank before acting. -- `ObjectHeader::parse` of 401 small headers (`local_metadata_bench`, - `object_header_parse_x401`): 24.13 → 25.11 µs median over 6 rounds - (+4.1%, about 2.4 ns per header; ranges 23.97–25.56 vs 24.73–25.21, five - of six base rounds at or below 24.29). The header parser has read its - continuation chunks from a queue since `a69c5be`; `4313917` removed its - per-header allocations but not all of the cost. Listing the 400-group - file through the facade, which parses those headers, is not slower - (−0.8%). -- One thread reading 256 x 256 hyperslabs of a contiguous dataset - (`concurrent_read`, `contiguous same`, 1 thread): 22682 → 21421 MB/s - median over 3 rounds (−5.6%; 22658–23411 vs 20358–21908), about 0.6 µs - more per 256 KiB slab. The multi-thread rows of that mode, which are - cache-bound and noisy, did not slow down; full reads did not either. +**Status:** open (speed only; values are correct). Measured on an idle tank +(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`) +against `main` `7a8fae0`, separate binaries alternating, 3 rounds +(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"): +`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges +23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read +continuation chunks from a bounded queue since `a69c5be` (bounding what a +crafted header can make it read); `4313917` removed its per-header +allocations but not all of the cost. Listing the 400-group file through the +facade, which parses the same headers, is 1.7% faster, so no user-visible +path is slower. + +An earlier run the same day under load also listed single-thread contiguous +hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with +overlapping ranges (noise), so that item is withdrawn. ## Files a SWMR writer had open could not be read past a stale end of file