diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 62beeb6..f6e8a3c 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -30,12 +30,15 @@ target: Criterion stretched it where 5 s could not hold the samples it needed > and in memory), Consolidation Efficiency, Ephemeral Tier, Multi-modal Search, > the Search and Read harnesses, and the "h5bench-Equivalent I/O Benchmarks" > and "Independent Validation: tank" sections. What does not yet meet that bar: -> the LongMemEval rows that need real embeddings (not re-run here, except the -> dated float16 comparison), the Consolidation Efficiency 100K cycle row and +> the Consolidation Efficiency 100K cycle row and > memory-reduction part (the 2026-09-24 run was stopped before it produced > them), the int8 side of "Quantising the index copy" (not re-run), and the > i7-12650H and macOS M3 Max rows under Cross-Platform Notes. That is a > known, tracked documentation gap, not a claim that those numbers are wrong. +> The LongMemEval rows that need real embeddings (vector-only, hybrid, RRF, +> stemmed hybrid, re-ranking, the weight sweep, the oracle variant) were +> re-run on 2026-09-27 on tank; see "Re-run with real embeddings" under +> LongMemEval Results. > > **Correctness note (2026-08-06).** Being dated and reproducible is necessary but > not sufficient — a number can be perfectly reproducible and still measure the @@ -325,6 +328,19 @@ which of two gold sessions ranks first, out of ~320. Those flips show the half-precision path was in effect; they do not change a single hit. The f32 run reproduces the published hybrid numbers exactly. +**Re-checked 2026-09-27** (tank, commit 7a8fae0, the same pair of runs on +`longmemeval_s_cleaned.json`, which is the same file; the machine was not +idle, which does not affect recall): the result is the same. The table above +reproduced exactly. Every Hit@k and MRR of the eight modes matched between f32 +and float16 at both levels, with two exceptions: RRF's session MRR (0.9253 vs +0.9254) and two per-type session MRRs in the fourth decimal. Three modes +differed by one question in the recency count. That re-run also corrects the +sentence above: two f32 runs on the same day differed by one question in +recency as well, so those flips are run-to-run variation and do not show that +the half-precision path was in effect. `--float16` is what shows that: the +harness prints "Stores use MemoryConfig::float16" and `MemoryConfig::float16` +is set on every store. + ### Opening a store (`read_from_disk`) `HDF5Memory::open` memory-mapped the file, copied the whole mapping into a @@ -482,7 +498,65 @@ The rows and columns of the uncompressed layouts are within 20% (chunked column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not explain the slower windows. -## Concurrent reads +## Local file speed after range reads + +### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) + +`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main` +`7a8fae0` (PRs #18 and #19), each built in its own worktree and run as +separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen +7 7800X3D, 16 threads). **Idle:** every round started with the 1-minute load +average below 2 (1.05–1.98; `target/ab-results2/load.log`). Criterion: +`taskset -c 5 local_metadata_bench --bench --warm-up-time 3 +--measurement-time 10`, 3 rounds each. Reads: `concurrent_read --dir +~/.cache/concurrent-read --decode-threads 1 --reps 3`, 3 rounds each. +Median (range) over the rounds. + +`local_metadata_bench` (the 400-group v1 fixture): + +| function | 8f59b2e | 7a8fae0 | change | +|---|---:|---:|---:| +| `object_header_parse_x401` | 23.86 µs (23.79–23.96) | 24.86 µs (24.69–24.98) | **+4.2%** | +| `snod_parse_all` | 1.842 µs (1.835–1.847) | 1.833 µs (1.833–1.854) | −0.5% | +| `btree_v1_walk` | 344 ns (338–346) | 348 ns (337–356) | +1.1% | +| `facade_list_400_groups` | 8.24 ms (8.04–8.28) | 8.10 ms (8.04–8.13) | −1.7% | + +`concurrent_read`, MB/s (64 datasets of 64 MiB `f32`; deflate chunks +256 x 256, level 4): + +| layout | mode | threads | 8f59b2e | 7a8fae0 | change | +|---|---|---:|---:|---:|---:| +| deflate | distinct | 1 | 891 (880–892) | 907 (906–908) | +1.7% | +| deflate | distinct | 2 | 1692 (1680–1699) | 1777 (1767–1779) | +5.0% | +| deflate | distinct | 4 | 3102 (3099–3102) | 3348 (3347–3350) | +8.0% | +| deflate | distinct | 8 | 5306 (5168–5345) | 6240 (6226–6248) | +17.6% | +| deflate | distinct | 16 | 6258 (6109–6442) | 8525 (8513–8561) | **+36.2%** | +| deflate | same | 1 | 234 (234–235) | 233 (233–234) | −0.5% | +| deflate | same | 16 | 2443 (2038–2444) | 2468 (2424–2473) | +1.0% | +| contiguous | distinct | 1 | 13477 (12949–13874) | 13302 (13287–13578) | −1.3% | +| contiguous | distinct | 16 | 12501 (12484–12530) | 12566 (12449–12567) | +0.5% | +| contiguous | same | 1 | 29866 (28832–30207) | 29364 (28838–29780) | −1.7% | +| contiguous | same | 16 | 163415 (161097–229146) | 233280 (159377–238440) | (noise) | + +What this shows: +- **Local metadata reads are at parity or faster.** Listing the 400-group + file through the facade is 1.7% faster than before M2/M3; the +7–10% + listing regression found while merging #18 is gone. +- **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header; + the base and candidate ranges do not overlap). It is the cost of reading + continuation chunks from a bounded queue (the fix for unbounded reads on + crafted headers) and does not show in the listing. Kept open in + `docs/known-issues.md`. +- **Full reads of deflate data got faster** after #18 (in-place chunk + decoding into the typed output and per-thread scratch buffers): +1.7% on + one thread, +36% at 16. +- Single-thread contiguous hyperslabs are within noise (−1.7%, overlapping + ranges). The multi-thread `contiguous same` rows read one 64 MiB dataset + out of the CPU caches and swing widely between rounds of the same build. + +An earlier run the same day at load 2.3–3.3 (two orphaned h5py processes, +since stopped, each using a core) reported that single-thread contiguous +hyperslab row as −5.6%; the idle rerun above does not reproduce it. ### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`) @@ -1410,7 +1484,67 @@ turn-level row plus its session Hit@1 in the tokenizer table. The run also produced figures this document does not publish (stemmed session Hit@5, Hit@10 and MRR, and per-type Hit@5/Hit@10/MRR for both modes), so there was nothing to compare them with. Rows that need real embeddings (vector-only, -hybrid, RRF, re-ranking, the weight sweep) were not re-run. +hybrid, RRF, re-ranking, the weight sweep) were not re-run then; they were on +2026-09-27 (next section). + +### Re-run with real embeddings (2026-09-27, tank) + +Every recall row in this section that needs real embeddings was measured +again on 2026-09-27 on tank (AMD Ryzen 7 7800X3D, 16 threads; MiniLM +embeddings on an RTX 5060 Ti, retrieval on the CPU), with the search code of +commit 7a8fae0. Six runs: + +```bash +cargo build --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings-cuda +B=target/release/longmemeval_bench W=weights/all-minilm-l6-v2 +$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --float16 +$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W --sweep +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --sweep +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --rerank-sweep +``` + +(`longmemeval_s.json`, used by the commands elsewhere in this section, is the +same file as `longmemeval_s_cleaned.json`.) + +**The machine was not idle.** The 1-minute load never fell below 2 in a +2-hour wait, because two stray test processes were each holding a core; it +was 2.1–2.7 when each run started and up to 8.4 while the runs were going. +Recall does not depend on load. Latency does, so no latency figure in this +file was updated from these runs. + +**What reproduced exactly:** every published full-haystack Hit@1/5/10 and MRR +at turn and session level for BM25, vector-only, hybrid 0.4/0.6, both stemmed +modes, and every re-ranking row; RRF except as below; the float16 table; 4 of +the 11 weight-sweep rows (0.0, 0.1, 0.3 and 0.9); and the oracle vector-only +figure. The headline, +hybrid 0.4/0.6 turn Hit@5 **81.4%**, is unchanged. + +**What changed.** Old values are kept here; the tables below show the new ones. + +| Figure | Published | 2026-09-27 | Why | +|---|---:|---:|---| +| Oracle, hybrid turn Hit@5 | 85.2% (2026-08-07, c913cd1) | **86.8%** | Weights. 85.2% was measured at 0.7/0.3, the default then. The default has been 0.4/0.6 since 29baabb. Today's oracle sweep gives 85.2% at 0.7/0.3 and 86.8% at 0.4/0.6. | +| Oracle, BM25-only turn Hit@5 with embeddings | 84.2% (2026-08-07) | 84.4% | Now equal to the zero-embedding figure, as it already was on the full haystack. At c913cd1, tied candidates were ordered by `HashMap` iteration. Since 3ed0489 (2026-09-19) they are broken by index. 3ed0489 is the likely cause; it was not bisected. | +| Ablation / sweep, hybrid 0.7/0.3, turn | 44.4% / 79.2% / 86.0% / 0.5868 | 44.2% / 79.2% / 85.8% / 0.5856 | The same tie-breaking change. 29baabb's re-run on 2026-09-19, made after 3ed0489 that same day, already had 44.2% / 85.8% / 0.5856. | +| Ablation, hybrid 0.7/0.3, session | 88.2% / 95.8% / 97.8% / 0.9158 | 88.0% / 95.8% / 97.6% / 0.9146 | Same. | +| Ablation / sweep, vector-only turn MRR | 0.5027 | 0.5031 | Same. The fusion and float16 tables already had 0.5031. | +| Sweep 0.2/0.8, turn | 53.6% / 78.2% / 0.6440 | 53.4% / 78.0% / 0.6429 | Same (the sweep was measured 2026-08-07, 1537a94). | +| Sweep 0.4/0.6, 0.5/0.5, 0.8/0.2 turn MRR | 0.6429 / 0.6234 / 0.5571 | 0.6430 / 0.6232 / 0.5574 | Same. | +| Sweep 0.6/0.4 | Hit@5 79.8%, MRR 0.6069, session Hit@5 96.6% | 79.6%, 0.6067, 96.4% | Same. | +| RRF turn MRR | 0.5967 (2026-09-19, aa92fef) | 0.5969 | Not explained. It may be a later search-path change, such as the int8 index becoming the default in 8b85d93. It may also be the run-to-run variation described next. It was not bisected. | + +**Not exactly deterministic.** Between two f32 runs today, the recency count +(the share of `knowledge-update` questions where the newest gold session +ranks first) differed by one question in two modes: hybrid 0.4/0.6 was +144/320 in one run and 145/320 in the other, and re-rank with a 1-day +half-life was 165/319 and 166/319. Each question gets a fresh store, so this +is not state carried between modes. The cause was not found; a candidate is +that the GPU embeddings are not bit-for-bit identical from run to run. Hit@k +and MRR agreed in every run that repeated a mode (hybrid 0.4/0.6 was measured +four times: three f32 runs and one float16 run), but a last-digit change in an +MRR, or a one-question change in recency, is within this variation. ### Full haystack — `longmemeval_s`, n=500 (the number to cite) @@ -1440,13 +1574,16 @@ the question. Real 384-d `all-MiniLM-L6-v2` embeddings, 190,015 unique texts encoded once on an RTX 5060 Ti (~13 min; the same work on the 8-core CPU was still unfinished after -30 minutes, so the GPU path is not a convenience here). Turn-level: +30 minutes, so the GPU path is not a convenience here). First measured +2026-08-07 (c913cd1); the values below are the 2026-09-27 re-run, which moved +the vector-only MRR and the `0.7/0.3` row (see the re-run section above). +Turn-level: | Mode | Hit@1 | Hit@5 | Hit@10 | MRR | |------|-------|-------|--------|-----| | BM25 only (`0.0`/`1.0`) | **53.8%** | 75.0% | 81.6% | **0.6320** | -| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5027 | -| Hybrid (`0.7`/`0.3`) | 44.4% | **79.2%** | **86.0%** | 0.5868 | +| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5031 | +| Hybrid (`0.7`/`0.3`) | 44.2% | **79.2%** | **85.8%** | 0.5856 | Session-level: @@ -1454,7 +1591,7 @@ Session-level: |------|-------|-------|--------|-----| | BM25 only | 86.2% | 93.6% | 96.6% | 0.8948 | | Vector only | 85.4% | 94.2% | 96.6% | 0.8901 | -| Hybrid | **88.2%** | **95.8%** | **97.8%** | **0.9158** | +| Hybrid | **88.0%** | **95.8%** | **97.6%** | **0.9146** | ### Fusion method — weighted vs. RRF, full haystack, n=500 @@ -1468,7 +1605,7 @@ takes a `Fusion`, and both run over the same HNSW + BM25 candidates: | BM25 only | **53.8%** | 75.0% | 81.6% | 0.6320 | 86.2% | 0.8948 | | Vector only | 36.0% | 71.8% | 81.6% | 0.5031 | 85.4% | 0.8901 | | **Weighted 0.4 / 0.6** | 51.6% | **81.4%** | **87.8%** | **0.6430** | **91.0%** | **0.9347** | -| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5967 | 89.6% | 0.9253 | +| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5969 | 89.6% | 0.9253 | **RRF loses to the tuned weighted sum** — 6.6pp of turn Hit@1 and 0.046 of MRR — and lands almost exactly where the old `0.7/0.3` weighting did (44.2% / @@ -1533,6 +1670,12 @@ one (see `newest_gold_first`); ~45% is chance. | + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% | | + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% | +Re-run on 2026-09-27 (`--rerank-sweep`, tank, commit 7a8fae0): every Hit@k and +MRR above reproduced exactly. The recency column came out 45.3%, 87.5%, +52.0%, 52.2%, 51.7% and 50.5%. Each of those is within one question of the +value in the table, which is the run-to-run variation described under +"Re-run with real embeddings" above, so the table was left as it was. + **The pre-fix row is the finding.** Ordering candidates by recency alone costs 40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in the pool rather than the ones that answer the question. It does ace the recency @@ -1546,9 +1689,9 @@ cannot reach the 87.5% the degenerate ordering gets. Those two rows are the ends of a trade-off, and the default sits deliberately near the relevance end. **Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency -moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because -the temporal term is capped by its weight (0.3) while relevance differences -between candidates are larger. The 24-hour default is kept; there is no +moves 1.4pp (1.7pp in the 2026-09-27 re-run) and MRR 0.003 — inside the +noise of a 500-question run — because the temporal term is capped by its +weight (0.3) while relevance differences between candidates are larger. The 24-hour default is kept; there is no measured reason to change it, and a corpus-matched value is not the lever it looks like. @@ -1556,24 +1699,28 @@ looks like. `0.7/0.3` was a documented default, never a searched one. Sweeping `vector_weight` from 0.0 to 1.0 (`--sweep`, reusing the one-time embedding -table) shows it is not merely suboptimal but **strictly dominated**: +table) shows it is not merely suboptimal but **strictly dominated**. First +measured 2026-08-07 (1537a94); the values below are the 2026-09-27 re-run, +which changed the 0.2, 0.4, 0.5, 0.6, 0.7, 0.8 and 1.0 rows in the last digit +or by one or two questions (see the re-run section above): | vector / keyword | Hit@1 | Hit@5 | Hit@10 | MRR | session Hit@5 | |---|---|---|---|---|---| | 0.0 / 1.0 (BM25) | **53.8%** | 75.0% | 81.6% | 0.6320 | 93.6% | | 0.1 / 0.9 | 53.2% | 77.4% | 83.8% | 0.6374 | 95.0% | -| 0.2 / 0.8 | 53.6% | 78.2% | 85.6% | 0.6440 | 95.4% | +| 0.2 / 0.8 | 53.4% | 78.0% | 85.6% | 0.6429 | 95.4% | | 0.3 / 0.7 | 53.2% | 78.8% | 87.2% | **0.6463** | 96.0% | -| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6429 | 96.8% | -| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6234 | **97.4%** | -| 0.6 / 0.4 | 46.6% | 79.8% | 87.4% | 0.6069 | 96.6% | -| 0.7 / 0.3 *(old default)* | 44.4% | 79.2% | 86.0% | 0.5868 | 95.8% | -| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5571 | 95.2% | +| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6430 | 96.8% | +| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6232 | **97.4%** | +| 0.6 / 0.4 | 46.6% | 79.6% | 87.4% | 0.6067 | 96.4% | +| 0.7 / 0.3 *(old default)* | 44.2% | 79.2% | 85.8% | 0.5856 | 95.8% | +| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5574 | 95.2% | | 0.9 / 0.1 | 37.8% | 73.4% | 84.6% | 0.5289 | 94.2% | -| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5027 | 94.2% | +| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5031 | 94.2% | **`0.4/0.6` beats `0.7/0.3` on every metric at both granularities** — Hit@1 -+7.2pp, Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. There is no trade being made; the ++7.4pp, Hit@5 +2.2, Hit@10 +2.0, MRR +0.057 (2026-09-27 figures; +7.2pp, ++2.2, +1.8 and +0.056 as measured on 2026-08-07). There is no trade being made; the old default was simply on the wrong side of the peak. **`0.4/0.6` is the recommended setting**, with `0.3/0.7` preferable if rank-1 precision matters most (it takes the best MRR in the sweep and gives up only 0.6pp of Hit@1 @@ -1634,11 +1781,27 @@ price of the harder corpus, and is the reason oracle-only numbers should not be presented as LongMemEval results. Session-level figures on this variant are degenerate — see below. -With real embeddings the same oracle corpus gives BM25-only 84.2% / vector-only -80.4% / hybrid **85.2%** Hit@5 turn-level — hybrid ahead at Hit@5 and Hit@10 and -behind at Hit@1, matching the full-haystack pattern above. (BM25-only reads 84.2% -here against 84.4% with zero embedding vectors: one question of 500 changes rank, -with MRR identical at 0.6597. On the full haystack the two agree exactly.) +With real embeddings, the same oracle corpus gives these turn-level figures +(2026-09-27, tank, commit 7a8fae0; command and load in "Re-run with real +embeddings" above): + +| Mode | Hit@1 | Hit@5 | Hit@10 | MRR | +|---|---:|---:|---:|---:| +| BM25 only | **52.6%** | 84.4% | 90.4% | 0.6597 | +| Vector only | 39.6% | 80.4% | 91.0% | 0.5605 | +| Hybrid 0.4 / 0.6 (default) | 52.4% | **86.8%** | **92.4%** | **0.6678** | +| Hybrid 0.7 / 0.3 (old default) | 48.4% | 85.2% | 92.2% | 0.6382 | + +Hybrid 0.4/0.6 leads at Hit@5, Hit@10 and MRR and is 0.2pp (one question) +behind BM25 at Hit@1, as on the full haystack. + +Until 2026-09-27 this paragraph gave BM25-only 84.2%, vector-only 80.4% and +hybrid **85.2%** Hit@5. Those were measured on 2026-08-07 (c913cd1), when the +hybrid default was 0.7/0.3; today's 0.7/0.3 row reproduces the 85.2%. The +0.4/0.6 default (29baabb) is what moves hybrid to 86.8%. The earlier BM25-only +84.2% with embeddings, one question below the zero-embedding 84.4%, predates +the index tie-break of 3ed0489; the two now agree, as they always did on the +full haystack. ### Retracted: session-level recall and the MemX comparison diff --git a/CHANGELOG.md b/CHANGELOG.md index f7d44f1..306dd13 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -477,6 +477,19 @@ Design: `docs/design/swmr.md`. `main` and it calls only format-crate code; its earlier +5–10% (323–355 vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured 357–374 ns against `main`'s 355–371 in the same rounds). + **Rechecked 2026-09-27 on an idle tank** (load below 2 at every round; + `8f59b2e` vs `main` `7a8fae0` as separate binaries, pinned to one CPU, 3 + alternating rounds, median (range); `BENCHMARKS.md`, "Local metadata and + data reads after range-read M2/M3"): facade listing 8.24 (8.04–8.28) → + 8.10 (8.04–8.13) ms, **−1.7%**, so the listing regression is gone; + symbol-table nodes −0.5%, B-tree walk +1.1%; `ObjectHeader::parse` ×401 + 23.86 (23.79–23.96) → 24.86 (24.69–24.98) µs, **+4.2%, open** + (`docs/known-issues.md`; about 2.5 ns per header, not visible in the + listing). Data reads (`concurrent_read --decode-threads 1`): full reads + of deflate data +1.7% (1 thread) to +36% (16 threads), contiguous reads + and deflate hyperslabs within ±2%. (A run earlier that day under load + 2.3–3.3 put single-thread contiguous hyperslabs at −5.6%; the idle rerun + does not reproduce it.) - Tests (2026-09-26, tank): - `clawhdf5-format/tests/storage_equivalence.rs` now also reads every dataset — whole, fill-aware, through a chunk cache (twice) and the @@ -631,6 +644,9 @@ Design: `docs/design/swmr.md`. instead of being inlined into the generic parser. Provisional (tank load 3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs 24.85–25.15 µs on `main`. + Rechecked 2026-09-27 on an idle tank (3 alternating rounds, load below + 2): 24.69–24.98 µs against 23.79–23.96 µs for `8f59b2e`, median +4.2%; a + residual cost remains (see the M2 speed note and `docs/known-issues.md`). - Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py `earliest`/`v110`/`latest` and clawhdf5-written files; structure comparisons with libhdf5 for version-2 B-trees, shrink on every index, diff --git a/ROADMAP.md b/ROADMAP.md index 914b02b..319a27e 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -135,7 +135,7 @@ **Crates:** `clawhdf5-agent`, `clawhdf5-bench` - [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547 -- [x] **8.2** LongMemEval benchmark — 500 queries, session Hit@1=100%, turn Hit@5=84.4% (beats MemX 51.6%), MRR=0.660 +- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results) - [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal - [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion - [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss @@ -168,7 +168,7 @@ Verified against current repo state on 2026-08-05 (see also `docs/superpowers/pl ### Recently closed out (2026-08-05, Tier 3–4 hardening pass) -- [x] Academic benchmark cross-validation — LongMemEval reproduced against MemX on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% vs MemX's 51.6%; recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05) +- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05) - [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer - [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories - [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open diff --git a/docs/known-issues.md b/docs/known-issues.md index 7eea8e7..98e1fac 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -7,6 +7,24 @@ deleting it. --- +## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27) + +**Status:** open (speed only; values are correct). Measured on an idle tank +(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`) +against `main` `7a8fae0`, separate binaries alternating, 3 rounds +(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"): +`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges +23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read +continuation chunks from a bounded queue since `a69c5be` (bounding what a +crafted header can make it read); `4313917` removed its per-header +allocations but not all of the cost. Listing the 400-group file through the +facade, which parses the same headers, is 1.7% faster, so no user-visible +path is slower. + +An earlier run the same day under load also listed single-thread contiguous +hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with +overlapping ranges (noise), so that item is withdrawn. + ## Files a SWMR writer had open could not be read past a stale end of file **Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any