From b39a705e77651da24815b63258f610710275046c Mon Sep 17 00:00:00 2001 From: osobh Date: Sun, 27 Sep 2026 19:25:36 -0500 Subject: [PATCH] docs: LongMemEval with real embeddings re-run on 2026-09-27 Oracle and full longmemeval_s haystack, f32 and --float16, plus --sweep (both corpora) and --rerank-sweep, on tank at 7a8fae0 (MiniLM on the RTX 5060 Ti). Not idle: two orphaned h5py test processes kept the 1-minute load at 2.1-2.7 (up to 8.4 during runs), so no latency figure was updated. The headline (hybrid 0.4/0.6 turn Hit@5 81.4%) and the float16 claim reproduce. Oracle hybrid is 86.8%, not 85.2%: the old figure was at the 0.7/0.3 default of the time (today's 0.7/0.3 gives 85.2%). The ablation's 0.7/0.3 row, vector-only MRR and seven sweep rows move in the last digit or by one or two questions, most likely from the index tie-break of 3ed0489. RRF turn MRR 0.5967 -> 0.5969 is not explained. Recency counts vary by one question between identical runs, so the float16 "flips" note is corrected. ROADMAP 8.2 no longer repeats the retracted session-level and MemX claims. Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 155 ++++++++++++++++++++++++++++++++++++++++++-------- ROADMAP.md | 4 +- 2 files changed, 132 insertions(+), 27 deletions(-) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 7937d0e..b5b3b07 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -30,12 +30,15 @@ target: Criterion stretched it where 5 s could not hold the samples it needed > and in memory), Consolidation Efficiency, Ephemeral Tier, Multi-modal Search, > the Search and Read harnesses, and the "h5bench-Equivalent I/O Benchmarks" > and "Independent Validation: tank" sections. What does not yet meet that bar: -> the LongMemEval rows that need real embeddings (not re-run here, except the -> dated float16 comparison), the Consolidation Efficiency 100K cycle row and +> the Consolidation Efficiency 100K cycle row and > memory-reduction part (the 2026-09-24 run was stopped before it produced > them), the int8 side of "Quantising the index copy" (not re-run), and the > i7-12650H and macOS M3 Max rows under Cross-Platform Notes. That is a > known, tracked documentation gap, not a claim that those numbers are wrong. +> The LongMemEval rows that need real embeddings (vector-only, hybrid, RRF, +> stemmed hybrid, re-ranking, the weight sweep, the oracle variant) were +> re-run on 2026-09-27 on tank; see "Re-run with real embeddings" under +> LongMemEval Results. > > **Correctness note (2026-08-06).** Being dated and reproducible is necessary but > not sufficient — a number can be perfectly reproducible and still measure the @@ -325,6 +328,19 @@ which of two gold sessions ranks first, out of ~320. Those flips show the half-precision path was in effect; they do not change a single hit. The f32 run reproduces the published hybrid numbers exactly. +**Re-checked 2026-09-27** (tank, commit 7a8fae0, the same pair of runs on +`longmemeval_s_cleaned.json`, which is the same file; the machine was not +idle, which does not affect recall): the result is the same. The table above +reproduced exactly. Every Hit@k and MRR of the eight modes matched between f32 +and float16 at both levels, with two exceptions: RRF's session MRR (0.9253 vs +0.9254) and two per-type session MRRs in the fourth decimal. Three modes +differed by one question in the recency count. That re-run also corrects the +sentence above: two f32 runs on the same day differed by one question in +recency as well, so those flips are run-to-run variation and do not show that +the half-precision path was in effect. `--float16` is what shows that: the +harness prints "Stores use MemoryConfig::float16" and `MemoryConfig::float16` +is set on every store. + ### Opening a store (`read_from_disk`) `HDF5Memory::open` memory-mapped the file, copied the whole mapping into a @@ -1490,7 +1506,67 @@ turn-level row plus its session Hit@1 in the tokenizer table. The run also produced figures this document does not publish (stemmed session Hit@5, Hit@10 and MRR, and per-type Hit@5/Hit@10/MRR for both modes), so there was nothing to compare them with. Rows that need real embeddings (vector-only, -hybrid, RRF, re-ranking, the weight sweep) were not re-run. +hybrid, RRF, re-ranking, the weight sweep) were not re-run then; they were on +2026-09-27 (next section). + +### Re-run with real embeddings (2026-09-27, tank) + +Every recall row in this section that needs real embeddings was measured +again on 2026-09-27 on tank (AMD Ryzen 7 7800X3D, 16 threads; MiniLM +embeddings on an RTX 5060 Ti, retrieval on the CPU), with the search code of +commit 7a8fae0. Six runs: + +```bash +cargo build --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings-cuda +B=target/release/longmemeval_bench W=weights/all-minilm-l6-v2 +$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --float16 +$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W --sweep +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --sweep +$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --rerank-sweep +``` + +(`longmemeval_s.json`, used by the commands elsewhere in this section, is the +same file as `longmemeval_s_cleaned.json`.) + +**The machine was not idle.** The 1-minute load never fell below 2 in a +2-hour wait, because two stray test processes were each holding a core; it +was 2.1–2.7 when each run started and up to 8.4 while the runs were going. +Recall does not depend on load. Latency does, so no latency figure in this +file was updated from these runs. + +**What reproduced exactly:** every published full-haystack Hit@1/5/10 and MRR +at turn and session level for BM25, vector-only, hybrid 0.4/0.6, both stemmed +modes, and every re-ranking row; RRF except as below; the float16 table; 4 of +the 11 weight-sweep rows (0.0, 0.1, 0.3 and 0.9); and the oracle vector-only +figure. The headline, +hybrid 0.4/0.6 turn Hit@5 **81.4%**, is unchanged. + +**What changed.** Old values are kept here; the tables below show the new ones. + +| Figure | Published | 2026-09-27 | Why | +|---|---:|---:|---| +| Oracle, hybrid turn Hit@5 | 85.2% (2026-08-07, c913cd1) | **86.8%** | Weights. 85.2% was measured at 0.7/0.3, the default then. The default has been 0.4/0.6 since 29baabb. Today's oracle sweep gives 85.2% at 0.7/0.3 and 86.8% at 0.4/0.6. | +| Oracle, BM25-only turn Hit@5 with embeddings | 84.2% (2026-08-07) | 84.4% | Now equal to the zero-embedding figure, as it already was on the full haystack. At c913cd1, tied candidates were ordered by `HashMap` iteration. Since 3ed0489 (2026-09-19) they are broken by index. 3ed0489 is the likely cause; it was not bisected. | +| Ablation / sweep, hybrid 0.7/0.3, turn | 44.4% / 79.2% / 86.0% / 0.5868 | 44.2% / 79.2% / 85.8% / 0.5856 | The same tie-breaking change. 29baabb's re-run on 2026-09-19, made after 3ed0489 that same day, already had 44.2% / 85.8% / 0.5856. | +| Ablation, hybrid 0.7/0.3, session | 88.2% / 95.8% / 97.8% / 0.9158 | 88.0% / 95.8% / 97.6% / 0.9146 | Same. | +| Ablation / sweep, vector-only turn MRR | 0.5027 | 0.5031 | Same. The fusion and float16 tables already had 0.5031. | +| Sweep 0.2/0.8, turn | 53.6% / 78.2% / 0.6440 | 53.4% / 78.0% / 0.6429 | Same (the sweep was measured 2026-08-07, 1537a94). | +| Sweep 0.4/0.6, 0.5/0.5, 0.8/0.2 turn MRR | 0.6429 / 0.6234 / 0.5571 | 0.6430 / 0.6232 / 0.5574 | Same. | +| Sweep 0.6/0.4 | Hit@5 79.8%, MRR 0.6069, session Hit@5 96.6% | 79.6%, 0.6067, 96.4% | Same. | +| RRF turn MRR | 0.5967 (2026-09-19, aa92fef) | 0.5969 | Not explained. It may be a later search-path change, such as the int8 index becoming the default in 8b85d93. It may also be the run-to-run variation described next. It was not bisected. | + +**Not exactly deterministic.** Between two f32 runs today, the recency count +(the share of `knowledge-update` questions where the newest gold session +ranks first) differed by one question in two modes: hybrid 0.4/0.6 was +144/320 in one run and 145/320 in the other, and re-rank with a 1-day +half-life was 165/319 and 166/319. Each question gets a fresh store, so this +is not state carried between modes. The cause was not found; a candidate is +that the GPU embeddings are not bit-for-bit identical from run to run. Hit@k +and MRR agreed in every run that repeated a mode (hybrid 0.4/0.6 was measured +four times: three f32 runs and one float16 run), but a last-digit change in an +MRR, or a one-question change in recency, is within this variation. ### Full haystack — `longmemeval_s`, n=500 (the number to cite) @@ -1520,13 +1596,16 @@ the question. Real 384-d `all-MiniLM-L6-v2` embeddings, 190,015 unique texts encoded once on an RTX 5060 Ti (~13 min; the same work on the 8-core CPU was still unfinished after -30 minutes, so the GPU path is not a convenience here). Turn-level: +30 minutes, so the GPU path is not a convenience here). First measured +2026-08-07 (c913cd1); the values below are the 2026-09-27 re-run, which moved +the vector-only MRR and the `0.7/0.3` row (see the re-run section above). +Turn-level: | Mode | Hit@1 | Hit@5 | Hit@10 | MRR | |------|-------|-------|--------|-----| | BM25 only (`0.0`/`1.0`) | **53.8%** | 75.0% | 81.6% | **0.6320** | -| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5027 | -| Hybrid (`0.7`/`0.3`) | 44.4% | **79.2%** | **86.0%** | 0.5868 | +| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5031 | +| Hybrid (`0.7`/`0.3`) | 44.2% | **79.2%** | **85.8%** | 0.5856 | Session-level: @@ -1534,7 +1613,7 @@ Session-level: |------|-------|-------|--------|-----| | BM25 only | 86.2% | 93.6% | 96.6% | 0.8948 | | Vector only | 85.4% | 94.2% | 96.6% | 0.8901 | -| Hybrid | **88.2%** | **95.8%** | **97.8%** | **0.9158** | +| Hybrid | **88.0%** | **95.8%** | **97.6%** | **0.9146** | ### Fusion method — weighted vs. RRF, full haystack, n=500 @@ -1548,7 +1627,7 @@ takes a `Fusion`, and both run over the same HNSW + BM25 candidates: | BM25 only | **53.8%** | 75.0% | 81.6% | 0.6320 | 86.2% | 0.8948 | | Vector only | 36.0% | 71.8% | 81.6% | 0.5031 | 85.4% | 0.8901 | | **Weighted 0.4 / 0.6** | 51.6% | **81.4%** | **87.8%** | **0.6430** | **91.0%** | **0.9347** | -| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5967 | 89.6% | 0.9253 | +| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5969 | 89.6% | 0.9253 | **RRF loses to the tuned weighted sum** — 6.6pp of turn Hit@1 and 0.046 of MRR — and lands almost exactly where the old `0.7/0.3` weighting did (44.2% / @@ -1613,6 +1692,12 @@ one (see `newest_gold_first`); ~45% is chance. | + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% | | + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% | +Re-run on 2026-09-27 (`--rerank-sweep`, tank, commit 7a8fae0): every Hit@k and +MRR above reproduced exactly. The recency column came out 45.3%, 87.5%, +52.0%, 52.2%, 51.7% and 50.5%. Each of those is within one question of the +value in the table, which is the run-to-run variation described under +"Re-run with real embeddings" above, so the table was left as it was. + **The pre-fix row is the finding.** Ordering candidates by recency alone costs 40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in the pool rather than the ones that answer the question. It does ace the recency @@ -1626,9 +1711,9 @@ cannot reach the 87.5% the degenerate ordering gets. Those two rows are the ends of a trade-off, and the default sits deliberately near the relevance end. **Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency -moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because -the temporal term is capped by its weight (0.3) while relevance differences -between candidates are larger. The 24-hour default is kept; there is no +moves 1.4pp (1.7pp in the 2026-09-27 re-run) and MRR 0.003 — inside the +noise of a 500-question run — because the temporal term is capped by its +weight (0.3) while relevance differences between candidates are larger. The 24-hour default is kept; there is no measured reason to change it, and a corpus-matched value is not the lever it looks like. @@ -1636,24 +1721,28 @@ looks like. `0.7/0.3` was a documented default, never a searched one. Sweeping `vector_weight` from 0.0 to 1.0 (`--sweep`, reusing the one-time embedding -table) shows it is not merely suboptimal but **strictly dominated**: +table) shows it is not merely suboptimal but **strictly dominated**. First +measured 2026-08-07 (1537a94); the values below are the 2026-09-27 re-run, +which changed the 0.2, 0.4, 0.5, 0.6, 0.7, 0.8 and 1.0 rows in the last digit +or by one or two questions (see the re-run section above): | vector / keyword | Hit@1 | Hit@5 | Hit@10 | MRR | session Hit@5 | |---|---|---|---|---|---| | 0.0 / 1.0 (BM25) | **53.8%** | 75.0% | 81.6% | 0.6320 | 93.6% | | 0.1 / 0.9 | 53.2% | 77.4% | 83.8% | 0.6374 | 95.0% | -| 0.2 / 0.8 | 53.6% | 78.2% | 85.6% | 0.6440 | 95.4% | +| 0.2 / 0.8 | 53.4% | 78.0% | 85.6% | 0.6429 | 95.4% | | 0.3 / 0.7 | 53.2% | 78.8% | 87.2% | **0.6463** | 96.0% | -| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6429 | 96.8% | -| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6234 | **97.4%** | -| 0.6 / 0.4 | 46.6% | 79.8% | 87.4% | 0.6069 | 96.6% | -| 0.7 / 0.3 *(old default)* | 44.4% | 79.2% | 86.0% | 0.5868 | 95.8% | -| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5571 | 95.2% | +| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6430 | 96.8% | +| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6232 | **97.4%** | +| 0.6 / 0.4 | 46.6% | 79.6% | 87.4% | 0.6067 | 96.4% | +| 0.7 / 0.3 *(old default)* | 44.2% | 79.2% | 85.8% | 0.5856 | 95.8% | +| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5574 | 95.2% | | 0.9 / 0.1 | 37.8% | 73.4% | 84.6% | 0.5289 | 94.2% | -| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5027 | 94.2% | +| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5031 | 94.2% | **`0.4/0.6` beats `0.7/0.3` on every metric at both granularities** — Hit@1 -+7.2pp, Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. There is no trade being made; the ++7.4pp, Hit@5 +2.2, Hit@10 +2.0, MRR +0.057 (2026-09-27 figures; +7.2pp, ++2.2, +1.8 and +0.056 as measured on 2026-08-07). There is no trade being made; the old default was simply on the wrong side of the peak. **`0.4/0.6` is the recommended setting**, with `0.3/0.7` preferable if rank-1 precision matters most (it takes the best MRR in the sweep and gives up only 0.6pp of Hit@1 @@ -1714,11 +1803,27 @@ price of the harder corpus, and is the reason oracle-only numbers should not be presented as LongMemEval results. Session-level figures on this variant are degenerate — see below. -With real embeddings the same oracle corpus gives BM25-only 84.2% / vector-only -80.4% / hybrid **85.2%** Hit@5 turn-level — hybrid ahead at Hit@5 and Hit@10 and -behind at Hit@1, matching the full-haystack pattern above. (BM25-only reads 84.2% -here against 84.4% with zero embedding vectors: one question of 500 changes rank, -with MRR identical at 0.6597. On the full haystack the two agree exactly.) +With real embeddings, the same oracle corpus gives these turn-level figures +(2026-09-27, tank, commit 7a8fae0; command and load in "Re-run with real +embeddings" above): + +| Mode | Hit@1 | Hit@5 | Hit@10 | MRR | +|---|---:|---:|---:|---:| +| BM25 only | **52.6%** | 84.4% | 90.4% | 0.6597 | +| Vector only | 39.6% | 80.4% | 91.0% | 0.5605 | +| Hybrid 0.4 / 0.6 (default) | 52.4% | **86.8%** | **92.4%** | **0.6678** | +| Hybrid 0.7 / 0.3 (old default) | 48.4% | 85.2% | 92.2% | 0.6382 | + +Hybrid 0.4/0.6 leads at Hit@5, Hit@10 and MRR and is 0.2pp (one question) +behind BM25 at Hit@1, as on the full haystack. + +Until 2026-09-27 this paragraph gave BM25-only 84.2%, vector-only 80.4% and +hybrid **85.2%** Hit@5. Those were measured on 2026-08-07 (c913cd1), when the +hybrid default was 0.7/0.3; today's 0.7/0.3 row reproduces the 85.2%. The +0.4/0.6 default (29baabb) is what moves hybrid to 86.8%. The earlier BM25-only +84.2% with embeddings, one question below the zero-embedding 84.4%, predates +the index tie-break of 3ed0489; the two now agree, as they always did on the +full haystack. ### Retracted: session-level recall and the MemX comparison diff --git a/ROADMAP.md b/ROADMAP.md index 914b02b..319a27e 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -135,7 +135,7 @@ **Crates:** `clawhdf5-agent`, `clawhdf5-bench` - [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547 -- [x] **8.2** LongMemEval benchmark — 500 queries, session Hit@1=100%, turn Hit@5=84.4% (beats MemX 51.6%), MRR=0.660 +- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results) - [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal - [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion - [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss @@ -168,7 +168,7 @@ Verified against current repo state on 2026-08-05 (see also `docs/superpowers/pl ### Recently closed out (2026-08-05, Tier 3–4 hardening pass) -- [x] Academic benchmark cross-validation — LongMemEval reproduced against MemX on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% vs MemX's 51.6%; recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05) +- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05) - [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer - [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories - [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open