docs: LongMemEval with real embeddings re-run on 2026-09-27

Oracle and full longmemeval_s haystack, f32 and --float16, plus --sweep
(both corpora) and --rerank-sweep, on tank at 7a8fae0 (MiniLM on the RTX
5060 Ti). Not idle: two orphaned h5py test processes kept the 1-minute load
at 2.1-2.7 (up to 8.4 during runs), so no latency figure was updated.

The headline (hybrid 0.4/0.6 turn Hit@5 81.4%) and the float16 claim
reproduce. Oracle hybrid is 86.8%, not 85.2%: the old figure was at the
0.7/0.3 default of the time (today's 0.7/0.3 gives 85.2%). The ablation's
0.7/0.3 row, vector-only MRR and seven sweep rows move in the last digit or
by one or two questions, most likely from the index tie-break of 3ed0489.
RRF turn MRR 0.5967 -> 0.5969 is not explained. Recency counts vary by one
question between identical runs, so the float16 "flips" note is corrected.
ROADMAP 8.2 no longer repeats the retracted session-level and MemX claims.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-27 19:25:36 -05:00
co-authored by Claude Opus 5.5
parent d239292655
commit b39a705e77
2 changed files with 132 additions and 27 deletions
+130 -25
View File
@@ -30,12 +30,15 @@ target: Criterion stretched it where 5 s could not hold the samples it needed
> and in memory), Consolidation Efficiency, Ephemeral Tier, Multi-modal Search,
> the Search and Read harnesses, and the "h5bench-Equivalent I/O Benchmarks"
> and "Independent Validation: tank" sections. What does not yet meet that bar:
> the LongMemEval rows that need real embeddings (not re-run here, except the
> dated float16 comparison), the Consolidation Efficiency 100K cycle row and
> the Consolidation Efficiency 100K cycle row and
> memory-reduction part (the 2026-09-24 run was stopped before it produced
> them), the int8 side of "Quantising the index copy" (not re-run), and the
> i7-12650H and macOS M3 Max rows under Cross-Platform Notes. That is a
> known, tracked documentation gap, not a claim that those numbers are wrong.
> The LongMemEval rows that need real embeddings (vector-only, hybrid, RRF,
> stemmed hybrid, re-ranking, the weight sweep, the oracle variant) were
> re-run on 2026-09-27 on tank; see "Re-run with real embeddings" under
> LongMemEval Results.
>
> **Correctness note (2026-08-06).** Being dated and reproducible is necessary but
> not sufficient — a number can be perfectly reproducible and still measure the
@@ -325,6 +328,19 @@ which of two gold sessions ranks first, out of ~320. Those flips show the
half-precision path was in effect; they do not change a single hit. The f32
run reproduces the published hybrid numbers exactly.
**Re-checked 2026-09-27** (tank, commit 7a8fae0, the same pair of runs on
`longmemeval_s_cleaned.json`, which is the same file; the machine was not
idle, which does not affect recall): the result is the same. The table above
reproduced exactly. Every Hit@k and MRR of the eight modes matched between f32
and float16 at both levels, with two exceptions: RRF's session MRR (0.9253 vs
0.9254) and two per-type session MRRs in the fourth decimal. Three modes
differed by one question in the recency count. That re-run also corrects the
sentence above: two f32 runs on the same day differed by one question in
recency as well, so those flips are run-to-run variation and do not show that
the half-precision path was in effect. `--float16` is what shows that: the
harness prints "Stores use MemoryConfig::float16" and `MemoryConfig::float16`
is set on every store.
### Opening a store (`read_from_disk`)
`HDF5Memory::open` memory-mapped the file, copied the whole mapping into a
@@ -1490,7 +1506,67 @@ turn-level row plus its session Hit@1 in the tokenizer table. The run also
produced figures this document does not publish (stemmed session Hit@5,
Hit@10 and MRR, and per-type Hit@5/Hit@10/MRR for both modes), so there was
nothing to compare them with. Rows that need real embeddings (vector-only,
hybrid, RRF, re-ranking, the weight sweep) were not re-run.
hybrid, RRF, re-ranking, the weight sweep) were not re-run then; they were on
2026-09-27 (next section).
### Re-run with real embeddings (2026-09-27, tank)
Every recall row in this section that needs real embeddings was measured
again on 2026-09-27 on tank (AMD Ryzen 7 7800X3D, 16 threads; MiniLM
embeddings on an RTX 5060 Ti, retrieval on the CPU), with the search code of
commit 7a8fae0. Six runs:
```bash
cargo build --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings-cuda
B=target/release/longmemeval_bench W=weights/all-minilm-l6-v2
$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --float16
$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W --sweep
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --sweep
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --rerank-sweep
```
(`longmemeval_s.json`, used by the commands elsewhere in this section, is the
same file as `longmemeval_s_cleaned.json`.)
**The machine was not idle.** The 1-minute load never fell below 2 in a
2-hour wait, because two stray test processes were each holding a core; it
was 2.1–2.7 when each run started and up to 8.4 while the runs were going.
Recall does not depend on load. Latency does, so no latency figure in this
file was updated from these runs.
**What reproduced exactly:** every published full-haystack Hit@1/5/10 and MRR
at turn and session level for BM25, vector-only, hybrid 0.4/0.6, both stemmed
modes, and every re-ranking row; RRF except as below; the float16 table; 4 of
the 11 weight-sweep rows (0.0, 0.1, 0.3 and 0.9); and the oracle vector-only
figure. The headline,
hybrid 0.4/0.6 turn Hit@5 **81.4%**, is unchanged.
**What changed.** Old values are kept here; the tables below show the new ones.
| Figure | Published | 2026-09-27 | Why |
|---|---:|---:|---|
| Oracle, hybrid turn Hit@5 | 85.2% (2026-08-07, c913cd1) | **86.8%** | Weights. 85.2% was measured at 0.7/0.3, the default then. The default has been 0.4/0.6 since 29baabb. Today's oracle sweep gives 85.2% at 0.7/0.3 and 86.8% at 0.4/0.6. |
| Oracle, BM25-only turn Hit@5 with embeddings | 84.2% (2026-08-07) | 84.4% | Now equal to the zero-embedding figure, as it already was on the full haystack. At c913cd1, tied candidates were ordered by `HashMap` iteration. Since 3ed0489 (2026-09-19) they are broken by index. 3ed0489 is the likely cause; it was not bisected. |
| Ablation / sweep, hybrid 0.7/0.3, turn | 44.4% / 79.2% / 86.0% / 0.5868 | 44.2% / 79.2% / 85.8% / 0.5856 | The same tie-breaking change. 29baabb's re-run on 2026-09-19, made after 3ed0489 that same day, already had 44.2% / 85.8% / 0.5856. |
| Ablation, hybrid 0.7/0.3, session | 88.2% / 95.8% / 97.8% / 0.9158 | 88.0% / 95.8% / 97.6% / 0.9146 | Same. |
| Ablation / sweep, vector-only turn MRR | 0.5027 | 0.5031 | Same. The fusion and float16 tables already had 0.5031. |
| Sweep 0.2/0.8, turn | 53.6% / 78.2% / 0.6440 | 53.4% / 78.0% / 0.6429 | Same (the sweep was measured 2026-08-07, 1537a94). |
| Sweep 0.4/0.6, 0.5/0.5, 0.8/0.2 turn MRR | 0.6429 / 0.6234 / 0.5571 | 0.6430 / 0.6232 / 0.5574 | Same. |
| Sweep 0.6/0.4 | Hit@5 79.8%, MRR 0.6069, session Hit@5 96.6% | 79.6%, 0.6067, 96.4% | Same. |
| RRF turn MRR | 0.5967 (2026-09-19, aa92fef) | 0.5969 | Not explained. It may be a later search-path change, such as the int8 index becoming the default in 8b85d93. It may also be the run-to-run variation described next. It was not bisected. |
**Not exactly deterministic.** Between two f32 runs today, the recency count
(the share of `knowledge-update` questions where the newest gold session
ranks first) differed by one question in two modes: hybrid 0.4/0.6 was
144/320 in one run and 145/320 in the other, and re-rank with a 1-day
half-life was 165/319 and 166/319. Each question gets a fresh store, so this
is not state carried between modes. The cause was not found; a candidate is
that the GPU embeddings are not bit-for-bit identical from run to run. Hit@k
and MRR agreed in every run that repeated a mode (hybrid 0.4/0.6 was measured
four times: three f32 runs and one float16 run), but a last-digit change in an
MRR, or a one-question change in recency, is within this variation.
### Full haystack — `longmemeval_s`, n=500 (the number to cite)
@@ -1520,13 +1596,16 @@ the question.
Real 384-d `all-MiniLM-L6-v2` embeddings, 190,015 unique texts encoded once on an
RTX 5060 Ti (~13 min; the same work on the 8-core CPU was still unfinished after
30 minutes, so the GPU path is not a convenience here). Turn-level:
30 minutes, so the GPU path is not a convenience here). First measured
2026-08-07 (c913cd1); the values below are the 2026-09-27 re-run, which moved
the vector-only MRR and the `0.7/0.3` row (see the re-run section above).
Turn-level:
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|------|-------|-------|--------|-----|
| BM25 only (`0.0`/`1.0`) | **53.8%** | 75.0% | 81.6% | **0.6320** |
| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5027 |
| Hybrid (`0.7`/`0.3`) | 44.4% | **79.2%** | **86.0%** | 0.5868 |
| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5031 |
| Hybrid (`0.7`/`0.3`) | 44.2% | **79.2%** | **85.8%** | 0.5856 |
Session-level:
@@ -1534,7 +1613,7 @@ Session-level:
|------|-------|-------|--------|-----|
| BM25 only | 86.2% | 93.6% | 96.6% | 0.8948 |
| Vector only | 85.4% | 94.2% | 96.6% | 0.8901 |
| Hybrid | **88.2%** | **95.8%** | **97.8%** | **0.9158** |
| Hybrid | **88.0%** | **95.8%** | **97.6%** | **0.9146** |
### Fusion method — weighted vs. RRF, full haystack, n=500
@@ -1548,7 +1627,7 @@ takes a `Fusion`, and both run over the same HNSW + BM25 candidates:
| BM25 only | **53.8%** | 75.0% | 81.6% | 0.6320 | 86.2% | 0.8948 |
| Vector only | 36.0% | 71.8% | 81.6% | 0.5031 | 85.4% | 0.8901 |
| **Weighted 0.4 / 0.6** | 51.6% | **81.4%** | **87.8%** | **0.6430** | **91.0%** | **0.9347** |
| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5967 | 89.6% | 0.9253 |
| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5969 | 89.6% | 0.9253 |
**RRF loses to the tuned weighted sum** — 6.6pp of turn Hit@1 and 0.046 of MRR
— and lands almost exactly where the old `0.7/0.3` weighting did (44.2% /
@@ -1613,6 +1692,12 @@ one (see `newest_gold_first`); ~45% is chance.
| + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% |
| + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% |
Re-run on 2026-09-27 (`--rerank-sweep`, tank, commit 7a8fae0): every Hit@k and
MRR above reproduced exactly. The recency column came out 45.3%, 87.5%,
52.0%, 52.2%, 51.7% and 50.5%. Each of those is within one question of the
value in the table, which is the run-to-run variation described under
"Re-run with real embeddings" above, so the table was left as it was.
**The pre-fix row is the finding.** Ordering candidates by recency alone costs
40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in
the pool rather than the ones that answer the question. It does ace the recency
@@ -1626,9 +1711,9 @@ cannot reach the 87.5% the degenerate ordering gets. Those two rows are the
ends of a trade-off, and the default sits deliberately near the relevance end.
**Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency
moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because
the temporal term is capped by its weight (0.3) while relevance differences
between candidates are larger. The 24-hour default is kept; there is no
moves 1.4pp (1.7pp in the 2026-09-27 re-run) and MRR 0.003 — inside the
noise of a 500-question run — because the temporal term is capped by its
weight (0.3) while relevance differences between candidates are larger. The 24-hour default is kept; there is no
measured reason to change it, and a corpus-matched value is not the lever it
looks like.
@@ -1636,24 +1721,28 @@ looks like.
`0.7/0.3` was a documented default, never a searched one. Sweeping
`vector_weight` from 0.0 to 1.0 (`--sweep`, reusing the one-time embedding
table) shows it is not merely suboptimal but **strictly dominated**:
table) shows it is not merely suboptimal but **strictly dominated**. First
measured 2026-08-07 (1537a94); the values below are the 2026-09-27 re-run,
which changed the 0.2, 0.4, 0.5, 0.6, 0.7, 0.8 and 1.0 rows in the last digit
or by one or two questions (see the re-run section above):
| vector / keyword | Hit@1 | Hit@5 | Hit@10 | MRR | session Hit@5 |
|---|---|---|---|---|---|
| 0.0 / 1.0 (BM25) | **53.8%** | 75.0% | 81.6% | 0.6320 | 93.6% |
| 0.1 / 0.9 | 53.2% | 77.4% | 83.8% | 0.6374 | 95.0% |
| 0.2 / 0.8 | 53.6% | 78.2% | 85.6% | 0.6440 | 95.4% |
| 0.2 / 0.8 | 53.4% | 78.0% | 85.6% | 0.6429 | 95.4% |
| 0.3 / 0.7 | 53.2% | 78.8% | 87.2% | **0.6463** | 96.0% |
| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6429 | 96.8% |
| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6234 | **97.4%** |
| 0.6 / 0.4 | 46.6% | 79.8% | 87.4% | 0.6069 | 96.6% |
| 0.7 / 0.3 *(old default)* | 44.4% | 79.2% | 86.0% | 0.5868 | 95.8% |
| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5571 | 95.2% |
| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6430 | 96.8% |
| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6232 | **97.4%** |
| 0.6 / 0.4 | 46.6% | 79.6% | 87.4% | 0.6067 | 96.4% |
| 0.7 / 0.3 *(old default)* | 44.2% | 79.2% | 85.8% | 0.5856 | 95.8% |
| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5574 | 95.2% |
| 0.9 / 0.1 | 37.8% | 73.4% | 84.6% | 0.5289 | 94.2% |
| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5027 | 94.2% |
| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5031 | 94.2% |
**`0.4/0.6` beats `0.7/0.3` on every metric at both granularities** — Hit@1
+7.2pp, Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. There is no trade being made; the
+7.4pp, Hit@5 +2.2, Hit@10 +2.0, MRR +0.057 (2026-09-27 figures; +7.2pp,
+2.2, +1.8 and +0.056 as measured on 2026-08-07). There is no trade being made; the
old default was simply on the wrong side of the peak. **`0.4/0.6` is the
recommended setting**, with `0.3/0.7` preferable if rank-1 precision matters
most (it takes the best MRR in the sweep and gives up only 0.6pp of Hit@1
@@ -1714,11 +1803,27 @@ price of the harder corpus, and is the reason oracle-only numbers should not be
presented as LongMemEval results. Session-level figures on this variant are
degenerate — see below.
With real embeddings the same oracle corpus gives BM25-only 84.2% / vector-only
80.4% / hybrid **85.2%** Hit@5 turn-level — hybrid ahead at Hit@5 and Hit@10 and
behind at Hit@1, matching the full-haystack pattern above. (BM25-only reads 84.2%
here against 84.4% with zero embedding vectors: one question of 500 changes rank,
with MRR identical at 0.6597. On the full haystack the two agree exactly.)
With real embeddings, the same oracle corpus gives these turn-level figures
(2026-09-27, tank, commit 7a8fae0; command and load in "Re-run with real
embeddings" above):
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|---|---:|---:|---:|---:|
| BM25 only | **52.6%** | 84.4% | 90.4% | 0.6597 |
| Vector only | 39.6% | 80.4% | 91.0% | 0.5605 |
| Hybrid 0.4 / 0.6 (default) | 52.4% | **86.8%** | **92.4%** | **0.6678** |
| Hybrid 0.7 / 0.3 (old default) | 48.4% | 85.2% | 92.2% | 0.6382 |
Hybrid 0.4/0.6 leads at Hit@5, Hit@10 and MRR and is 0.2pp (one question)
behind BM25 at Hit@1, as on the full haystack.
Until 2026-09-27 this paragraph gave BM25-only 84.2%, vector-only 80.4% and
hybrid **85.2%** Hit@5. Those were measured on 2026-08-07 (c913cd1), when the
hybrid default was 0.7/0.3; today's 0.7/0.3 row reproduces the 85.2%. The
0.4/0.6 default (29baabb) is what moves hybrid to 86.8%. The earlier BM25-only
84.2% with embeddings, one question below the zero-embedding 84.4%, predates
the index tie-break of 3ed0489; the two now agree, as they always did on the
full haystack.
### Retracted: session-level recall and the MemX comparison
+2 -2
View File
@@ -135,7 +135,7 @@
**Crates:** `clawhdf5-agent`, `clawhdf5-bench`
- [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547
- [x] **8.2** LongMemEval benchmark — 500 queries, session Hit@1=100%, turn Hit@5=84.4% (beats MemX 51.6%), MRR=0.660
- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results)
- [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal
- [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion
- [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss
@@ -168,7 +168,7 @@ Verified against current repo state on 2026-08-05 (see also `docs/superpowers/pl
### Recently closed out (2026-08-05, Tier 3–4 hardening pass)
- [x] Academic benchmark cross-validation — LongMemEval reproduced against MemX on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% vs MemX's 51.6%; recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
- [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer
- [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories
- [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open