Benchmarks re-measured: LongMemEval with real embeddings, local reads on an idle machine #20

Merged
osobh merged 3 commits from docs/bench-refresh-2026-09-27 into main 2026-09-28 02:38:09 +00:00
4 changed files with 225 additions and 28 deletions
+189 -26
View File
@@ -30,12 +30,15 @@ target: Criterion stretched it where 5 s could not hold the samples it needed
> and in memory), Consolidation Efficiency, Ephemeral Tier, Multi-modal Search,
> the Search and Read harnesses, and the "h5bench-Equivalent I/O Benchmarks"
> and "Independent Validation: tank" sections. What does not yet meet that bar:
> the LongMemEval rows that need real embeddings (not re-run here, except the
> dated float16 comparison), the Consolidation Efficiency 100K cycle row and
> the Consolidation Efficiency 100K cycle row and
> memory-reduction part (the 2026-09-24 run was stopped before it produced
> them), the int8 side of "Quantising the index copy" (not re-run), and the
> i7-12650H and macOS M3 Max rows under Cross-Platform Notes. That is a
> known, tracked documentation gap, not a claim that those numbers are wrong.
> The LongMemEval rows that need real embeddings (vector-only, hybrid, RRF,
> stemmed hybrid, re-ranking, the weight sweep, the oracle variant) were
> re-run on 2026-09-27 on tank; see "Re-run with real embeddings" under
> LongMemEval Results.
>
> **Correctness note (2026-08-06).** Being dated and reproducible is necessary but
> not sufficient — a number can be perfectly reproducible and still measure the
@@ -325,6 +328,19 @@ which of two gold sessions ranks first, out of ~320. Those flips show the
half-precision path was in effect; they do not change a single hit. The f32
run reproduces the published hybrid numbers exactly.
**Re-checked 2026-09-27** (tank, commit 7a8fae0, the same pair of runs on
`longmemeval_s_cleaned.json`, which is the same file; the machine was not
idle, which does not affect recall): the result is the same. The table above
reproduced exactly. Every Hit@k and MRR of the eight modes matched between f32
and float16 at both levels, with two exceptions: RRF's session MRR (0.9253 vs
0.9254) and two per-type session MRRs in the fourth decimal. Three modes
differed by one question in the recency count. That re-run also corrects the
sentence above: two f32 runs on the same day differed by one question in
recency as well, so those flips are run-to-run variation and do not show that
the half-precision path was in effect. `--float16` is what shows that: the
harness prints "Stores use MemoryConfig::float16" and `MemoryConfig::float16`
is set on every store.
### Opening a store (`read_from_disk`)
`HDF5Memory::open` memory-mapped the file, copied the whole mapping into a
@@ -482,7 +498,65 @@ The rows and columns of the uncompressed layouts are within 20% (chunked
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
explain the slower windows.
## Concurrent reads
## Local file speed after range reads
### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank)
`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main`
`7a8fae0` (PRs #18 and #19), each built in its own worktree and run as
separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen
7 7800X3D, 16 threads). **Idle:** every round started with the 1-minute load
average below 2 (1.05–1.98; `target/ab-results2/load.log`). Criterion:
`taskset -c 5 local_metadata_bench --bench --warm-up-time 3
--measurement-time 10`, 3 rounds each. Reads: `concurrent_read --dir
~/.cache/concurrent-read --decode-threads 1 --reps 3`, 3 rounds each.
Median (range) over the rounds.
`local_metadata_bench` (the 400-group v1 fixture):
| function | 8f59b2e | 7a8fae0 | change |
|---|---:|---:|---:|
| `object_header_parse_x401` | 23.86 µs (23.79–23.96) | 24.86 µs (24.69–24.98) | **+4.2%** |
| `snod_parse_all` | 1.842 µs (1.835–1.847) | 1.833 µs (1.833–1.854) | −0.5% |
| `btree_v1_walk` | 344 ns (338–346) | 348 ns (337–356) | +1.1% |
| `facade_list_400_groups` | 8.24 ms (8.04–8.28) | 8.10 ms (8.04–8.13) | −1.7% |
`concurrent_read`, MB/s (64 datasets of 64 MiB `f32`; deflate chunks
256 x 256, level 4):
| layout | mode | threads | 8f59b2e | 7a8fae0 | change |
|---|---|---:|---:|---:|---:|
| deflate | distinct | 1 | 891 (880–892) | 907 (906–908) | +1.7% |
| deflate | distinct | 2 | 1692 (1680–1699) | 1777 (1767–1779) | +5.0% |
| deflate | distinct | 4 | 3102 (3099–3102) | 3348 (3347–3350) | +8.0% |
| deflate | distinct | 8 | 5306 (5168–5345) | 6240 (6226–6248) | +17.6% |
| deflate | distinct | 16 | 6258 (6109–6442) | 8525 (8513–8561) | **+36.2%** |
| deflate | same | 1 | 234 (234–235) | 233 (233–234) | −0.5% |
| deflate | same | 16 | 2443 (2038–2444) | 2468 (2424–2473) | +1.0% |
| contiguous | distinct | 1 | 13477 (12949–13874) | 13302 (13287–13578) | −1.3% |
| contiguous | distinct | 16 | 12501 (12484–12530) | 12566 (12449–12567) | +0.5% |
| contiguous | same | 1 | 29866 (28832–30207) | 29364 (28838–29780) | −1.7% |
| contiguous | same | 16 | 163415 (161097–229146) | 233280 (159377–238440) | (noise) |
What this shows:
- **Local metadata reads are at parity or faster.** Listing the 400-group
file through the facade is 1.7% faster than before M2/M3; the +7–10%
listing regression found while merging #18 is gone.
- **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header;
the base and candidate ranges do not overlap). It is the cost of reading
continuation chunks from a bounded queue (the fix for unbounded reads on
crafted headers) and does not show in the listing. Kept open in
`docs/known-issues.md`.
- **Full reads of deflate data got faster** after #18 (in-place chunk
decoding into the typed output and per-thread scratch buffers): +1.7% on
one thread, +36% at 16.
- Single-thread contiguous hyperslabs are within noise (−1.7%, overlapping
ranges). The multi-thread `contiguous same` rows read one 64 MiB dataset
out of the CPU caches and swing widely between rounds of the same build.
An earlier run the same day at load 2.3–3.3 (two orphaned h5py processes,
since stopped, each using a core) reported that single-thread contiguous
hyperslab row as −5.6%; the idle rerun above does not reproduce it.
### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`)
@@ -1410,7 +1484,67 @@ turn-level row plus its session Hit@1 in the tokenizer table. The run also
produced figures this document does not publish (stemmed session Hit@5,
Hit@10 and MRR, and per-type Hit@5/Hit@10/MRR for both modes), so there was
nothing to compare them with. Rows that need real embeddings (vector-only,
hybrid, RRF, re-ranking, the weight sweep) were not re-run.
hybrid, RRF, re-ranking, the weight sweep) were not re-run then; they were on
2026-09-27 (next section).
### Re-run with real embeddings (2026-09-27, tank)
Every recall row in this section that needs real embeddings was measured
again on 2026-09-27 on tank (AMD Ryzen 7 7800X3D, 16 threads; MiniLM
embeddings on an RTX 5060 Ti, retrieval on the CPU), with the search code of
commit 7a8fae0. Six runs:
```bash
cargo build --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings-cuda
B=target/release/longmemeval_bench W=weights/all-minilm-l6-v2
$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --float16
$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W --sweep
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --sweep
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --rerank-sweep
```
(`longmemeval_s.json`, used by the commands elsewhere in this section, is the
same file as `longmemeval_s_cleaned.json`.)
**The machine was not idle.** The 1-minute load never fell below 2 in a
2-hour wait, because two stray test processes were each holding a core; it
was 2.1–2.7 when each run started and up to 8.4 while the runs were going.
Recall does not depend on load. Latency does, so no latency figure in this
file was updated from these runs.
**What reproduced exactly:** every published full-haystack Hit@1/5/10 and MRR
at turn and session level for BM25, vector-only, hybrid 0.4/0.6, both stemmed
modes, and every re-ranking row; RRF except as below; the float16 table; 4 of
the 11 weight-sweep rows (0.0, 0.1, 0.3 and 0.9); and the oracle vector-only
figure. The headline,
hybrid 0.4/0.6 turn Hit@5 **81.4%**, is unchanged.
**What changed.** Old values are kept here; the tables below show the new ones.
| Figure | Published | 2026-09-27 | Why |
|---|---:|---:|---|
| Oracle, hybrid turn Hit@5 | 85.2% (2026-08-07, c913cd1) | **86.8%** | Weights. 85.2% was measured at 0.7/0.3, the default then. The default has been 0.4/0.6 since 29baabb. Today's oracle sweep gives 85.2% at 0.7/0.3 and 86.8% at 0.4/0.6. |
| Oracle, BM25-only turn Hit@5 with embeddings | 84.2% (2026-08-07) | 84.4% | Now equal to the zero-embedding figure, as it already was on the full haystack. At c913cd1, tied candidates were ordered by `HashMap` iteration. Since 3ed0489 (2026-09-19) they are broken by index. 3ed0489 is the likely cause; it was not bisected. |
| Ablation / sweep, hybrid 0.7/0.3, turn | 44.4% / 79.2% / 86.0% / 0.5868 | 44.2% / 79.2% / 85.8% / 0.5856 | The same tie-breaking change. 29baabb's re-run on 2026-09-19, made after 3ed0489 that same day, already had 44.2% / 85.8% / 0.5856. |
| Ablation, hybrid 0.7/0.3, session | 88.2% / 95.8% / 97.8% / 0.9158 | 88.0% / 95.8% / 97.6% / 0.9146 | Same. |
| Ablation / sweep, vector-only turn MRR | 0.5027 | 0.5031 | Same. The fusion and float16 tables already had 0.5031. |
| Sweep 0.2/0.8, turn | 53.6% / 78.2% / 0.6440 | 53.4% / 78.0% / 0.6429 | Same (the sweep was measured 2026-08-07, 1537a94). |
| Sweep 0.4/0.6, 0.5/0.5, 0.8/0.2 turn MRR | 0.6429 / 0.6234 / 0.5571 | 0.6430 / 0.6232 / 0.5574 | Same. |
| Sweep 0.6/0.4 | Hit@5 79.8%, MRR 0.6069, session Hit@5 96.6% | 79.6%, 0.6067, 96.4% | Same. |
| RRF turn MRR | 0.5967 (2026-09-19, aa92fef) | 0.5969 | Not explained. It may be a later search-path change, such as the int8 index becoming the default in 8b85d93. It may also be the run-to-run variation described next. It was not bisected. |
**Not exactly deterministic.** Between two f32 runs today, the recency count
(the share of `knowledge-update` questions where the newest gold session
ranks first) differed by one question in two modes: hybrid 0.4/0.6 was
144/320 in one run and 145/320 in the other, and re-rank with a 1-day
half-life was 165/319 and 166/319. Each question gets a fresh store, so this
is not state carried between modes. The cause was not found; a candidate is
that the GPU embeddings are not bit-for-bit identical from run to run. Hit@k
and MRR agreed in every run that repeated a mode (hybrid 0.4/0.6 was measured
four times: three f32 runs and one float16 run), but a last-digit change in an
MRR, or a one-question change in recency, is within this variation.
### Full haystack — `longmemeval_s`, n=500 (the number to cite)
@@ -1440,13 +1574,16 @@ the question.
Real 384-d `all-MiniLM-L6-v2` embeddings, 190,015 unique texts encoded once on an
RTX 5060 Ti (~13 min; the same work on the 8-core CPU was still unfinished after
30 minutes, so the GPU path is not a convenience here). Turn-level:
30 minutes, so the GPU path is not a convenience here). First measured
2026-08-07 (c913cd1); the values below are the 2026-09-27 re-run, which moved
the vector-only MRR and the `0.7/0.3` row (see the re-run section above).
Turn-level:
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|------|-------|-------|--------|-----|
| BM25 only (`0.0`/`1.0`) | **53.8%** | 75.0% | 81.6% | **0.6320** |
| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5027 |
| Hybrid (`0.7`/`0.3`) | 44.4% | **79.2%** | **86.0%** | 0.5868 |
| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5031 |
| Hybrid (`0.7`/`0.3`) | 44.2% | **79.2%** | **85.8%** | 0.5856 |
Session-level:
@@ -1454,7 +1591,7 @@ Session-level:
|------|-------|-------|--------|-----|
| BM25 only | 86.2% | 93.6% | 96.6% | 0.8948 |
| Vector only | 85.4% | 94.2% | 96.6% | 0.8901 |
| Hybrid | **88.2%** | **95.8%** | **97.8%** | **0.9158** |
| Hybrid | **88.0%** | **95.8%** | **97.6%** | **0.9146** |
### Fusion method — weighted vs. RRF, full haystack, n=500
@@ -1468,7 +1605,7 @@ takes a `Fusion`, and both run over the same HNSW + BM25 candidates:
| BM25 only | **53.8%** | 75.0% | 81.6% | 0.6320 | 86.2% | 0.8948 |
| Vector only | 36.0% | 71.8% | 81.6% | 0.5031 | 85.4% | 0.8901 |
| **Weighted 0.4 / 0.6** | 51.6% | **81.4%** | **87.8%** | **0.6430** | **91.0%** | **0.9347** |
| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5967 | 89.6% | 0.9253 |
| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5969 | 89.6% | 0.9253 |
**RRF loses to the tuned weighted sum** — 6.6pp of turn Hit@1 and 0.046 of MRR
— and lands almost exactly where the old `0.7/0.3` weighting did (44.2% /
@@ -1533,6 +1670,12 @@ one (see `newest_gold_first`); ~45% is chance.
| + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% |
| + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% |
Re-run on 2026-09-27 (`--rerank-sweep`, tank, commit 7a8fae0): every Hit@k and
MRR above reproduced exactly. The recency column came out 45.3%, 87.5%,
52.0%, 52.2%, 51.7% and 50.5%. Each of those is within one question of the
value in the table, which is the run-to-run variation described under
"Re-run with real embeddings" above, so the table was left as it was.
**The pre-fix row is the finding.** Ordering candidates by recency alone costs
40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in
the pool rather than the ones that answer the question. It does ace the recency
@@ -1546,9 +1689,9 @@ cannot reach the 87.5% the degenerate ordering gets. Those two rows are the
ends of a trade-off, and the default sits deliberately near the relevance end.
**Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency
moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because
the temporal term is capped by its weight (0.3) while relevance differences
between candidates are larger. The 24-hour default is kept; there is no
moves 1.4pp (1.7pp in the 2026-09-27 re-run) and MRR 0.003 — inside the
noise of a 500-question run — because the temporal term is capped by its
weight (0.3) while relevance differences between candidates are larger. The 24-hour default is kept; there is no
measured reason to change it, and a corpus-matched value is not the lever it
looks like.
@@ -1556,24 +1699,28 @@ looks like.
`0.7/0.3` was a documented default, never a searched one. Sweeping
`vector_weight` from 0.0 to 1.0 (`--sweep`, reusing the one-time embedding
table) shows it is not merely suboptimal but **strictly dominated**:
table) shows it is not merely suboptimal but **strictly dominated**. First
measured 2026-08-07 (1537a94); the values below are the 2026-09-27 re-run,
which changed the 0.2, 0.4, 0.5, 0.6, 0.7, 0.8 and 1.0 rows in the last digit
or by one or two questions (see the re-run section above):
| vector / keyword | Hit@1 | Hit@5 | Hit@10 | MRR | session Hit@5 |
|---|---|---|---|---|---|
| 0.0 / 1.0 (BM25) | **53.8%** | 75.0% | 81.6% | 0.6320 | 93.6% |
| 0.1 / 0.9 | 53.2% | 77.4% | 83.8% | 0.6374 | 95.0% |
| 0.2 / 0.8 | 53.6% | 78.2% | 85.6% | 0.6440 | 95.4% |
| 0.2 / 0.8 | 53.4% | 78.0% | 85.6% | 0.6429 | 95.4% |
| 0.3 / 0.7 | 53.2% | 78.8% | 87.2% | **0.6463** | 96.0% |
| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6429 | 96.8% |
| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6234 | **97.4%** |
| 0.6 / 0.4 | 46.6% | 79.8% | 87.4% | 0.6069 | 96.6% |
| 0.7 / 0.3 *(old default)* | 44.4% | 79.2% | 86.0% | 0.5868 | 95.8% |
| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5571 | 95.2% |
| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6430 | 96.8% |
| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6232 | **97.4%** |
| 0.6 / 0.4 | 46.6% | 79.6% | 87.4% | 0.6067 | 96.4% |
| 0.7 / 0.3 *(old default)* | 44.2% | 79.2% | 85.8% | 0.5856 | 95.8% |
| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5574 | 95.2% |
| 0.9 / 0.1 | 37.8% | 73.4% | 84.6% | 0.5289 | 94.2% |
| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5027 | 94.2% |
| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5031 | 94.2% |
**`0.4/0.6` beats `0.7/0.3` on every metric at both granularities** — Hit@1
+7.2pp, Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. There is no trade being made; the
+7.4pp, Hit@5 +2.2, Hit@10 +2.0, MRR +0.057 (2026-09-27 figures; +7.2pp,
+2.2, +1.8 and +0.056 as measured on 2026-08-07). There is no trade being made; the
old default was simply on the wrong side of the peak. **`0.4/0.6` is the
recommended setting**, with `0.3/0.7` preferable if rank-1 precision matters
most (it takes the best MRR in the sweep and gives up only 0.6pp of Hit@1
@@ -1634,11 +1781,27 @@ price of the harder corpus, and is the reason oracle-only numbers should not be
presented as LongMemEval results. Session-level figures on this variant are
degenerate — see below.
With real embeddings the same oracle corpus gives BM25-only 84.2% / vector-only
80.4% / hybrid **85.2%** Hit@5 turn-level — hybrid ahead at Hit@5 and Hit@10 and
behind at Hit@1, matching the full-haystack pattern above. (BM25-only reads 84.2%
here against 84.4% with zero embedding vectors: one question of 500 changes rank,
with MRR identical at 0.6597. On the full haystack the two agree exactly.)
With real embeddings, the same oracle corpus gives these turn-level figures
(2026-09-27, tank, commit 7a8fae0; command and load in "Re-run with real
embeddings" above):
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|---|---:|---:|---:|---:|
| BM25 only | **52.6%** | 84.4% | 90.4% | 0.6597 |
| Vector only | 39.6% | 80.4% | 91.0% | 0.5605 |
| Hybrid 0.4 / 0.6 (default) | 52.4% | **86.8%** | **92.4%** | **0.6678** |
| Hybrid 0.7 / 0.3 (old default) | 48.4% | 85.2% | 92.2% | 0.6382 |
Hybrid 0.4/0.6 leads at Hit@5, Hit@10 and MRR and is 0.2pp (one question)
behind BM25 at Hit@1, as on the full haystack.
Until 2026-09-27 this paragraph gave BM25-only 84.2%, vector-only 80.4% and
hybrid **85.2%** Hit@5. Those were measured on 2026-08-07 (c913cd1), when the
hybrid default was 0.7/0.3; today's 0.7/0.3 row reproduces the 85.2%. The
0.4/0.6 default (29baabb) is what moves hybrid to 86.8%. The earlier BM25-only
84.2% with embeddings, one question below the zero-embedding 84.4%, predates
the index tie-break of 3ed0489; the two now agree, as they always did on the
full haystack.
### Retracted: session-level recall and the MemX comparison
+16
View File
@@ -477,6 +477,19 @@ Design: `docs/design/swmr.md`.
`main` and it calls only format-crate code; its earlier +5–10% (323–355
vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured
357–374 ns against `main`'s 355–371 in the same rounds).
**Rechecked 2026-09-27 on an idle tank** (load below 2 at every round;
`8f59b2e` vs `main` `7a8fae0` as separate binaries, pinned to one CPU, 3
alternating rounds, median (range); `BENCHMARKS.md`, "Local metadata and
data reads after range-read M2/M3"): facade listing 8.24 (8.04–8.28) →
8.10 (8.04–8.13) ms, **−1.7%**, so the listing regression is gone;
symbol-table nodes −0.5%, B-tree walk +1.1%; `ObjectHeader::parse` ×401
23.86 (23.79–23.96) → 24.86 (24.69–24.98) µs, **+4.2%, open**
(`docs/known-issues.md`; about 2.5 ns per header, not visible in the
listing). Data reads (`concurrent_read --decode-threads 1`): full reads
of deflate data +1.7% (1 thread) to +36% (16 threads), contiguous reads
and deflate hyperslabs within ±2%. (A run earlier that day under load
2.3–3.3 put single-thread contiguous hyperslabs at −5.6%; the idle rerun
does not reproduce it.)
- Tests (2026-09-26, tank):
- `clawhdf5-format/tests/storage_equivalence.rs` now also reads every
dataset — whole, fill-aware, through a chunk cache (twice) and the
@@ -631,6 +644,9 @@ Design: `docs/design/swmr.md`.
instead of being inlined into the generic parser. Provisional (tank load
3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs
24.85–25.15 µs on `main`.
Rechecked 2026-09-27 on an idle tank (3 alternating rounds, load below
2): 24.69–24.98 µs against 23.79–23.96 µs for `8f59b2e`, median +4.2%; a
residual cost remains (see the M2 speed note and `docs/known-issues.md`).
- Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py
`earliest`/`v110`/`latest` and clawhdf5-written files; structure
comparisons with libhdf5 for version-2 B-trees, shrink on every index,
+2 -2
View File
@@ -135,7 +135,7 @@
**Crates:** `clawhdf5-agent`, `clawhdf5-bench`
- [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547
- [x] **8.2** LongMemEval benchmark — 500 queries, session Hit@1=100%, turn Hit@5=84.4% (beats MemX 51.6%), MRR=0.660
- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results)
- [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal
- [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion
- [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss
@@ -168,7 +168,7 @@ Verified against current repo state on 2026-08-05 (see also `docs/superpowers/pl
### Recently closed out (2026-08-05, Tier 3–4 hardening pass)
- [x] Academic benchmark cross-validation — LongMemEval reproduced against MemX on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% vs MemX's 51.6%; recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
- [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer
- [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories
- [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open
+18
View File
@@ -7,6 +7,24 @@ deleting it.
---
## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27)
**Status:** open (speed only; values are correct). Measured on an idle tank
(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`)
against `main` `7a8fae0`, separate binaries alternating, 3 rounds
(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"):
`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges
23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read
continuation chunks from a bounded queue since `a69c5be` (bounding what a
crafted header can make it read); `4313917` removed its per-header
allocations but not all of the cost. Listing the 400-group file through the
facade, which parses the same headers, is 1.7% faster, so no user-visible
path is slower.
An earlier run the same day under load also listed single-thread contiguous
hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with
overlapping ranges (noise), so that item is withdrawn.
## Files a SWMR writer had open could not be read past a stale end of file
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any