Benchmarks re-measured: LongMemEval with real embeddings, local reads on an idle machine #20
+189
-26
@@ -30,12 +30,15 @@ target: Criterion stretched it where 5 s could not hold the samples it needed
|
||||
> and in memory), Consolidation Efficiency, Ephemeral Tier, Multi-modal Search,
|
||||
> the Search and Read harnesses, and the "h5bench-Equivalent I/O Benchmarks"
|
||||
> and "Independent Validation: tank" sections. What does not yet meet that bar:
|
||||
> the LongMemEval rows that need real embeddings (not re-run here, except the
|
||||
> dated float16 comparison), the Consolidation Efficiency 100K cycle row and
|
||||
> the Consolidation Efficiency 100K cycle row and
|
||||
> memory-reduction part (the 2026-09-24 run was stopped before it produced
|
||||
> them), the int8 side of "Quantising the index copy" (not re-run), and the
|
||||
> i7-12650H and macOS M3 Max rows under Cross-Platform Notes. That is a
|
||||
> known, tracked documentation gap, not a claim that those numbers are wrong.
|
||||
> The LongMemEval rows that need real embeddings (vector-only, hybrid, RRF,
|
||||
> stemmed hybrid, re-ranking, the weight sweep, the oracle variant) were
|
||||
> re-run on 2026-09-27 on tank; see "Re-run with real embeddings" under
|
||||
> LongMemEval Results.
|
||||
>
|
||||
> **Correctness note (2026-08-06).** Being dated and reproducible is necessary but
|
||||
> not sufficient — a number can be perfectly reproducible and still measure the
|
||||
@@ -325,6 +328,19 @@ which of two gold sessions ranks first, out of ~320. Those flips show the
|
||||
half-precision path was in effect; they do not change a single hit. The f32
|
||||
run reproduces the published hybrid numbers exactly.
|
||||
|
||||
**Re-checked 2026-09-27** (tank, commit 7a8fae0, the same pair of runs on
|
||||
`longmemeval_s_cleaned.json`, which is the same file; the machine was not
|
||||
idle, which does not affect recall): the result is the same. The table above
|
||||
reproduced exactly. Every Hit@k and MRR of the eight modes matched between f32
|
||||
and float16 at both levels, with two exceptions: RRF's session MRR (0.9253 vs
|
||||
0.9254) and two per-type session MRRs in the fourth decimal. Three modes
|
||||
differed by one question in the recency count. That re-run also corrects the
|
||||
sentence above: two f32 runs on the same day differed by one question in
|
||||
recency as well, so those flips are run-to-run variation and do not show that
|
||||
the half-precision path was in effect. `--float16` is what shows that: the
|
||||
harness prints "Stores use MemoryConfig::float16" and `MemoryConfig::float16`
|
||||
is set on every store.
|
||||
|
||||
### Opening a store (`read_from_disk`)
|
||||
|
||||
`HDF5Memory::open` memory-mapped the file, copied the whole mapping into a
|
||||
@@ -482,7 +498,65 @@ The rows and columns of the uncompressed layouts are within 20% (chunked
|
||||
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
|
||||
explain the slower windows.
|
||||
|
||||
## Concurrent reads
|
||||
## Local file speed after range reads
|
||||
|
||||
### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank)
|
||||
|
||||
`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main`
|
||||
`7a8fae0` (PRs #18 and #19), each built in its own worktree and run as
|
||||
separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen
|
||||
7 7800X3D, 16 threads). **Idle:** every round started with the 1-minute load
|
||||
average below 2 (1.05–1.98; `target/ab-results2/load.log`). Criterion:
|
||||
`taskset -c 5 local_metadata_bench --bench --warm-up-time 3
|
||||
--measurement-time 10`, 3 rounds each. Reads: `concurrent_read --dir
|
||||
~/.cache/concurrent-read --decode-threads 1 --reps 3`, 3 rounds each.
|
||||
Median (range) over the rounds.
|
||||
|
||||
`local_metadata_bench` (the 400-group v1 fixture):
|
||||
|
||||
| function | 8f59b2e | 7a8fae0 | change |
|
||||
|---|---:|---:|---:|
|
||||
| `object_header_parse_x401` | 23.86 µs (23.79–23.96) | 24.86 µs (24.69–24.98) | **+4.2%** |
|
||||
| `snod_parse_all` | 1.842 µs (1.835–1.847) | 1.833 µs (1.833–1.854) | −0.5% |
|
||||
| `btree_v1_walk` | 344 ns (338–346) | 348 ns (337–356) | +1.1% |
|
||||
| `facade_list_400_groups` | 8.24 ms (8.04–8.28) | 8.10 ms (8.04–8.13) | −1.7% |
|
||||
|
||||
`concurrent_read`, MB/s (64 datasets of 64 MiB `f32`; deflate chunks
|
||||
256 x 256, level 4):
|
||||
|
||||
| layout | mode | threads | 8f59b2e | 7a8fae0 | change |
|
||||
|---|---|---:|---:|---:|---:|
|
||||
| deflate | distinct | 1 | 891 (880–892) | 907 (906–908) | +1.7% |
|
||||
| deflate | distinct | 2 | 1692 (1680–1699) | 1777 (1767–1779) | +5.0% |
|
||||
| deflate | distinct | 4 | 3102 (3099–3102) | 3348 (3347–3350) | +8.0% |
|
||||
| deflate | distinct | 8 | 5306 (5168–5345) | 6240 (6226–6248) | +17.6% |
|
||||
| deflate | distinct | 16 | 6258 (6109–6442) | 8525 (8513–8561) | **+36.2%** |
|
||||
| deflate | same | 1 | 234 (234–235) | 233 (233–234) | −0.5% |
|
||||
| deflate | same | 16 | 2443 (2038–2444) | 2468 (2424–2473) | +1.0% |
|
||||
| contiguous | distinct | 1 | 13477 (12949–13874) | 13302 (13287–13578) | −1.3% |
|
||||
| contiguous | distinct | 16 | 12501 (12484–12530) | 12566 (12449–12567) | +0.5% |
|
||||
| contiguous | same | 1 | 29866 (28832–30207) | 29364 (28838–29780) | −1.7% |
|
||||
| contiguous | same | 16 | 163415 (161097–229146) | 233280 (159377–238440) | (noise) |
|
||||
|
||||
What this shows:
|
||||
- **Local metadata reads are at parity or faster.** Listing the 400-group
|
||||
file through the facade is 1.7% faster than before M2/M3; the +7–10%
|
||||
listing regression found while merging #18 is gone.
|
||||
- **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header;
|
||||
the base and candidate ranges do not overlap). It is the cost of reading
|
||||
continuation chunks from a bounded queue (the fix for unbounded reads on
|
||||
crafted headers) and does not show in the listing. Kept open in
|
||||
`docs/known-issues.md`.
|
||||
- **Full reads of deflate data got faster** after #18 (in-place chunk
|
||||
decoding into the typed output and per-thread scratch buffers): +1.7% on
|
||||
one thread, +36% at 16.
|
||||
- Single-thread contiguous hyperslabs are within noise (−1.7%, overlapping
|
||||
ranges). The multi-thread `contiguous same` rows read one 64 MiB dataset
|
||||
out of the CPU caches and swing widely between rounds of the same build.
|
||||
|
||||
An earlier run the same day at load 2.3–3.3 (two orphaned h5py processes,
|
||||
since stopped, each using a core) reported that single-thread contiguous
|
||||
hyperslab row as −5.6%; the idle rerun above does not reproduce it.
|
||||
|
||||
### Results after in-place chunk decoding (2026-09-26, tank, `c5334b1`)
|
||||
|
||||
@@ -1410,7 +1484,67 @@ turn-level row plus its session Hit@1 in the tokenizer table. The run also
|
||||
produced figures this document does not publish (stemmed session Hit@5,
|
||||
Hit@10 and MRR, and per-type Hit@5/Hit@10/MRR for both modes), so there was
|
||||
nothing to compare them with. Rows that need real embeddings (vector-only,
|
||||
hybrid, RRF, re-ranking, the weight sweep) were not re-run.
|
||||
hybrid, RRF, re-ranking, the weight sweep) were not re-run then; they were on
|
||||
2026-09-27 (next section).
|
||||
|
||||
### Re-run with real embeddings (2026-09-27, tank)
|
||||
|
||||
Every recall row in this section that needs real embeddings was measured
|
||||
again on 2026-09-27 on tank (AMD Ryzen 7 7800X3D, 16 threads; MiniLM
|
||||
embeddings on an RTX 5060 Ti, retrieval on the CPU), with the search code of
|
||||
commit 7a8fae0. Six runs:
|
||||
|
||||
```bash
|
||||
cargo build --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings-cuda
|
||||
B=target/release/longmemeval_bench W=weights/all-minilm-l6-v2
|
||||
$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W
|
||||
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W
|
||||
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --float16
|
||||
$B benchmarks/longmemeval/longmemeval_oracle.json --embeddings $W --sweep
|
||||
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --sweep
|
||||
$B benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings $W --rerank-sweep
|
||||
```
|
||||
|
||||
(`longmemeval_s.json`, used by the commands elsewhere in this section, is the
|
||||
same file as `longmemeval_s_cleaned.json`.)
|
||||
|
||||
**The machine was not idle.** The 1-minute load never fell below 2 in a
|
||||
2-hour wait, because two stray test processes were each holding a core; it
|
||||
was 2.1–2.7 when each run started and up to 8.4 while the runs were going.
|
||||
Recall does not depend on load. Latency does, so no latency figure in this
|
||||
file was updated from these runs.
|
||||
|
||||
**What reproduced exactly:** every published full-haystack Hit@1/5/10 and MRR
|
||||
at turn and session level for BM25, vector-only, hybrid 0.4/0.6, both stemmed
|
||||
modes, and every re-ranking row; RRF except as below; the float16 table; 4 of
|
||||
the 11 weight-sweep rows (0.0, 0.1, 0.3 and 0.9); and the oracle vector-only
|
||||
figure. The headline,
|
||||
hybrid 0.4/0.6 turn Hit@5 **81.4%**, is unchanged.
|
||||
|
||||
**What changed.** Old values are kept here; the tables below show the new ones.
|
||||
|
||||
| Figure | Published | 2026-09-27 | Why |
|
||||
|---|---:|---:|---|
|
||||
| Oracle, hybrid turn Hit@5 | 85.2% (2026-08-07, c913cd1) | **86.8%** | Weights. 85.2% was measured at 0.7/0.3, the default then. The default has been 0.4/0.6 since 29baabb. Today's oracle sweep gives 85.2% at 0.7/0.3 and 86.8% at 0.4/0.6. |
|
||||
| Oracle, BM25-only turn Hit@5 with embeddings | 84.2% (2026-08-07) | 84.4% | Now equal to the zero-embedding figure, as it already was on the full haystack. At c913cd1, tied candidates were ordered by `HashMap` iteration. Since 3ed0489 (2026-09-19) they are broken by index. 3ed0489 is the likely cause; it was not bisected. |
|
||||
| Ablation / sweep, hybrid 0.7/0.3, turn | 44.4% / 79.2% / 86.0% / 0.5868 | 44.2% / 79.2% / 85.8% / 0.5856 | The same tie-breaking change. 29baabb's re-run on 2026-09-19, made after 3ed0489 that same day, already had 44.2% / 85.8% / 0.5856. |
|
||||
| Ablation, hybrid 0.7/0.3, session | 88.2% / 95.8% / 97.8% / 0.9158 | 88.0% / 95.8% / 97.6% / 0.9146 | Same. |
|
||||
| Ablation / sweep, vector-only turn MRR | 0.5027 | 0.5031 | Same. The fusion and float16 tables already had 0.5031. |
|
||||
| Sweep 0.2/0.8, turn | 53.6% / 78.2% / 0.6440 | 53.4% / 78.0% / 0.6429 | Same (the sweep was measured 2026-08-07, 1537a94). |
|
||||
| Sweep 0.4/0.6, 0.5/0.5, 0.8/0.2 turn MRR | 0.6429 / 0.6234 / 0.5571 | 0.6430 / 0.6232 / 0.5574 | Same. |
|
||||
| Sweep 0.6/0.4 | Hit@5 79.8%, MRR 0.6069, session Hit@5 96.6% | 79.6%, 0.6067, 96.4% | Same. |
|
||||
| RRF turn MRR | 0.5967 (2026-09-19, aa92fef) | 0.5969 | Not explained. It may be a later search-path change, such as the int8 index becoming the default in 8b85d93. It may also be the run-to-run variation described next. It was not bisected. |
|
||||
|
||||
**Not exactly deterministic.** Between two f32 runs today, the recency count
|
||||
(the share of `knowledge-update` questions where the newest gold session
|
||||
ranks first) differed by one question in two modes: hybrid 0.4/0.6 was
|
||||
144/320 in one run and 145/320 in the other, and re-rank with a 1-day
|
||||
half-life was 165/319 and 166/319. Each question gets a fresh store, so this
|
||||
is not state carried between modes. The cause was not found; a candidate is
|
||||
that the GPU embeddings are not bit-for-bit identical from run to run. Hit@k
|
||||
and MRR agreed in every run that repeated a mode (hybrid 0.4/0.6 was measured
|
||||
four times: three f32 runs and one float16 run), but a last-digit change in an
|
||||
MRR, or a one-question change in recency, is within this variation.
|
||||
|
||||
### Full haystack — `longmemeval_s`, n=500 (the number to cite)
|
||||
|
||||
@@ -1440,13 +1574,16 @@ the question.
|
||||
|
||||
Real 384-d `all-MiniLM-L6-v2` embeddings, 190,015 unique texts encoded once on an
|
||||
RTX 5060 Ti (~13 min; the same work on the 8-core CPU was still unfinished after
|
||||
30 minutes, so the GPU path is not a convenience here). Turn-level:
|
||||
30 minutes, so the GPU path is not a convenience here). First measured
|
||||
2026-08-07 (c913cd1); the values below are the 2026-09-27 re-run, which moved
|
||||
the vector-only MRR and the `0.7/0.3` row (see the re-run section above).
|
||||
Turn-level:
|
||||
|
||||
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|
||||
|------|-------|-------|--------|-----|
|
||||
| BM25 only (`0.0`/`1.0`) | **53.8%** | 75.0% | 81.6% | **0.6320** |
|
||||
| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5027 |
|
||||
| Hybrid (`0.7`/`0.3`) | 44.4% | **79.2%** | **86.0%** | 0.5868 |
|
||||
| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5031 |
|
||||
| Hybrid (`0.7`/`0.3`) | 44.2% | **79.2%** | **85.8%** | 0.5856 |
|
||||
|
||||
Session-level:
|
||||
|
||||
@@ -1454,7 +1591,7 @@ Session-level:
|
||||
|------|-------|-------|--------|-----|
|
||||
| BM25 only | 86.2% | 93.6% | 96.6% | 0.8948 |
|
||||
| Vector only | 85.4% | 94.2% | 96.6% | 0.8901 |
|
||||
| Hybrid | **88.2%** | **95.8%** | **97.8%** | **0.9158** |
|
||||
| Hybrid | **88.0%** | **95.8%** | **97.6%** | **0.9146** |
|
||||
|
||||
### Fusion method — weighted vs. RRF, full haystack, n=500
|
||||
|
||||
@@ -1468,7 +1605,7 @@ takes a `Fusion`, and both run over the same HNSW + BM25 candidates:
|
||||
| BM25 only | **53.8%** | 75.0% | 81.6% | 0.6320 | 86.2% | 0.8948 |
|
||||
| Vector only | 36.0% | 71.8% | 81.6% | 0.5031 | 85.4% | 0.8901 |
|
||||
| **Weighted 0.4 / 0.6** | 51.6% | **81.4%** | **87.8%** | **0.6430** | **91.0%** | **0.9347** |
|
||||
| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5967 | 89.6% | 0.9253 |
|
||||
| RRF (k=60) | 45.0% | 78.8% | 87.6% | 0.5969 | 89.6% | 0.9253 |
|
||||
|
||||
**RRF loses to the tuned weighted sum** — 6.6pp of turn Hit@1 and 0.046 of MRR
|
||||
— and lands almost exactly where the old `0.7/0.3` weighting did (44.2% /
|
||||
@@ -1533,6 +1670,12 @@ one (see `newest_gold_first`); ~45% is chance.
|
||||
| + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% |
|
||||
| + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% |
|
||||
|
||||
Re-run on 2026-09-27 (`--rerank-sweep`, tank, commit 7a8fae0): every Hit@k and
|
||||
MRR above reproduced exactly. The recency column came out 45.3%, 87.5%,
|
||||
52.0%, 52.2%, 51.7% and 50.5%. Each of those is within one question of the
|
||||
value in the table, which is the run-to-run variation described under
|
||||
"Re-run with real embeddings" above, so the table was left as it was.
|
||||
|
||||
**The pre-fix row is the finding.** Ordering candidates by recency alone costs
|
||||
40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in
|
||||
the pool rather than the ones that answer the question. It does ace the recency
|
||||
@@ -1546,9 +1689,9 @@ cannot reach the 87.5% the degenerate ordering gets. Those two rows are the
|
||||
ends of a trade-off, and the default sits deliberately near the relevance end.
|
||||
|
||||
**Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency
|
||||
moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because
|
||||
the temporal term is capped by its weight (0.3) while relevance differences
|
||||
between candidates are larger. The 24-hour default is kept; there is no
|
||||
moves 1.4pp (1.7pp in the 2026-09-27 re-run) and MRR 0.003 — inside the
|
||||
noise of a 500-question run — because the temporal term is capped by its
|
||||
weight (0.3) while relevance differences between candidates are larger. The 24-hour default is kept; there is no
|
||||
measured reason to change it, and a corpus-matched value is not the lever it
|
||||
looks like.
|
||||
|
||||
@@ -1556,24 +1699,28 @@ looks like.
|
||||
|
||||
`0.7/0.3` was a documented default, never a searched one. Sweeping
|
||||
`vector_weight` from 0.0 to 1.0 (`--sweep`, reusing the one-time embedding
|
||||
table) shows it is not merely suboptimal but **strictly dominated**:
|
||||
table) shows it is not merely suboptimal but **strictly dominated**. First
|
||||
measured 2026-08-07 (1537a94); the values below are the 2026-09-27 re-run,
|
||||
which changed the 0.2, 0.4, 0.5, 0.6, 0.7, 0.8 and 1.0 rows in the last digit
|
||||
or by one or two questions (see the re-run section above):
|
||||
|
||||
| vector / keyword | Hit@1 | Hit@5 | Hit@10 | MRR | session Hit@5 |
|
||||
|---|---|---|---|---|---|
|
||||
| 0.0 / 1.0 (BM25) | **53.8%** | 75.0% | 81.6% | 0.6320 | 93.6% |
|
||||
| 0.1 / 0.9 | 53.2% | 77.4% | 83.8% | 0.6374 | 95.0% |
|
||||
| 0.2 / 0.8 | 53.6% | 78.2% | 85.6% | 0.6440 | 95.4% |
|
||||
| 0.2 / 0.8 | 53.4% | 78.0% | 85.6% | 0.6429 | 95.4% |
|
||||
| 0.3 / 0.7 | 53.2% | 78.8% | 87.2% | **0.6463** | 96.0% |
|
||||
| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6429 | 96.8% |
|
||||
| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6234 | **97.4%** |
|
||||
| 0.6 / 0.4 | 46.6% | 79.8% | 87.4% | 0.6069 | 96.6% |
|
||||
| 0.7 / 0.3 *(old default)* | 44.4% | 79.2% | 86.0% | 0.5868 | 95.8% |
|
||||
| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5571 | 95.2% |
|
||||
| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6430 | 96.8% |
|
||||
| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6232 | **97.4%** |
|
||||
| 0.6 / 0.4 | 46.6% | 79.6% | 87.4% | 0.6067 | 96.4% |
|
||||
| 0.7 / 0.3 *(old default)* | 44.2% | 79.2% | 85.8% | 0.5856 | 95.8% |
|
||||
| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5574 | 95.2% |
|
||||
| 0.9 / 0.1 | 37.8% | 73.4% | 84.6% | 0.5289 | 94.2% |
|
||||
| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5027 | 94.2% |
|
||||
| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5031 | 94.2% |
|
||||
|
||||
**`0.4/0.6` beats `0.7/0.3` on every metric at both granularities** — Hit@1
|
||||
+7.2pp, Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. There is no trade being made; the
|
||||
+7.4pp, Hit@5 +2.2, Hit@10 +2.0, MRR +0.057 (2026-09-27 figures; +7.2pp,
|
||||
+2.2, +1.8 and +0.056 as measured on 2026-08-07). There is no trade being made; the
|
||||
old default was simply on the wrong side of the peak. **`0.4/0.6` is the
|
||||
recommended setting**, with `0.3/0.7` preferable if rank-1 precision matters
|
||||
most (it takes the best MRR in the sweep and gives up only 0.6pp of Hit@1
|
||||
@@ -1634,11 +1781,27 @@ price of the harder corpus, and is the reason oracle-only numbers should not be
|
||||
presented as LongMemEval results. Session-level figures on this variant are
|
||||
degenerate — see below.
|
||||
|
||||
With real embeddings the same oracle corpus gives BM25-only 84.2% / vector-only
|
||||
80.4% / hybrid **85.2%** Hit@5 turn-level — hybrid ahead at Hit@5 and Hit@10 and
|
||||
behind at Hit@1, matching the full-haystack pattern above. (BM25-only reads 84.2%
|
||||
here against 84.4% with zero embedding vectors: one question of 500 changes rank,
|
||||
with MRR identical at 0.6597. On the full haystack the two agree exactly.)
|
||||
With real embeddings, the same oracle corpus gives these turn-level figures
|
||||
(2026-09-27, tank, commit 7a8fae0; command and load in "Re-run with real
|
||||
embeddings" above):
|
||||
|
||||
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|
||||
|---|---:|---:|---:|---:|
|
||||
| BM25 only | **52.6%** | 84.4% | 90.4% | 0.6597 |
|
||||
| Vector only | 39.6% | 80.4% | 91.0% | 0.5605 |
|
||||
| Hybrid 0.4 / 0.6 (default) | 52.4% | **86.8%** | **92.4%** | **0.6678** |
|
||||
| Hybrid 0.7 / 0.3 (old default) | 48.4% | 85.2% | 92.2% | 0.6382 |
|
||||
|
||||
Hybrid 0.4/0.6 leads at Hit@5, Hit@10 and MRR and is 0.2pp (one question)
|
||||
behind BM25 at Hit@1, as on the full haystack.
|
||||
|
||||
Until 2026-09-27 this paragraph gave BM25-only 84.2%, vector-only 80.4% and
|
||||
hybrid **85.2%** Hit@5. Those were measured on 2026-08-07 (c913cd1), when the
|
||||
hybrid default was 0.7/0.3; today's 0.7/0.3 row reproduces the 85.2%. The
|
||||
0.4/0.6 default (29baabb) is what moves hybrid to 86.8%. The earlier BM25-only
|
||||
84.2% with embeddings, one question below the zero-embedding 84.4%, predates
|
||||
the index tie-break of 3ed0489; the two now agree, as they always did on the
|
||||
full haystack.
|
||||
|
||||
### Retracted: session-level recall and the MemX comparison
|
||||
|
||||
|
||||
@@ -477,6 +477,19 @@ Design: `docs/design/swmr.md`.
|
||||
`main` and it calls only format-crate code; its earlier +5–10% (323–355
|
||||
vs 353–363 ns) was run-to-run layout noise (at `2893b6c` it measured
|
||||
357–374 ns against `main`'s 355–371 in the same rounds).
|
||||
**Rechecked 2026-09-27 on an idle tank** (load below 2 at every round;
|
||||
`8f59b2e` vs `main` `7a8fae0` as separate binaries, pinned to one CPU, 3
|
||||
alternating rounds, median (range); `BENCHMARKS.md`, "Local metadata and
|
||||
data reads after range-read M2/M3"): facade listing 8.24 (8.04–8.28) →
|
||||
8.10 (8.04–8.13) ms, **−1.7%**, so the listing regression is gone;
|
||||
symbol-table nodes −0.5%, B-tree walk +1.1%; `ObjectHeader::parse` ×401
|
||||
23.86 (23.79–23.96) → 24.86 (24.69–24.98) µs, **+4.2%, open**
|
||||
(`docs/known-issues.md`; about 2.5 ns per header, not visible in the
|
||||
listing). Data reads (`concurrent_read --decode-threads 1`): full reads
|
||||
of deflate data +1.7% (1 thread) to +36% (16 threads), contiguous reads
|
||||
and deflate hyperslabs within ±2%. (A run earlier that day under load
|
||||
2.3–3.3 put single-thread contiguous hyperslabs at −5.6%; the idle rerun
|
||||
does not reproduce it.)
|
||||
- Tests (2026-09-26, tank):
|
||||
- `clawhdf5-format/tests/storage_equivalence.rs` now also reads every
|
||||
dataset — whole, fill-aware, through a chunk cache (twice) and the
|
||||
@@ -631,6 +644,9 @@ Design: `docs/design/swmr.md`.
|
||||
instead of being inlined into the generic parser. Provisional (tank load
|
||||
3–4; Criterion, separate binaries, 2 alternating rounds): 24.96–25.12 vs
|
||||
24.85–25.15 µs on `main`.
|
||||
Rechecked 2026-09-27 on an idle tank (3 alternating rounds, load below
|
||||
2): 24.69–24.98 µs against 23.79–23.96 µs for `8f59b2e`, median +4.2%; a
|
||||
residual cost remains (see the M2 speed note and `docs/known-issues.md`).
|
||||
- Tests: `crates/clawhdf5-tools/tests/edit_coverage_interop.rs` (h5py
|
||||
`earliest`/`v110`/`latest` and clawhdf5-written files; structure
|
||||
comparisons with libhdf5 for version-2 B-trees, shrink on every index,
|
||||
|
||||
+2
-2
@@ -135,7 +135,7 @@
|
||||
**Crates:** `clawhdf5-agent`, `clawhdf5-bench`
|
||||
|
||||
- [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547
|
||||
- [x] **8.2** LongMemEval benchmark — 500 queries, session Hit@1=100%, turn Hit@5=84.4% (beats MemX 51.6%), MRR=0.660
|
||||
- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results)
|
||||
- [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal
|
||||
- [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion
|
||||
- [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss
|
||||
@@ -168,7 +168,7 @@ Verified against current repo state on 2026-08-05 (see also `docs/superpowers/pl
|
||||
|
||||
### Recently closed out (2026-08-05, Tier 3–4 hardening pass)
|
||||
|
||||
- [x] Academic benchmark cross-validation — LongMemEval reproduced against MemX on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% vs MemX's 51.6%; recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
|
||||
- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
|
||||
- [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer
|
||||
- [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories
|
||||
- [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open
|
||||
|
||||
@@ -7,6 +7,24 @@ deleting it.
|
||||
|
||||
---
|
||||
|
||||
## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27)
|
||||
|
||||
**Status:** open (speed only; values are correct). Measured on an idle tank
|
||||
(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`)
|
||||
against `main` `7a8fae0`, separate binaries alternating, 3 rounds
|
||||
(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"):
|
||||
`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges
|
||||
23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read
|
||||
continuation chunks from a bounded queue since `a69c5be` (bounding what a
|
||||
crafted header can make it read); `4313917` removed its per-header
|
||||
allocations but not all of the cost. Listing the 400-group file through the
|
||||
facade, which parses the same headers, is 1.7% faster, so no user-visible
|
||||
path is slower.
|
||||
|
||||
An earlier run the same day under load also listed single-thread contiguous
|
||||
hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with
|
||||
overlapping ranges (noise), so that item is withdrawn.
|
||||
|
||||
## Files a SWMR writer had open could not be read past a stale end of file
|
||||
|
||||
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any
|
||||
|
||||
Reference in New Issue
Block a user