Benchmarks re-measured: LongMemEval with real embeddings, local reads on an idle machine #20

Merged
osobh merged 3 commits from docs/bench-refresh-2026-09-27 into main 2026-09-28 02:38:09 +00:00
Owner

Docs only. This re-measures two sets of published numbers that were stale or taken under load, using the MiniLM model and LongMemEval data restored on 2026-09-27. No code changes.

LongMemEval re-run with real embeddings (tank, RTX 5060 Ti, main 7a8fae0)

  • The headline is unchanged. Full-haystack hybrid 0.4/0.6, turn-level Hit@5, is still 81.4%, so the README badge, the README tables and CLAUDE.md stay as they were.
  • The float16 claim holds: f32 and float16 match on every Hit@k and MRR.
  • Unchanged figures: every full-haystack Hit@k and MRR for BM25, vector-only, hybrid, the stemmed modes and every re-ranking row reproduced. So did oracle BM25 (84.4%, MRR 0.6597) and oracle vector-only (80.4%, MRR 0.5605).
  • Figures updated, with their causes and the old values kept as history:
    • Oracle hybrid Hit@5: 85.2% → 86.8%. The 2026-08-07 figure was measured at 0.7/0.3, the default then, and today's sweep gives exactly 85.2% at 0.7/0.3. The default has been 0.4/0.6 since 29baabb.
    • Oracle BM25 with embeddings: 84.2% → 84.4%. Likely 3ed0489, which breaks tied scores by index instead of hash order; not bisected.
    • Ablation and sweep rows: fourth-decimal MRR and 0.2-point Hit@k changes. These come from the same tie-break change, and 29baabb's commit message already recorded them.
    • RRF turn MRR: 0.5967 → 0.5969, not explained.
  • Run-to-run variation, found and documented: one recency metric differed by one question between two f32 runs. The float16 section's claim that recency flips proved the half-precision path was active is corrected.
  • ROADMAP.md: it still cited the session Hit@1 = 100% and "beats MemX" comparisons that BENCHMARKS.md had already retracted. It now gives current numbers and points to the retraction.

Local reads after range-read M2/M3, measured idle

  • Setup: 8f59b2e (just before #18) against main 7a8fae0, run as separate binaries alternating between the two. Every round started with the load below 2; the first attempt had been spoiled by two orphaned h5py test processes, since stopped.
  • Facade listing: −1.7%. The +7–10% regression found while merging #18 is gone.
  • Full deflate reads: +1.7% at 1 thread up to +36% at 16 threads.
  • ObjectHeader::parse: +4.2%. This is real (the ranges don't overlap): about 2.5 ns per header, from the bounded continuation-chunk queue. It isn't visible in the listing, and it stays open in docs/known-issues.md.
  • Single-thread contiguous hyperslabs: the −5.6% from the loaded run was noise. Idle it measures −1.7% with overlapping ranges, so it's withdrawn from known-issues.

Commands, commits, load logs and raw outputs are recorded in BENCHMARKS.md. The raw files are in target/ab-results2/ and target/lme-2026-09-27/ on tank.

🤖 Generated with Claude Code

Docs only. This re-measures two sets of published numbers that were stale or taken under load, using the MiniLM model and LongMemEval data restored on 2026-09-27. No code changes. ## LongMemEval re-run with real embeddings (tank, RTX 5060 Ti, `main` `7a8fae0`) - **The headline is unchanged.** Full-haystack hybrid 0.4/0.6, turn-level Hit@5, is still **81.4%**, so the README badge, the README tables and CLAUDE.md stay as they were. - **The float16 claim holds:** f32 and float16 match on every Hit@k and MRR. - **Unchanged figures:** every full-haystack Hit@k and MRR for BM25, vector-only, hybrid, the stemmed modes and every re-ranking row reproduced. So did oracle BM25 (84.4%, MRR 0.6597) and oracle vector-only (80.4%, MRR 0.5605). - **Figures updated, with their causes and the old values kept as history:** - **Oracle hybrid Hit@5:** 85.2% → **86.8%**. The 2026-08-07 figure was measured at 0.7/0.3, the default then, and today's sweep gives exactly 85.2% at 0.7/0.3. The default has been 0.4/0.6 since `29baabb`. - **Oracle BM25 with embeddings:** 84.2% → 84.4%. Likely `3ed0489`, which breaks tied scores by index instead of hash order; not bisected. - **Ablation and sweep rows:** fourth-decimal MRR and 0.2-point Hit@k changes. These come from the same tie-break change, and `29baabb`'s commit message already recorded them. - **RRF turn MRR:** 0.5967 → 0.5969, not explained. - **Run-to-run variation, found and documented:** one recency metric differed by one question between two f32 runs. The float16 section's claim that recency flips proved the half-precision path was active is corrected. - **ROADMAP.md:** it still cited the session Hit@1 = 100% and "beats MemX" comparisons that BENCHMARKS.md had already retracted. It now gives current numbers and points to the retraction. ## Local reads after range-read M2/M3, measured idle - **Setup:** `8f59b2e` (just before #18) against `main` `7a8fae0`, run as separate binaries alternating between the two. Every round started with the load below 2; the first attempt had been spoiled by two orphaned h5py test processes, since stopped. - **Facade listing:** −1.7%. The +7–10% regression found while merging #18 is gone. - **Full deflate reads:** +1.7% at 1 thread up to +36% at 16 threads. - **`ObjectHeader::parse`:** **+4.2%**. This is real (the ranges don't overlap): about 2.5 ns per header, from the bounded continuation-chunk queue. It isn't visible in the listing, and it stays open in `docs/known-issues.md`. - **Single-thread contiguous hyperslabs:** the −5.6% from the loaded run was noise. Idle it measures −1.7% with overlapping ranges, so it's withdrawn from known-issues. Commands, commits, load logs and raw outputs are recorded in BENCHMARKS.md. The raw files are in `target/ab-results2/` and `target/lme-2026-09-27/` on tank. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
osobh added 3 commits 2026-09-28 00:43:29 +00:00
8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.

The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Oracle and full longmemeval_s haystack, f32 and --float16, plus --sweep
(both corpora) and --rerank-sweep, on tank at 7a8fae0 (MiniLM on the RTX
5060 Ti). Not idle: two orphaned h5py test processes kept the 1-minute load
at 2.1-2.7 (up to 8.4 during runs), so no latency figure was updated.

The headline (hybrid 0.4/0.6 turn Hit@5 81.4%) and the float16 claim
reproduce. Oracle hybrid is 86.8%, not 85.2%: the old figure was at the
0.7/0.3 default of the time (today's 0.7/0.3 gives 85.2%). The ablation's
0.7/0.3 row, vector-only MRR and seven sweep rows move in the last digit or
by one or two questions, most likely from the index tie-break of 3ed0489.
RRF turn MRR 0.5967 -> 0.5969 is not explained. Recency counts vary by one
question between identical runs, so the float16 "flips" note is corrected.
ROADMAP 8.2 no longer repeats the retracted session-level and MemX claims.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
docs: local read A/B re-run on an idle machine
CI / test-arm64 (pull_request) Successful in 1m25s
CI / test (pull_request) Successful in 15m57s
7179006aee
The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
osobh merged commit 425585ee71 into main 2026-09-28 02:38:09 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: quantumclaw/clawhdf5#20