Docs only. This re-measures two sets of published numbers that were stale or taken under load, using the MiniLM model and LongMemEval data restored on 2026-09-27. No code changes.
LongMemEval re-run with real embeddings (tank, RTX 5060 Ti, main7a8fae0)
The headline is unchanged. Full-haystack hybrid 0.4/0.6, turn-level Hit@5, is still 81.4%, so the README badge, the README tables and CLAUDE.md stay as they were.
The float16 claim holds: f32 and float16 match on every Hit@k and MRR.
Unchanged figures: every full-haystack Hit@k and MRR for BM25, vector-only, hybrid, the stemmed modes and every re-ranking row reproduced. So did oracle BM25 (84.4%, MRR 0.6597) and oracle vector-only (80.4%, MRR 0.5605).
Figures updated, with their causes and the old values kept as history:
Oracle hybrid Hit@5: 85.2% → 86.8%. The 2026-08-07 figure was measured at 0.7/0.3, the default then, and today's sweep gives exactly 85.2% at 0.7/0.3. The default has been 0.4/0.6 since 29baabb.
Oracle BM25 with embeddings: 84.2% → 84.4%. Likely 3ed0489, which breaks tied scores by index instead of hash order; not bisected.
Ablation and sweep rows: fourth-decimal MRR and 0.2-point Hit@k changes. These come from the same tie-break change, and 29baabb's commit message already recorded them.
RRF turn MRR: 0.5967 → 0.5969, not explained.
Run-to-run variation, found and documented: one recency metric differed by one question between two f32 runs. The float16 section's claim that recency flips proved the half-precision path was active is corrected.
ROADMAP.md: it still cited the session Hit@1 = 100% and "beats MemX" comparisons that BENCHMARKS.md had already retracted. It now gives current numbers and points to the retraction.
Local reads after range-read M2/M3, measured idle
Setup:8f59b2e (just before #18) against main7a8fae0, run as separate binaries alternating between the two. Every round started with the load below 2; the first attempt had been spoiled by two orphaned h5py test processes, since stopped.
Facade listing: −1.7%. The +7–10% regression found while merging #18 is gone.
Full deflate reads: +1.7% at 1 thread up to +36% at 16 threads.
ObjectHeader::parse:+4.2%. This is real (the ranges don't overlap): about 2.5 ns per header, from the bounded continuation-chunk queue. It isn't visible in the listing, and it stays open in docs/known-issues.md.
Single-thread contiguous hyperslabs: the −5.6% from the loaded run was noise. Idle it measures −1.7% with overlapping ranges, so it's withdrawn from known-issues.
Commands, commits, load logs and raw outputs are recorded in BENCHMARKS.md. The raw files are in target/ab-results2/ and target/lme-2026-09-27/ on tank.
Docs only. This re-measures two sets of published numbers that were stale or taken under load, using the MiniLM model and LongMemEval data restored on 2026-09-27. No code changes.
## LongMemEval re-run with real embeddings (tank, RTX 5060 Ti, `main` `7a8fae0`)
- **The headline is unchanged.** Full-haystack hybrid 0.4/0.6, turn-level Hit@5, is still **81.4%**, so the README badge, the README tables and CLAUDE.md stay as they were.
- **The float16 claim holds:** f32 and float16 match on every Hit@k and MRR.
- **Unchanged figures:** every full-haystack Hit@k and MRR for BM25, vector-only, hybrid, the stemmed modes and every re-ranking row reproduced. So did oracle BM25 (84.4%, MRR 0.6597) and oracle vector-only (80.4%, MRR 0.5605).
- **Figures updated, with their causes and the old values kept as history:**
- **Oracle hybrid Hit@5:** 85.2% → **86.8%**. The 2026-08-07 figure was measured at 0.7/0.3, the default then, and today's sweep gives exactly 85.2% at 0.7/0.3. The default has been 0.4/0.6 since `29baabb`.
- **Oracle BM25 with embeddings:** 84.2% → 84.4%. Likely `3ed0489`, which breaks tied scores by index instead of hash order; not bisected.
- **Ablation and sweep rows:** fourth-decimal MRR and 0.2-point Hit@k changes. These come from the same tie-break change, and `29baabb`'s commit message already recorded them.
- **RRF turn MRR:** 0.5967 → 0.5969, not explained.
- **Run-to-run variation, found and documented:** one recency metric differed by one question between two f32 runs. The float16 section's claim that recency flips proved the half-precision path was active is corrected.
- **ROADMAP.md:** it still cited the session Hit@1 = 100% and "beats MemX" comparisons that BENCHMARKS.md had already retracted. It now gives current numbers and points to the retraction.
## Local reads after range-read M2/M3, measured idle
- **Setup:** `8f59b2e` (just before #18) against `main` `7a8fae0`, run as separate binaries alternating between the two. Every round started with the load below 2; the first attempt had been spoiled by two orphaned h5py test processes, since stopped.
- **Facade listing:** −1.7%. The +7–10% regression found while merging #18 is gone.
- **Full deflate reads:** +1.7% at 1 thread up to +36% at 16 threads.
- **`ObjectHeader::parse`:** **+4.2%**. This is real (the ranges don't overlap): about 2.5 ns per header, from the bounded continuation-chunk queue. It isn't visible in the listing, and it stays open in `docs/known-issues.md`.
- **Single-thread contiguous hyperslabs:** the −5.6% from the loaded run was noise. Idle it measures −1.7% with overlapping ranges, so it's withdrawn from known-issues.
Commands, commits, load logs and raw outputs are recorded in BENCHMARKS.md. The raw files are in `target/ab-results2/` and `target/lme-2026-09-27/` on tank.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.
The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Oracle and full longmemeval_s haystack, f32 and --float16, plus --sweep
(both corpora) and --rerank-sweep, on tank at 7a8fae0 (MiniLM on the RTX
5060 Ti). Not idle: two orphaned h5py test processes kept the 1-minute load
at 2.1-2.7 (up to 8.4 during runs), so no latency figure was updated.
The headline (hybrid 0.4/0.6 turn Hit@5 81.4%) and the float16 claim
reproduce. Oracle hybrid is 86.8%, not 85.2%: the old figure was at the
0.7/0.3 default of the time (today's 0.7/0.3 gives 85.2%). The ablation's
0.7/0.3 row, vector-only MRR and seven sweep rows move in the last digit or
by one or two questions, most likely from the index tie-break of 3ed0489.
RRF turn MRR 0.5967 -> 0.5969 is not explained. Recency counts vary by one
question between identical runs, so the float16 "flips" note is corrected.
ROADMAP 8.2 no longer repeats the retracted session-level and MemX claims.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
osobh
merged commit 425585ee71 into main2026-09-28 02:38:09 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Docs only. This re-measures two sets of published numbers that were stale or taken under load, using the MiniLM model and LongMemEval data restored on 2026-09-27. No code changes.
LongMemEval re-run with real embeddings (tank, RTX 5060 Ti,
main7a8fae0)29baabb.3ed0489, which breaks tied scores by index instead of hash order; not bisected.29baabb's commit message already recorded them.Local reads after range-read M2/M3, measured idle
8f59b2e(just before #18) againstmain7a8fae0, run as separate binaries alternating between the two. Every round started with the load below 2; the first attempt had been spoiled by two orphaned h5py test processes, since stopped.ObjectHeader::parse: +4.2%. This is real (the ranges don't overlap): about 2.5 ns per header, from the bounded continuation-chunk queue. It isn't visible in the listing, and it stays open indocs/known-issues.md.Commands, commits, load logs and raw outputs are recorded in BENCHMARKS.md. The raw files are in
target/ab-results2/andtarget/lme-2026-09-27/on tank.🤖 Generated with Claude Code