Benchmarks re-measured: LongMemEval with real embeddings, local reads on an idle machine #20

Merged
osobh merged 3 commits from docs/bench-refresh-2026-09-27 into main 2026-09-28 02:38:09 +00:00
3 Commits
Author SHA1 Message Date
osobhandClaude Opus 5.5 7179006aee docs: local read A/B re-run on an idle machine
CI / test-arm64 (pull_request) Successful in 1m25s
CI / test (pull_request) Successful in 15m57s
The first run of the day could not get an idle tank: two orphaned h5py
SWMR reader processes from earlier interop tests (since stopped) kept a
core each busy. Re-run with the load below 2 at every round:
ObjectHeader::parse is +4.2% (real: the ranges do not overlap; about 2.5
ns per header, not visible in the facade listing, which is -1.7%); full
deflate reads +1.7% (1 thread) to +36% (16 threads); the -5.6% single-
thread contiguous hyperslab result from the loaded run is noise (-1.7%,
overlapping ranges) and is withdrawn from known-issues.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 19:43:07 -05:00
osobhandClaude Opus 5.5 b39a705e77 docs: LongMemEval with real embeddings re-run on 2026-09-27
Oracle and full longmemeval_s haystack, f32 and --float16, plus --sweep
(both corpora) and --rerank-sweep, on tank at 7a8fae0 (MiniLM on the RTX
5060 Ti). Not idle: two orphaned h5py test processes kept the 1-minute load
at 2.1-2.7 (up to 8.4 during runs), so no latency figure was updated.

The headline (hybrid 0.4/0.6 turn Hit@5 81.4%) and the float16 claim
reproduce. Oracle hybrid is 86.8%, not 85.2%: the old figure was at the
0.7/0.3 default of the time (today's 0.7/0.3 gives 85.2%). The ablation's
0.7/0.3 row, vector-only MRR and seven sweep rows move in the last digit or
by one or two questions, most likely from the index tie-break of 3ed0489.
RRF turn MRR 0.5967 -> 0.5969 is not explained. Recency counts vary by one
question between identical runs, so the float16 "flips" note is corrected.
ROADMAP 8.2 no longer repeats the retracted session-level and MemX claims.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 19:25:36 -05:00
osobhandClaude Opus 5.5 d239292655 docs: local metadata and data reads after range-read M2/M3, rechecked
8f59b2e (main before PR #18) against 7a8fae0, separate binaries,
alternating rounds on tank (6 Criterion rounds of local_metadata_bench,
3 of concurrent_read --decode-threads 1). Still provisional: two orphaned
h5py test processes held the load at 2.1-2.6 and the load < 2 gate was
not met in 2 hours.

The facade listing regression is gone (-0.8%). ObjectHeader::parse x401
is +4.1% and one-thread contiguous hyperslab reads -5.6%; both listed as
open in known-issues. Full deflate reads are 7-23% faster.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-27 15:35:59 -05:00