bench(longmemeval): --float16, and float16 measured on real embeddings
`longmemeval_bench --float16` builds every per-question store with MemoryConfig::float16, so the vector stage searches half-rounded embeddings exactly as such a store holds them. Full longmemeval_s (500 questions, ~494 turns each) with real all-MiniLM-L6-v2 embeddings, f32 vs float16, on tank (CUDA): identical at every Hit@k and MRR, turn and session level, in all eight modes — bar RRF session MRR 0.9253 vs 0.9254 and one or two flips out of ~320 in which gold session ranks first. The f32 run reproduces the published hybrid numbers exactly. The earlier float16 evidence was synthetic clustered data only; this is the real-embedding check. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
+3
-1
@@ -89,7 +89,9 @@
|
||||
precision.** At 100K x 384 the file goes from 154.0 to 80.8 MiB (−48%), a
|
||||
checkpoint from 752 to 512 ms and open from 300 to 252 ms, with the same
|
||||
vector recall@10 against an exact scan (0.999 vs 0.994) and the same
|
||||
`hybrid_search` latency; at 10K open is 3 ms slower. The cache rounds each
|
||||
`hybrid_search` latency; at 10K open is 3 ms slower. On the full
|
||||
LongMemEval haystack with real MiniLM embeddings every retrieval metric is
|
||||
identical to `f32` (`longmemeval_bench --float16`). The cache rounds each
|
||||
embedding as it is saved, so memory and file agree bit for bit and a store
|
||||
returns the same results before and after a reopen (tested). Out-of-range
|
||||
values are refused with `MemoryError::InvalidEntry` rather than stored as
|
||||
|
||||
Reference in New Issue
Block a user