perf(agent): persistent incremental BM25 index; no store rewrite per query
hybrid_search rebuilt the BM25 index from scratch (re-tokenising every record) and rewrote the whole .h5 file on every single query, so a query cost O(store size) in both CPU and disk I/O. Steady-state p50 per the search harness: 5.5 -> 0.24 ms (1K), 49 -> 2.1 ms (10K), 884 -> 23 ms (100K). - BM25Index is incremental: add_document / remove_document keep it exactly equivalent to a fresh build over the same live documents (property test: 60 random op sequences compared against BM25Index::build after every step). IDF moves to query time since it depends on the live document count. Top-k uses a bounded heap, ties break by doc id (results were HashMap-ordered), and the "WAND" code that computed a bound and then discarded it is removed. - HDF5Memory keeps one index for its lifetime, built lazily. Appends are picked up by ensure_bm25_fresh whatever path added them; delete and in-place update report themselves; compaction drops the index. A test drives every mutation and compares against a fresh build. - A query no longer calls flush(). Activation boosts are marked dirty and persisted by the next checkpoint, including a best-effort one on drop so a search-only session keeps them (approved behaviour change). Activation weights are capped at 16.0; they previously grew without bound. The archived mission branch's BM25 cache was reviewed and not used: it was invalidated by every write, so interleaved save/search still rebuilt per query, and it changed the default fusion weights. Co-Authored-By: Claude Fable 5.1 <[email protected]>
This commit is contained in:
co-authored by
Claude Fable 5.1
parent
61424d1418
commit
2bfbb7fb4b
@@ -139,6 +139,26 @@ build: 36472.8 ms (2742 vectors/s) · exact scan: 40 QPS, p50 24644 µs
|
||||
| 128 | 0.9990 | 7633 | 126 | 248 |
|
||||
| 256 | 0.9990 | 2823 | 352 | 510 |
|
||||
|
||||
### After: persistent keyword index, no store rewrite per query
|
||||
|
||||
`hybrid_search` used to rebuild the BM25 index from scratch (re-tokenising every
|
||||
record) and rewrite the whole `.h5` file on **every query**. The index is now
|
||||
kept for the life of the store and updated incrementally, and activation boosts
|
||||
are persisted by the next checkpoint instead of inside the query. Steady-state
|
||||
p50: **5.5 → 0.24 ms** (1K), **49 → 2.1 ms** (10K), **884 → 23 ms** (100K).
|
||||
|
||||
The first query after `open()` is slower than before (it pays for the better —
|
||||
slower — HNSW build plus the one-off keyword index build); persisting the HNSW
|
||||
index removes that.
|
||||
|
||||
### End to end: `HDF5Memory::hybrid_search` (k = 10, weights 0.7 / 0.3)
|
||||
|
||||
| N | ingest ms | checkpoint ms | open ms | first query ms | p50 ms | p99 ms | QPS |
|
||||
|---:|---:|---:|---:|---:|---:|---:|---:|
|
||||
| 1000 | 11 | 3.8 | 0.9 | 195.9 | 0.24 | 0.27 | 4130.4 |
|
||||
| 10000 | 104 | 31.1 | 10.9 | 2627.1 | 2.09 | 2.11 | 479.5 |
|
||||
| 100000 | 1436 | 684.7 | 278.0 | 36308.1 | 22.90 | 25.46 | 43.5 |
|
||||
|
||||
## Vector Search Latency
|
||||
|
||||
Brute-force cosine similarity over 384-dimensional embeddings (OpenAI text-embedding-3-small size).
|
||||
|
||||
Reference in New Issue
Block a user