fix(agent): re-ranking threw away the retrieval score
reranker::rerank built its combined score from temporal decay, source authority and Hebbian activation. RerankInput carried no relevance score, so it could not have used one: re-ranking a candidate pool reordered it purely by age and discarded the retriever's ordering. The OpenClaw backend re-ranks every search, so that was its shipping behaviour. Measured over the full LongMemEval haystack (500 questions, real MiniLM embeddings), ordering by metadata alone costs 40.6pp of Hit@1 (11.0% vs 51.6%) and two thirds of MRR (0.1829 vs 0.6430) — the results are the newest memories in the pool rather than the ones answering the question. RerankInput::relevance and ReRankConfig::relevance_weight (1.0 by default) make relevance lead, with the metadata signals breaking near-ties. Retrieval is preserved (Hit@1 52.0%, +0.4pp against no re-ranking; MRR -0.003) and recency discrimination improves 6-7pp, from chance to ~52%. A half-life sweep (1, 7, 30, 90 days) moves recency 1.4pp and MRR 0.003 — inside the noise — because the temporal term is capped by its weight while relevance gaps are larger. The 24-hour default is kept: there is no measured reason to change it. The two ends of the trade-off are recorded in BENCHMARKS.md rather than just the good news. Breaking: RerankInput and ReRankConfig gained fields. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
@@ -653,6 +653,46 @@ add. There is no case here for changing the default; `TokenFilter::Stemmed`
|
||||
is available via `HDF5Memory::set_token_filter` for callers who want Hit@5/@10
|
||||
over rank-1 precision.
|
||||
|
||||
### Re-ranking and recency — full haystack, n=500
|
||||
|
||||
`reranker::rerank` combines temporal decay, source authority and Hebbian
|
||||
activation. Until now its combined score contained **no relevance term at
|
||||
all** — `RerankInput` did not carry the retrieval score — so a caller that
|
||||
re-ranked its candidates threw the retriever's ordering away and returned them
|
||||
ordered by age. The OpenClaw backend did exactly that on every search.
|
||||
|
||||
Measuring that is unambiguous. "Recency" below is the share of
|
||||
`knowledge-update` questions where the newest gold session outranked the stale
|
||||
one (see `newest_gold_first`); ~45% is chance.
|
||||
|
||||
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR | recency |
|
||||
|---|---|---|---|---|---|
|
||||
| Hybrid 0.4/0.6, no re-rank | 51.6% | **81.4%** | 87.8% | 0.6430 | 45.0% |
|
||||
| + re-rank, **metadata only** (pre-fix) | 11.0% | 24.8% | 43.8% | 0.1829 | **87.5%** |
|
||||
| + re-rank, relevance-led, half-life 1 day | **52.0%** | 79.8% | 87.8% | 0.6403 | 51.7% |
|
||||
| + re-rank, relevance-led, half-life 7 days | 51.8% | 80.8% | 87.6% | **0.6437** | **52.2%** |
|
||||
| + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% |
|
||||
| + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% |
|
||||
|
||||
**The pre-fix row is the finding.** Ordering candidates by recency alone costs
|
||||
40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in
|
||||
the pool rather than the ones that answer the question. It does ace the recency
|
||||
metric, which is exactly what makes that metric worth having — a number that
|
||||
only goes up when a change is good would not have caught this.
|
||||
|
||||
With relevance leading, retrieval is preserved (Hit@1 +0.4pp, MRR −0.003
|
||||
against no re-ranking) and recency discrimination gains 6–7pp. That is a real
|
||||
improvement but not a solved problem: recency only breaks near-ties, so it
|
||||
cannot reach the 87.5% the degenerate ordering gets. Those two rows are the
|
||||
ends of a trade-off, and the default sits deliberately near the relevance end.
|
||||
|
||||
**Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency
|
||||
moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because
|
||||
the temporal term is capped by its weight (0.3) while relevance differences
|
||||
between candidates are larger. The 24-hour default is kept; there is no
|
||||
measured reason to change it, and a corpus-matched value is not the lever it
|
||||
looks like.
|
||||
|
||||
### Weight sweep — full haystack, n=500
|
||||
|
||||
`0.7/0.3` was a documented default, never a searched one. Sweeping
|
||||
|
||||
Reference in New Issue
Block a user