diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 806fe14..7f8944e 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -653,6 +653,46 @@ add. There is no case here for changing the default; `TokenFilter::Stemmed` is available via `HDF5Memory::set_token_filter` for callers who want Hit@5/@10 over rank-1 precision. +### Re-ranking and recency — full haystack, n=500 + +`reranker::rerank` combines temporal decay, source authority and Hebbian +activation. Until now its combined score contained **no relevance term at +all** — `RerankInput` did not carry the retrieval score — so a caller that +re-ranked its candidates threw the retriever's ordering away and returned them +ordered by age. The OpenClaw backend did exactly that on every search. + +Measuring that is unambiguous. "Recency" below is the share of +`knowledge-update` questions where the newest gold session outranked the stale +one (see `newest_gold_first`); ~45% is chance. + +| Mode | Hit@1 | Hit@5 | Hit@10 | MRR | recency | +|---|---|---|---|---|---| +| Hybrid 0.4/0.6, no re-rank | 51.6% | **81.4%** | 87.8% | 0.6430 | 45.0% | +| + re-rank, **metadata only** (pre-fix) | 11.0% | 24.8% | 43.8% | 0.1829 | **87.5%** | +| + re-rank, relevance-led, half-life 1 day | **52.0%** | 79.8% | 87.8% | 0.6403 | 51.7% | +| + re-rank, relevance-led, half-life 7 days | 51.8% | 80.8% | 87.6% | **0.6437** | **52.2%** | +| + re-rank, relevance-led, half-life 30 days | 51.8% | 81.0% | 87.8% | 0.6427 | 51.4% | +| + re-rank, relevance-led, half-life 90 days | **52.0%** | 80.4% | 87.8% | 0.6425 | 50.8% | + +**The pre-fix row is the finding.** Ordering candidates by recency alone costs +40.6pp of Hit@1 and two thirds of MRR: the results are the newest memories in +the pool rather than the ones that answer the question. It does ace the recency +metric, which is exactly what makes that metric worth having — a number that +only goes up when a change is good would not have caught this. + +With relevance leading, retrieval is preserved (Hit@1 +0.4pp, MRR −0.003 +against no re-ranking) and recency discrimination gains 6–7pp. That is a real +improvement but not a solved problem: recency only breaks near-ties, so it +cannot reach the 87.5% the degenerate ordering gets. Those two rows are the +ends of a trade-off, and the default sits deliberately near the relevance end. + +**Half-life is not a sensitive knob.** Across 1, 7, 30 and 90 days recency +moves 1.4pp and MRR 0.003 — inside the noise of a 500-question run — because +the temporal term is capped by its weight (0.3) while relevance differences +between candidates are larger. The 24-hour default is kept; there is no +measured reason to change it, and a corpus-matched value is not the lever it +looks like. + ### Weight sweep — full haystack, n=500 `0.7/0.3` was a documented default, never a searched one. Sweeping diff --git a/CHANGELOG.md b/CHANGELOG.md index 4b4da7f..f3e103d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,28 @@ # Changelog +## Unreleased + +### Retrieval quality +- `clawhdf5-agent`: **re-ranking discarded the retrieval score.** + `reranker::rerank` built its combined score from temporal decay, source + authority and Hebbian activation only — `RerankInput` had no relevance field + — so re-ranking a candidate pool reordered it by age and threw the + retriever's ordering away. The OpenClaw backend re-ranked every search, so + this was its shipping behaviour: measured over the full LongMemEval haystack + it cost **40.6pp of Hit@1** (11.0% vs 51.6%) and two thirds of MRR (0.183 vs + 0.643). `RerankInput::relevance` and `ReRankConfig::relevance_weight` (1.0 by + default) fix it: relevance leads and the metadata signals break near-ties, + which restores retrieval (Hit@1 +0.4pp vs no re-ranking) and improves + recency discrimination by 6–7pp. **Breaking:** `RerankInput` and + `ReRankConfig` gained fields, so literal constructions need updating; + `..Default::default()` does not. +- `clawhdf5-bench`: the LongMemEval harness feeds the dataset's real session + dates to the store instead of a synthetic counter (decay needs true + intervals, not just the right order), and reports `newest_gold_first` — on a + `knowledge-update` question, did the newest gold session outrank the stale + one it supersedes? Plain recall cannot see this, because both are labelled + gold. New `--rerank-sweep`. + ## v2.5.0 (2026-09-19) ### Upgrade Notes