docs: the saturated-score finding, and its fix verified on a fresh harvest
deploy / build (push) Canceled after 0s
deploy / test (push) Canceled after 1m4s

01a0c940 (old question): relevance 2.93-3.00, spread 0.07. 01a0c950
(kind+evidence): spread 2.98, a survey at 0.0 and a measured method paper
at 2.98, 8 of 9 tagged. The research agent reading the old manifest had
already called the score overscored, unprompted.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
This commit is contained in:
Omar Sobh
2026-09-22 08:34:06 -05:00
co-authored by Claude Opus 5
parent 9628b26795
commit 4aafbca2c5
+24
View File
@@ -434,6 +434,30 @@ existed, is filled per paper by a Choice over the mission's topics, plus
a four-level `relevance` Score the ranking phase can start from; PORTICO's
abstract scored 3.0 at 1.0.
**The paper triage ran live, and the first thing it proved was its own
defect** (`48606f2`). Mission `01a0c940`, 10 papers: 10 tagged, 10 scored,
every answer confident — and the relevance scores were **2.933.00, a
spread of 0.07** on a 4-level scale. The harvest runs the operator's own
arXiv topic queries, so "is this relevant to an agent platform" is true by
construction and the scale had no room. The research agent reading that
manifest said so itself, unprompted, about BabelArena: *"`score: 2.98` —
overscored. A score around 1.01.5 would be honest."*
Measured on the same ten abstracts: actionability saturated too (0.20);
**evidence strength** — position piece → measured with ablations — spread
**1.64**. The manifest now carries `evidence` and `kind` (method /
benchmark / measurement / survey / position, with the confidence of that
call) and no relevance number. Verified on a fresh harvest (`01a0c950`,
9 papers, new topics): **spread 2.98** — a survey at 0.0 labelled `survey`
at 0.93, a measured method paper at 2.98 — and 8 of 9 tagged, the untagged
one genuinely fitting neither topic.
Generalised so it cannot recur quietly: `patterns::spread()` +
`SATURATED_BELOW`, and `triage_papers` warns when a live harvest's own
scores collapse. **A score that returns the same number for everything is
a defect in the question, not a fact about the population** — and it looks
exactly like a working feature: N confident numbers, no information.
**Next for this tier:** accumulate skill-triage agreement rows from real
missions before any promotion from shadow to selection; a local backend
only if the vendor dependency bites (logit read-out over the 9B fleet