docs: the saturated-score finding, and its fix verified on a fresh harvest
01a0c940 (old question): relevance 2.93-3.00, spread 0.07. 01a0c950 (kind+evidence): spread 2.98, a survey at 0.0 and a measured method paper at 2.98, 8 of 9 tagged. The research agent reading the old manifest had already called the score overscored, unprompted. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
This commit is contained in:
co-authored by
Claude Opus 5
parent
9628b26795
commit
4aafbca2c5
@@ -434,6 +434,30 @@ existed, is filled per paper by a Choice over the mission's topics, plus
|
||||
a four-level `relevance` Score the ranking phase can start from; PORTICO's
|
||||
abstract scored 3.0 at 1.0.
|
||||
|
||||
**The paper triage ran live, and the first thing it proved was its own
|
||||
defect** (`48606f2`). Mission `01a0c940`, 10 papers: 10 tagged, 10 scored,
|
||||
every answer confident — and the relevance scores were **2.93–3.00, a
|
||||
spread of 0.07** on a 4-level scale. The harvest runs the operator's own
|
||||
arXiv topic queries, so "is this relevant to an agent platform" is true by
|
||||
construction and the scale had no room. The research agent reading that
|
||||
manifest said so itself, unprompted, about BabelArena: *"`score: 2.98` —
|
||||
overscored. A score around 1.0–1.5 would be honest."*
|
||||
|
||||
Measured on the same ten abstracts: actionability saturated too (0.20);
|
||||
**evidence strength** — position piece → measured with ablations — spread
|
||||
**1.64**. The manifest now carries `evidence` and `kind` (method /
|
||||
benchmark / measurement / survey / position, with the confidence of that
|
||||
call) and no relevance number. Verified on a fresh harvest (`01a0c950`,
|
||||
9 papers, new topics): **spread 2.98** — a survey at 0.0 labelled `survey`
|
||||
at 0.93, a measured method paper at 2.98 — and 8 of 9 tagged, the untagged
|
||||
one genuinely fitting neither topic.
|
||||
|
||||
Generalised so it cannot recur quietly: `patterns::spread()` +
|
||||
`SATURATED_BELOW`, and `triage_papers` warns when a live harvest's own
|
||||
scores collapse. **A score that returns the same number for everything is
|
||||
a defect in the question, not a fact about the population** — and it looks
|
||||
exactly like a working feature: N confident numbers, no information.
|
||||
|
||||
**Next for this tier:** accumulate skill-triage agreement rows from real
|
||||
missions before any promotion from shadow to selection; a local backend
|
||||
only if the vendor dependency bites (logit read-out over the 9B fleet
|
||||
|
||||
Reference in New Issue
Block a user