docs: the saturated-score finding, and its fix verified on a fresh harvest
01a0c940 (old question): relevance 2.93-3.00, spread 0.07. 01a0c950 (kind+evidence): spread 2.98, a survey at 0.0 and a measured method paper at 2.98, 8 of 9 tagged. The research agent reading the old manifest had already called the score overscored, unprompted. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
This commit is contained in:
co-authored by
Claude Opus 5
parent
9628b26795
commit
4aafbca2c5
@@ -434,6 +434,30 @@ existed, is filled per paper by a Choice over the mission's topics, plus
|
|||||||
a four-level `relevance` Score the ranking phase can start from; PORTICO's
|
a four-level `relevance` Score the ranking phase can start from; PORTICO's
|
||||||
abstract scored 3.0 at 1.0.
|
abstract scored 3.0 at 1.0.
|
||||||
|
|
||||||
|
**The paper triage ran live, and the first thing it proved was its own
|
||||||
|
defect** (`48606f2`). Mission `01a0c940`, 10 papers: 10 tagged, 10 scored,
|
||||||
|
every answer confident — and the relevance scores were **2.93–3.00, a
|
||||||
|
spread of 0.07** on a 4-level scale. The harvest runs the operator's own
|
||||||
|
arXiv topic queries, so "is this relevant to an agent platform" is true by
|
||||||
|
construction and the scale had no room. The research agent reading that
|
||||||
|
manifest said so itself, unprompted, about BabelArena: *"`score: 2.98` —
|
||||||
|
overscored. A score around 1.0–1.5 would be honest."*
|
||||||
|
|
||||||
|
Measured on the same ten abstracts: actionability saturated too (0.20);
|
||||||
|
**evidence strength** — position piece → measured with ablations — spread
|
||||||
|
**1.64**. The manifest now carries `evidence` and `kind` (method /
|
||||||
|
benchmark / measurement / survey / position, with the confidence of that
|
||||||
|
call) and no relevance number. Verified on a fresh harvest (`01a0c950`,
|
||||||
|
9 papers, new topics): **spread 2.98** — a survey at 0.0 labelled `survey`
|
||||||
|
at 0.93, a measured method paper at 2.98 — and 8 of 9 tagged, the untagged
|
||||||
|
one genuinely fitting neither topic.
|
||||||
|
|
||||||
|
Generalised so it cannot recur quietly: `patterns::spread()` +
|
||||||
|
`SATURATED_BELOW`, and `triage_papers` warns when a live harvest's own
|
||||||
|
scores collapse. **A score that returns the same number for everything is
|
||||||
|
a defect in the question, not a fact about the population** — and it looks
|
||||||
|
exactly like a working feature: N confident numbers, no information.
|
||||||
|
|
||||||
**Next for this tier:** accumulate skill-triage agreement rows from real
|
**Next for this tier:** accumulate skill-triage agreement rows from real
|
||||||
missions before any promotion from shadow to selection; a local backend
|
missions before any promotion from shadow to selection; a local backend
|
||||||
only if the vendor dependency bites (logit read-out over the 9B fleet
|
only if the vendor dependency bites (logit read-out over the 9B fleet
|
||||||
|
|||||||
Reference in New Issue
Block a user