diff --git a/docs/NEXT-SESSION.md b/docs/NEXT-SESSION.md index 5adcdcc..94f7eb2 100644 --- a/docs/NEXT-SESSION.md +++ b/docs/NEXT-SESSION.md @@ -434,6 +434,30 @@ existed, is filled per paper by a Choice over the mission's topics, plus a four-level `relevance` Score the ranking phase can start from; PORTICO's abstract scored 3.0 at 1.0. +**The paper triage ran live, and the first thing it proved was its own +defect** (`48606f2`). Mission `01a0c940`, 10 papers: 10 tagged, 10 scored, +every answer confident — and the relevance scores were **2.93–3.00, a +spread of 0.07** on a 4-level scale. The harvest runs the operator's own +arXiv topic queries, so "is this relevant to an agent platform" is true by +construction and the scale had no room. The research agent reading that +manifest said so itself, unprompted, about BabelArena: *"`score: 2.98` — +overscored. A score around 1.0–1.5 would be honest."* + +Measured on the same ten abstracts: actionability saturated too (0.20); +**evidence strength** — position piece → measured with ablations — spread +**1.64**. The manifest now carries `evidence` and `kind` (method / +benchmark / measurement / survey / position, with the confidence of that +call) and no relevance number. Verified on a fresh harvest (`01a0c950`, +9 papers, new topics): **spread 2.98** — a survey at 0.0 labelled `survey` +at 0.93, a measured method paper at 2.98 — and 8 of 9 tagged, the untagged +one genuinely fitting neither topic. + +Generalised so it cannot recur quietly: `patterns::spread()` + +`SATURATED_BELOW`, and `triage_papers` warns when a live harvest's own +scores collapse. **A score that returns the same number for everything is +a defect in the question, not a fact about the population** — and it looks +exactly like a working feature: N confident numbers, no information. + **Next for this tier:** accumulate skill-triage agreement rows from real missions before any promotion from shadow to selection; a local backend only if the vendor dependency bites (logit read-out over the 9B fleet