diff --git a/docs/NEXT-SESSION.md b/docs/NEXT-SESSION.md index 94f7eb2..48a80f4 100644 --- a/docs/NEXT-SESSION.md +++ b/docs/NEXT-SESSION.md @@ -458,6 +458,15 @@ scores collapse. **A score that returns the same number for everything is a defect in the question, not a fact about the population** — and it looks exactly like a working feature: N confident numbers, no information. +**Third skill-triage agreement row, and the first substantive one** +(`01a0c940`, a real continuous-research mission): of 7 observable skills +the oracle said 5 apply; the agent read 4 of those and **0 it did not +expect**. The single disagreement is a miss by the AGENT, not the oracle: +`executive-summary-writing` at 0.77, unread, on a mission whose whole +output is digest entries. `duplicate-detection` sat correctly below the +line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the +oracle is right where the agent is wrong, rather than merely plausible. + **Next for this tier:** accumulate skill-triage agreement rows from real missions before any promotion from shadow to selection; a local backend only if the vendor dependency bites (logit read-out over the 9B fleet