diff --git a/docs/NEXT-SESSION.md b/docs/NEXT-SESSION.md index 475b158..8fa1d44 100644 --- a/docs/NEXT-SESSION.md +++ b/docs/NEXT-SESSION.md @@ -467,6 +467,17 @@ output is digest entries. `duplicate-detection` sat correctly below the line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the oracle is right where the agent is wrong, rather than merely plausible. +**Four rows now, and they say the same thing.** `01a0c950` (the second +research mission): 7 observable, oracle says 5 apply, agent read 2, **0 +unexpected**. Across all four rows the oracle's **precision is 4/4 — the +agent has never once read a skill the oracle scored below the line** — +while recall varies (4/5, then 2/5) and every gap is a skill the agent +skipped that the mission's own output argues it wanted +(`executive-summary-writing` on a digest mission, `podcast-dialogue-writing` +on one that writes a script). Zero observed cases of the oracle scoring +low something the agent needed. That is the number that would have to go +wrong for selection to be unsafe, and it has not yet. + **Next for this tier:** accumulate skill-triage agreement rows from real missions before any promotion from shadow to selection; a local backend only if the vendor dependency bites (logit read-out over the 9B fleet