From ad67ce2fde572df2faf9f79639b83b24c4ef2a51 Mon Sep 17 00:00:00 2001 From: Omar Sobh Date: Tue, 22 Sep 2026 09:02:06 -0500 Subject: [PATCH] =?UTF-8?q?docs:=20fourth=20skill-triage=20row=20=E2=80=94?= =?UTF-8?q?=20precision=204/4=20across=20every=20run=20so=20far?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz --- docs/NEXT-SESSION.md | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/docs/NEXT-SESSION.md b/docs/NEXT-SESSION.md index 475b158..8fa1d44 100644 --- a/docs/NEXT-SESSION.md +++ b/docs/NEXT-SESSION.md @@ -467,6 +467,17 @@ output is digest entries. `duplicate-detection` sat correctly below the line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the oracle is right where the agent is wrong, rather than merely plausible. +**Four rows now, and they say the same thing.** `01a0c950` (the second +research mission): 7 observable, oracle says 5 apply, agent read 2, **0 +unexpected**. Across all four rows the oracle's **precision is 4/4 — the +agent has never once read a skill the oracle scored below the line** — +while recall varies (4/5, then 2/5) and every gap is a skill the agent +skipped that the mission's own output argues it wanted +(`executive-summary-writing` on a digest mission, `podcast-dialogue-writing` +on one that writes a script). Zero observed cases of the oracle scoring +low something the agent needed. That is the number that would have to go +wrong for selection to be unsafe, and it has not yet. + **Next for this tier:** accumulate skill-triage agreement rows from real missions before any promotion from shadow to selection; a local backend only if the vendor dependency bites (logit read-out over the 9B fleet