docs: fourth skill-triage row — precision 4/4 across every run so far
deploy / test (push) Successful in 4m59s
deploy / build (push) Successful in 59s

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
This commit is contained in:
Omar Sobh
2026-09-22 09:02:06 -05:00
co-authored by Claude Opus 5
parent 7b28950da1
commit ad67ce2fde
+11
View File
@@ -467,6 +467,17 @@ output is digest entries. `duplicate-detection` sat correctly below the
line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the
oracle is right where the agent is wrong, rather than merely plausible. oracle is right where the agent is wrong, rather than merely plausible.
**Four rows now, and they say the same thing.** `01a0c950` (the second
research mission): 7 observable, oracle says 5 apply, agent read 2, **0
unexpected**. Across all four rows the oracle's **precision is 4/4 — the
agent has never once read a skill the oracle scored below the line** —
while recall varies (4/5, then 2/5) and every gap is a skill the agent
skipped that the mission's own output argues it wanted
(`executive-summary-writing` on a digest mission, `podcast-dialogue-writing`
on one that writes a script). Zero observed cases of the oracle scoring
low something the agent needed. That is the number that would have to go
wrong for selection to be unsafe, and it has not yet.
**Next for this tier:** accumulate skill-triage agreement rows from real **Next for this tier:** accumulate skill-triage agreement rows from real
missions before any promotion from shadow to selection; a local backend missions before any promotion from shadow to selection; a local backend
only if the vendor dependency bites (logit read-out over the 9B fleet only if the vendor dependency bites (logit read-out over the 9B fleet