docs: fourth skill-triage row — precision 4/4 across every run so far
Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
This commit is contained in:
co-authored by
Claude Opus 5
parent
7b28950da1
commit
ad67ce2fde
@@ -467,6 +467,17 @@ output is digest entries. `duplicate-detection` sat correctly below the
|
|||||||
line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the
|
line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the
|
||||||
oracle is right where the agent is wrong, rather than merely plausible.
|
oracle is right where the agent is wrong, rather than merely plausible.
|
||||||
|
|
||||||
|
**Four rows now, and they say the same thing.** `01a0c950` (the second
|
||||||
|
research mission): 7 observable, oracle says 5 apply, agent read 2, **0
|
||||||
|
unexpected**. Across all four rows the oracle's **precision is 4/4 — the
|
||||||
|
agent has never once read a skill the oracle scored below the line** —
|
||||||
|
while recall varies (4/5, then 2/5) and every gap is a skill the agent
|
||||||
|
skipped that the mission's own output argues it wanted
|
||||||
|
(`executive-summary-writing` on a digest mission, `podcast-dialogue-writing`
|
||||||
|
on one that writes a script). Zero observed cases of the oracle scoring
|
||||||
|
low something the agent needed. That is the number that would have to go
|
||||||
|
wrong for selection to be unsafe, and it has not yet.
|
||||||
|
|
||||||
**Next for this tier:** accumulate skill-triage agreement rows from real
|
**Next for this tier:** accumulate skill-triage agreement rows from real
|
||||||
missions before any promotion from shadow to selection; a local backend
|
missions before any promotion from shadow to selection; a local backend
|
||||||
only if the vendor dependency bites (logit read-out over the 9B fleet
|
only if the vendor dependency bites (logit read-out over the 9B fleet
|
||||||
|
|||||||
Reference in New Issue
Block a user