docs: the first substantive skill-triage agreement row
deploy / test (push) Successful in 5m13s
deploy / build (push) Successful in 1m4s

01a0c940: 7 observable skills, oracle says 5 apply, agent read 4 of those
and 0 unexpected. The one disagreement is the agent's miss —
executive-summary-writing at 0.77, unread, on a mission whose output is
digest entries.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz
This commit is contained in:
Omar Sobh
2026-09-22 08:35:09 -05:00
co-authored by Claude Opus 5
parent 4aafbca2c5
commit 716044833b
+9
View File
@@ -458,6 +458,15 @@ scores collapse. **A score that returns the same number for everything is
a defect in the question, not a fact about the population** — and it looks a defect in the question, not a fact about the population** — and it looks
exactly like a working feature: N confident numbers, no information. exactly like a working feature: N confident numbers, no information.
**Third skill-triage agreement row, and the first substantive one**
(`01a0c940`, a real continuous-research mission): of 7 observable skills
the oracle said 5 apply; the agent read 4 of those and **0 it did not
expect**. The single disagreement is a miss by the AGENT, not the oracle:
`executive-summary-writing` at 0.77, unread, on a mission whose whole
output is digest entries. `duplicate-detection` sat correctly below the
line at 0.31. Precision 4/4, recall 4/5 — the first row that argues the
oracle is right where the agent is wrong, rather than merely plausible.
**Next for this tier:** accumulate skill-triage agreement rows from real **Next for this tier:** accumulate skill-triage agreement rows from real
missions before any promotion from shadow to selection; a local backend missions before any promotion from shadow to selection; a local backend
only if the vendor dependency bites (logit read-out over the 9B fleet only if the vendor dependency bites (logit read-out over the 9B fleet