From 17e6d08cd119e87724d06beac05bd19a504e27cf Mon Sep 17 00:00:00 2001 From: Omar Sobh Date: Tue, 22 Sep 2026 15:17:02 -0500 Subject: [PATCH] =?UTF-8?q?docs:=20papers=5Fresearch=20evidenced=20?= =?UTF-8?q?=E2=80=94=206=20of=2012,=20and=20its=20citations=20are=20real?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit First run of the papers_research team: a bounded library on microVM and sandbox isolation for agents, delivered as a README index plus four notes with complete frontmatter, judged met. The judge could only confirm the frontmatter was present. A paper team whose container has no pdf-to-text tool is exactly where invented citations appear, so every arXiv id was checked against the arXiv API: all four exist with exact title matches. One, 2603.02277 (SandboxEscapeBench), was already cited in the operator's own research pass — independent evidence of on-topic work, not plausible filler. It also confirmed ToolSearch by a real run, one of the allowlisted tools the task-permission shadow had not yet seen. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz --- docs/TEMPLATE-MATURITY.md | 17 ++++++++++++++--- 1 file changed, 14 insertions(+), 3 deletions(-) diff --git a/docs/TEMPLATE-MATURITY.md b/docs/TEMPLATE-MATURITY.md index b00f8b0..1e17a74 100644 --- a/docs/TEMPLATE-MATURITY.md +++ b/docs/TEMPLATE-MATURITY.md @@ -63,11 +63,11 @@ holding (it was 55 of 85 dangling once). | `gpu` | arch_analyst kernel_author bench_engineer coder committer | coding_readwrite | 13 | no | | `threejs` | scene_designer coder shader_author perf_engineer committer | coding_readwrite | 13 | no | | `codebase_research` | code_archeologist architecture_mapper flow_tracer vault_scribe | research_readonly | 11 | **yes** — first run 2026-09-22 | -| `papers_research` | domain_scout paper_reader library_curator | research_web_readonly | 8 | no | +| `papers_research` | domain_scout paper_reader library_curator | research_web_readonly | 8 | **yes** — first run 2026-09-22 | | `insight_research` | implementation_tracker novelty_hunter publication_drafter | research_readonly | 8 | no | | `continuous_improvement` | brain_inspector improvement_proposer improvement_evaluator | research_readonly | 7 | no | -**Five of twelve are evidenced.** `backend` was exercised on 2026-09-22, +**Six of twelve are evidenced.** `backend` was exercised on 2026-09-22, the first time anything had been staffed from it: all five roles provisioned with real agents (api_designer, db_engineer, coder, tester, committer), and the mission delivered cursor pagination — `Paged`, @@ -87,7 +87,18 @@ citations verify exactly** (`hook_script_with` 562, `NODE_EXTRACT` 789, `TaskPolicy` 297, `ROLE_POLICIES` 261, `settings_hook` 809, `install_command_with` 819). Nothing hallucinated. **Five of twelve.** -**Three of the remaining seven are still unevidenced.** The other nine are well-formed scaffolding: +`papers_research` followed: asked for a bounded library on microVM and +sandbox isolation for agents, it delivered a README index and four notes +with complete frontmatter, judged met. The judge could only check the +frontmatter was *present*; a paper team whose container has no +pdf-to-text tool is exactly where invented citations live, so every +arXiv id was checked against arXiv itself — **all four exist, with exact +title matches**. One of them, `2603.02277` (SandboxEscapeBench), is a paper +the operator's own research pass had already cited, which is independent +evidence it found on-topic work rather than plausible filler. **Six of +twelve.** + +**Two of the remaining six are still unevidenced.** The other nine are well-formed scaffolding: roles, prompts, brain seeds and resolving skills, and no run behind any of them. They will probably work — they are structurally identical to the three that do — but "probably" is the word, and this codebase has a name for the gap