A review of what our agents can actually be asked to do, graded by evidence rather than by what the TOML declares. Recipes: research_and_code (13 harness scenarios) and research_only (staffing measured and corrected — 1 of 9 applicable skills under the old rust_sdlc default, 4 of 4 under topic_research) are proven. security_hardening, benchmark and refactor are exercised once each. continuous_research is the outlier: the most moving parts of any recipe, ran twice today, and NOTHING in the harness would notice if it broke. Dead keys the recipes lean on, from phase_config::DECLARED_BUT_UNREAD: loop, produces, input_from_phase, and mcp_bundles at phase level. Not hidden — security_hardening.toml annotates its own decoration inline, and the other five should copy that. The risk is a reader taking loop = "until_no_more_int_items" for a loop. Teams: all 12 resolve every declared skill (0 unbound, the skill-binding-repair work holding), but only rust_sdlc, topic_research and continuous_research have a run behind them. The other nine are well-formed scaffolding. Five are blocked on a target stack we do not have a repo for; four could be exercised against repositories we already have, and nobody currently knows whether they run at all. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01WZb5A2kfVfjpdwSochkuHz