Files
clawmates/skills/research/signal-to-noise-ranking.md
T
Omar SobhandClaude Opus 5 a2d7e3ea92
deploy / test (push) Successful in 4m4s
deploy / build (push) Successful in 1m45s
feat(skills): author the Continuous Research skills, and make its roles honest
The boot log has said `continuous_research — 4 role skills bound, 8 unresolved`
on every start. Those eight roles ran without the instructions their template
promised them, silently: `skills_loader` reports the miss and carries on.

Worse than the missing files was what the prompts described. v1 told the
harvester to sweep "RSS feeds, GitHub trending, HN front page, YouTube /
podcast RSS" — none of which exist. `harvest.rs` searches arXiv and nothing
else. A role prompt describing a machine that was never built is the failure
this codebase keeps paying for, so v2 describes the machine that exists.

Roles now match the pipeline: paper_reader (the harvest already ran; read the
manifest and the papers), signal_ranker (unchanged in spirit), script_writer
(the podcast half, which had no role at all).

Seven skills authored under skills/research/, kebab-case to match the loader —
team_template_loader.rs:177-181 documents the snake_case/kebab-case trap that
already unbinds skills elsewhere:

  arxiv-daily                  what the harvest guarantees, so an agent does
                               NOT re-search arXiv and corrupt the seen-set
  paper-to-project-relevance   name a file or roadmap item, or say "no bearing"
  duplicate-detection          the seen-set catches identity; this catches the
                               same work under a different id
  signal-to-noise-ranking      novelty/relevance/depth, and the two biases to
                               resist (recency up, inconvenience down)
  executive-summary-writing    what it is / why it matters / what to do — decide,
                               do not hedge
  obsidian-vault-conventions   the vault is a human's live workspace; never
                               main, never reorganise, hash the body not the file
  podcast-dialogue-writing     write for someone on a treadmill; the 10-70 char
                               highlight bound is the API's, not a style rule

`web_fetch` dropped from mcp_bundles: runtime_provision.rs binds every mission
claw to `["clawmates_door"]` and never reads that field, so declaring it
instructed roles to use a tool that never arrived. The prompts say `curl` via
Bash, which is what they actually have.

Boot now reports `continuous_research — 11 role skills bound`, with no
unresolved clause. 344 tests pass.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-17 14:44:03 -07:00

53 lines
2.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: signal-to-noise-ranking
description: Scoring harvested items on novelty, relevance and depth so a digest surfaces the few that matter.
when_to_use: You are ranking harvested research items before deciding which reach the digest or the podcast.
tags: [research, judgement]
---
# Ranking signal against noise
Score each item on three axes, 03. Sum them. The sum decides where it goes, and
the axes exist so that a high score has to be earned in more than one way.
## The axes
**Novelty (03)** — how different is this from what we have already seen on this
topic in the last 30 days? A fourth paper making the same claim is a 0 even if
the claim is important. Recency is not novelty: something published today that
restates last month's result is still a 0.
**Relevance (03)** — how directly does it connect to a named project or open
question? Not the field — the project. See `paper-to-project-relevance`. If you
cannot name the thing it touches, this is at most a 1.
**Depth (03)** — is this a primary source with an evaluation, or the fifth blog
post about someone else's paper? A primary source with a weak evaluation can
still beat a good summary of it, because you can check the former.
## Thresholds
```
>= 6 goes in the digest / the episode
>= 4 goes to watchlist.md — worth knowing, not worth airtime
< 4 dropped from the digest, KEPT in the raw manifest
```
Dropped is not deleted. The manifest is the audit trail: "we saw it and passed"
is a different and more useful record than silence, and it is what lets you
notice later that the ranking was wrong.
## Two failure modes to resist
- **Up-ranking the recent.** Time is not a quality signal. A preprint from this
morning gets no credit for being from this morning.
- **Down-ranking the inconvenient.** A paper that contradicts the approach we
are currently taking is *high* signal, not low. The temptation to score it
down is exactly why it needs airtime.
## Escape hatch
Anything the operator has explicitly flagged bypasses scoring entirely. A human
saying "look at this" outranks the rubric; the rubric exists for the other
several hundred items.