feat(papers): find papers on arXiv, shelve the PDF, catalogue the note
Corrects a misread of the design. I had built this as "read the vault to find papers"; the vault is the CARD CATALOGUE, not the source. Papers are found on arXiv, the PDF is pulled down and shelved in our own library, and a note recording it goes in the vault. Three parts, and which is which matters: arXiv — where papers are found blob store — the shelf; the PDF lives there (cm-files, local + S3) the vault — the catalogue; one note per paper, pointing at the shelf The checkmark list (corpus, 0064) is what makes this continuous rather than a job that redoes itself every week — the failure that killed the previous attempt (0030-0044, dropped in 0053). The load-bearing detail: every catalogue note carries `source_id: arxiv:NNNN.NNNNN` in frontmatter, which is exactly the key corpus::parse_note reads. So the checkmark list is rebuildable FROM the vault. If the database were lost, re-indexing restores what we have — the catalogue is authoritative, the index is derived. A test asserts that round trip rather than trusting the two halves to agree. Version suffixes are stripped (2401.12345v3 -> 2401.12345) or a weekly job re-downloads a paper every time authors post a revision. Fetches are rejected unless the bytes start with %PDF: arXiv serves an HTML holding page while a PDF renders, and shelving that leaves a file that looks present and is unreadable. Verified against live arXiv, not fixtures: arxiv:2607.29678 TokTier: Exact Stateful Tokenization for Agentic LLM… arxiv:2607.29677 ExtractBench: A Benchmark for Schema-Guided Enterpri… arxiv:2607.29658 Reusing Past Repairs Through Hierarchical Trajectory… pdf: 1,361,770 bytes, %PDF verified 388 tests, clippy clean. Co-Authored-By: Claude Opus 5 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
6e5ccc25a6
commit
e4a395b72e
@@ -18,6 +18,7 @@ pub mod mission_orchestrator;
|
|||||||
pub mod mission_refiner;
|
pub mod mission_refiner;
|
||||||
pub mod corpus;
|
pub mod corpus;
|
||||||
pub mod mission_delivery;
|
pub mod mission_delivery;
|
||||||
|
pub mod papers;
|
||||||
pub mod phase_config;
|
pub mod phase_config;
|
||||||
pub mod runtime_preflight;
|
pub mod runtime_preflight;
|
||||||
pub mod mission_runtime;
|
pub mod mission_runtime;
|
||||||
|
|||||||
@@ -0,0 +1,345 @@
|
|||||||
|
//! Finding papers, shelving them, and cataloguing them.
|
||||||
|
//!
|
||||||
|
//! The library has three parts and it matters which is which:
|
||||||
|
//!
|
||||||
|
//! - **arXiv** is where papers are *found*.
|
||||||
|
//! - **The blob store** is the *shelf* — the PDF itself lives there.
|
||||||
|
//! - **The vault** is the *card catalogue* — a markdown note per paper, with
|
||||||
|
//! the metadata and a pointer to the shelf.
|
||||||
|
//!
|
||||||
|
//! Plus [`crate::corpus`], which is the list of checkmarks: it is what stops
|
||||||
|
//! the same paper being fetched twice across weekly runs. That list is the
|
||||||
|
//! reason this can be a *continuous* job rather than one that redoes itself
|
||||||
|
//! forever — the failure that killed the previous attempt at this (migrations
|
||||||
|
//! 0030-0044, dropped in 0053).
|
||||||
|
//!
|
||||||
|
//! # The contract that ties it together
|
||||||
|
//!
|
||||||
|
//! Every note this module writes carries `source_id: arxiv:NNNN.NNNNN` in its
|
||||||
|
//! frontmatter. `corpus::parse_note` reads exactly that key, so re-indexing
|
||||||
|
//! the vault re-derives the checkmark list from the notes themselves. The
|
||||||
|
//! catalogue is authoritative; the index is rebuildable from it. If the
|
||||||
|
//! database were lost, a re-index of the vault would restore what we have.
|
||||||
|
|
||||||
|
use serde::{Deserialize, Serialize};
|
||||||
|
|
||||||
|
/// One paper as arXiv describes it.
|
||||||
|
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
|
||||||
|
pub struct Paper {
|
||||||
|
/// Bare arXiv id, e.g. `2401.12345` — no version suffix.
|
||||||
|
pub arxiv_id: String,
|
||||||
|
pub title: String,
|
||||||
|
pub authors: Vec<String>,
|
||||||
|
pub summary: String,
|
||||||
|
pub published: String,
|
||||||
|
pub pdf_url: String,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Paper {
|
||||||
|
/// The checkmark key. Version suffixes are stripped upstream so `v1` and
|
||||||
|
/// `v2` of the same paper are one entry, not two.
|
||||||
|
pub fn source_id(&self) -> String {
|
||||||
|
format!("arxiv:{}", self.arxiv_id)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Where the PDF is shelved in the blob store.
|
||||||
|
pub fn blob_key(&self) -> String {
|
||||||
|
format!("papers/arxiv/{}.pdf", self.arxiv_id)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Where the catalogue note goes in the vault.
|
||||||
|
///
|
||||||
|
/// Under a dedicated folder so the library never collides with the
|
||||||
|
/// hand-written parts of the vault (`30 Resources`, `40 Projects`, and so
|
||||||
|
/// on). A human should always be able to tell which notes a machine wrote.
|
||||||
|
pub fn note_path(&self) -> String {
|
||||||
|
format!("60 Papers/arxiv-{}.md", self.arxiv_id)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Strip an arXiv version suffix: `2401.12345v3` -> `2401.12345`.
|
||||||
|
///
|
||||||
|
/// Without this a weekly job re-downloads a paper every time the authors post
|
||||||
|
/// a revision, and the checkmark list quietly fills with near-duplicates.
|
||||||
|
pub fn normalize_arxiv_id(raw: &str) -> String {
|
||||||
|
let id = raw.rsplit('/').next().unwrap_or(raw);
|
||||||
|
match id.find('v') {
|
||||||
|
// Only a trailing `vN` counts; the `v` in a word must not truncate.
|
||||||
|
Some(i) if id[i + 1..].chars().all(|c| c.is_ascii_digit()) && i + 1 < id.len() => {
|
||||||
|
id[..i].to_string()
|
||||||
|
}
|
||||||
|
_ => id.to_string(),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Parse arXiv's Atom feed.
|
||||||
|
///
|
||||||
|
/// Hand-rolled rather than pulling an XML crate: the feed is a fixed, simple
|
||||||
|
/// shape and this reads five fields from it. If arXiv's format ever drifts,
|
||||||
|
/// `entries_are_parsed_from_a_real_feed` fails loudly rather than silently
|
||||||
|
/// returning zero papers — which is the failure mode that matters, because a
|
||||||
|
/// search returning nothing looks exactly like "no new papers this week".
|
||||||
|
pub fn parse_atom(xml: &str) -> Vec<Paper> {
|
||||||
|
let mut out = Vec::new();
|
||||||
|
for chunk in xml.split("<entry>").skip(1) {
|
||||||
|
let entry = chunk.split("</entry>").next().unwrap_or(chunk);
|
||||||
|
let field = |tag: &str| -> Option<String> {
|
||||||
|
let open = format!("<{tag}>");
|
||||||
|
let close = format!("</{tag}>");
|
||||||
|
let start = entry.find(&open)? + open.len();
|
||||||
|
let end = entry[start..].find(&close)? + start;
|
||||||
|
Some(unescape(entry[start..end].trim()))
|
||||||
|
};
|
||||||
|
|
||||||
|
let Some(raw_id) = field("id") else { continue };
|
||||||
|
let arxiv_id = normalize_arxiv_id(&raw_id);
|
||||||
|
if arxiv_id.is_empty() {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let Some(title) = field("title") else { continue };
|
||||||
|
|
||||||
|
let authors = entry
|
||||||
|
.split("<author>")
|
||||||
|
.skip(1)
|
||||||
|
.filter_map(|a| {
|
||||||
|
let start = a.find("<name>")? + 6;
|
||||||
|
let end = a[start..].find("</name>")? + start;
|
||||||
|
Some(unescape(a[start..end].trim()))
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
|
||||||
|
// The PDF link is an attribute, not an element.
|
||||||
|
let pdf_url = entry
|
||||||
|
.split("<link")
|
||||||
|
.find(|l| l.contains("title=\"pdf\""))
|
||||||
|
.and_then(|l| {
|
||||||
|
let start = l.find("href=\"")? + 6;
|
||||||
|
let end = l[start..].find('"')? + start;
|
||||||
|
Some(l[start..end].to_string())
|
||||||
|
})
|
||||||
|
.unwrap_or_else(|| format!("https://arxiv.org/pdf/{arxiv_id}"));
|
||||||
|
|
||||||
|
out.push(Paper {
|
||||||
|
title: title.split_whitespace().collect::<Vec<_>>().join(" "),
|
||||||
|
summary: field("summary")
|
||||||
|
.unwrap_or_default()
|
||||||
|
.split_whitespace()
|
||||||
|
.collect::<Vec<_>>()
|
||||||
|
.join(" "),
|
||||||
|
published: field("published").unwrap_or_default(),
|
||||||
|
authors,
|
||||||
|
pdf_url,
|
||||||
|
arxiv_id,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
out
|
||||||
|
}
|
||||||
|
|
||||||
|
fn unescape(s: &str) -> String {
|
||||||
|
s.replace("&", "&")
|
||||||
|
.replace("<", "<")
|
||||||
|
.replace(">", ">")
|
||||||
|
.replace(""", "\"")
|
||||||
|
.replace("'", "'")
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Search arXiv. `max_results` is capped to keep one run bounded.
|
||||||
|
pub async fn search(query: &str, max_results: usize) -> Result<Vec<Paper>, String> {
|
||||||
|
let max = max_results.clamp(1, 50);
|
||||||
|
let url = format!(
|
||||||
|
"https://export.arxiv.org/api/query?search_query={}&start=0&max_results={max}\
|
||||||
|
&sortBy=submittedDate&sortOrder=descending",
|
||||||
|
urlencoding(query)
|
||||||
|
);
|
||||||
|
let body = reqwest::Client::new()
|
||||||
|
.get(&url)
|
||||||
|
.header("User-Agent", "clawmates-papers/0.1 (research library)")
|
||||||
|
.timeout(std::time::Duration::from_secs(60))
|
||||||
|
.send()
|
||||||
|
.await
|
||||||
|
.map_err(|e| format!("arxiv query: {e}"))?
|
||||||
|
.text()
|
||||||
|
.await
|
||||||
|
.map_err(|e| format!("arxiv body: {e}"))?;
|
||||||
|
Ok(parse_atom(&body))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Download the PDF. Returns the bytes; the caller decides where to shelve it.
|
||||||
|
pub async fn fetch_pdf(paper: &Paper) -> Result<Vec<u8>, String> {
|
||||||
|
let bytes = reqwest::Client::new()
|
||||||
|
.get(&paper.pdf_url)
|
||||||
|
.header("User-Agent", "clawmates-papers/0.1 (research library)")
|
||||||
|
.timeout(std::time::Duration::from_secs(180))
|
||||||
|
.send()
|
||||||
|
.await
|
||||||
|
.map_err(|e| format!("fetch pdf {}: {e}", paper.arxiv_id))?
|
||||||
|
.bytes()
|
||||||
|
.await
|
||||||
|
.map_err(|e| format!("read pdf {}: {e}", paper.arxiv_id))?;
|
||||||
|
|
||||||
|
// A PDF starts with `%PDF`. arXiv serves an HTML holding page when a PDF
|
||||||
|
// is still rendering, and shelving that would leave a file that looks
|
||||||
|
// present and is unreadable.
|
||||||
|
if !bytes.starts_with(b"%PDF") {
|
||||||
|
return Err(format!(
|
||||||
|
"{} did not return a PDF ({} bytes, starts {:?})",
|
||||||
|
paper.pdf_url,
|
||||||
|
bytes.len(),
|
||||||
|
String::from_utf8_lossy(&bytes[..bytes.len().min(16)])
|
||||||
|
));
|
||||||
|
}
|
||||||
|
Ok(bytes.to_vec())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The catalogue note for a shelved paper.
|
||||||
|
///
|
||||||
|
/// `source_id` in the frontmatter is the load-bearing part — it is what
|
||||||
|
/// `corpus::parse_note` reads to rebuild the checkmark list from the vault.
|
||||||
|
pub fn catalogue_note(paper: &Paper, blob_key: &str) -> String {
|
||||||
|
let authors = if paper.authors.is_empty() {
|
||||||
|
"unknown".to_string()
|
||||||
|
} else {
|
||||||
|
paper.authors.join(", ")
|
||||||
|
};
|
||||||
|
format!(
|
||||||
|
"---\n\
|
||||||
|
source_id: arxiv:{id}\n\
|
||||||
|
arxiv: {id}\n\
|
||||||
|
title: \"{title}\"\n\
|
||||||
|
authors: \"{authors}\"\n\
|
||||||
|
published: {published}\n\
|
||||||
|
pdf: {blob_key}\n\
|
||||||
|
url: https://arxiv.org/abs/{id}\n\
|
||||||
|
added: {added}\n\
|
||||||
|
tags: [paper, arxiv]\n\
|
||||||
|
---\n\
|
||||||
|
\n\
|
||||||
|
# {title}\n\
|
||||||
|
\n\
|
||||||
|
**Authors:** {authors} \n\
|
||||||
|
**arXiv:** [{id}](https://arxiv.org/abs/{id}) \n\
|
||||||
|
**PDF:** `{blob_key}`\n\
|
||||||
|
\n\
|
||||||
|
## Abstract\n\
|
||||||
|
\n\
|
||||||
|
{summary}\n\
|
||||||
|
\n\
|
||||||
|
## Notes\n\
|
||||||
|
\n\
|
||||||
|
_Catalogued automatically. Add your own notes below._\n",
|
||||||
|
id = paper.arxiv_id,
|
||||||
|
title = paper.title.replace('"', "'"),
|
||||||
|
authors = authors,
|
||||||
|
published = paper.published,
|
||||||
|
blob_key = blob_key,
|
||||||
|
added = paper.published,
|
||||||
|
summary = paper.summary,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn urlencoding(s: &str) -> String {
|
||||||
|
s.bytes()
|
||||||
|
.map(|b| match b {
|
||||||
|
b'A'..=b'Z' | b'a'..=b'z' | b'0'..=b'9' | b'-' | b'_' | b'.' | b'~' => {
|
||||||
|
(b as char).to_string()
|
||||||
|
}
|
||||||
|
b' ' => "+".to_string(),
|
||||||
|
_ => format!("%{b:02X}"),
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
#[cfg(test)]
|
||||||
|
mod tests {
|
||||||
|
use super::*;
|
||||||
|
|
||||||
|
/// A revision must not read as a new paper.
|
||||||
|
#[test]
|
||||||
|
fn version_suffixes_are_stripped() {
|
||||||
|
assert_eq!(normalize_arxiv_id("http://arxiv.org/abs/2401.12345v3"), "2401.12345");
|
||||||
|
assert_eq!(normalize_arxiv_id("2401.12345v1"), "2401.12345");
|
||||||
|
assert_eq!(normalize_arxiv_id("2401.12345"), "2401.12345");
|
||||||
|
// Old-style ids contain letters and a slash.
|
||||||
|
assert_eq!(normalize_arxiv_id("http://arxiv.org/abs/cs/0701001"), "0701001");
|
||||||
|
// A trailing `v` with no digits is part of the id, not a version.
|
||||||
|
assert_eq!(normalize_arxiv_id("2401.1234v"), "2401.1234v");
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Parsed against the real shape of arXiv's Atom feed. If this fails the
|
||||||
|
/// format drifted — which otherwise shows up as "no new papers", which is
|
||||||
|
/// indistinguishable from a quiet week.
|
||||||
|
#[test]
|
||||||
|
fn entries_are_parsed_from_a_real_feed() {
|
||||||
|
let xml = r#"<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<feed xmlns="http://www.w3.org/2005/Atom">
|
||||||
|
<entry>
|
||||||
|
<id>http://arxiv.org/abs/2401.12345v2</id>
|
||||||
|
<published>2026-01-15T10:00:00Z</published>
|
||||||
|
<title>Attention Is All You Need Again</title>
|
||||||
|
<summary> We show that
|
||||||
|
attention still works. </summary>
|
||||||
|
<author><name>Ada Lovelace</name></author>
|
||||||
|
<author><name>Alan Turing</name></author>
|
||||||
|
<link href="http://arxiv.org/abs/2401.12345v2" rel="alternate" type="text/html"/>
|
||||||
|
<link title="pdf" href="http://arxiv.org/pdf/2401.12345v2" rel="related" type="application/pdf"/>
|
||||||
|
</entry>
|
||||||
|
</feed>"#;
|
||||||
|
let papers = parse_atom(xml);
|
||||||
|
assert_eq!(papers.len(), 1);
|
||||||
|
let p = &papers[0];
|
||||||
|
assert_eq!(p.arxiv_id, "2401.12345", "version stripped");
|
||||||
|
assert_eq!(p.title, "Attention Is All You Need Again", "whitespace collapsed");
|
||||||
|
assert_eq!(p.summary, "We show that attention still works.");
|
||||||
|
assert_eq!(p.authors, vec!["Ada Lovelace", "Alan Turing"]);
|
||||||
|
assert_eq!(p.pdf_url, "http://arxiv.org/pdf/2401.12345v2");
|
||||||
|
assert_eq!(p.source_id(), "arxiv:2401.12345");
|
||||||
|
assert_eq!(p.blob_key(), "papers/arxiv/2401.12345.pdf");
|
||||||
|
assert_eq!(p.note_path(), "60 Papers/arxiv-2401.12345.md");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn an_empty_feed_yields_no_papers_rather_than_panicking() {
|
||||||
|
assert!(parse_atom("<feed></feed>").is_empty());
|
||||||
|
assert!(parse_atom("").is_empty());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn xml_entities_are_unescaped() {
|
||||||
|
let xml = r#"<feed><entry><id>http://arxiv.org/abs/1v1</id>
|
||||||
|
<title>Cats & Dogs <3</title><summary>a "quote"</summary>
|
||||||
|
</entry></feed>"#;
|
||||||
|
let p = &parse_atom(xml)[0];
|
||||||
|
assert_eq!(p.title, "Cats & Dogs <3");
|
||||||
|
assert_eq!(p.summary, "a \"quote\"");
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The note must carry the identity `corpus::parse_note` reads, or the
|
||||||
|
/// catalogue cannot rebuild the checkmark list and the library forgets
|
||||||
|
/// itself the moment the database is lost.
|
||||||
|
#[test]
|
||||||
|
fn a_catalogue_note_round_trips_through_the_corpus_parser() {
|
||||||
|
let paper = Paper {
|
||||||
|
arxiv_id: "2401.12345".into(),
|
||||||
|
title: "A \"Quoted\" Title".into(),
|
||||||
|
authors: vec!["Ada Lovelace".into()],
|
||||||
|
summary: "Summary text.".into(),
|
||||||
|
published: "2026-01-15T10:00:00Z".into(),
|
||||||
|
pdf_url: "http://arxiv.org/pdf/2401.12345".into(),
|
||||||
|
};
|
||||||
|
let note = catalogue_note(&paper, &paper.blob_key());
|
||||||
|
|
||||||
|
let parsed = crate::corpus::parse_note(&paper.note_path(), ¬e);
|
||||||
|
assert_eq!(
|
||||||
|
parsed.declared_source_id.as_deref(),
|
||||||
|
Some("arxiv:2401.12345"),
|
||||||
|
"the corpus parser must recover the identity from the note"
|
||||||
|
);
|
||||||
|
assert_eq!(parsed.title.as_deref(), Some("A 'Quoted' Title"));
|
||||||
|
assert!(note.contains("papers/arxiv/2401.12345.pdf"), "note points at the shelf");
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn queries_are_url_encoded() {
|
||||||
|
assert_eq!(urlencoding("all:agent topologies"), "all%3Aagent+topologies");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -199,3 +199,30 @@ async fn index_the_real_vault() {
|
|||||||
assert_eq!(b.updated, 0);
|
assert_eq!(b.updated, 0);
|
||||||
assert_eq!(b.unchanged, a.scanned);
|
assert_eq!(b.unchanged, a.scanned);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Live arXiv check. Ignored by default (needs network); run with
|
||||||
|
/// `cargo test -p cm-api --test corpus_vault live_arxiv -- --ignored --nocapture`.
|
||||||
|
///
|
||||||
|
/// Guards the one failure that hides: if arXiv's feed format drifts, parsing
|
||||||
|
/// returns zero papers, which looks exactly like "no new papers this week".
|
||||||
|
#[tokio::test]
|
||||||
|
#[ignore]
|
||||||
|
async fn live_arxiv_search_and_fetch() {
|
||||||
|
let papers = cm_api::papers::search("all:agentic topologies", 3)
|
||||||
|
.await
|
||||||
|
.expect("arxiv search");
|
||||||
|
println!("found {} papers", papers.len());
|
||||||
|
assert!(!papers.is_empty(), "arXiv returned nothing — format drift?");
|
||||||
|
|
||||||
|
for p in &papers {
|
||||||
|
println!(" {} | {}", p.source_id(), &p.title[..p.title.len().min(60)]);
|
||||||
|
assert!(!p.arxiv_id.is_empty());
|
||||||
|
assert!(!p.title.is_empty());
|
||||||
|
assert!(!p.arxiv_id.contains('v'), "version must be stripped: {}", p.arxiv_id);
|
||||||
|
}
|
||||||
|
|
||||||
|
let pdf = cm_api::papers::fetch_pdf(&papers[0]).await.expect("fetch pdf");
|
||||||
|
println!("pdf bytes: {}", pdf.len());
|
||||||
|
assert!(pdf.starts_with(b"%PDF"));
|
||||||
|
assert!(pdf.len() > 10_000, "suspiciously small pdf: {}", pdf.len());
|
||||||
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user