Files
clawhdf5/docs/agent-memory.md
T
osobhandClaude Opus 5.5 4d1a43fd7a docs: agent memory guide in docs/agent-memory.md
The agent-memory detail that lived only in README.md (architecture,
modules, performance and footprint tables, LongMemEval, feature flags and
settings, file schema, SQLite migration, research foundation) moves to its
own page, so the README can lead with the HDF5 library. Code examples are
updated to the current API (MemoryConfig::new takes a PathBuf,
HDF5Memory::search with SearchOptions, consolidation with timestamps) and
were compiled and run against the workspace; the CLI section was run
against the built `clawhdf5` binary. New: the CLI's search defaults to
0.7/0.3, not the library's 0.4/0.6; the /integrity group of signed stores.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:10:22 -05:00

25 KiB
Raw Blame History

Agent memory (clawhdf5-agent)

clawhdf5-agent is a persistent, searchable memory store for AI agents, built on clawhdf5's HDF5 writer: records (text, embedding, source channel, timestamp, session, tags), sessions and a knowledge graph in one .h5 file, with a write-ahead log beside it. This page is the long form of the agent part of the README; every number on it comes from BENCHMARKS.md, where the commands and machines are.

Integration status: ClawBrainHub's CLI uses this crate's bm25::BM25Index; no agent framework uses the store. clawhdf5 is not an OpenClaw memory plugin (openclaw.md), and ZeroClaw does not use it.

Quick start

[dependencies]
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }  # not on crates.io yet
use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions};

// A new store: 384-dim embeddings (float16 on disk and an int8 HNSW index by default).
let mut memory = HDF5Memory::create(MemoryConfig::new("agent.h5".into(), "my-agent", 384))?;

memory.save(MemoryEntry {
    chunk: "User prefers dark mode and vim keybindings.".into(),
    embedding: embed("User prefers dark mode and vim keybindings."), // your embedder
    source_channel: "chat".into(),
    timestamp: now,
    session_id: "session-001".into(),
    tags: "preference".into(),
})?;

// Hybrid search: HNSW vector + BM25 keyword, fused 0.4 / 0.6 (the measured default).
let query = embed("what editor does the user like?");
for r in memory.search(&query, "editor preferences", &SearchOptions::new(5)) {
    println!("[{:.3}] {}", r.score, r.chunk);
}
memory.flush_wal()?; // checkpoint the WAL into agent.h5

clawhdf5 stores embeddings; it does not compute them. Any dimension works, fixed when the store is created. HDF5Memory::open(path) reopens a store (holding its single-writer lock); HDF5Memory::open_read_only(path) gives a lock-free point-in-time view.

HDF5Memory::search(query_emb, text, &SearchOptions) is the full search path; hybrid_search(query_emb, text, vector_weight, keyword_weight, k) and hybrid_search_with are thin wrappers over it.

use clawhdf5_agent::confidence::ConfidenceConfig;
use clawhdf5_agent::reranker::ReRankConfig;

// Only memories from these source channels; still a full page of k results.
let work = memory.search(&query, "deadline", &SearchOptions::new(5).with_sources(["slack", "email"]));
// Re-rank (relevance, recency, source authority, activation), then drop
// low-confidence results: the pipeline ClawhdfBackend runs.
let careful = memory.search(
    &query,
    "user preferences",
    &SearchOptions::new(5)
        .with_rerank(ReRankConfig::default())
        .with_confidence(ConfidenceConfig::default()),
);

The source-channel filter is applied before ranking: an exact scan of the allowed records whenever that is cheaper than the index would be, and as the fallback when the index returns a short pool. Hebbian activation boosts are persisted by the next checkpoint (or on drop), not per query; search never writes the store.

Signed checkpoints

use clawhdf5_agent::signing;

let key = signing::generate_key();          // keep the secret key; publish the public one
let public = key.verifying_key();
memory.set_signing_key(key);                 // never written to disk
memory.flush_wal()?;                         // this checkpoint is signed
let report = HDF5Memory::verify(std::path::Path::new("agent.h5"), &public)?;
assert!(report.is_valid());                  // report.changed_records names edited records

The Ed25519 signature covers every record (text, embedding as stored, channel, timestamp, session, tags, deleted flag, activation) through a SHA-256 Merkle tree, plus the store's settings, sessions and knowledge graph, so a change made with any tool is caught and located. It covers checkpoints, not saves still in the WAL (report.wal_entries_unsigned counts those). A signed store refuses to checkpoint without the key (MemoryError::SigningKeyRequired). CLI: clawhdf5 keygen, --signing-key <file> on writing commands, and verify --public-key. Signing adds about 20% to a checkpoint and 32 bytes per record to the file (BENCHMARKS.md § Signed checkpoints).

Architecture

                         ┌─────────────────┐
                         │   Agent Query   │
                         └────────┬────────┘
                                  │
                ┌─────────────────▼──────────────────┐
                │ HDF5Memory::search                 │
                │ optional source-channel filter     │
                │ HNSW vector + BM25 keyword         │
                │ weighted fusion (0.4 / 0.6)        │
                │ × √(Hebbian activation)            │
                └─────────────────┬──────────────────┘
                                  │  opt-in (SearchOptions);
                                  │  ClawhdfBackend turns both on
                ┌─────────────────▼──────────────────┐
                │ Multi-factor re-ranking            │
                │ relevance · recency · authority ·  │
                │ activation                         │
                ├────────────────────────────────────┤
                │ Confidence rejection               │
                │ (suppress bad matches)             │
                └─────────────────┬──────────────────┘
                                  │
     ┌────────────────────────────▼────────────────────────────┐
     │ In memory                                               │
     │  cache (embeddings) · BM25 index · HNSW index           │
     │  provenance ledger + anomaly alerts (session-scoped)    │
     └────────────────────────────┬────────────────────────────┘
                                  │ WAL append; checkpoint
     ┌────────────────────────────▼────────────────────────────┐
     │ agent_memory.h5   /meta · /memory · /sessions ·         │
     │                   /knowledge_graph                      │
     │ agent_memory.h5.wal   chained-CRC write-ahead log       │
     │ agent_memory.h5.ann   HNSW graph (derived, rebuildable) │
     │ agent_memory.h5.lock  single-writer lock                │
     └─────────────────────────────────────────────────────────┘

Durability. Every WAL entry carries a CRC32 chained to the previous entry's, so a corrupted, reordered, duplicated or spliced entry stops replay instead of loading bad data. Each checkpoint records a WAL mark in /meta, so a crash between a checkpoint and the WAL truncate never applies an entry twice. Checkpoints and snapshots are made durable as a unit (temp file synced, renamed, directory synced). Individual WAL appends are not fsynced (a latency trade-off): saves since the last checkpoint can be lost on power failure or a kernel panic, not on a process crash. An unreadable WAL is quarantined to <store>.h5.wal.corrupt-<ts> rather than blocking open().

Single writer. create/open take an exclusive advisory lock on <store>.h5.lock; a second opener gets MemoryError::Locked.

Write bookkeeping. save/save_batch/save_or_update run each write through an in-memory (session-scoped, not persisted) provenance ledger — an unkeyed content hash per record, for detecting accidental corruption, not tampering — and a write-anomaly detector (rate limits, injection patterns, source distribution). Alerts never block a save; drain them with take_anomaly_alerts. The source classification is inferred from the caller's source_channel string, a heuristic, not an authenticated trust boundary.

Modules

Module What it does
hybrid Vector + BM25 fusion: min-max-normalised weighted sum, vector 0.4 / keyword 0.6 by default (hybrid::DEFAULT_FUSION, tuned on LongMemEval); RRF via Fusion::Rrf / hybrid_search_with (measured worse)
reranker Re-ranking by retrieval relevance (leads, weight 1.0), recency, source authority, activation. Opt-in via SearchOptions::with_rerank; on in ClawhdfBackend
confidence Low-confidence rejection. Opt-in via SearchOptions::with_confidence; on in ClawhdfBackend
bm25 Incremental Okapi BM25 index kept for the life of the store; optional stemming
signing Ed25519-signed checkpoints (above)
wal Write-ahead log, format v4, chained CRC32 per entry; reads v2 and v3 (v1 only through the one-time migration in open)
knowledge Entity/relation graph: BFS, spreading activation, fuzzy (Levenshtein) entity resolution
consolidation Three tiers (Working → Episodic → Semantic): importance, novelty, time decay
temporal Sorted timestamp index, session DAG, entity timeline
multimodal Cross-modal search over text/image/audio/video embeddings (exact scan)
provenance, anomaly Session-scoped write bookkeeping (above)
openclaw ClawhdfBackend, a Markdown-oriented backend (below). Named for OpenClaw, but not an OpenClaw plugin (openclaw.md)
vector_search Flat cosine search paths: pre-normed, SIMD, BLAS, GPU, parallel
ivf / pq Standalone IVF and IVF-PQ indexes; not used by HDF5Memory, whose index is HNSW
query_expand, entity_extract Synonym/acronym/temporal query expansion; rule-based entity extraction into the graph
memory_strategy, decision_gate When to save: save-every, semantic shift, user correction; trivial/substantive classification
ephemeral In-memory TTL/LFU working tier
async_memory Tokio wrapper over the store (async feature)

Library components

The consolidation tiers, the graph algorithms and the temporal and multi-modal indexes are components you drive directly; the store persists the records, sessions and graph they work over.

use clawhdf5_agent::knowledge::KnowledgeCache;

let mut kg = KnowledgeCache::new();
let alice = kg.add_entity("Alice", "person", -1);
let bob = kg.add_entity("Bob", "person", -1);
let acme = kg.add_entity("Acme Corp", "company", -1);
kg.add_relation(alice, acme, "works_at", 1.0);
kg.add_relation(alice, bob, "manages", 0.8);

let neighbors = kg.bfs_neighbors(alice, 2);                     // 2-hop neighbourhood
let activated = kg.spreading_activation(&[alice], 0.5, 0.01, 5); // related entities
let (id, created) = kg.resolve_or_create("alice", "person", -1, 2); // fuzzy (Levenshtein <= 2)
assert_eq!((id, created), (alice, false));
use clawhdf5_agent::consolidation::{ConsolidationConfig, ConsolidationEngine, UntrustedSource};

let mut engine = ConsolidationEngine::new(ConsolidationConfig {
    working_capacity: 100,
    ..Default::default()
});
let id = engine.add_memory("User prefers dark mode".into(), embed("dark mode"), UntrustedSource::User, now);
engine.access_memory(id, now + 60.0);   // reactivates it
engine.consolidate(now + 3600.0);        // promote (Working -> Episodic -> Semantic) and evict
let stats = engine.get_stats();
println!("working {} episodic {} semantic {}", stats.working_count, stats.episodic_count, stats.semantic_count);

System and correction sources get elevated importance and go through a separate entry point, add_trusted_memory(.., TrustedSource::System, ..), so untrusted content cannot claim them.

use clawhdf5_agent::temporal::TemporalIndex;

let mut index = TemporalIndex::new();
index.insert(1, 1_700_000_000.0);
index.insert(2, 1_700_003_600.0);                               // an hour later
let in_range = index.range_query(1_700_000_000.0, 1_700_010_800.0);
let recent = index.latest(10);

Markdown backend

ClawhdfBackend ingests Markdown by section and searches it with the full pipeline. It is a library API, not an OpenClaw plugin.

use clawhdf5_agent::openclaw::{ClawhdfBackend, MemoryBackend};

let mut backend = ClawhdfBackend::create(std::path::Path::new("memory.h5"), 384)?;
let md = std::fs::read_to_string("MEMORY.md")?;
let sections = backend.ingest_markdown("MEMORY.md", &md)?;   // one record per heading
for r in backend.search("dark mode", &embed("dark mode"), 5) {
    println!("[{:.3}] {} ({})", r.score, r.text, r.path);
}
let exported = backend.export_markdown("MEMORY.md")?;

Limits: ingested sections carry no embedding, so their search is keyword-only unless you save records with vectors through save_entry; ingesting a file again adds its sections again; export_markdown writes every heading as ##, so it is not a lossless round trip.

Performance

Unless marked otherwise, measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D, 8C/16T), commit 5c8323c, 384-dim embeddings; commands in BENCHMARKS.md.

HNSW (the default vector stage) — search_harness, clustered data, N = 100K, M = 16, ef_construction = 64, ef = 64, recall against an exact scan (§ Quantising the index copy):

index recall@10 QPS build
f32 0.9945 13 399 3.2 s
i8 + exact re-score (default for new stores) 0.9940 21 848 1.8 s

A paired comparison (medians of alternating runs, same binary), not re-run on 2026-09-24: a single f32 run that day measured recall 0.9945, 19 001 QPS and a 2.7 s build, so the 1.63x ratio has not been re-checked. On a Raspberry Pi 5 (NEON SDOT) the int8 index is 1.18x the f32 QPS at equal recall. Before the v2.4.0 neighbour-selection fix, recall@10 at 100K was 0.31.

Operations:

Operation Latency Scale
hybrid_search p50 0.07 ms / 0.49 ms / 4.69 ms 1K / 10K / 100K records
BM25 keyword search 20.4 µs 1K records
Knowledge graph BFS 23.1 µs 1K entities
Spreading activation 10.1 µs 100 entities
Temporal range query 622 ns 10K timestamps
Consolidation cycle 115.2 µs 1K records
Cross-modal search (exact scan, 2 embeddings per record) 842.0 µs / 8.44 ms 1K / 10K records
Memory write (WAL append) 26.1 µs per record

float16 stores (the default) add about 2 µs per write for rounding (§ Write Path).

Brute-force and IVF (Criterion; not used by HDF5Memory):

Scale Flat IVF (nprobe=10) IVF-PQ
1K 47.4 µs — —
10K 500.5 µs 24.8 µs —
100K 6.58 ms 592 µs 869 µs

No comparison with MemX is made: its published figure is end-to-end and ours is one component (BENCHMARKS.md).

Consolidation — 1,000 records (10 signal + 990 noise), working_capacity = 100: the store goes from 1,000 to 100 records with Hit@1 on the signal records staying at 100%, and search from 2.22 ms to 0.24 ms (§ Consolidation Efficiency).

LongMemEval retrieval recall

Full longmemeval_s haystack, all 500 questions (47.7 sessions and 493.5 turns each; 4.0% of sessions are evidence), real all-MiniLM-L6-v2 embeddings, k = 10. Re-run 2026-09-27 on tank; the headline reproduced exactly (§ LongMemEval Results):

Mode Turn-level Hit@5 Session-level Hit@5
BM25 only 75.0% 93.6%
Vector only (MiniLM) 71.8% 94.2%
Hybrid 0.4 / 0.6 (default) 81.4% 96.8%

This is retrieval recall (did a gold turn appear in the top k), not the official LongMemEval QA accuracy; the two are not comparable. A weight sweep found the old 0.7 / 0.3 default strictly dominated by 0.4 / 0.6, the default since v2.5.0; use 0.3 / 0.7 if rank-1 precision matters most. Earlier session-level figures of 100% and a claimed win over MemX were retracted (BENCHMARKS.md). The benchmark's vector stage needs clawhdf5-bench's embeddings feature.

Memory footprint

On disk — float16 embeddings (the default), 200-character synthetic text, footprint_bench: 810.4 KB at 1K records, 7.8 MB at 10K, 76.7 MB at 100K (803–829 bytes per record). The synthetic text is far more repetitive than real text (40 distinct strings, deflated), so real records will be larger; the embeddings alone are 768 B per record. On the same data, 100K × 384 takes 80.8 MiB as float16 and 154.0 MiB as f32 (§ Memory Footprint).

In memory — a store reopened from disk, counting allocator (§ Memory footprint):

Records Raw vectors f32 index i8 index (default)
1K 1 MiB 4 MiB (2.40x) 2 MiB (1.64x)
10K 15 MiB 44 MiB (3.03x) 27 MiB (1.81x)
100K 146 MiB 399 MiB (2.72x) 256 MiB (1.74x)

The f32 column was re-measured on 2026-09-24; the i8 column was not.

Feature flags and settings

clawhdf5-agent flag Default Description
float16 yes Half-precision cosine kernel. Half-precision storage is the MemoryConfig::float16 setting, not this feature
hnsw yes HNSW index for the vector stage (clawhdf5-ann); without it, an exact linear scan
parallel yes Parallel HNSW bulk build (identical graph) and Rayon search strategies
zstd no Zstd instead of deflate for embeddings when MemoryConfig::compression is on (links libzstd)
fast-math / openblas / accelerate no BLAS matrix-vector multiply (generic / OpenBLAS / Apple Accelerate)
gpu no GPU distance computation via wgpu (clawhdf5-gpu)
async no Tokio async wrapper with background flush

For an exact linear scan: --no-default-features --features float16.

Settings stored in the file (MemoryConfig):

  • float16 (on for new stores): embeddings on disk as IEEE half precision, rounded as they enter the cache so memory and file agree; values must lie within ±65504. On LongMemEval with real MiniLM embeddings every retrieval metric matches f32. Opt out with float16 = false or clawhdf5 create --f32. Existing stores keep their setting.
  • quantized_index (on for new stores): the HNSW index's copy of the embeddings as i8, re-scored against the exact embeddings; see the table above. Opt out with quantized_index = false or create --f32-index.
  • hnsw_m, hnsw_ef_construction, hnsw_ef_search: 16 / 64 / scaled with k by default.
  • compression (off): deflate (or Zstd) for embeddings; text of 4 KiB or more is always deflated.
  • wal_enabled (on), wal_max_entries, hebbian_boost, decay_factor.

File schema

agent_memory.h5
├── /meta                          (attributes)
│   ├── schema_version, edgehdf5_version (writer tag, kept for compatibility)
│   ├── agent_id, embedder, embedding_dim, chunk_size, overlap, created_at
│   ├── float16, compression, compression_level, compact_threshold,
│   │   hebbian_boost, decay_factor, wal_enabled, wal_max_entries
│   ├── quantized_index, hnsw_m, hnsw_ef_construction, hnsw_ef_search
│   ├── wal_applied_len, wal_applied_crc   (WAL mark of the last checkpoint)
│   └── ann_generation                     (ties the .ann sidecar to this checkpoint)
├── /memory
│   ├── chunks:              string[N]
│   ├── embeddings:          f32[N × D], or f16 for a `float16` store (chunked)
│   ├── source_channel, session_ids, tags:  string[N]
│   ├── timestamps:          f64[N]
│   ├── tombstones:          u8[N]
│   ├── norms:               f32[N]      (pre-computed L2)
│   └── activation_weights:  f32[N]      (Hebbian)
├── /sessions
│   ├── ids, channels, summaries:  string[S]
│   ├── start_idxs, end_idxs:      i64[S]
│   └── timestamps:                f64[S]
├── /knowledge_graph
│   ├── entity_ids, entity_emb_idxs: i64[E]; entity_names, entity_types: string[E]
│   ├── relation_srcs, relation_tgts: i64[R]; relation_types: string[R]
│   ├── relation_weights: f32[R]; relation_ts: f64[R]
│   └── alias_strings: string[A]; alias_entity_ids: i64[A]   (when aliases exist)
└── /integrity                     (signed stores: per-record hashes and the signed manifest)

A store is an ordinary HDF5 file: h5py, h5dump and h5rs read it (the agent's h5py_interop test checks a whole store). Beside it: <store>.h5.wal, <store>.h5.ann (HNSW graph; derived, safe to delete) and <store>.h5.lock.

CLI

clawhdf5-cli installs a binary named clawhdf5:

cargo install --path crates/clawhdf5-cli
clawhdf5 --path agent.h5 create --agent-id my-agent --dim 384 --wal
echo '{"chunk":"User prefers dark mode","embedding":[0.1, ...],"source_channel":"chat","timestamp":1700000000.0,"session_id":"s1","tags":"pref"}' \
  | clawhdf5 --path agent.h5 save
clawhdf5 --path agent.h5 search --embedding '[0.1, ...]' --query 'dark mode preferences' \
  --top-k 5 --vector-weight 0.4 --keyword-weight 0.6
clawhdf5 --path agent.h5 stats          # also: recall <index>, export, agents-md, flush-wal
clawhdf5 --path agent.h5 snapshot backup.h5
clawhdf5 keygen --out signing.key       # then --signing-key signing.key; verify --public-key <hex>

Output is JSON. The CLI's search defaults to weights 0.7 / 0.3, not the library's 0.4 / 0.6, so pass them. recall, stats, agents-md and export open the store read-only.

Migrating from SQLite

cargo install --path crates/clawhdf5-migrate
clawhdf5-migrate --sqlite old.db --hdf5 memory.h5 --agent-id my-agent --embedder minilm

The output is an ordinary agent store, written through the agent's API. The source must use the memory_chunks / sessions / entities / relations layout (names configurable with --*-table); this is not ZeroClaw's schema, and ZeroClaw does not use clawhdf5. What carries over:

SQLite Agent store
memory_chunks records (text, embedding, source channel, timestamp, session id, tags); rows with deleted = 1 become deleted records, or are left out with --skip-deleted
sessions sessions (id, start/end index, channel, summary, timestamp)
entities, relations knowledge-graph entities and relations; entities get new ids and relations are re-pointed

Records are written in id order and numbered from 0. Embeddings are stored as float16 like any new store; --f32 keeps full precision (and is required for values beyond ±65504). The dimension is detected from the first row unless --embedding-dim is given, and a row of another length is an error, never truncated or padded; a source with no records needs --embedding-dim. Every row is checked before the output is created. --incremental adds only rows the store does not hold (records already in it take the source's deleted flag). The tool reads the result back read-only, compares it with the source (every row with --validate-full) and checks that a migrated record is found by search; --dry-run only counts rows. clawhdf5-migrate bundles SQLite, so it compiles C.

Older crate names: rustyhdf5* is now clawhdf5*, edgehdf5-memory is clawhdf5-agent, and the edgehdf5 CLI is clawhdf5-cli.

Research foundation

The design draws on recent papers on agent memory:

Paper Idea Module
MemX (2026) Hybrid fusion + multi-factor re-ranking hybrid, reranker
Graph-Native Cognitive Memory (2026) Weighted, timestamped relations; entity timelines knowledge, temporal
CraniMem (2026) Bounded hippocampal memory consolidation
D-MEM (2026) Surprise-gated storage (as a novelty score) consolidation
SYNAPSE (2025) Spreading activation for recall knowledge
RAGdb (2025) Zero-dependency edge RAG architecture
MemoryGraft (2025) Memory poisoning attacks anomaly, provenance
MemoryArena (2026) Multi-session benchmark temporal
AI Hippocampus (2026) Memory taxonomy survey overall design