Numbers, API names, feature defaults and PR references checked against CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history. - int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives them (2026-09-19/20, machine not recorded, not re-run) instead of none; the Pi 5 1.18x carries 2026-09-21. - BENCHMARKS headline: the libhdf5 chunked-write figure is the newest measurement (35x, 2026-09-23), not 45.3x (2026-08-03). - Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug) in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to the bad_nbit_parms_walk.h5 flip. - README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/ time datasets too; zlib-rs byte-identity scoped to what was measured; macOS default links the system libz for inflate. - Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4 unlimited-dimension size warning. - agent-memory.md: string-dataset compression threshold, agents-md prints Markdown, float16 file sizes linked to their study. - known-issues.md: contiguous selection reads, 1.21x vs h5py threads. - docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify, dated figures, fast-math is not BLAS. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
26 KiB
Agent memory (clawhdf5-agent)
clawhdf5-agent is a persistent, searchable memory store for AI agents,
built on clawhdf5's HDF5 writer: records (text, embedding, source channel,
timestamp, session, tags), sessions and a knowledge graph in one .h5 file,
with a write-ahead log beside it. This page is the long form of the agent
part of the README; every number on it comes from
BENCHMARKS.md, where the commands and machines are.
- Quick start · Search · Signed checkpoints
- Architecture · Modules · Library components
- Performance · LongMemEval · Footprint
- Feature flags and settings · File schema
- CLI · Migrating from SQLite · Research foundation
Integration status: ClawBrainHub's CLI uses this crate's bm25::BM25Index;
no agent framework uses the store. clawhdf5 is not an OpenClaw memory
plugin (openclaw.md), and ZeroClaw does not use it.
Quick start
[dependencies]
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } # not on crates.io yet
use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions};
// A new store: 384-dim embeddings (float16 on disk and an int8 HNSW index by default).
let mut memory = HDF5Memory::create(MemoryConfig::new("agent.h5".into(), "my-agent", 384))?;
memory.save(MemoryEntry {
chunk: "User prefers dark mode and vim keybindings.".into(),
embedding: embed("User prefers dark mode and vim keybindings."), // your embedder
source_channel: "chat".into(),
timestamp: now,
session_id: "session-001".into(),
tags: "preference".into(),
})?;
// Hybrid search: HNSW vector + BM25 keyword, fused 0.4 / 0.6 (the measured default).
let query = embed("what editor does the user like?");
for r in memory.search(&query, "editor preferences", &SearchOptions::new(5)) {
println!("[{:.3}] {}", r.score, r.chunk);
}
memory.flush_wal()?; // checkpoint the WAL into agent.h5
clawhdf5 stores embeddings; it does not compute them. Any dimension works,
fixed when the store is created. HDF5Memory::open(path) reopens a store
(holding its single-writer lock); HDF5Memory::open_read_only(path) gives a
lock-free point-in-time view.
Search
HDF5Memory::search(query_emb, text, &SearchOptions) is the full search
path; hybrid_search(query_emb, text, vector_weight, keyword_weight, k) and
hybrid_search_with are thin wrappers over it.
use clawhdf5_agent::confidence::ConfidenceConfig;
use clawhdf5_agent::reranker::ReRankConfig;
// Only memories from these source channels; still a full page of k results.
let work = memory.search(&query, "deadline", &SearchOptions::new(5).with_sources(["slack", "email"]));
// Re-rank (relevance, recency, source authority, activation), then drop
// low-confidence results: the pipeline ClawhdfBackend runs.
let careful = memory.search(
&query,
"user preferences",
&SearchOptions::new(5)
.with_rerank(ReRankConfig::default())
.with_confidence(ConfidenceConfig::default()),
);
The source-channel filter is applied before ranking: an exact scan of the allowed records whenever that is cheaper than the index would be, and as the fallback when the index returns a short pool. Hebbian activation boosts are persisted by the next checkpoint (or on drop), not per query; search never writes the store.
Signed checkpoints
use clawhdf5_agent::signing;
let key = signing::generate_key(); // keep the secret key; publish the public one
let public = key.verifying_key();
memory.set_signing_key(key); // never written to disk
memory.flush_wal()?; // this checkpoint is signed
let report = HDF5Memory::verify(std::path::Path::new("agent.h5"), &public)?;
assert!(report.is_valid()); // report.changed_records names edited records
The Ed25519 signature covers every record (text, embedding as stored,
channel, timestamp, session, tags, deleted flag, activation) through a
SHA-256 Merkle tree, plus the store's settings, sessions and knowledge
graph, so a change made with any tool is caught and located. It covers
checkpoints, not saves still in the WAL (report.wal_entries_unsigned
counts those). A signed store refuses to checkpoint without the key
(MemoryError::SigningKeyRequired). CLI: clawhdf5 keygen,
--signing-key <file> on writing commands, and verify --public-key.
Signing adds about 20% to a checkpoint and 32 bytes per record to the file
(BENCHMARKS.md § Signed checkpoints).
Architecture
┌─────────────────┐
│ Agent Query │
└────────┬────────┘
│
┌─────────────────▼──────────────────┐
│ HDF5Memory::search │
│ optional source-channel filter │
│ HNSW vector + BM25 keyword │
│ weighted fusion (0.4 / 0.6) │
│ × √(Hebbian activation) │
└─────────────────┬──────────────────┘
│ opt-in (SearchOptions);
│ ClawhdfBackend turns both on
┌─────────────────▼──────────────────┐
│ Multi-factor re-ranking │
│ relevance · recency · authority · │
│ activation │
├────────────────────────────────────┤
│ Confidence rejection │
│ (suppress bad matches) │
└─────────────────┬──────────────────┘
│
┌────────────────────────────▼────────────────────────────┐
│ In memory │
│ cache (embeddings) · BM25 index · HNSW index │
│ provenance ledger + anomaly alerts (session-scoped) │
└────────────────────────────┬────────────────────────────┘
│ WAL append; checkpoint
┌────────────────────────────▼────────────────────────────┐
│ agent_memory.h5 /meta · /memory · /sessions · │
│ /knowledge_graph │
│ agent_memory.h5.wal chained-CRC write-ahead log │
│ agent_memory.h5.ann HNSW graph (derived, rebuildable) │
│ agent_memory.h5.lock single-writer lock │
└─────────────────────────────────────────────────────────┘
Durability. Every WAL entry carries a CRC32 chained to the previous
entry's, so a corrupted, reordered, duplicated or spliced entry stops replay
instead of loading bad data. Each checkpoint records a WAL mark in /meta,
so a crash between a checkpoint and the WAL truncate never applies an entry
twice. Checkpoints and snapshots are made durable as a unit (temp file
synced, renamed, directory synced). Individual WAL appends are not
fsynced (a latency trade-off): saves since the last checkpoint can be lost
on power failure or a kernel panic, not on a process crash. An unreadable
WAL is quarantined to <store>.h5.wal.corrupt-<ts> rather than blocking
open().
Single writer. create/open take an exclusive advisory lock on
<store>.h5.lock; a second opener gets MemoryError::Locked.
Write bookkeeping. save/save_batch/save_or_update run each write
through an in-memory (session-scoped, not persisted) provenance ledger — an
unkeyed content hash per record, for detecting accidental corruption, not
tampering — and a write-anomaly detector (rate limits, injection patterns,
source distribution). Alerts never block a save; drain them with
take_anomaly_alerts. The source classification is inferred from the
caller's source_channel string, a heuristic, not an authenticated trust
boundary.
Modules
| Module | What it does |
|---|---|
hybrid |
Vector + BM25 fusion: min-max-normalised weighted sum, vector 0.4 / keyword 0.6 by default (hybrid::DEFAULT_FUSION, tuned on LongMemEval); RRF via Fusion::Rrf / hybrid_search_with (measured worse) |
reranker |
Re-ranking by retrieval relevance (leads, weight 1.0), recency, source authority, activation. Opt-in via SearchOptions::with_rerank; on in ClawhdfBackend |
confidence |
Low-confidence rejection. Opt-in via SearchOptions::with_confidence; on in ClawhdfBackend |
bm25 |
Incremental Okapi BM25 index kept for the life of the store; optional stemming |
signing |
Ed25519-signed checkpoints (above) |
wal |
Write-ahead log, format v4, chained CRC32 per entry; reads v2 and v3 (v1 only through the one-time migration in open) |
knowledge |
Entity/relation graph: BFS, spreading activation, fuzzy (Levenshtein) entity resolution |
consolidation |
Three tiers (Working → Episodic → Semantic): importance, novelty, time decay |
temporal |
Sorted timestamp index, session DAG, entity timeline |
multimodal |
Cross-modal search over text/image/audio/video embeddings (exact scan) |
provenance, anomaly |
Session-scoped write bookkeeping (above) |
openclaw |
ClawhdfBackend, a Markdown-oriented backend (below). Named for OpenClaw, but not an OpenClaw plugin (openclaw.md) |
vector_search |
Flat cosine search paths: pre-normed, SIMD, BLAS, GPU, parallel |
ivf / pq |
Standalone IVF and IVF-PQ indexes; not used by HDF5Memory, whose index is HNSW |
query_expand, entity_extract |
Synonym/acronym/temporal query expansion; rule-based entity extraction into the graph |
memory_strategy, decision_gate |
When to save: save-every, semantic shift, user correction; trivial/substantive classification |
ephemeral |
In-memory TTL/LFU working tier |
async_memory |
Tokio wrapper over the store (async feature) |
Library components
The consolidation tiers, the graph algorithms and the temporal and multi-modal indexes are components you drive directly; the store persists the records, sessions and graph they work over.
use clawhdf5_agent::knowledge::KnowledgeCache;
let mut kg = KnowledgeCache::new();
let alice = kg.add_entity("Alice", "person", -1);
let bob = kg.add_entity("Bob", "person", -1);
let acme = kg.add_entity("Acme Corp", "company", -1);
kg.add_relation(alice, acme, "works_at", 1.0);
kg.add_relation(alice, bob, "manages", 0.8);
let neighbors = kg.bfs_neighbors(alice, 2); // 2-hop neighbourhood
let activated = kg.spreading_activation(&[alice], 0.5, 0.01, 5); // related entities
let (id, created) = kg.resolve_or_create("alice", "person", -1, 2); // fuzzy (Levenshtein <= 2)
assert_eq!((id, created), (alice, false));
use clawhdf5_agent::consolidation::{ConsolidationConfig, ConsolidationEngine, UntrustedSource};
let mut engine = ConsolidationEngine::new(ConsolidationConfig {
working_capacity: 100,
..Default::default()
});
let id = engine.add_memory("User prefers dark mode".into(), embed("dark mode"), UntrustedSource::User, now);
engine.access_memory(id, now + 60.0); // reactivates it
engine.consolidate(now + 3600.0); // promote (Working -> Episodic -> Semantic) and evict
let stats = engine.get_stats();
println!("working {} episodic {} semantic {}", stats.working_count, stats.episodic_count, stats.semantic_count);
System and correction sources get elevated importance and go through a
separate entry point, add_trusted_memory(.., TrustedSource::System, ..),
so untrusted content cannot claim them.
use clawhdf5_agent::temporal::TemporalIndex;
let mut index = TemporalIndex::new();
index.insert(1, 1_700_000_000.0);
index.insert(2, 1_700_003_600.0); // an hour later
let in_range = index.range_query(1_700_000_000.0, 1_700_010_800.0);
let recent = index.latest(10);
Markdown backend
ClawhdfBackend ingests Markdown by section and searches it with the full
pipeline. It is a library API, not an OpenClaw plugin.
use clawhdf5_agent::openclaw::{ClawhdfBackend, MemoryBackend};
let mut backend = ClawhdfBackend::create(std::path::Path::new("memory.h5"), 384)?;
let md = std::fs::read_to_string("MEMORY.md")?;
let sections = backend.ingest_markdown("MEMORY.md", &md)?; // one record per heading
for r in backend.search("dark mode", &embed("dark mode"), 5) {
println!("[{:.3}] {} ({})", r.score, r.text, r.path);
}
let exported = backend.export_markdown("MEMORY.md")?;
Limits: ingested sections carry no embedding, so their search is
keyword-only unless you save records with vectors through save_entry;
ingesting a file again adds its sections again; export_markdown writes
every heading as ##, so it is not a lossless round trip.
Performance
Unless marked otherwise, measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D,
8C/16T), commit 5c8323c, 384-dim embeddings; commands in
BENCHMARKS.md.
HNSW (the default vector stage) — search_harness, clustered data,
N = 100K, M = 16, ef_construction = 64, ef = 64, recall against an exact scan
(§ Quantising the index copy):
| index | recall@10 | QPS | build |
|---|---|---|---|
f32 |
0.9945 | 13 399 | 3.2 s |
i8 + exact re-score (default for new stores) |
0.9940 | 21 848 | 1.8 s |
A paired comparison (medians of alternating runs, same binary: the int8
index answers 1.63x the queries per second at equal recall), recorded
2026-09-20 with the machine not recorded, and not re-run since: a single
f32 run on 2026-09-24 (tank) measured recall 0.9945, 19 001 QPS and a
2.7 s build, so the 1.63x ratio has not been re-checked. On a Raspberry
Pi 5 (NEON SDOT) the int8 index is 1.18x the f32 QPS at equal recall
(2026-09-21; § On ARM).
Before the v2.4.0 neighbour-selection fix, recall@10 at 100K was
0.31.
Operations:
| Operation | Latency | Scale |
|---|---|---|
hybrid_search p50 |
0.07 ms / 0.49 ms / 4.69 ms | 1K / 10K / 100K records |
| BM25 keyword search | 20.4 µs | 1K records |
| Knowledge graph BFS | 23.1 µs | 1K entities |
| Spreading activation | 10.1 µs | 100 entities |
| Temporal range query | 622 ns | 10K timestamps |
| Consolidation cycle | 115.2 µs | 1K records |
| Cross-modal search (exact scan, 2 embeddings per record) | 842.0 µs / 8.44 ms | 1K / 10K records |
| Memory write (WAL append) | 26.1 µs | per record |
float16 stores (the default) add about 2 µs per write for rounding
(§ Write Path).
Brute-force and IVF (Criterion; not used by HDF5Memory):
| Scale | Flat | IVF (nprobe=10) | IVF-PQ |
|---|---|---|---|
| 1K | 47.4 µs | — | — |
| 10K | 500.5 µs | 24.8 µs | — |
| 100K | 6.58 ms | 592 µs | 869 µs |
No comparison with MemX is made: its published figure is end-to-end and ours is one component (BENCHMARKS.md).
Consolidation — 1,000 records (10 signal + 990 noise),
working_capacity = 100: the store goes from 1,000 to 100 records with
Hit@1 on the signal records staying at 100%, and search from 2.22 ms to
0.24 ms (§ Consolidation Efficiency).
LongMemEval retrieval recall
Full longmemeval_s haystack, all 500 questions (47.7 sessions and 493.5
turns each; 4.0% of sessions are evidence), real all-MiniLM-L6-v2
embeddings, k = 10. Re-run 2026-09-27 on tank; the headline reproduced
exactly (§ LongMemEval Results):
| Mode | Turn-level Hit@5 | Session-level Hit@5 |
|---|---|---|
| BM25 only | 75.0% | 93.6% |
| Vector only (MiniLM) | 71.8% | 94.2% |
| Hybrid 0.4 / 0.6 (default) | 81.4% | 96.8% |
This is retrieval recall (did a gold turn appear in the top k), not the
official LongMemEval QA accuracy; the two are not comparable. A weight sweep
found the old 0.7 / 0.3 default strictly dominated by 0.4 / 0.6, the default
since v2.5.0; use 0.3 / 0.7 if rank-1 precision matters most. Earlier
session-level figures of 100% and a claimed win over MemX were retracted
(BENCHMARKS.md).
The benchmark's vector stage needs clawhdf5-bench's embeddings feature.
Memory footprint
On disk — float16 embeddings (the default), 200-character synthetic
text, footprint_bench: 810.4 KB at 1K records, 7.8 MB at 10K, 76.7 MB at
100K (803–829 bytes per record). The synthetic text is far more repetitive
than real text (40 distinct strings, deflated), so real records will be
larger; the embeddings alone are 768 B per record
(§ Memory Footprint, 2026-09-24). In the
float16 study (clustered data, 2026-09-23), 100K × 384 takes 80.8 MiB as
float16 and 154.0 MiB as f32
(§ float16 embedding storage).
In memory — a store reopened from disk, counting allocator (§ Memory footprint):
| Records | Raw vectors | f32 index |
i8 index (default) |
|---|---|---|---|
| 1K | 1 MiB | 4 MiB (2.40x) | 2 MiB (1.64x) |
| 10K | 15 MiB | 44 MiB (3.03x) | 27 MiB (1.81x) |
| 100K | 146 MiB | 399 MiB (2.72x) | 256 MiB (1.74x) |
The f32 column was re-measured on 2026-09-24 (tank); the i8 column was
first measured 2026-09-19 (commit c0a9206, machine not recorded) and not
re-run (§ Quantising the index copy).
Feature flags and settings
clawhdf5-agent flag |
Default | Description |
|---|---|---|
float16 |
yes | Half-precision cosine kernel. Half-precision storage is the MemoryConfig::float16 setting, not this feature |
hnsw |
yes | HNSW index for the vector stage (clawhdf5-ann); without it, an exact linear scan |
parallel |
yes | Parallel HNSW bulk build (identical graph) and Rayon search strategies |
zstd |
no | Zstd instead of deflate for embeddings when MemoryConfig::compression is on (links libzstd) |
fast-math / openblas / accelerate |
no | BLAS matrix-vector multiply (generic / OpenBLAS / Apple Accelerate) |
gpu |
no | GPU distance computation via wgpu (clawhdf5-gpu) |
async |
no | Tokio async wrapper with background flush |
For an exact linear scan: --no-default-features --features float16.
Settings stored in the file (MemoryConfig):
float16(on for new stores): embeddings on disk as IEEE half precision, rounded as they enter the cache so memory and file agree; values must lie within ±65504. On LongMemEval with real MiniLM embeddings every retrieval metric matchesf32. Opt out withfloat16 = falseorclawhdf5 create --f32. Existing stores keep their setting.quantized_index(on for new stores): the HNSW index's copy of the embeddings asi8, re-scored against the exact embeddings; see the table above. Opt out withquantized_index = falseorcreate --f32-index.hnsw_m,hnsw_ef_construction,hnsw_ef_search: 16 / 64 / scaled withkby default.compression(off): deflate (or Zstd) for embeddings; string datasets (text, channels, tags, ...) of 4 KiB or more are always deflated.wal_enabled(on),wal_max_entries,hebbian_boost,decay_factor.
File schema
agent_memory.h5
├── /meta (attributes)
│ ├── schema_version, edgehdf5_version (writer tag, kept for compatibility)
│ ├── agent_id, embedder, embedding_dim, chunk_size, overlap, created_at
│ ├── float16, compression, compression_level, compact_threshold,
│ │ hebbian_boost, decay_factor, wal_enabled, wal_max_entries
│ ├── quantized_index, hnsw_m, hnsw_ef_construction, hnsw_ef_search
│ ├── wal_applied_len, wal_applied_crc (WAL mark of the last checkpoint)
│ └── ann_generation (ties the .ann sidecar to this checkpoint)
├── /memory
│ ├── chunks: string[N]
│ ├── embeddings: f32[N × D], or f16 for a `float16` store (chunked)
│ ├── source_channel, session_ids, tags: string[N]
│ ├── timestamps: f64[N]
│ ├── tombstones: u8[N]
│ ├── norms: f32[N] (pre-computed L2)
│ └── activation_weights: f32[N] (Hebbian)
├── /sessions
│ ├── ids, channels, summaries: string[S]
│ ├── start_idxs, end_idxs: i64[S]
│ └── timestamps: f64[S]
├── /knowledge_graph
│ ├── entity_ids, entity_emb_idxs: i64[E]; entity_names, entity_types: string[E]
│ ├── relation_srcs, relation_tgts: i64[R]; relation_types: string[R]
│ ├── relation_weights: f32[R]; relation_ts: f64[R]
│ └── alias_strings: string[A]; alias_entity_ids: i64[A] (when aliases exist)
└── /integrity (signed stores: per-record hashes and the signed manifest)
A store is an ordinary HDF5 file: h5py, h5dump and h5rs read it (the
agent's h5py_interop test checks a whole store). Beside it:
<store>.h5.wal, <store>.h5.ann (HNSW graph; derived, safe to delete)
and <store>.h5.lock.
CLI
clawhdf5-cli installs a binary named clawhdf5:
cargo install --path crates/clawhdf5-cli
clawhdf5 --path agent.h5 create --agent-id my-agent --dim 384 --wal
echo '{"chunk":"User prefers dark mode","embedding":[0.1, ...],"source_channel":"chat","timestamp":1700000000.0,"session_id":"s1","tags":"pref"}' \
| clawhdf5 --path agent.h5 save
clawhdf5 --path agent.h5 search --embedding '[0.1, ...]' --query 'dark mode preferences' \
--top-k 5 --vector-weight 0.4 --keyword-weight 0.6
clawhdf5 --path agent.h5 stats # also: recall <index>, export, agents-md, flush-wal
clawhdf5 --path agent.h5 snapshot backup.h5
clawhdf5 keygen --out signing.key # then --signing-key signing.key; verify --public-key <hex>
Output is JSON (Markdown for agents-md). The CLI's search defaults to
weights 0.7 / 0.3, not the library's 0.4 / 0.6, so pass them. recall,
stats, agents-md and export open the store read-only.
Migrating from SQLite
cargo install --path crates/clawhdf5-migrate
clawhdf5-migrate --sqlite old.db --hdf5 memory.h5 --agent-id my-agent --embedder minilm
The output is an ordinary agent store, written through the agent's API. The
source must use the memory_chunks / sessions / entities / relations
layout (names configurable with --*-table); this is not ZeroClaw's schema,
and ZeroClaw does not use clawhdf5. What carries over:
| SQLite | Agent store |
|---|---|
memory_chunks |
records (text, embedding, source channel, timestamp, session id, tags); rows with deleted = 1 become deleted records, or are left out with --skip-deleted |
sessions |
sessions (id, start/end index, channel, summary, timestamp) |
entities, relations |
knowledge-graph entities and relations; entities get new ids and relations are re-pointed |
Records are written in id order and numbered from 0. Embeddings are
stored as float16 like any new store; --f32 keeps full precision (and is
required for values beyond ±65504). The dimension is detected from the
first row unless --embedding-dim is given, and a row of another length is
an error, never truncated or padded; a source with no records needs
--embedding-dim. Every row is checked before the output is created.
--incremental adds only rows the store does not hold (records already in
it take the source's deleted flag). The tool reads the result back
read-only, compares it with the source (every row with --validate-full)
and checks that a migrated record is found by search; --dry-run only
counts rows. clawhdf5-migrate bundles SQLite, so it compiles C.
Older crate names: rustyhdf5* is now clawhdf5*, edgehdf5-memory is
clawhdf5-agent, and the edgehdf5 CLI is clawhdf5-cli.
Research foundation
The design draws on recent papers on agent memory:
| Paper | Idea | Module |
|---|---|---|
| MemX (2026) | Hybrid fusion + multi-factor re-ranking | hybrid, reranker |
| Graph-Native Cognitive Memory (2026) | Weighted, timestamped relations; entity timelines | knowledge, temporal |
| CraniMem (2026) | Bounded hippocampal memory | consolidation |
| D-MEM (2026) | Surprise-gated storage (as a novelty score) | consolidation |
| SYNAPSE (2025) | Spreading activation for recall | knowledge |
| RAGdb (2025) | Zero-dependency edge RAG | architecture |
| MemoryGraft (2025) | Memory poisoning attacks | anomaly, provenance |
| MemoryArena (2026) | Multi-session benchmark | temporal |
| AI Hippocampus (2026) | Memory taxonomy survey | overall design |