diff --git a/README.md b/README.md index 348f832..889e631 100644 --- a/README.md +++ b/README.md @@ -1,1120 +1,444 @@ -# ClawhDF5 +# clawhdf5 -**The memory layer AI agents deserve. One file. Pure Rust. Zero C dependencies.** +**A pure-Rust HDF5 reader, writer and in-place editor — no libhdf5, and no +C by default — with an agent-memory store built on it.** [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![Rust](https://img.shields.io/badge/rust-1.92%2B-orange.svg)](https://www.rust-lang.org) -[![Tests](https://img.shields.io/badge/tests-1850%2B-brightgreen.svg)](#building) -[![LongMemEval](https://img.shields.io/badge/LongMemEval__s-Turn--Level%20Hit@5%2081.4%25%20hybrid-blue.svg)](BENCHMARKS.md#longmemeval-results) -[![Footprint](https://img.shields.io/badge/on--disk-~820%20B%2Frecord%20float16%2C%20synthetic%20text-lightgrey.svg)](BENCHMARKS.md#memory-footprint-1) +[![Conformance](https://img.shields.io/badge/conformance-602%2F697%20files%20identical%20to%20h5py-brightgreen.svg)](CONFORMANCE.md) +[![LongMemEval](https://img.shields.io/badge/LongMemEval__s-turn%20Hit@5%2081.4%25%20hybrid-blue.svg)](BENCHMARKS.md#longmemeval-results) -ClawHDF5 is a pure-Rust HDF5 implementation combined with a research-grade agent memory engine. It gives AI agents persistent, searchable, cryptographically verifiable memory (Ed25519-signed checkpoints) — all stored in a single portable file. +clawhdf5 implements the HDF5 file format from the specification, in Rust. +It reads superblocks v0–3, every group and chunk-index structure libhdf5 +writes, the standard filters and the common plugin filters, +variable-length data and virtual datasets, and follows files a SWMR writer +is appending to. It reads files from libhdf5, h5py and netCDF-4 and writes +files they read. The same library opens files over HTTP and in object +stores by range requests, runs in the browser as WebAssembly, and has +Python bindings with an h5py-shaped API. -> **Two things live here:** -> - **A general-purpose, pure-Rust HDF5 library** — zero C dependencies, NetCDF-4 support, SIMD/GPU acceleration. See the **[Crate Map](#crate-map)** and **[BENCHMARKS.md](BENCHMARKS.md)** for the libhdf5 head-to-head numbers. -> - **An agent memory layer built on top of it** — vector search, knowledge graph, hippocampal-style consolidation, in `clawhdf5-agent`. +Two things live in this repository: -The crates are not on crates.io yet, so depend on them from git: +- **The HDF5 library** — the `clawhdf5` crate and its parts, `h5rs` + command-line tools, Python, WebAssembly and NetCDF-4 layers. +- **Agent memory** (`clawhdf5-agent`) — a single-file store for AI agents + (HNSW + BM25 hybrid search, write-ahead log, signed checkpoints) whose + files are ordinary HDF5. See [docs/agent-memory.md](docs/agent-memory.md). + +Nothing is published to crates.io or PyPI yet: use it +[from git or a checkout](#install). + +## Contents + +- [Evidence](#evidence) — conformance, robustness on hostile files, speed +- [What is supported](#what-is-supported) — the feature matrix +- [Install](#install) · [Quick start: Rust](#quick-start-rust) · [Quick start: Python](#quick-start-python) +- [Remote files, the browser, SWMR](#remote-files-the-browser-swmr) · [h5rs tools](#h5rs-tools) +- [Agent memory](#agent-memory) · [Crate map](#crate-map) · [Building and testing](#building-and-testing) +- [Documentation](#documentation) · [Who uses it](#who-uses-it) + +## Evidence + +**Conformance.** Every file of eight public corpora — the libhdf5 source +tree's test files, the HDF Group's +[CVE reproducer corpus](https://github.com/HDFGroup/cve_hdf5), netcdf-c, +netcdf4-python, pyfive, h5wasm, h5py's and xarray's data files, 697 files +in all, pinned by commit — is read by clawhdf5 and by h5py/libhdf5 and +compared object by object (object set, shapes, a SHA-256 of every +dataset's and attribute's values). Run of 2026-09-28 on tank, h5py 3.16 / +HDF5 2.0 ([CONFORMANCE.md](CONFORMANCE.md)): + +| files | ok (identical to h5py) | mismatch | libhdf5 cannot open | ref-bug¹ | our-error¹ | panic / hang / crash / OOM | +|---:|---:|---:|---:|---:|---:|---:| +| 697 | **602** | **0** | 92 | 2 | 1 | **0** | + +¹ The three remaining objects are corrupt data (scale-offset codes past +the end of a chunk, short unfiltered chunks, an N-Bit parameter list one +value short) that HDF5 2.0 returns only by reading past a buffer; +clawhdf5 refuses them, as libhdf5's development branch and its own +`test_filter_bad_params` do. Details and evidence in +[CONFORMANCE.md § Reference bugs](CONFORMANCE.md#reference-bugs). The +run is a nightly CI job (`.gitea/workflows/conformance.yml`) that fails on +any panic, hang or crash, or on an ok file that stops being ok. + +**Robustness on hostile files.** On the 147 CVE and fuzzer files +([CONFORMANCE.md § CVE corpus](CONFORMANCE.md#cve-corpus-clawhdf5-vs-h5dump-vs-h5py)): + +| tool | panic | crash | hang | OOM | +|---|---:|---:|---:|---:| +| clawhdf5 | 0 | 0 | 0 | 0 | +| h5dump 1.14.6 | 0 | 2 | 0 | 0 | +| h5py 3.16.0 / HDF5 2.0.0 | 0 | 1 | 0 | 0 | + +Sizes and addresses read from a file are checked before use +(overflow-checked arithmetic, fallible allocation on the chunked read +paths, bounded recursion in B-trees and object-header chains), and +`scripts/h5rs-fuzz.sh` runs every `h5rs` subcommand over the corpus +looking for panics, crashes and hangs. + +**Reads from many threads.** A `File` is `Send + Sync` and there is no +library-wide lock, so one open file serves many threads. Full reads of 64 +deflate-compressed 64 MiB datasets, each read decoding on its calling +thread (`concurrent_read --decode-threads 1`), tank (Ryzen 7 7800X3D, +16 threads), 2026-09-26, commit `c5334b1` +([BENCHMARKS.md](BENCHMARKS.md#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)): + +| threads | clawhdf5, one `File` | h5py, threads | h5py, processes | clawhdf5 / h5py processes | +|---:|---:|---:|---:|---:| +| 1 | 670 MB/s | 410 MB/s | 397 MB/s | 1.69x | +| 16 | 4944 MB/s | 390 MB/s | 3135 MB/s | 1.58x | + +That run was noisier than others on the same machine, so compare ratios +within it rather than MB/s across runs. A contiguous (uncompressed) full +read on one thread ran at 6718 MB/s against h5py's 5545 in the same run. + +**Against libhdf5 1.14.6 from Rust**, tank, 2026-08-03 +([BENCHMARKS.md § Independent Validation](BENCHMARKS.md#independent-validation-tank-ryzen-7-7800x3d-2026-08-03)): +sequential read of 100K `f32` 23.3 µs vs 63.6 µs (2.7x); 128 attribute +writes 85.2 µs vs 877 µs (10.3x); 64 group creates 130 µs vs 1.37 ms +(10.6x); a 512×512 `f32` chunked deflate-6 write 1.44 ms vs 65.0 ms +(re-measured 2026-09-23 with the pure-Rust deflate: 1.46 ms vs 51.4 ms, +35x); a 100K `f32` sequential write is a tie. The writer (`FileBuilder`) +assembles a file in memory and writes it once, which is part of that +difference; read the caveats in [BENCHMARKS.md](BENCHMARKS.md#caveats) +before quoting these. + +## What is supported + +Limits and open issues, with dates, are in +[docs/known-issues.md](docs/known-issues.md). + +| Area | Supported | Read only | Not supported | +|---|---|---|---| +| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers | Metadata cache images | Writing files HDF5 1.8 can read | +| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | +| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex (HDF5 2.0 class 11) | Variable-length strings and sequences, object references | Writing variable-length data; decoding region and attribute references; x87 long double and binary128 | +| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does); fill values; resizable datasets; virtual datasets (read limits in known-issues) | Chunk indexes v1 B-tree and implicit (the editor also changes them) | External raw data files (explicit error) | +| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4, Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | +| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; filters this build cannot encode (refused before any write) | +| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | +| **Bindings** | Python (read, `'w'` for numeric arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound or VL-sequence datasets | Node.js (the package does not work; see known-issues) | + +Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`, +`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure +Rust; h5py + hdf5plugin read what clawhdf5 writes with them, and ZFP decodes +bit-exact against hdf5plugin 7.1. `pcodec` (opt-in) uses a private filter ID +that only clawhdf5 reads. + +**C dependencies, precisely.** The core crates build no C by default: no +libhdf5, and deflate is [zlib-rs](https://github.com/trifectatechfoundation/zlib-rs), +which produces output byte-identical to zlib-ng and matches its HDF5 read and +write speed within 6% +([BENCHMARKS.md](BENCHMARKS.md#deflate-backend-zlib-rs-vs-zlib-ng)). CI +fails if a C-building crate enters their default dependency tree. C comes in +only when you ask: `fast-deflate` (zlib-ng, needs cmake), `zstd`, `szip`, +`https` and the cloud stores (ring / aws-lc-rs), the BLAS backends, +`clawhdf5-migrate` (bundled SQLite) and the Node.js bindings. + +## Install + +The crates are not on crates.io; depend on the repository (MSRV 1.92): ```toml [dependencies] -clawhdf5 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } # core HDF5 read/write -clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } # + agent memory layer +clawhdf5 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +# optional parts +clawhdf5-remote = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } # HTTP / object stores +clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } # agent memory ``` -> **C dependencies, precisely:** the core crates (`clawhdf5`, `clawhdf5-agent`, -> `-format`, `-io`, `-filters`, `-ann`, `-accel`, `-netcdf4`, `-cli`) build no C -> code by default — no libhdf5, and deflate is the pure-Rust -> [zlib-rs](https://github.com/trifectatechfoundation/zlib-rs), which matches -> zlib-ng on HDF5 reads and writes and produces byte-identical output -> ([BENCHMARKS.md § Deflate backend](BENCHMARKS.md#deflate-backend-zlib-rs-vs-zlib-ng)). -> CI fails if a C-building crate enters their default dependency tree. C comes -> in only when you ask for it: `fast-deflate` (zlib-ng, needs cmake), `zstd`, -> `szip`, the BLAS backends, `clawhdf5-migrate` (bundled SQLite) and the -> Node.js bindings. +or, with a checkout, `clawhdf5 = { path = "../clawhdf5/crates/clawhdf5" }`. +Add `features = ["plugin-filters"]` for every plugin filter. -> **New here?** Start with the **[Quickstart Guide](docs/QUICKSTART.md)** · See **[Use Cases](docs/USE_CASES.md)** · Read **[Benchmarks](BENCHMARKS.md)** - -## What's new (v2.2 → v2.7, and unreleased) - -Five releases in September 2026. Details, including upgrade notes and every -breaking change, are in [CHANGELOG.md](CHANGELOG.md). - -**HDF5 correctness (read these if you read files with an earlier release)** -- **Extensible Array chunk indexes returned wrong data** past the 36th chunk — - any dataset with one unlimited dimension. Silent: plausible numbers from the - wrong chunks. Fixed in v2.7.0; re-read affected data. -- Fixed and Extensible Array checksums are now verified, so a corrupt chunk - index is `ChecksumMismatch` instead of wrong data (v2.7.0). -- Compound datatypes written with default libver bounds (plain - `h5py.File(path, 'w')`) were mis-parsed; HDF5 2.0 compound v5 and native - complex (class 11) types now parse (v2.2.0–v2.3.0). -- Committed datatypes, fill values, soft links and `H5T_STD_REF` references now - read correctly; external links and external raw data are explicit errors; - `attrs()` no longer silently drops attributes (v2.3.0–v2.5.0). -- Datasets indexed by a version-2 B-tree now read (v2.5.0). - -**Security and robustness** -- A crafted file could abort any reader via B-tree v2 recursion or explode it - via shared children; both are now fast errors (v2.7.0). -- Virtual-dataset source paths are confined to the file's directory; chunked - reads use overflow-checked sizes and fallible allocation, and the facade - writes files atomically (v2.3.0). -- Agent store: single-writer lock plus `open_read_only`; a crash between - checkpoint and WAL truncate no longer duplicates entries; unreadable WALs are - quarantined instead of blocking `open()` (v2.3.0). - -**Search quality and speed** -- HNSW neighbour selection now uses the paper's diversity heuristic: recall@10 - at 100K went from 0.31 to 0.98 (v2.4.0). -- `hybrid_search` is 79–190× faster than v2.3.0 (p50 0.07 ms at 1K, 4.65 ms at - 100K). It no longer rebuilds BM25 or rewrites the store per query, and the - HNSW graph is persisted (v2.4.0). -- Default fusion weights are now the measured 0.4 / 0.6 (v2.5.0). Re-ranking had - been discarding the retrieval score, costing the Markdown backend 40.6pp of - Hit@1; fixed in v2.6.0. -- Selection reads whose bounding box covers at most half the dataset decode - only the chunks they touch (a 64×64 window: 105 ms to 0.39 ms), and full - reads are 1.2–1.9× faster (v2.5.0). - -**Memory** -- A loaded store holds ~30% less (embeddings stored once, v2.6.0), and the - int8 HNSW index, **on by default for new stores** (unreleased), brings a - 100K × 384 store to 1.74× the raw vectors. At equal recall it is also faster - than `f32`: 1.63× QPS on AVX2, 1.18× on a Raspberry Pi 5 (NEON `SDOT`). - -**Interop and search (unreleased)** -- **Files we write now open in h5py and libhdf5.** Every `f32` dataset — - including every agent store's embeddings — and every empty dataset was - refused by libhdf5. Both were write-side bugs in every release; agent stores - fix themselves at their next checkpoint. See - [docs/known-issues.md](docs/known-issues.md). -- `MemoryConfig::float16` now stores half-precision embeddings (it was - ignored), and is on by default for new stores: 48% smaller files, and - identical LongMemEval retrieval on real embeddings. -- `HDF5Memory::search` with `SearchOptions`: filter by source channel (exact - filtered top-k, never slower than unfiltered), and opt-in re-ranking and - confidence rejection, which used to be reachable only through `ClawhdfBackend`. - -**Remote files (unreleased)** -- New crate `clawhdf5-remote`: `open_url("http://…")` reads a file on an - HTTP server (or in S3/GCS/Azure, opt-in) by range requests through a - block cache, without downloading it; `h5rs` takes URLs with its `remote` - feature. See [Reading remote files](#reading-remote-files). -- `File::open_swmr` follows a file an h5py/libhdf5 SWMR writer is still - appending to (`Dataset::refresh`, bounded retries); copies of such files - taken mid-write read with every open path. See - [Following a file a SWMR writer is appending to](#following-a-file-a-swmr-writer-is-appending-to). - -**Tooling** -- CI now runs the h5py/netCDF4 interop suites for real (they had been skipping - silently) and runs an aarch64 job for the NEON kernels. - ---- - -## Why ClawhDF5? - -Every AI agent needs memory. Today that means scattered Markdown files, SQLite databases, cloud-hosted vector stores, and glue code. ClawhDF5 replaces all of it: - -| Problem | Status Quo | ClawhDF5 | -|---------|-----------|----------| -| Vector search | External DB (Pinecone, Qdrant) | Built-in, sub-millisecond | -| Keyword search | Separate FTS engine | Integrated BM25 | -| Knowledge graph | Neo4j or none | In-file graph with spreading activation | -| Memory consolidation | Manual pruning | Hippocampal-inspired automatic tiers | -| Temporal queries | Custom code | Native temporal index (622 ns range query over 10K) | -| Multi-modal | Multiple stores | Unified cross-modal search (exact scan: 842 µs over 1K records) | -| Integrity | Hope for the best | Ed25519-signed checkpoints that pinpoint any edited record, chained-CRC WAL, checksummed chunk indexes, write-anomaly alerts | -| Portability | Config + DB + files | **One `.h5` file. Copy it anywhere.** | - ---- - -## Performance - -The brute-force/IVF vector search, agent-memory, on-disk footprint and consolidation figures below were measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D, 8C/16T), commit 5c8323c, 384-dim embeddings; the commands are in [BENCHMARKS.md](BENCHMARKS.md). Exceptions are marked where they appear: the HDF5 Core I/O table immediately below is from a separate, independently reproduced run (see its own hardware note), and the HNSW `f32`/`i8` table and the in-memory `i8` column were not re-measured on 2026-09-24. - -### HDF5 Core I/O (vs libhdf5 1.14.6) - -*Benchmark numbers are being validated in collaboration with engineers from the HDF5 Group to confirm methodology and reproducibility.* - -Figures below are from an independent reproduction run on a second machine (AMD Ryzen 7 7800X3D, 2026-08-03). Full methodology, the original i7-12650H run, and two additional benchmarks added to close prior coverage gaps (an I/O-inclusive metadata-open comparison and an honest zero-copy-mmap measurement) are in [BENCHMARKS.md § Independent Validation](BENCHMARKS.md#independent-validation-tank-ryzen-7-7800x3d-2026-08-03). - -| Operation | ClawhDF5 | libhdf5 | Speedup | -|-----------|----------|---------|---------| -| Attribute write (128 attrs) | 85.2 µs | 877 µs | **10.3×** | -| Group create (64 groups) | 130 µs | 1.37 ms | **10.6×** | -| Chunked write, deflate-6 (512×512 f32) | 1.44 ms | 65.0 ms | **45.3×** | -| Sequential read (100K f32) | 23.3 µs | 63.6 µs | **2.7×** | -| Sequential write (100K f32) | 210 µs | 189 µs | **≈ tie** | - -The chunked-write row was re-measured on the same machine on 2026-09-23, after -the default deflate backend became pure-Rust zlib-rs: 1.46 ms against -libhdf5's 51.4 ms (**35×**), and 1.48 ms with zlib-ng. libhdf5's own time on -that machine moved from 65.0 to 51.4 ms between the two dates, which is most -of the difference from 45×; compare same-day numbers only. - -### Vector Search - -**HNSW (the default backend for `hybrid_search`)** — `search_harness`, clustered -384-dim data, M = 16, ef_construction = 64, recall measured against an exact scan. -See [BENCHMARKS.md § Search harness](BENCHMARKS.md#search-harness-baseline-v230) -and [§ Quantising the index copy](BENCHMARKS.md#quantising-the-index-copy-quantized_index): - -| N = 100K, ef = 64 | recall@10 | QPS | build | -|---|---:|---:|---:| -| `f32` index | 0.9945 | 13 399 | 3.2 s | -| `i8` index + exact re-score (**default for new stores**) | 0.9940 | **21 848** | **1.8 s** | - -Before the v2.4.0 neighbour-selection fix, recall@10 at 100K was 0.31. These -two rows are a paired comparison (medians of alternating runs, same binary). -A single `f32` run on 2026-09-24 measured recall 0.9945, 19 001 QPS and a -2.7 s build; the int8 row was not re-run, so the pair has not been re-checked -([§ Quantising the index copy](BENCHMARKS.md#quantising-the-index-copy-quantized_index)). - -**Brute-force and IVF paths** (Criterion, tank, 2026-09-24): - -| Scale | Flat | IVF (nprobe=10) | IVF-PQ | MemX¹ (claimed, end-to-end) | -|-------|------|-----------------|--------|----------| -| 1K | **47.4 µs** | — | — | — | -| 10K | 500.5 µs | **24.8 µs** | — | — | -| 100K | 6.58 ms | 592 µs | **869 µs** | <90 ms | - -> These replace figures from the original i7-12650H run (flat 54 µs / 753 µs / -> 11.4 ms); a 2026-08-05 run on tank had already matched the new ones — see -> [BENCHMARKS.md § Vector Search Latency](BENCHMARKS.md#vector-search-latency). - -### Agent Memory Operations - -| Operation | Latency | Scale | -|-----------|---------|-------| -| Hybrid search (`HDF5Memory::hybrid_search`, p50) | **0.07 ms** / 0.49 ms / 4.69 ms | 1K / 10K / 100K records | -| BM25 keyword search | **20.4 µs** | 1K records | -| Knowledge graph BFS | **23.1 µs** | 1K entities | -| Spreading activation | **10.1 µs** | 100 entities | -| Temporal range query | **622 ns** | 10K timestamps | -| Consolidation cycle | **115.2 µs** | 1K records | -| Cross-modal search (exact scan, 2 embeddings per record) | **842.0 µs** / 8.44 ms | 1K / 10K records | -| Memory write (WAL) | **26.1 µs** | per record (group-commit append; HDF5 batched at flush) | -| Importance gate | **57.6 ns** | per record (trivial skip) | - -The old 18 µs WAL write was undated, from another machine: v2.3.0 measures -24.3 µs on the same hardware as this table, the same as an `f32` store today. -`float16` stores (the new default) add ~2 µs for rounding; the int8 index adds -nothing. See [BENCHMARKS.md § Write Path](BENCHMARKS.md#write-path). -Knowledge-graph traversal was briefly 6.5x slower (155 µs) until this re-run -found and fixed an adjacency index rebuilt on every traversal; see -[§ Knowledge Graph](BENCHMARKS.md#knowledge-graph). - -### Chunked Write Throughput (codec comparison) - -Measured with Criterion on f32 matrices. Auto-shuffle is applied before all compression codecs -by default (AoS→SoA byte transpose, +157–204% throughput for float data): - -| Codec | 128×128 f32 | 512×512 f32 | Notes | -|-------|-------------|-------------|-------| -| Zstd level 3 | **148 µs / 422 MiB/s** | **1.34 ms / 748 MiB/s** | With auto-shuffle | -| Deflate level 6 | 153 µs / 407 MiB/s | 1.39 ms / 719 MiB/s | With auto-shuffle | -| Pcodec | 528 µs / 118 MiB/s | 1.69 ms / 591 MiB/s | Best compression ratio | - -Use `.with_zstd(3)` or `.with_deflate(6)` for write-heavy workloads — both now perform at ~720–750 MiB/s on large matrices. Use `.with_pcodec()` for write-once/read-many workloads where compression ratio matters more than encode speed. Disable auto-shuffle with `.without_shuffle()` for byte arrays that don't benefit from AoS→SoA transposition. - -> ¹ MemX ([arxiv:2603.16171](https://arxiv.org/abs/2603.16171), March 2026): Rust + libSQL, claims <90ms at 100K records. **Not like-for-like:** MemX's figure is *end-to-end* (embeddings + FTS5 + four-factor re-ranking); ours is a *single component* (raw vector search), so the two columns are not comparable and no ratio is given. See [BENCHMARKS.md](BENCHMARKS.md#comparison-to-memx-arxiv260316171). - -### LongMemEval Retrieval Recall - -Evaluated against the full **`longmemeval_s`** haystack — all 500 questions, 47.7 -sessions and 493.5 turns each, with only 4.0% of haystack sessions being evidence -sessions. See [BENCHMARKS.md § LongMemEval -Results](BENCHMARKS.md#longmemeval-results) for the full scoring-target -declaration: - -| Mode | Turn-Level Hit@5 | Session-Level Hit@5 | -|------|------------------|---------------------| -| BM25 only | 75.0% | 93.6% | -| Vector only (MiniLM) | 71.8% | 94.2% | -| Hybrid (0.4/0.6, tuned) | **81.4%** | **96.8%** | - -Hybrid is the strongest configuration, which is what running two retrieval stages -is for. The weights matter more than the stages: a sweep of `vector_weight` from -0.0 to 1.0 found the old `0.7/0.3` default is **strictly dominated** by -`0.4/0.6` — better on Hit@1, Hit@5, Hit@10 and MRR at both granularities. Since -v2.5.0 `0.4/0.6` is the default (`hybrid::DEFAULT_FUSION`, used by -`unified_search`, `hybrid_search_with` and `ClawhdfBackend`); callers that -pass weights to `hybrid_search` explicitly choose their own. Use `0.3/0.7` if -rank-1 precision matters most. Reciprocal rank fusion is selectable -(`hybrid::Fusion::Rrf`) but measured worse than the weighted sum. See -[BENCHMARKS.md § Weight sweep](BENCHMARKS.md#weight-sweep--full-haystack-n500). - -The benchmark's vector stage requires `clawhdf5-bench`'s `embeddings` feature -(real MiniLM embeddings); without it the vector stage is inert and only the BM25 row is produced, which is what every previously published -number here measured. - -On the easier `longmemeval_oracle` variant (evidence sessions only) the same -harness scores 84.4% turn-level Hit@5 / MRR 0.6597, reproduced identically on a -second machine. The 9.4-point gap is the cost of the real haystack, and is why the -full-haystack number is the one quoted here. - -This is **retrieval recall** (did the gold memory appear in the top-k), not the -official LongMemEval QA-accuracy metric — the two are not comparable, and -retrieval recall reported as QA accuracy typically overstates by 20–30 points. - -> **Previously reported here and now retracted:** session-level Hit@5 of 100.0% / -> MRR 1.0000, and a claim of beating MemX's 51.6%. Those session-level figures were -> degenerate on the oracle variant (any returned document is a hit by -> construction); the 93.6% above is a different, real measurement on a corpus where -> evidence sessions are 4.0% of the haystack. The MemX comparison stays withdrawn — -> MemX measures fact-level granularity over 220,349 records, which running the full -> haystack does not fix. Details in -> [BENCHMARKS.md](BENCHMARKS.md#retracted-session-level-recall-and-the-memx-comparison). - -> Enable embeddings via `hybrid_search(query_emb, text, 0.4, 0.6, k)` for substantially higher recall. The vector stage is served by the HNSW index by default (the `hnsw` feature is on by default); build with `--no-default-features --features float16` to fall back to an exact linear cosine scan. - -### Memory Footprint - -**On disk** — 384-dim `float16` embeddings (the default for new stores), -200-char text, `footprint_bench` -([BENCHMARKS.md § Memory Footprint](BENCHMARKS.md#memory-footprint-1)): - -| Records | File Size | Bytes/Record | Gzip-6 compressed | -|---------|-----------|--------------|-------------------| -| 1K | 810.4 KB | 829 B | 56.4 KB | -| 10K | 7.8 MB | 820 B | 471.3 KB | -| 100K | 76.7 MB | 803 B | 4.5 MB | - -The benchmark's synthetic embeddings and text are far more repetitive than -real data (only 40 distinct texts), so no column here is an expectation for -real data. The compressed column is an upper bound, and the Bytes/Record -column is optimistic too: it is not an uncompressed figure, because the store -always deflates its text (any string dataset of 4 KiB or more) whatever -`MemoryConfig::compression` says. The `float16` embeddings alone are 768 B per -record, so 200 characters of real text would take a record above 820 B. -This table used to show `f32` stores (1.7 KB per record, 169.8 MB at 100K); -those were not re-measured. The float16 study compares the two on the same -data: 100K × 384 records take 80.8 MiB as `float16` and 154.0 MiB as `f32`. - -**In memory** — a store reopened from disk, 384-dim `f32`, measured with a -counting allocator ([BENCHMARKS.md § Memory footprint](BENCHMARKS.md#memory-footprint)): - -| Records | Raw vectors | Reopened, `f32` index | Reopened, `i8` index (default) | -|---------|-------------|-----------------------|--------------------------------| -| 1K | 1 MiB | 4 MiB (2.40x) | 2 MiB (1.64x) | -| 10K | 15 MiB | 44 MiB (3.03x) | 27 MiB (1.81x) | -| 100K | 146 MiB | 399 MiB (2.72x) | **256 MiB (1.74x)** | - -Down from 505 MiB (3.44x) at 100K before v2.6.0, when the cache held every -embedding twice. The `f32` column was re-measured on 2026-09-24 and reproduced -exactly; the `i8` column was not re-run. - -### Consolidation Efficiency - -1,000 records (10 signal + 990 noise), `working_capacity = 100` -([BENCHMARKS.md § Consolidation Efficiency](BENCHMARKS.md#consolidation-efficiency)): - -| Metric | Before | After | Delta | -|--------|--------|-------|-------| -| Records in store | 1,000 | 100 | −90% | -| Hit@1 recall (signal records) | 100% | 100% | no loss | -| Search latency (avg) | 2.22 ms | 0.24 ms | **9.3x faster** | - -The consolidation cycle that does this took 0.13 ms; a cycle over 10K records -takes 2.81 ms and over 100K 46.7 ms. - -**Full benchmark details: [BENCHMARKS.md](BENCHMARKS.md)** - ---- - -## Agent Memory Architecture - -ClawhDF5's agent memory engine draws on 15+ recent papers on agentic memory systems (see [Research Foundation](#research-foundation)). - -``` - ┌─────────────────┐ - │ Agent Query │ - └────────┬────────┘ - │ - ┌─────────────────▼──────────────────┐ - │ HDF5Memory::search │ - │ optional source-channel filter │ - │ HNSW vector + BM25 keyword │ - │ weighted fusion (0.4 / 0.6) │ - │ × √(Hebbian activation) │ - └─────────────────┬──────────────────┘ - │ opt-in (SearchOptions); - │ ClawhdfBackend turns both on - ┌─────────────────▼──────────────────┐ - │ Multi-factor re-ranking │ - │ relevance · recency · authority · │ - │ activation │ - ├────────────────────────────────────┤ - │ Confidence rejection │ - │ (suppress bad matches) │ - └─────────────────┬──────────────────┘ - │ - ┌────────────────────────────▼────────────────────────────┐ - │ In memory │ - │ cache (flat f32 embeddings) · BM25 index · HNSW index │ - │ provenance ledger + anomaly alerts (session-scoped) │ - └────────────────────────────┬────────────────────────────┘ - │ WAL append; checkpoint - ┌────────────────────────────▼────────────────────────────┐ - │ agent_memory.h5 /meta · /memory · /sessions · │ - │ /knowledge_graph │ - │ agent_memory.h5.wal chained-CRC write-ahead log │ - │ agent_memory.h5.ann HNSW graph (derived, rebuildable) │ - │ agent_memory.h5.lock single-writer lock │ - └─────────────────────────────────────────────────────────┘ -``` - -Consolidation tiers (Working → Episodic → Semantic), the knowledge-graph -algorithms, temporal and multi-modal indexes are library components you drive -directly; the store persists the records, sessions and graph they work over. - -### Module Overview - -| Module | What It Does | -|--------|-------------| -| **`knowledge`** | Entity/relation graph with BFS traversal, spreading activation, fuzzy (Levenshtein) entity resolution | -| **`consolidation`** | Three-tier memory (Working → Episodic → Semantic) with importance scoring, novelty, and time-decay | -| **`hybrid`** | Vector + BM25 fusion. Default is a min-max-normalised weighted sum, vector 0.4 / keyword 0.6 (`hybrid::DEFAULT_FUSION`, tuned on LongMemEval); RRF is available via `Fusion::Rrf` / `hybrid_search_with`. The vector stage uses the HNSW index by default (`hnsw` feature); disable with `--no-default-features --features float16` for an exact linear scan | -| **`reranker`** | Multi-factor re-ranking: retrieval relevance (leads, weight 1.0), temporal recency, source authority, activation weight. Opt-in via `SearchOptions::with_rerank`; on in `ClawhdfBackend` | -| **`confidence`** | Low-confidence rejection — suppresses spurious recalls when nothing matches. Opt-in via `SearchOptions::with_confidence`; on in `ClawhdfBackend` | -| **`temporal`** | Sorted timestamp index, session DAG, entity timeline, temporal query hints | -| **`multimodal`** | Cross-modal search across text/image/audio/video embeddings | -| **`signing`** | Ed25519-signed checkpoints: SHA-256 per record in a Merkle tree, plus hashes of settings, sessions and the knowledge graph; `HDF5Memory::verify` names any edited record | -| **`provenance`** | Source attribution and an unkeyed FNV-1a content hash per record, held in memory for the session, for detecting accidental corruption (not tamper-proof) | -| **`anomaly`** | Write rate limiting, 15 injection-pattern detectors, source-distribution analysis. Alerts never block a save; drain them with `take_anomaly_alerts` | -| **`openclaw`** | `ClawhdfBackend`: a Markdown-oriented backend (ingest by section, search, read back by path, export). Named for OpenClaw, but **not an OpenClaw plugin** — see [docs/openclaw.md](docs/openclaw.md) | -| **`vector_search`** | Flat cosine, pre-normed, SIMD, BLAS, GPU, parallel search paths | -| **`ivf` / `pq`** | Standalone IVF and IVF-PQ indexes (benchmarked to 100K vectors); not used by `HDF5Memory`, whose ANN index is HNSW | -| **`bm25`** | Incremental Okapi BM25 inverted index, kept for the life of the store; optional stemming | -| **`query_expand`** | Synonym / acronym / temporal query expansion | -| **`entity_extract`** | Rule-based entity extraction from text chunks into the knowledge graph | -| **`wal`** | Write-ahead log (v4) with a chained CRC32 per entry, so a corrupted, reordered, duplicated or spliced entry stops replay; checkpoints record a WAL mark so nothing is applied twice. Appends are not fsynced | -| **`memory_strategy`** | Pluggable strategies: save-every, semantic-shift, user-correction detection | -| **`decision_gate`** | Sub-microsecond trivial/substantive classification | -| **`ephemeral`** | In-memory TTL/LFU working tier | -| **`async_memory`** | Tokio-based async wrapper over the memory store (`async` feature) | - ---- - -## Quick Start - -### HDF5 File I/O - -```rust -use clawhdf5::{File, FileBuilder, AttrValue}; - -// Write -let mut builder = FileBuilder::new(); -builder.create_dataset("temperatures") - .with_f64_data(&[22.5, 23.1, 21.8]) - .with_shape(&[3]); -builder.write("output.h5")?; - -// Read -let file = File::open("output.h5")?; -let ds = file.dataset("temperatures")?; -let values = ds.read_f64()?; -assert_eq!(values, vec![22.5, 23.1, 21.8]); -``` - -### Groups and links - -```rust -use clawhdf5::{AttrValue, FileBuilder}; - -let mut b = FileBuilder::new(); -// A path creates its missing intermediate groups, as in h5py. -b.create_dataset("run/2026/temps").with_f64_data(&[22.5, 23.1]); -// Builders nest; a group added at an existing path is merged into it. -let mut run = b.create_group("run"); -run.set_attr("operator", AttrValue::String("ana".into())); -let mut cal = run.create_group("calibration"); -cal.track_order(true); // h5py lists members in insertion order -cal.create_dataset("offset").with_f64_data(&[0.1]); -run.add_group(cal.finish()); -b.add_group(run.finish()); -b.add_soft_link("latest", "/run/2026"); // h5py.SoftLink -b.add_hard_link("temps", "/run/2026/temps"); // f["temps"] = f["run/2026/temps"] -b.add_external_link("raw", "raw.h5", "/data"); -b.write("groups.h5")?; -``` - -A group holds at most 65 535 links; more is an error, as is a link over -65 515 bytes (a very long soft-link target) in a group of more than 8 links. - -### Modifying an existing file - -```rust -use clawhdf5::{AttrValue, FileEditor, Selection}; - -// A file from h5py or clawhdf5, dataset "x" chunked with maxshape=(None,). -let mut ed = FileEditor::open("data.h5")?; // exclusive lock, like libhdf5 -ed.resize("x", &[1100])?; // h5py: ds.resize((1100,)) -let sel = Selection::Hyperslab { start: vec![1000], stride: vec![1], count: vec![100], block: vec![1] }; -ed.write_values("x", &sel, &[0.5f64; 100])?; // ds[1000:1100] = 0.5 -ed.set_attr("x", "units", &AttrValue::String("m/s".into()))?; -ed.resize("x", &[900])?; // shrinking prunes chunks, like h5py -``` - -Each call changes the file in place (no rewrite) and syncs it. Any chunk -index (version-2 B-trees for several unlimited dimensions included) and -attributes in compact or dense storage are handled as libhdf5 handles -them; space an edit frees is reused by later edits of the same editor. What it -cannot change safely is refused before anything is written; see -[known issues](docs/known-issues.md) for the limits. - -### Reading remote files - -[`clawhdf5-remote`](crates/clawhdf5-remote/README.md) opens a file on an -HTTP server (or, with its `s3`/`gcs`/`azure` features, in an object store) -without downloading it: the read API is the same `clawhdf5::File`, and -only the bytes an operation needs are fetched, by `Range` requests through -a block cache (1 MiB blocks; opening fetches the first one). A file that -changes on the server while it is open is an error, never a mix of old and -new bytes. - -```rust -let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/tall.h5")?; -let values = file.dataset("/g2/dset2.1")?.read_f64()?; -``` - -To try it without a server of your own, the crate's test server serves a -directory with range support: - -```bash -cargo run -p clawhdf5-remote --example range_server -- crates/clawhdf5/tests/fixtures 127.0.0.1:8000 -# in another shell: list the file, read one dataset, print what it cost -cargo run -p clawhdf5-remote --example read_url -- http://127.0.0.1:8000/tall.h5 /g2/dset2.1 -``` - -```text -/g1 group -/g2 group -/g2/dset2.1 dataset [10] F32 -/g2/dset2.2 dataset [3, 5] F32 -/g1/g1.1 group -/g1/g1.2 group -/g1/g1.2/g1.2.1 group -/g1/g1.1/dset1.1.1 dataset [10, 10] I32 -/g1/g1.1/dset1.1.2 dataset [20] I32 -/g2/dset2.1: 10 values, first [1.0, 1.100000023841858, 1.2000000476837158, ...] -1 range requests (the one at open included), 9968 bytes fetched, 9968 bytes cached -``` - -(`tall.h5` is 9 968 bytes, so the first block holds all of it.) `h5rs` -built with `--features remote` takes the same URLs: -`h5rs ls -r http://127.0.0.1:8000/tall.h5`. Plain HTTP builds no C; -`https://` is the `https` feature (rustls with ring, which compiles C). -Limits are in [known issues](docs/known-issues.md). - -### Following a file a SWMR writer is appending to - -`File::open_swmr` reads a file that a libhdf5 writer in SWMR mode (h5py -`f.swmr_mode = True`) is still appending to, as h5py's -`File(path, "r", swmr=True)` does: `Dataset::refresh()` picks up the new -extent, every read reads the chunk index as it is now, and a read that -races the writer (a checksum that fails mid-flush) is retried, up to 100 -attempts as in libhdf5, and never returned torn. - -```rust -use std::time::{Duration, Instant}; - -let file = clawhdf5::File::open_swmr("live.h5")?; -let mut ds = file.dataset("samples")?; -let (mut seen, mut last_growth) = (0, Instant::now()); -// Stop when the writer closes the file, or when the dataset has not grown -// for a minute: a writer that crashed or was killed never clears the -// SWMR-write flag, so `swmr_writer_active()` alone can stay true forever. -while file.swmr_writer_active()? && last_growth.elapsed() < Duration::from_secs(60) { - ds.refresh()?; - let n = ds.shape()?[0]; - if n > seen { - // read rows seen..n ... - (seen, last_growth) = (n, Instant::now()); - } - std::thread::sleep(Duration::from_millis(100)); -} -ds.refresh()?; // the final extent -``` - -`swmr_writer_active()` reads the superblock's SWMR-write flag, which libhdf5 -clears only when the writer closes the file; a file whose writer died keeps -it set (as the mid-write copy in `tests/fixtures/swmr_mid_write.h5` does), -so a follower needs its own stop condition, like the idle timeout above. - -Design and limits: [docs/design/swmr.md](docs/design/swmr.md). - -### Python - -`crates/clawhdf5-py` is a Python package (PyO3 + numpy) that reads HDF5 with -an h5py-shaped API and no libhdf5. It is not on PyPI; build it with +The Python package is not on PyPI; build it with [maturin](https://www.maturin.rs) into a virtualenv: ```bash python -m venv .venv && . .venv/bin/activate pip install maturin numpy -maturin develop --release -m crates/clawhdf5-py/Cargo.toml +maturin develop --release -m crates/clawhdf5-py/Cargo.toml # add --features https for https:// python -c "import clawhdf5; print(clawhdf5.__version__)" ``` +`h5rs`: `cargo install --path crates/clawhdf5-tools` (add +`--features remote` for URLs). + +## Quick start: Rust + +```rust +use clawhdf5::{AttrValue, File, FileBuilder, Selection}; + +// Write: a chunked, deflate-compressed 2-D dataset that can grow along axis 0. +let data: Vec = (0..1000 * 64).map(|i| i as f64).collect(); +let mut b = FileBuilder::new(); +b.create_dataset("run/temps") // intermediate groups are created, as in h5py + .with_f64_data(&data) + .with_shape(&[1000, 64]) + .with_maxshape(&[u64::MAX, 64]) // u64::MAX = unlimited + .with_chunks(&[100, 64]) + .with_deflate(6) + .set_attr("units", AttrValue::String("K".into())); +b.write("data.h5")?; + +// Read: whole datasets, or a hyperslab (only the chunks it touches are decoded). +let file = File::open("data.h5")?; +let ds = file.dataset("run/temps")?; +assert_eq!(ds.shape()?, vec![1000, 64]); +let all = ds.read_f64()?; +let rows = ds.read_f64_selection(&Selection::Hyperslab { + start: vec![10, 0], stride: vec![1, 1], count: vec![2, 64], block: vec![1, 1], +})?; +assert_eq!(rows.len(), 128); +println!("{:?} {:?}", ds.attr("units")?, file.root().groups()?); +``` + +Edit that file in place — no rewrite; each call is written and synced +before it returns, and anything the editor cannot do safely is refused +before a byte is written: + +```rust +use clawhdf5::{AttrValue, FileEditor, Selection}; + +let mut ed = FileEditor::open("data.h5")?; // exclusive lock, as libhdf5 takes +ed.resize("run/temps", &[1100, 64])?; // h5py: ds.resize((1100, 64)) +let sel = Selection::Hyperslab { + start: vec![1000, 0], stride: vec![1, 1], count: vec![100, 64], block: vec![1, 1], +}; +ed.write_values("run/temps", &sel, &vec![0.5f64; 100 * 64])?; // ds[1000:1100] = 0.5 +ed.set_attr("run/temps", "calibrated", &AttrValue::I64(1))?; +``` + +The editor changes chunk indexes and heaps as libhdf5 does (the tests +compare index shapes and heap bookkeeping with libhdf5's, and check every +edited file with h5py, h5dump and `h5rs check`). More in +[docs/QUICKSTART.md](docs/QUICKSTART.md): groups and links, filters, +strings and variable-length data, NetCDF-4. + +## Quick start: Python + ```python import numpy as np import clawhdf5 with clawhdf5.File("data.h5", "r") as f: - print(list(f.keys())) # sorted member names, like h5py - ds = f["group/temperatures"] # relative or absolute ("/group/...") paths - print(ds.shape, ds.dtype) # dtype is the numpy dtype h5py reports - block = ds[100:200, ::4] # a small selection reads only its chunks + print(list(f.keys())) # member names, like h5py + ds = f["group/temperatures"] # relative or absolute paths + print(ds.shape, ds.dtype, ds.chunks) + block = ds[100:200, ::4] # a small selection decodes only its chunks row = ds[-1] # integers drop the axis picked = ds[[1, 5, 9], :] # one increasing index list per key units = ds.attrs["units"] # attributes come back as h5py returns them everything = np.asarray(ds) + ids = f["table"]["id"] # compound -> structured array; one field - records = f["table"] # compound -> numpy structured array - ids = records["id"] # one field - -# A file on a web server: range requests through a block cache, nothing -# downloaded up front; the same read API. The GIL is released while waiting. -with clawhdf5.File("http://data.example.org/run42.h5") as f: - first = f["group/temperatures"][0] -f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024, - headers={"Authorization": "Bearer ..."}) -``` - -An existing file opened with `"r+"` is edited in place (through -`clawhdf5::FileEditor`), with h5py's indexing, broadcasting and numeric -conversion; each edit is on disk when the statement returns: - -```python -with clawhdf5.File("data.h5", "r+") as f: +with clawhdf5.File("data.h5", "r+") as f: # edited in place, h5py semantics f["group/temperatures"][100:200, ::4] = 0.0 f["series"].resize(5000, axis=0) # chunked datasets, within maxshape - f["series"][4000:] = new_values + f["series"][4000:] = np.ones(1000) f["group"].attrs["calibrated"] = True + +with clawhdf5.File("http://data.example.org/run42.h5") as f: # range requests, no download + first = f["group/temperatures"][0] ``` -Creating or deleting datasets, groups and attributes in an existing file is -not supported (`NotImplementedError`); limits are in -[known issues](docs/known-issues.md). +Reads release the GIL, so Python threads read in parallel. The test suite +(`crates/clawhdf5-py/tests`) compares every read and every edit with h5py, +locally and over HTTP. Types, keys, writing (`'w'`: numeric arrays) and +limits: [crates/clawhdf5-py/README.md](crates/clawhdf5-py/README.md). -The default build reads `http://` URLs only; build with -`maturin develop --release --features https` (rustls with ring, which -compiles C) for `https://`, and `--features s3` (or `gcs`, `azure`) for -object-store URLs. +## Remote files, the browser, SWMR -Reads cover integers and IEEE floats of every width in either byte order, -`bool`, enums, complex, fixed and variable-length strings, variable-length -sequences, opaque, HDF5 array types and compounds; other types (references, -bitfields, ...) raise `TypeError` instead of returning guessed data. Keys -follow h5py (negative steps, `None` and boolean masks are refused). The -read itself runs with the GIL released, so Python threads read in parallel. -A selection whose bounding box covers at most half the dataset decodes only -the chunks (or contiguous rows) that box overlaps; a larger one — including -a strided slice across the whole dataset — decodes the whole dataset, as -do datasets that are compact, virtual, unwritten, or chunked with a -non-default fill value (`docs/known-issues.md`). An index list is read one -group of neighbouring chunks at a time. -Writing (`File(path, "w")`, `create_dataset`, `create_group`, `attrs[...] =`) -covers `float64`, `float32`, `int64`, `int32` and `uint8` arrays. The tests -in `crates/clawhdf5-py/tests` compare every read and every in-place edit -with h5py; run them with -`pip install pytest h5py && pytest crates/clawhdf5-py/tests`. - -### Agent Memory +**Remote files** ([`clawhdf5-remote`](crates/clawhdf5-remote/README.md), +design: [docs/design/range-reads.md](docs/design/range-reads.md)). The +same `clawhdf5::File`, over HTTP range requests or an object store, through +a block cache (1 MiB blocks, LRU budget, concurrent requests deduplicated, +runs coalesced into parallel requests). The file is pinned by ETag / +Last-Modified and length: a file that changes on the server is an error, +never a mix of old and new bytes. ```rust -use clawhdf5_agent::{HDF5Memory, MemoryConfig, MemoryEntry, AgentMemory}; +let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/tall.h5")?; +let values = file.dataset("/g2/dset2.1")?.read_f64()?; +``` -// Create memory store -let config = MemoryConfig::new("agent.h5".into(), "my-agent", 384); -let mut memory = HDF5Memory::create(config)?; +Try it with the crate's test server: -// Save a memory -memory.save(MemoryEntry { - chunk: "User prefers dark mode and vim keybindings.".into(), - embedding: embed("User prefers dark mode..."), // your embedder - source_channel: "chat".into(), - timestamp: now(), - session_id: "session-001".into(), - tags: "preference".into(), -})?; +```bash +cargo run -p clawhdf5-remote --example range_server -- crates/clawhdf5/tests/fixtures 127.0.0.1:8000 +cargo run -p clawhdf5-remote --example read_url -- http://127.0.0.1:8000/tall.h5 /g2/dset2.1 +``` -// Hybrid search: vector + BM25, weighted 0.4 / 0.6 (the measured default) -let results = memory.hybrid_search(&query_embedding, "user preferences", 0.4, 0.6, 5); -for result in results { - println!("[{:.3}] {}", result.score, result.chunk); +Plain HTTP builds no C; `https` (rustls + ring) and `s3` / `gcs` / `azure` +are opt-in features. The cloud backends are built and their URL handling +tested, but have not been run against a real bucket. + +**In the browser** (`clawhdf5-wasm`, demo and API in +[examples/wasm-viewer](examples/wasm-viewer/README.md)): `open(bytes)` reads +a file held in memory; `openUrl(url)` reads a file on a web server by range +requests, fetching only what each call needs, on the main thread (no +worker, no synchronous XHR). In a 200 MB h5py file, listing the root, +reading two small datasets, a group's attributes, the large dataset's shape +and a 10-value window of it took 5 requests and 6 MiB (tank, 2026-09-27, +the viewer's Node + Chromium test suite). + +**SWMR reading** (design: [docs/design/swmr.md](docs/design/swmr.md)). +`File::open_swmr` follows a file a libhdf5 SWMR writer (h5py +`f.swmr_mode = True`) is still appending to, as h5py's +`File(path, "r", swmr=True)` does: `Dataset::refresh()` picks up the new +extent, and a read that races the writer (a checksum failing mid-flush) is +retried, up to 100 attempts as in libhdf5, never returned torn. Tested live +against an h5py writer appending to Extensible-Array and v2-B-tree indexed +datasets for 2 500 steps (20 000 in a release build), beside h5py's own +SWMR reader. clawhdf5 does not write SWMR files. + +```rust +let file = clawhdf5::File::open_swmr("live.h5")?; +let mut ds = file.dataset("samples")?; +while file.swmr_writer_active()? { // add your own timeout: a writer that died keeps the flag set + ds.refresh()?; + let n = ds.shape()?[0]; + // read the new rows ... + std::thread::sleep(std::time::Duration::from_millis(100)); } ``` -### Search Options +## h5rs tools -```rust -use clawhdf5_agent::SearchOptions; -use clawhdf5_agent::confidence::ConfidenceConfig; -use clawhdf5_agent::reranker::ReRankConfig; - -// Only memories from these source channels; still a full page of k results. -let work = memory.search( - &query_embedding, - "deadline", - &SearchOptions::new(5).with_sources(["slack", "email"]), -); - -// Re-rank by relevance, recency, source authority and activation, then drop -// low-confidence results — the pipeline ClawhdfBackend runs. -let careful = memory.search( - &query_embedding, - "user preferences", - &SearchOptions::new(5) - .with_rerank(ReRankConfig::default()) - .with_confidence(ConfidenceConfig::default()), -); -``` - -### Signed Checkpoints - -```rust -use clawhdf5_agent::signing; - -// Once, somewhere safe: keep the secret key, publish the public key. -let key = signing::generate_key(); -let public = key.verifying_key(); - -// Every checkpoint is signed from now on. The key is never written to disk; -// a signed store refuses to checkpoint without it. -memory.set_signing_key(key); -memory.flush_wal()?; - -// Anyone holding the public key can check the file, e.g. after copying it. -let report = HDF5Memory::verify(std::path::Path::new("agent.h5"), &public)?; -assert!(report.is_valid()); -// On a tampered file: report.changed_records lists the records that differ. -``` - -The signature covers every record (text, embedding as stored, channel, -timestamp, session, tags, deleted flag, activation), the store's settings, -its sessions and its knowledge graph — a change made with any tool is caught. -It covers checkpoints, not saves still in the WAL -(`report.wal_entries_unsigned` counts those). CLI: `clawhdf5-cli keygen`, -`--signing-key ` on writing commands, and `verify --public-key`. -Signing adds about 20% to a checkpoint and 32 bytes per record to the file -([BENCHMARKS.md § Signed checkpoints](BENCHMARKS.md#signed-checkpoints)). - -### Knowledge Graph - -```rust -use clawhdf5_agent::knowledge::KnowledgeCache; - -let mut kg = KnowledgeCache::new(); - -// Add entities -let alice = kg.add_entity("Alice", "person", -1); -let bob = kg.add_entity("Bob", "person", -1); -let acme = kg.add_entity("Acme Corp", "company", -1); - -// Add relations -kg.add_relation(alice, acme, "works_at", 1.0); -kg.add_relation(bob, acme, "works_at", 1.0); -kg.add_relation(alice, bob, "manages", 0.8); - -// Traverse -let neighbors = kg.bfs_neighbors(alice, 2); // 2-hop neighborhood - -// Spreading activation — find related entities -let activated = kg.spreading_activation(&[alice], 0.5, 0.01, 5); - -// Entity resolution — fuzzy matching -let (id, created) = kg.resolve_or_create("alice", "person", -1, 2); -// id == alice, created == false: matched the existing entity (Levenshtein distance ≤ 2) -``` - -### Memory Consolidation - -```rust -use clawhdf5_agent::consolidation::*; - -let config = ConsolidationConfig::default(); -let mut engine = ConsolidationEngine::new(config); - -let now = 1_700_000_000.0; // seconds since the epoch - -// Add memories — automatically scored for importance. -// Elevated sources (System, …) go through a separate, explicit API. -let id = engine.add_memory("User prefers dark mode".into(), vec![0.1, 0.2, ...], UntrustedSource::User, now); -engine.add_trusted_memory("ok".into(), vec![0.0, 0.0, ...], TrustedSource::System, now); - -// Access a memory (reactivates it) -engine.access_memory(id, now); - -// Run consolidation cycle -engine.consolidate(now); -let stats = engine.get_stats(); -// Working memories promote to Episodic (if important enough) -// Episodic memories promote to Semantic (if accessed enough) -// Low-decay memories get evicted when tiers are full -``` - -### Temporal Queries - -```rust -use clawhdf5_agent::temporal::*; - -let mut index = TemporalIndex::new(); -index.insert(1, 1700000000.0); // record 1 at timestamp -index.insert(2, 1700003600.0); // record 2, 1 hour later - -// Range query — "what happened between 2pm and 5pm?" -let ids = index.range_query(1700000000.0, 1700010800.0); - -// Latest 10 memories -let recent = index.latest(10); -``` - -### Markdown Backend - -`ClawhdfBackend` ingests Markdown by section and searches it with the full -pipeline. It is a library API — clawhdf5 is **not** an OpenClaw memory plugin -([docs/openclaw.md](docs/openclaw.md)). Sections stored this way carry no -embedding, so their search is keyword-only unless you save records with -vectors through `save_entry`. - -```rust -use clawhdf5_agent::openclaw::*; - -// Create backend -let mut backend = ClawhdfBackend::create(std::path::Path::new("memory.h5"), 384)?; - -// Ingest existing Markdown memory files -let md = std::fs::read_to_string("MEMORY.md")?; -let count = backend.ingest_markdown("MEMORY.md", &md)?; - -// Search (full pipeline: weighted vector + BM25 fusion → re-rank → confidence filter) -let results = backend.search("user preferences", &query_embedding, 5); - -// Export back to Markdown -let exported = backend.export_markdown("MEMORY.md")?; -``` - ---- - -## Crate Map - -``` -clawhdf5 workspace (19 crates, ~86K lines of Rust in src/, ~104K with tests - and benches; plus libaec-sys, an internal FFI bindings - crate for the optional szip feature) -│ -├── Core HDF5 -│ ├── clawhdf5-format — Binary parser/writer (no_std-capable), shared type definitions -│ ├── clawhdf5-io — I/O abstraction (file/memory readers; optional mmap, async, HSDS, MPI) -│ ├── clawhdf5-filters — Fast deflate path (zlib-ng); the filter registry and the lz4/zstd/pcodec/szip/LZF/bitshuffle/bzip2/Blosc/Blosc2 filters live in clawhdf5-format -│ ├── clawhdf5-derive — Proc macros -│ ├── clawhdf5 — High-level API -│ ├── clawhdf5-netcdf4 — NetCDF-4 support -│ ├── clawhdf5-accel — SIMD (AVX2, NEON incl. SDOT int8; AVX-512 behind `avx512`) -│ ├── clawhdf5-gpu — GPU compute (wgpu, hand-written WGSL compute shaders) -│ └── clawhdf5-remote — Remote files: HTTP(S) range requests, object stores, block cache -│ -├── Agent Memory -│ ├── clawhdf5-agent — Memory engine (24.7K lines, 32 modules; chained-CRC WAL) -│ ├── clawhdf5-ann — HNSW approximate nearest neighbor (default backend; f32 or int8 storage; `parallel` build) -│ ├── clawhdf5-migrate — SQLite → HDF5 migration -│ ├── clawhdf5-android — Android JNI bridge -│ └── clawhdf5-cli — CLI tool -│ -├── Bindings -│ ├── clawhdf5-py — Python (PyO3) -│ ├── clawhdf5-napi — Node.js (napi-rs) -│ └── clawhdf5-wasm — Browser (WebAssembly, wasm-bindgen; read-only; remote files by HTTP range requests) -│ -└── Tooling - ├── clawhdf5-tools — h5rs: ls, dump, stat, diff, check - └── clawhdf5-bench — Benchmark suite -``` - ---- - -## Research Foundation - -ClawhDF5's agent memory design draws from 15+ recent papers: - -| Paper | Key Insight | ClawhDF5 Module | -|-------|-------------|-----------------| -| **MemX** (2026) | Hybrid fusion + multi-factor re-ranking | `hybrid`, `reranker` | -| **Graph-Native Cognitive Memory** (2026) | Graph-structured memory (weighted, timestamped relations; entity timelines) | `knowledge`, `temporal` | -| **CraniMem** (2026) | Bounded hippocampal memory | `consolidation` | -| **D-MEM** (2026) | Surprise-gated storage (implemented as a novelty score) | `consolidation` | -| **SYNAPSE** (2025) | Spreading activation for recall | `knowledge` | -| **RAGdb** (2025) | Zero-dependency edge RAG | Architecture | -| **MemoryGraft** (2025) | Memory poisoning attacks | `anomaly`, `provenance` | -| **MemoryArena** (2026) | Multi-session benchmark | `temporal` | -| **AI Hippocampus** (2026) | Memory taxonomy survey | Overall design | - ---- - -## Feature Flags - -### `clawhdf5-agent` - -| Flag | Default | Description | -|------|---------|-------------| -| `float16` | **yes** | Half-precision cosine kernel (`cosine_similarity_f16`). Half-precision *storage* is the `MemoryConfig::float16` setting below, and needs no feature | -| `hnsw` | **yes** | HNSW approximate vector index for `hybrid_search` (via `clawhdf5-ann`); disable for an exact linear scan | -| `parallel` | **yes** | Parallel HNSW bulk build (same graph, ~3× faster on 16 cores) and Rayon brute-force search strategies | -| `zstd` | no | Compress embeddings with Zstd instead of deflate when `MemoryConfig::compression` is on (links libzstd) | -| `fast-math` | no | BLAS matrix-vector multiply | -| `accelerate` | no | Apple Accelerate / AMX (macOS) | -| `openblas` | no | OpenBLAS (Linux) | -| `gpu` | no | GPU search via wgpu | -| `async` | no | Tokio async with background flush | - -To opt out of the parallel build: `--no-default-features --features float16,hnsw`. -For an exact linear cosine scan instead of HNSW: `--no-default-features --features float16`. - -`MemoryConfig::hnsw_m`, `hnsw_ef_construction` and `hnsw_ef_search` tune the -vector index (16 / 64 / scale-with-`k` by default) and are stored with the -file. - -`MemoryConfig::quantized_index` (**on by default** for new stores) holds the -HNSW index's own copy of the embeddings as `i8`, roughly halving a loaded -store's memory (2.72x -> 1.74x the raw vectors at 100k x 384). Quantised -distances are approximate, so the query path re-scores the candidate pool -against the exact embeddings the store already holds, which keeps recall at the -`f32` index's level. It is also **faster**: 1.63x the queries per second at -equal recall on x86-64 (AVX2) and 1.18x on a Raspberry Pi 5 (NEON `SDOT`), with -index builds 1.8x and 2.3x faster respectively. Stores created before the -setting existed keep their `f32` index; opt out for new stores with -`quantized_index = false` or `clawhdf5-cli create --f32-index`. See -[BENCHMARKS.md § Quantising the index copy](BENCHMARKS.md#quantising-the-index-copy-quantized_index). - -`MemoryConfig::float16` (**on by default** for new stores) stores the -embeddings on disk as IEEE half precision (numpy `float16`): at 100K × 384 the -file drops from 154 to 81 MiB, checkpoints and opens get faster, and on the -full LongMemEval haystack with real MiniLM embeddings every retrieval metric -matches `f32`. Embeddings are rounded as they are saved, so the store searches -the same before and after a reopen; values must lie within ±65504. Existing -stores keep their setting. Opt out with `float16 = false` or -`clawhdf5-cli create --f32` — e.g. for unnormalised vectors. See -[BENCHMARKS.md § float16 embedding storage](BENCHMARKS.md#float16-embedding-storage-memoryconfigfloat16). - -### `clawhdf5-format` - -| Flag | Default | Description | -|------|---------|-------------| -| `std` | yes | Standard library (disable for `no_std`) | -| `deflate` | yes | Deflate compression | -| `checksum` | yes | Jenkins lookup3 verification | -| `provenance` | yes | SHA-256 provenance attributes | -| `zlib-rs` | **yes** | Pure-Rust deflate backend ([zlib-rs](https://github.com/trifectatechfoundation/zlib-rs)) | -| `fast-deflate` | no | zlib-ng deflate backend instead (C; needs `cmake`). Overrides `zlib-rs` when both are on | -| `system-zlib-decompress` | **yes** | Use Apple's system libz for decompression (macOS only; no effect elsewhere) | -| `parallel` | no | Parallel chunk encoding + compression (rayon) | -| `fast-checksum` | no | crc32fast-accelerated checksums | -| `lz4` | no | LZ4 block compression filter (id 32004) | -| `zstd` | no | Zstandard compression filter (id 32015) | -| `pcodec` | no | Pcodec lossless numerical codec (via `pco` crate). Private, unregistered filter id 480: **only clawhdf5 can read these datasets** (h5py/libhdf5 cannot). Files from clawhdf5 <= 2.7.0 used id 32023, which is registered to Granular BitRound; they still read. | -| `system-zlib` | no | System zlib backend for deflate (C) | -| `blake3_hash` | no | BLAKE3 content hashing for provenance | -| `szip` | no | SZIP filter (id 4) via libaec (C, through the internal `libaec-sys` crate) | -| `lzf` | **yes** | LZF filter (id 32000), h5py's built-in `compression="lzf"`: read and write. No dependencies | -| `bitshuffle` | no | Bitshuffle filter (id 32008) with its LZ4 and Zstandard modes: read and write. Pure Rust (lz4_flex, ruzstd) | -| `bzip2` | no | bzip2 filter (id 307): read and write. Pure Rust (the `bzip2` crate's libbz2-rs-sys backend compiles no C) | -| `blosc` | no | Blosc 1 filter (id 32001): reads BloscLZ, LZ4/LZ4HC, Snappy, Zlib and Zstandard frames with byte or bit shuffle; writes LZ4, Snappy, Zlib or Zstandard (not BloscLZ). Pure Rust | -| `blosc2` | no | Blosc2 filter (id 32026), read only: hdf5plugin's frames and B2ND (n-D) chunks, BloscLZ, LZ4/LZ4HC, Zlib and Zstandard, with shuffle, bit shuffle, delta or truncated precision. Pure Rust | -| `zfp` | no | ZFP filter (id 32013, H5Z-ZFP), read only: every mode (rate, precision, accuracy, reversible, expert) for int32, int64, float and double, 1-4-D, returning exactly libzfp's values. Pure Rust, no dependencies | -| `plugin-filters` | no | All six above | - -clawhdf5 cannot write Blosc2 or ZFP. Any other -filter can be supplied at run time with `filter_registry::register_filter` (a -decoder closure, or a `FilterCodec` that also encodes). The facade -(`clawhdf5`) forwards `lzf`, `bitshuffle`, `bzip2`, `blosc`, `blosc2`, `zfp` -and `plugin-filters`. Write -with `DatasetBuilder::with_lzf()`, `with_bitshuffle(..)`, `with_bzip2(..)` -and `with_blosc(..)`; h5py + hdf5plugin read the result (tested both ways in -`crates/clawhdf5/tests/plugin_filters_interop.rs`). The pure-Rust Zstandard -encoder has one level (about zstd's level 1); no speed or ratio claims are -made for these codecs. - -### `clawhdf5-ann` - -| Flag | Default | Description | -|------|---------|-------------| -| `parallel` | no | Batched bulk build runs neighbour planning and back-link pruning on a Rayon pool; the graph is identical with or without it (enabled by `clawhdf5-agent`'s default `parallel`) | - -### `clawhdf5-io` - -| Flag | Default | Description | -|------|---------|-------------| -| `mmap` | no | Memory-mapped reads (`memmap2`) | -| `async` | no | Tokio-based async I/O | -| `hsds` | no | HSDS (HDF REST service) client | -| `mpi-io` | no | MPI-backed I/O via the `mpi` crate | - -> **Parallel I/O (MPI) limitation:** `mpi-io`'s read path is a root-rank read -> followed by a broadcast, and its write path gathers all ranks' shards to -> rank 0 before writing — not true collective I/O -> (`MPI_File_read_at_all`/`write_at_all`). It does not provide I/O bandwidth -> that scales with rank count; true collective I/O is tracked as future work. - ---- - -## Building +`h5rs` (crate `clawhdf5-tools`) is a pure-Rust counterpart of the HDF5 +command-line tools: ```bash -# Default (pure Rust: no cmake or C compiler needed) -cargo build --workspace +h5rs ls -r file.h5 # like h5ls +h5rs dump file.h5 # like h5dump: DDL, or --json (hdf5-json) +h5rs stat file.h5 # like h5stat +h5rs diff a.h5 b.h5 # like h5diff +h5rs check --data file.h5 # structural and checksum validator +``` -# Agent memory with all accelerations (Linux) -cargo build -p clawhdf5-agent --features fast-math +`dump` output is byte-identical to h5dump's on the interop test files, and +the `ls`/`stat`/`diff` tests compare with h5ls, h5stat and h5diff. `check` +walks the file's structures, verifies their checksums (superblock, object +headers, v2 B-trees, fractal heaps, chunk indexes) and with `--data` +decodes every dataset; it validates with the library's own parsers, so it +accepts what they accept. With `--features remote` every subcommand takes a +URL. Details: [crates/clawhdf5-tools/README.md](crates/clawhdf5-tools/README.md). -# Agent memory with Apple Accelerate (macOS) -cargo build -p clawhdf5-agent --features "accelerate,gpu" +## Agent memory -# Tests -cargo test --workspace # all 1,850+ tests -cargo test -p clawhdf5-agent # agent memory tests -scripts/ci-test.sh # what CI runs: fmt, clippy matrix, tests, - # h5py/netCDF4 interop, no_std +`clawhdf5-agent` stores an agent's memories — text, embeddings, sessions, +a knowledge graph — in one HDF5 file (readable by h5py), with: -# The interop suites need a Python with h5py; on a PEP 668 system that has to -# be a virtualenv. `ci-test.sh` finds `.venv` on its own, or set -# CLAWHDF5_PYTHON. Without one they skip — set CLAWHDF5_REQUIRE_INTEROP=1 to -# make that a failure instead. +- **Hybrid search**: HNSW (clawhdf5-ann) vector + BM25 keyword, weighted + 0.4 / 0.6, optional source filter, re-ranking and confidence rejection. + On the full LongMemEval `longmemeval_s` haystack (500 questions, real + MiniLM embeddings) turn-level Hit@5 is **81.4%** — retrieval recall, not + the official QA-accuracy metric (tank, re-run 2026-09-27, + [BENCHMARKS.md](BENCHMARKS.md#longmemeval-results)). +- **Compact by default**: float16 embeddings on disk (48% smaller at 100K) + and an int8 index copy with exact re-scoring — 1.74x the raw vectors in + memory at 100K instead of 2.72x, and 1.63x the QPS at equal recall on + AVX2 (paired runs; see BENCHMARKS.md for which rows were re-run). +- **Durability**: a write-ahead log with a chained CRC per entry, + crash-safe checkpoints, a single-writer lock and a read-only open. WAL + appends are not fsynced: saves since the last checkpoint can be lost on + power failure. +- **Signed checkpoints**: Ed25519 over a SHA-256 Merkle tree of the + records, settings, sessions and graph; `HDF5Memory::verify` names the + edited records. + +```rust +use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions}; + +let mut memory = HDF5Memory::create(MemoryConfig::new("agent.h5".into(), "my-agent", 384))?; +memory.save(MemoryEntry { + chunk: "User prefers dark mode and vim keybindings.".into(), + embedding: embed("User prefers dark mode and vim keybindings."), // your embedder + source_channel: "chat".into(), + timestamp: now, + session_id: "session-001".into(), + tags: "preference".into(), +})?; +for r in memory.search(&embed("what editor?"), "editor preferences", &SearchOptions::new(5)) { + println!("[{:.3}] {}", r.score, r.chunk); +} +``` + +Architecture, every module, performance tables, feature flags, file +schema, CLI and SQLite migration: [docs/agent-memory.md](docs/agent-memory.md). + +## Crate map + +19 crates under `crates/`, plus `libaec-sys` (FFI for the optional SZIP +filter). + +| Crate | Role | +|---|---| +| **HDF5** | | +| `clawhdf5` | The facade: `File`, `FileBuilder`, `FileEditor`, `Dataset`, `Group`, SWMR reading | +| `clawhdf5-format` | The format itself (superblock, headers, B-trees, heaps, datatypes), the filter pipeline and registry, every codec but the deflate backends; `no_std`-capable | +| `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) | +| `clawhdf5-io` | I/O helpers: mmap, async, an HSDS client, `mpi-io` (not collective I/O) | +| `clawhdf5-remote` | HTTP(S) and object-store files through a block cache | +| `clawhdf5-netcdf4` | NetCDF-4 dimensions, variables, CF attributes | +| `clawhdf5-derive` | Derive macros for HDF5-serialisable structs | +| `clawhdf5-tools` | `h5rs`: `ls`, `dump`, `stat`, `diff`, `check` | +| **Bindings** | | +| `clawhdf5-py` | Python (PyO3 + numpy) | +| `clawhdf5-wasm` | Browser (wasm-bindgen), read-only | +| `clawhdf5-napi` | Node.js (unpublished; does not work, see known-issues) | +| `clawhdf5-android` | Android JNI bindings for the agent store | +| **Agent memory** | | +| `clawhdf5-agent` | The memory store | +| `clawhdf5-ann` | HNSW index (`f32` or `i8` storage) | +| `clawhdf5-accel` | SIMD kernels (AVX2, NEON incl. `SDOT`; AVX-512 behind a feature) | +| `clawhdf5-gpu` | Vector distance computation on the GPU (wgpu, WGSL); HDF5 I/O is CPU-only | +| `clawhdf5-migrate` | SQLite → agent store migration | +| `clawhdf5-cli` | The `clawhdf5` agent-memory CLI | +| `clawhdf5-bench` | Benchmarks and harnesses | + +## Building and testing + +```bash +cargo build --workspace # pure Rust: no cmake or C compiler needed +cargo test --workspace +scripts/ci-test.sh # what CI runs: fmt, clippy matrix, tests, interop, no_std, no-C check +conformance/run.sh # the conformance report (needs h5py, hdf5plugin, h5dump) +``` + +The interop suites need a Python with h5py (and netCDF4, xarray); on a +PEP 668 system that has to be a virtualenv, which `ci-test.sh` finds as +`.venv` or through `CLAWHDF5_PYTHON`. Without one they skip; set +`CLAWHDF5_REQUIRE_INTEROP=1` to make that a failure, as CI does: + +```bash python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 xarray - -# Benchmarks -cargo bench -p clawhdf5-agent # agent memory suite -cargo bench -p clawhdf5-bench # h5bench-equivalent I/O suite ``` ---- +CI (`.gitea/workflows/`) runs `ci-test.sh` on x86-64, lints and tests the +NEON code on aarch64, and runs the conformance corpus nightly. -## HDF5 File Schema +## Documentation -``` -agent_memory.h5 -├── /meta (attributes) -│ ├── schema_version: "1.0", edgehdf5_version -│ ├── agent_id, embedder, embedding_dim, chunk_size, overlap, created_at -│ ├── float16, compression, compression_level, compact_threshold, -│ │ hebbian_boost, decay_factor, wal_enabled, wal_max_entries -│ ├── quantized_index, hnsw_m, hnsw_ef_construction, hnsw_ef_search -│ ├── wal_applied_len, wal_applied_crc (WAL mark of the last checkpoint) -│ └── ann_generation (ties the .ann sidecar to this checkpoint) -├── /memory -│ ├── chunks: string[N] -│ ├── embeddings: f32[N × D], or f16 for a `float16` store -│ │ (chunked; deflate, or Zstd with the `zstd` -│ │ feature, when compression is on) -│ ├── source_channel: string[N] -│ ├── timestamps: f64[N] -│ ├── session_ids: string[N] -│ ├── tags: string[N] -│ ├── tombstones: u8[N] -│ ├── norms: f32[N] (pre-computed L2) -│ └── activation_weights: f32[N] (Hebbian) -├── /sessions -│ ├── ids, channels, summaries: string[S] -│ ├── start_idxs, end_idxs: i64[S] -│ └── timestamps: f64[S] -└── /knowledge_graph - ├── entity_ids, entity_emb_idxs: i64[E]; entity_names, entity_types: string[E] - ├── relation_srcs, relation_tgts: i64[R]; relation_types: string[R] - ├── relation_weights: f32[R]; relation_ts: f64[R] - └── alias_strings: string[A]; alias_entity_ids: i64[A] (when aliases exist) -``` +| | | +|---|---| +| [docs/QUICKSTART.md](docs/QUICKSTART.md) | Longer quick starts: HDF5 in Rust and Python, NetCDF-4, agent memory, CLI | +| [docs/USE_CASES.md](docs/USE_CASES.md) | Where clawhdf5 fits, and where it does not | +| [docs/agent-memory.md](docs/agent-memory.md) | The agent-memory store in full | +| [CONFORMANCE.md](CONFORMANCE.md) | The conformance report, generated by `conformance/run.sh` | +| [BENCHMARKS.md](BENCHMARKS.md) | Every measurement, with date, machine and command | +| [docs/known-issues.md](docs/known-issues.md) | Open limits and fixed bugs, with dates | +| [CHANGELOG.md](CHANGELOG.md) | Changes, including everything since v2.7.0 | +| [docs/README.md](docs/README.md) | Index of every document | -Alongside the store: `.h5.wal` (write-ahead log), `.h5.ann` -(HNSW graph; derived, safe to delete) and `.h5.lock` (single-writer -lock). A second writer gets `MemoryError::Locked`; use -`HDF5Memory::open_read_only` for a lock-free point-in-time view. +## Who uses it ---- - -## Migration - -### From rustyhdf5 / edgehdf5 - -Replace in `Cargo.toml` and source: - -| Old | New | -|-----|-----| -| `rustyhdf5*` | `clawhdf5*` | -| `edgehdf5-memory` | `clawhdf5-agent` | -| `edgehdf5` (CLI) | `clawhdf5-cli` | - -### From SQLite - -```bash -cargo install --path crates/clawhdf5-migrate -clawhdf5-migrate --sqlite old.db --hdf5 memory.h5 --agent-id my-agent --embedder minilm -``` - -The output is an ordinary `clawhdf5-agent` store, written through the agent's -own API: open it with `HDF5Memory::open` (or `clawhdf5-cli --path memory.h5 …`) -and search it straight away. The source must use the `memory_chunks` / `sessions` / `entities` / `relations` layout (names are -configurable with `--*-table`); note that this is not ZeroClaw's schema, and -ZeroClaw does not use clawhdf5. What carries over: - -| SQLite | Agent store | -|--------|-------------| -| `memory_chunks` | memory records (text, embedding, source channel, timestamp, session id, tags); rows with `deleted = 1` become deleted records, or are left out with `--skip-deleted` | -| `sessions` | sessions (id, start/end index, channel, summary, timestamp) | -| `entities`, `relations` | knowledge graph entities and relations; entities get new ids and relations are re-pointed at them | - -The chunk `id` column has no counterpart in the agent store, so records are -written in `id` order and numbered from 0. Embeddings are stored as float16 -like any new store; `--f32` keeps full precision (and is required for values -beyond ±65504). The embedding dimension is detected from the first row unless -`--embedding-dim` is given, and every row must have it: a row of another length -is an error, never truncated or padded. A source with no memory records (only -sessions or the graph) needs `--embedding-dim`, since a store's dimension is -fixed when it is created. Every row is checked before the output is created, -so a source that cannot be migrated leaves an existing store at `--hdf5` as it -was. `--incremental` adds to an existing store only the rows it does not -already hold; the source must have the store's dimension, and records already -in the store take the source's deleted flag (a row deleted in SQLite since the -last run is deleted in the store; one un-deleted there is written again, as -the agent has no un-delete). The tool reads the result back with -`HDF5Memory::open_read_only`, compares it with the source (every row with -`--validate-full`) and checks that a migrated record is found by search; -`--dry-run` only counts the rows. - ---- - -## Roadmap - -See [ROADMAP.md](ROADMAP.md) for the full implementation tracker. - -**Phase 1 complete** — all 8 tracks delivered: -- ✅ Knowledge Graph with spreading activation -- ✅ Hippocampal memory consolidation -- ✅ RRF hybrid retrieval + re-ranking + confidence rejection -- ✅ Temporal reasoning with sub-µs queries -- ✅ Memory security + anomaly detection -- ✅ Multi-modal memory (text/image/audio/video) -- ✅ Markdown ingest/export backend (`ClawhdfBackend`); an OpenClaw plugin was never built — see [docs/openclaw.md](docs/openclaw.md) -- ✅ Comprehensive Criterion benchmarks - -**Phase 2** — MemoryArena and LongMemEval academic benchmarks are done (see [BENCHMARKS.md](BENCHMARKS.md), reproduced on a second machine); remaining: crates.io/PyPI publishing. The Node bindings are unpublished and known to be broken ([known issues](docs/known-issues.md)). - ---- - -## Part of the RedClaw Ecosystem - -ClawhDF5 powers the `.brain` format for [ClawBrainHub](https://clawbrainhub.com) — the brain registry for AI agents. One file that packages identity, skills, memory, knowledge, and cryptographic provenance. - ---- +[ClawBrainHub](https://clawbrainhub.com) is the one verified consumer: its +`.brain` files are HDF5 files it reads and writes through the facade +(`File`, `FileBuilder`, `AttrValue`, `Selection`), and its CLI uses +`clawhdf5_agent::bm25::BM25Index` (builds and passes its tests against +`main`, checked 2026-09-25). clawhdf5 is **not** an OpenClaw memory plugin +([docs/openclaw.md](docs/openclaw.md)), and ZeroClaw does not use it. ## License -MIT - ---- - -

- Built by RedClaw Systems
- ~86,000 lines of Rust. Zero C dependencies. One file to remember everything. -

+MIT — see [LICENSE](LICENSE).