docs: QUICKSTART and USE_CASES on the current APIs

QUICKSTART used APIs that do not exist (file.dataset_names(),
File::attr, AttrValue::Str, memory.search(&q, 5), MemoryConfig::new with
a &str, consolidation without timestamps), `clawhdf5 = "2.0"` from
crates.io, and "3-45x faster than libhdf5". It now covers HDF5 in Rust
(write, read, strings, in-place append, remote, SWMR), Python (read, r+,
w, URLs), NetCDF-4, h5rs, agent memory and the CLI, every snippet
compiled and run (Python against a wheel built from the tree).

USE_CASES dropped claims with no source (the agent crate adds ~2MB,
IVF-PQ under 1.2 ms on modest hardware, an OpenClaw scenario, a .brain
layout and `clawhub publish` commands) and now covers the HDF5 cases
(no-C builds, threads, remote data, untrusted files, SWMR, in-place
edits), the agent cases with measured numbers, and when to use
something else.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 11:10:22 -05:00
co-authored by Claude Opus 5.5
parent a90373ca84
commit b0b4018919
2 changed files with 449 additions and 671 deletions
+160 -189
View File
@@ -1,209 +1,180 @@
# ClawhDF5 Use Cases
# Where clawhdf5 fits
Real-world scenarios where ClawhDF5 solves problems that other approaches can't.
Situations clawhdf5 was built for, what it gives you in each, and — at the
end — when to use something else. Code for each is in
[QUICKSTART.md](QUICKSTART.md); limits are in [known-issues.md](known-issues.md).
---
## 1. Personal AI Assistant
## HDF5 data
**Scenario:** You run a personal AI assistant (like OpenClaw, MemGPT, or a custom agent) that accumulates knowledge about you over weeks and months — preferences, decisions, context from past conversations.
### Reading HDF5 where libhdf5 is a burden
**Problem:** Most assistants either forget everything between sessions (stateless) or dump everything into a growing context window (expensive, eventually hits token limits).
You ship a Rust service, a CLI, a static binary, a WebAssembly page or a
cross-compiled ARM build, and linking libhdf5 (and its C toolchain,
threadsafe-build and version questions) is the hard part.
**ClawhDF5 solution:**
- The default build compiles no C at all, including deflate (pure-Rust
zlib-rs); `scripts/ci-test.sh` fails if a C-building crate enters the core
crates' default dependency tree.
- Reads are checked against h5py object by object on 697 public files;
602 are identical and none mismatches ([CONFORMANCE.md](../CONFORMANCE.md)).
- The common plugin filters (LZF, bitshuffle, bzip2, Blosc, Blosc2, ZFP)
are pure Rust too, so files written with hdf5plugin read without
installing plugins.
```
conversation → embedding → save to agent.h5
│
┌─────────────┤
│ │
Working Knowledge
Memory Graph
(recent) (entities)
│ │
consolidate traverse
│ │
Episodic "Who is
Memory Alice's
(important) manager?"
│
Semantic
Memory
(core facts)
```
### Many threads reading one file
- **Daily conversations** enter Working memory (bounded, auto-evicts old/trivial stuff)
- **Important facts** promote to Episodic ("User got promoted to VP on March 5th")
- **Core preferences** solidify in Semantic ("User is vegan, lives in SF, uses dark mode")
- **Entity tracking** via knowledge graph ("Alice → manages → Bob", "User → works_at → Acme")
- **One file** — back it up, move it to a new machine, it travels with the agent
A service answers requests from one large HDF5 file, and h5py threads do
not scale (libhdf5 serialises API calls; h5py users fall back to process
pools).
**What you'd need without ClawhDF5:** SQLite for structured data + Pinecone for vectors + a separate entity store + custom consolidation logic + Markdown files + glue code.
- A `clawhdf5::File` is `Send + Sync` with no library-wide lock: open it
once and share it.
- Full reads of deflate data from 16 threads through one `File` ran at
1.58x the throughput of 16 h5py processes on tank on 2026-09-26
([BENCHMARKS.md](../BENCHMARKS.md#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)).
- The Python bindings release the GIL for every read, so Python threads
get the same.
### Data on a web server or in object storage
The file is on HTTP, S3, GCS or Azure, and you need a few datasets from it,
not the whole download.
- `clawhdf5_remote::open_url` (Rust), `clawhdf5.File(url)` (Python) and
`h5rs` with `--features remote` read by range requests through a block
cache, with the file pinned by ETag/Last-Modified so a changed file is an
error rather than mixed data.
- In the browser, `clawhdf5-wasm`'s `openUrl` does the same from the page's
main thread; the [viewer](../examples/wasm-viewer/README.md) is a working
example. Opening one dataset of a 3000-dataset, 198 MB h5py file took
5 requests and 5.2 MB at 1 MiB blocks (tank, 2026-09-27, CHANGELOG).
- Design and measured request counts: [design/range-reads.md](design/range-reads.md).
### Files you did not write and do not trust
User uploads, files from instruments or old archives, fuzzed inputs.
- On the HDF Group's CVE corpus clawhdf5 has no panic, crash, hang or
runaway allocation, where h5dump 1.14.6 crashes on 2 files and h5py on 1
([CONFORMANCE.md](../CONFORMANCE.md#cve-corpus-clawhdf5-vs-h5dump-vs-h5py)).
- `h5rs check --data file.h5` validates the structures and checksums and
decodes every dataset; it uses the library's parsers, so it accepts what
they accept, not everything libhdf5 would reject.
### Watching a running experiment
An acquisition process writes with libhdf5 in SWMR mode and a dashboard or
monitor follows it.
- `File::open_swmr` + `Dataset::refresh()` follow the writer as h5py's
SWMR reader does, retrying reads that race a flush and never returning
torn data. Tested live against an h5py writer.
- clawhdf5 does not write SWMR files; the writer stays libhdf5.
### Patching files in place
Fix a calibration constant, append to a time series, grow a dataset: files
too large to rewrite, or written by someone else.
- `FileEditor` (Rust) and `clawhdf5.File(path, 'r+')` (Python) overwrite
values, resize chunked datasets and set attributes without rewriting the
file, changing indexes and heaps as libhdf5 does; everything is checked
against h5py and h5dump in the tests.
- Anything it cannot do safely is refused before a byte is written.
---
## 2. OpenClaw
## Agent memory
Not supported: clawhdf5 is not an OpenClaw memory plugin, and the config this
section used to show was never valid. See [openclaw.md](openclaw.md).
### A personal assistant that remembers
An assistant accumulates preferences, decisions and context over months.
- `clawhdf5-agent` keeps records, sessions and a knowledge graph in one
`.h5` file with a write-ahead log: back it up or move it with the agent.
- Hybrid search (HNSW + BM25) reaches 81.4% turn-level Hit@5 on the full
LongMemEval haystack — retrieval recall, not QA accuracy
([BENCHMARKS.md](../BENCHMARKS.md#longmemeval-results)).
- The consolidation engine (Working → Episodic → Semantic) and the
knowledge graph are library components you drive; see
[agent-memory.md](agent-memory.md#library-components).
### Several agents, kept apart
A coding agent, a research agent and a scheduler should not read each
other's memories.
- One store per agent; each store has a single writer (an exclusive lock),
and other processes can open it read-only.
- `SearchOptions::with_sources` restricts a search to chosen source
channels.
- The write-anomaly detector flags injection patterns and write bursts
(alerts, never blocks); its source classification is a heuristic on the
`source_channel` string, not an authenticated boundary.
- There is no built-in way to share a graph between stores; export and
import it yourself.
### On a small device
A Raspberry Pi or another ARM board, no server, no network.
- Pure Rust, no database server, one file.
- The int8 index uses NEON `SDOT` on cores with the dot-product extension
(plain NEON elsewhere); on a Raspberry Pi 5 it
was 1.18x the `f32` index's QPS at equal recall
([BENCHMARKS.md](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)).
CI builds and tests the aarch64 code on an ARM runner.
- WAL appends are not fsynced: on power loss, saves since the last
checkpoint can be lost, while checkpoints themselves are made durable as
a unit. Checkpoint (`flush_wal`) as often as you need.
- `clawhdf5-android` has JNI bindings for the store.
### Tamper-evident memory
You need to know whether a store was edited outside your agent.
- With a signing key, every checkpoint stores an Ed25519-signed manifest
(SHA-256 per record in a Merkle tree, plus settings, sessions and graph);
`HDF5Memory::verify` names the records that changed. Saves still in the
WAL are not covered until the next checkpoint.
### `.brain` files (ClawBrainHub)
[ClawBrainHub](https://clawbrainhub.com) packages agents as `.brain` files,
which are HDF5 files its `cbh-core` crate reads and writes through
clawhdf5's facade (`File`, `FileBuilder`, `AttrValue`, `Selection`). It is
the one verified consumer of clawhdf5.
---
## 3. Multi-Agent System
## When to use something else
**Scenario:** You have multiple specialized agents — a coding agent, a research agent, a scheduling agent — that need to share knowledge without sharing everything.
- **Parallel writes from MPI ranks**: `clawhdf5-io`'s `mpi-io` gathers
writes to rank 0 and reads on one rank then broadcasts; it is not
collective I/O. Use libhdf5 with MPI-IO.
- **Writing SWMR files**, **creating or deleting objects in an existing
file**, **writing variable-length data**, **writing Blosc2 or ZFP**: not
supported.
- **Files that must open in HDF5 1.8**: clawhdf5's output is not tested
there.
- **Node.js**: the package does not work
([known-issues.md](known-issues.md#the-nodejs-package-packagesclawhdf5-node-does-not-work)).
- **An OpenClaw or ZeroClaw memory backend**: clawhdf5 is neither
([openclaw.md](openclaw.md)).
**Problem:** Giving agents a shared database creates security issues (coding agent shouldn't see personal data) and conflicts (agents overwrite each other's memories).
## Choosing features
**ClawhDF5 solution:**
```
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Coding Agent │ │Research Agent│ │Schedule Agent│
│ coding.h5 │ │ research.h5 │ │ schedule.h5 │
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘
│ │ │
└────────┬────────┘ │
│ │
┌───────▼────────┐ │
│ Shared KG only │◄────────────────┘
│ (export/import)│
└────────────────┘
```
- Each agent has its own `.h5` file (full isolation)
- Knowledge graph entities/relations can be exported and imported between agents
- **Source isolation** in the provenance system prevents user-sourced memories from contaminating system memories within a single agent
- **Anomaly detection** catches if one agent is writing suspiciously (injection attack via tool output)
---
## 4. Edge / Embedded AI
**Scenario:** You're building an AI agent that runs on a Raspberry Pi, phone, or embedded device with limited resources. No cloud database. No internet for vector DB queries.
**Problem:** Most memory solutions require a server (Pinecone, Qdrant) or heavy dependencies (Python, CUDA).
**ClawhDF5 solution:**
- **Pure Rust** — compiles to a single static binary, no C dependencies
- **Single file** — all memory in one `.h5` file, no database server
- **Small footprint** — the agent crate adds ~2MB to your binary
- **ARM support** — runs on ARM64 (Raspberry Pi, phones) natively
- **Android bridge** — `clawhdf5-android` provides JNI bindings for Android apps
- **IVF-PQ** for ANN search keeps latency under 1.2ms even at 100K vectors on modest hardware
- **WAL** for crash safety — if the device loses power, no data corruption
```rust
// Same API whether you're on a server or a Pi
let config = MemoryConfig::new("/data/agent.h5", "edge-agent", 384);
let mut memory = HDF5Memory::create(config)?;
```
---
## 5. Scientific Data + AI Memory
**Scenario:** You work with HDF5 files (common in physics, climate science, genomics) and want to add AI-powered search over your datasets.
**Problem:** Existing HDF5 libraries (h5py, HDF5 C library) don't have vector search. You'd need a separate tool.
**ClawhDF5 solution:**
ClawhDF5 is a full HDF5 implementation that *also* has agent memory. You can:
- **Read existing HDF5 files** from CERN, NASA, NOAA — no C library needed
- **Add vector search** to your datasets by embedding them and storing in the agent memory layer
- **Query across datasets** using hybrid search (find the experiment that matches your description)
- **Track data provenance** with the built-in provenance system
```rust
use clawhdf5::File;
use clawhdf5_agent::{HDF5Memory, MemoryConfig};
// Read your scientific data
let data = File::open("experiment_results.h5")?;
let measurements = data.dataset("sensor_readings")?.read_f64()?;
// Create a searchable memory alongside it
let mut memory = HDF5Memory::create(
MemoryConfig::new("experiment_memory.h5", "lab-assistant", 384)
)?;
// Embed and index experiment descriptions
memory.save(MemoryEntry {
chunk: "Experiment 47: Temperature response at 350K with catalyst B".into(),
embedding: embed("Temperature response..."),
source_channel: "lab-notebook".into(),
..default()
})?;
// Later: "which experiments used catalyst B above 300K?"
let results = memory.hybrid_search(&query_emb, "catalyst B temperature", 0.6, 0.4, 10);
```
---
## 6. The `.brain` Format (ClawBrainHub)
**Scenario:** You've built an amazing AI agent with custom personality, skills, and accumulated knowledge. You want to package it and distribute it.
**Problem:** Agent identity is scattered across config files, prompt templates, skill definitions, vector stores, and various databases. There's no standard format.
**ClawhDF5 solution — the `.brain` file:**
```
agent.brain (HDF5)
├── /meta — schema version, author, license
├── /identity — system prompt, personality, avatar
├── /skills — tool definitions, MCP configs
├── /memory — vector embeddings, knowledge graph
├── /media — voice samples, images
├── /runtime — model preferences, resource limits
└── /provenance — SHA-256 hashes, Ed25519 signatures
```
One file. Cryptographically signed. Publishable to [ClawBrainHub](https://clawbrainhub.com).
```bash
# Create a brain file
clawhdf5 --path agent.brain create --agent-id my-agent --dim 384
# Publish to ClawBrainHub (coming soon)
clawhub publish agent.brain
# Pull a brain
clawhub pull redclawsystems/research-assistant
```
This is the container image for intelligence.
---
## Choosing the Right Features
| Your Situation | Features to Enable | Why |
|----------------|-------------------|-----|
| **Quick prototype** | Default | Vector search works out of the box |
| **Production agent** | defaults (`float16`, `hnsw`, `parallel`) | HNSW search and a parallel index build; half-precision *storage* is `MemoryConfig::float16`, on by default for new stores |
| **macOS** | + `accelerate` | Apple AMX coprocessor for matrix ops |
| **Linux server** | + `openblas` or `fast-math` | BLAS acceleration |
| **GPU available** | + `gpu` | wgpu-based search, wins at 100K+ scale |
| **Long-running agent** | + `async` | Tokio async with background flush |
| **Edge device** | Default only | Minimal dependencies, smallest binary |
```toml
# Not on crates.io yet: depend on the repository.
# Production agent on Linux
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5", features = ["fast-math"] }
# Edge device
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
# macOS with GPU
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5", features = ["accelerate", "gpu", "async"] }
```
---
<p align="center"><em>Built by <a href="https://git.redclaw.dev/quantumclaw">RedClaw Systems</a></em></p>
| Situation | Crate / features |
|---|---|
| Read and write HDF5 | `clawhdf5` (defaults: `mmap`, `provenance`, `lzf`) |
| Plugin-filtered files (hdf5plugin) | `clawhdf5`, `features = ["plugin-filters"]` |
| Zstd, LZ4 | `zstd` (links libzstd), `lz4` |
| SZIP | `clawhdf5-format`'s `szip` (libaec, C) |
| zlib-ng instead of zlib-rs | `fast-deflate` (needs cmake) |
| Remote files | `clawhdf5-remote` (`http` default; `https`, `s3`, `gcs`, `azure`) |
| Agent memory | `clawhdf5-agent` (defaults: `float16`, `hnsw`, `parallel`) |
| BLAS for the agent's brute-force paths | `fast-math`, `openblas`, or `accelerate` (macOS) |
| GPU distance computation | `gpu` (wgpu) |
| Async wrapper | `async` (Tokio) |