Files
clawhdf5/docs/USE_CASES.md
T
osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00

183 lines
7.9 KiB
Markdown

# Where clawhdf5 fits
Situations clawhdf5 was built for, what it gives you in each, and — at the
end — when to use something else. Code for each is in
[QUICKSTART.md](QUICKSTART.md); limits are in [known-issues.md](known-issues.md).
---
## HDF5 data
### Reading HDF5 where libhdf5 is a burden
You ship a Rust service, a CLI, a static binary, a WebAssembly page or a
cross-compiled ARM build, and linking libhdf5 (and its C toolchain,
threadsafe-build and version questions) is the hard part.
- The default build compiles no C at all, including deflate (pure-Rust
zlib-rs); `scripts/ci-test.sh` fails if a C-building crate enters the core
crates' default dependency tree.
- Reads are checked against h5py object by object on 697 public files;
602 are identical and none mismatches ([CONFORMANCE.md](../CONFORMANCE.md)).
- The common plugin filters (LZF, bitshuffle, bzip2, Blosc, Blosc2, ZFP)
are pure Rust too, so files written with hdf5plugin read without
installing plugins.
### Many threads reading one file
A service answers requests from one large HDF5 file, and h5py threads do
not scale (libhdf5 serialises API calls; h5py users fall back to process
pools).
- A `clawhdf5::File` is `Send + Sync` with no library-wide lock: open it
once and share it.
- Full reads of deflate data from 16 threads through one `File` ran at
1.58x the throughput of 16 h5py processes on tank on 2026-09-26
([BENCHMARKS.md](../BENCHMARKS.md#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)).
- The Python bindings release the GIL for every read, so Python threads
get the same.
### Data on a web server or in object storage
The file is on HTTP, S3, GCS or Azure, and you need a few datasets from it,
not the whole download.
- `clawhdf5_remote::open_url` (Rust), `clawhdf5.File(url)` (Python) and
`h5rs` with `--features remote` read by range requests through a block
cache, with the file pinned by ETag/Last-Modified so a changed file is an
error rather than mixed data.
- In the browser, `clawhdf5-wasm`'s `openUrl` does the same from the page's
main thread; the [viewer](../examples/wasm-viewer/README.md) is a working
example. Opening and reading one dataset of a 3000-dataset, 198 MB h5py
file took 5 requests and 5.2 MB at 1 MiB blocks (h5py's default
`libver="earliest"`; 7 requests and 6.7 MB with `"latest"`) (tank,
2026-09-27, CHANGELOG "Unreleased").
- Design and measured request counts: [design/range-reads.md](design/range-reads.md).
### Files you did not write and do not trust
User uploads, files from instruments or old archives, fuzzed inputs.
- On the HDF Group's CVE corpus clawhdf5 has no panic, crash, hang or
runaway allocation, where h5dump 1.14.6 crashes on 2 files and h5py on 1
([CONFORMANCE.md](../CONFORMANCE.md#cve-corpus-clawhdf5-vs-h5dump-vs-h5py)).
- `h5rs check --data file.h5` validates the structures and checksums and
decodes every dataset; it uses the library's parsers, so it accepts what
they accept, not everything libhdf5 would reject.
### Watching a running experiment
An acquisition process writes with libhdf5 in SWMR mode and a dashboard or
monitor follows it.
- `File::open_swmr` + `Dataset::refresh()` follow the writer as h5py's
SWMR reader does, retrying reads that race a flush and never returning
torn data. Tested live against an h5py writer.
- clawhdf5 does not write SWMR files; the writer stays libhdf5.
### Patching files in place
Fix a calibration constant, append to a time series, grow a dataset: files
too large to rewrite, or written by someone else.
- `FileEditor` (Rust) and `clawhdf5.File(path, 'r+')` (Python) overwrite
values, resize chunked datasets and set attributes without rewriting the
file, changing indexes and heaps as libhdf5 does; everything is checked
against h5py and h5dump in the tests.
- Anything it cannot do safely is refused before a byte is written.
---
## Agent memory
### A personal assistant that remembers
An assistant accumulates preferences, decisions and context over months.
- `clawhdf5-agent` keeps records, sessions and a knowledge graph in one
`.h5` file with a write-ahead log: back it up or move it with the agent.
- Hybrid search (HNSW + BM25) reaches 81.4% turn-level Hit@5 on the full
LongMemEval haystack with real MiniLM embeddings — retrieval recall, not
QA accuracy (tank, 2026-09-27; [BENCHMARKS.md](../BENCHMARKS.md#longmemeval-results)).
- The consolidation engine (Working → Episodic → Semantic) and the
knowledge graph are library components you drive; see
[agent-memory.md](agent-memory.md#library-components).
### Several agents, kept apart
A coding agent, a research agent and a scheduler should not read each
other's memories.
- One store per agent; each store has a single writer (an exclusive lock),
and other processes can open it read-only.
- `SearchOptions::with_sources` restricts a search to chosen source
channels.
- The write-anomaly detector flags injection patterns and write bursts
(alerts, never blocks); its source classification is a heuristic on the
`source_channel` string, not an authenticated boundary.
- There is no built-in way to share a graph between stores; export and
import it yourself.
### On a small device
A Raspberry Pi or another ARM board, no server, no network.
- Pure Rust, no database server, one file.
- The int8 index uses NEON `SDOT` on cores with the dot-product extension
(plain NEON elsewhere); on a Raspberry Pi 5 it
was 1.18x the `f32` index's QPS at equal recall (2026-09-21, `114a2df`,
not re-run since; [BENCHMARKS.md](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)).
CI builds and tests the aarch64 code on an ARM runner.
- WAL appends are not fsynced: on power loss, saves since the last
checkpoint can be lost, while checkpoints themselves are made durable as
a unit. Checkpoint (`flush_wal`) as often as you need.
- `clawhdf5-android` has JNI bindings for the store.
### Tamper-evident memory
You need to know whether a store was edited outside your agent.
- With a signing key, every checkpoint stores an Ed25519-signed manifest
(SHA-256 per record in a Merkle tree, plus settings, sessions and graph);
`HDF5Memory::verify` names the records that changed. Saves still in the
WAL are not covered until the next checkpoint.
### `.brain` files (ClawBrainHub)
[ClawBrainHub](https://clawbrainhub.com) packages agents as `.brain` files,
which are HDF5 files its `cbh-core` crate reads and writes through
clawhdf5's facade (`File`, `FileBuilder`, `AttrValue`, `Selection`). It is
the one verified consumer of clawhdf5.
---
## When to use something else
- **Parallel writes from MPI ranks**: `clawhdf5-io`'s `mpi-io` gathers
writes to rank 0 and reads on one rank then broadcasts; it is not
collective I/O. Use libhdf5 with MPI-IO.
- **Writing SWMR files**, **creating or deleting objects in an existing
file**, **writing variable-length data**, **writing Blosc2 or ZFP**: not
supported.
- **Files that must open in HDF5 1.8**: clawhdf5's output is not tested
there.
- **Node.js**: the package does not work
([known-issues.md](known-issues.md#the-nodejs-package-packagesclawhdf5-node-does-not-work)).
- **An OpenClaw or ZeroClaw memory backend**: clawhdf5 is neither
([openclaw.md](openclaw.md)).
## Choosing features
| Situation | Crate / features |
|---|---|
| Read and write HDF5 | `clawhdf5` (defaults: `mmap`, `provenance`, `lzf`) |
| Plugin-filtered files (hdf5plugin) | `clawhdf5`, `features = ["plugin-filters"]` |
| Zstd, LZ4 | `zstd` (links libzstd), `lz4` |
| SZIP | `clawhdf5-format`'s `szip` (libaec, C) |
| zlib-ng instead of zlib-rs | `fast-deflate` (needs cmake) |
| Remote files | `clawhdf5-remote` (`http` default; `https`, `s3`, `gcs`, `azure`) |
| Agent memory | `clawhdf5-agent` (defaults: `float16`, `hnsw`, `parallel`) |
| Faster brute-force paths in the agent | `clawhdf5-agent`'s `fast-math` (matrixmultiply, pure Rust), or BLAS: `openblas`, `accelerate` (macOS) |
| GPU distance computation | `clawhdf5-agent`'s `gpu` (wgpu) |
| Async wrapper | `clawhdf5-agent`'s `async` (Tokio) |