QUICKSTART used APIs that do not exist (file.dataset_names(), File::attr, AttrValue::Str, memory.search(&q, 5), MemoryConfig::new with a &str, consolidation without timestamps), `clawhdf5 = "2.0"` from crates.io, and "3-45x faster than libhdf5". It now covers HDF5 in Rust (write, read, strings, in-place append, remote, SWMR), Python (read, r+, w, URLs), NetCDF-4, h5rs, agent memory and the CLI, every snippet compiled and run (Python against a wheel built from the tree). USE_CASES dropped claims with no source (the agent crate adds ~2MB, IVF-PQ under 1.2 ms on modest hardware, an OpenClaw scenario, a .brain layout and `clawhub publish` commands) and now covers the HDF5 cases (no-C builds, threads, remote data, untrusted files, SWMR, in-place edits), the agent cases with measured numbers, and when to use something else. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
12 KiB
clawhdf5 quick start
Short, working examples for each way in. Every snippet here was compiled
and run against the repository (2026-09-28); the Rust ones assume a
function returning Result<_, Box<dyn std::error::Error>>.
| You want to | Go to |
|---|---|
| Read or write HDF5 from Rust | HDF5 in Rust |
| Read or edit HDF5 from Python without libhdf5 | Python |
| Read NetCDF-4 files | NetCDF-4 |
| Inspect or validate files on the command line | h5rs |
| Give an AI agent a memory store | Agent memory |
What is and is not supported: the feature matrix and known-issues.md.
1. HDF5 in Rust
Install
Not on crates.io yet; depend on the repository (MSRV 1.92):
[dependencies]
clawhdf5 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
# every plugin filter (bitshuffle, bzip2, Blosc, Blosc2, ZFP; LZF is on by default):
# clawhdf5 = { git = "...", features = ["plugin-filters"] }
Write a file
use clawhdf5::{AttrValue, FileBuilder};
let mut b = FileBuilder::new();
b.set_attr("title", AttrValue::String("run 42".into())); // a root attribute
b.create_dataset("temperatures") // 1-D f64, contiguous
.with_f64_data(&[22.5, 23.1, 21.8, 24.0]);
b.create_dataset("grid") // 2-D f32, chunked + gzip
.with_f32_data(&vec![1.5f32; 256 * 256])
.with_shape(&[256, 256])
.with_chunks(&[64, 64])
.with_deflate(4)
.with_fletcher32();
b.create_dataset("counts") // LZF (default feature), as h5py's compression="lzf"
.with_i32_data(&(0..10_000).collect::<Vec<i32>>())
.with_chunks(&[1000])
.with_lzf();
b.create_dataset("log") // appendable: unlimited first axis
.with_f64_data(&[])
.with_shape(&[0])
.with_maxshape(&[u64::MAX])
.with_chunks(&[1024]);
let mut sensors = b.create_group("sensors"); // groups nest; paths work too
sensors.set_attr("site", AttrValue::String("north".into()));
sensors.create_dataset("ids").with_i32_data(&[7, 8, 9]);
b.add_group(sensors.finish());
b.add_soft_link("latest", "/sensors");
b.write("example.h5")?;
h5py, h5dump and h5rs check --data read the result. FileBuilder holds
the file in memory and writes it once (atomically). Other data:
with_f16_data, with_i64_data, with_u64_data, with_u8_data,
with_compound_data (with CompoundTypeBuilder), enums, array types;
filters with_shuffle, with_zstd, with_lz4, with_bitshuffle,
with_bzip2, with_blosc (behind features); with_fill_value,
track_order, hard and external links, virtual datasets. The writer does
not write variable-length data.
Read a file
use clawhdf5::{File, Selection};
let file = File::open("example.h5")?;
let root = file.root();
println!("datasets {:?}, groups {:?}", root.datasets()?, root.groups()?);
println!("attrs {:?}", root.attrs()?);
let grid = file.dataset("grid")?;
println!("{:?} {:?} {:?}", grid.shape()?, grid.dtype()?, grid.max_dimensions()?);
let values: Vec<f32> = grid.read_f32()?; // integers/floats convert as libhdf5 does
let window = grid.read_f32_selection(&Selection::Hyperslab {
start: vec![0, 0], stride: vec![2, 2], count: vec![16, 16], block: vec![1, 1],
})?; // every other element of a 32x32 corner
let ids = file.group("sensors")?.dataset("ids")?.read_i64()?;
let same = file.dataset("latest/ids")?.read_i32()?; // through the soft link
A selection whose bounding box covers at most half the dataset decodes only
the chunks it touches; a larger one decodes the whole dataset
(known-issues.md).
File::open maps the file (mmap feature, default); File::open_buffered
reads it into memory, File::from_bytes takes a buffer, and
File::open_storage any Storage backend. A File is Send + Sync:
share it between threads.
Strings and variable-length data:
let file = clawhdf5::File::open("strings.h5")?; // written by h5py
let names: Vec<String> = file.dataset("names")?.read_string()?; // fixed- or variable-length
read_vlen::<T>() reads variable-length sequences, and
File::decode_strings / decode_vlen decode such values inside compounds
and raw attributes.
Edit a file in place
FileEditor changes an existing file (from h5py or clawhdf5) without
rewriting it: values, dataset extents, attributes. Here, appending batches
to the unlimited log dataset written above:
use clawhdf5::{FileEditor, Selection};
let mut ed = FileEditor::open("example.h5")?;
for batch in 0..3u64 {
let rows = vec![batch as f64; 500];
ed.resize("log", &[(batch + 1) * 500])?;
let sel = Selection::Hyperslab {
start: vec![batch * 500], stride: vec![1], count: vec![500], block: vec![1],
};
ed.write_values("log", &sel, &rows)?;
}
Each call is written and synced before it returns. The editor holds an exclusive lock and has no journal: a crash in the middle of an edit can leave the file inconsistent. What it refuses (before writing anything): known-issues.md § In-place modification.
Remote files and SWMR
// clawhdf5-remote = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
let file = clawhdf5_remote::open_url("http://127.0.0.1:8000/tall.h5")?;
let values = file.dataset("/g2/dset2.1")?.read_f64()?;
Serve a directory with range support to try it:
cargo run -p clawhdf5-remote --example range_server -- crates/clawhdf5/tests/fixtures 127.0.0.1:8000.
https:// needs the https feature; s3://, gs://, az:// the s3,
gcs, azure features (credentials from the environment).
See crates/clawhdf5-remote/README.md.
A file an h5py/libhdf5 SWMR writer is still appending to:
use std::time::{Duration, Instant};
let file = clawhdf5::File::open_swmr("live.h5")?;
let mut ds = file.dataset("samples")?;
let (mut seen, mut last_growth) = (0, Instant::now());
// Stop when the writer closes the file, or when the dataset has not grown for
// a minute (a writer that died never clears the SWMR-write flag).
while file.swmr_writer_active()? && last_growth.elapsed() < Duration::from_secs(60) {
ds.refresh()?; // h5py: ds.refresh()
let n = ds.shape()?[0];
if n > seen {
// read rows seen..n ...
(seen, last_growth) = (n, Instant::now());
}
std::thread::sleep(Duration::from_millis(100));
}
Design and limits: design/swmr.md.
2. Python
Not on PyPI yet; build the package with maturin into a virtualenv:
python -m venv .venv && . .venv/bin/activate
pip install maturin numpy
maturin develop --release -m crates/clawhdf5-py/Cargo.toml
Reading follows h5py:
import numpy as np
import clawhdf5
with clawhdf5.File("data.h5", "r") as f:
print(list(f.keys())) # member names, like h5py
ds = f["group/temperatures"] # relative or absolute paths
print(ds.shape, ds.dtype, ds.chunks)
block = ds[100:200, ::4] # a small selection decodes only its chunks
row = ds[-1] # integers drop the axis
picked = ds[[1, 5, 9], :] # one increasing index list per key
units = ds.attrs["units"] # attributes come back as h5py returns them
everything = np.asarray(ds)
ids = f["table"]["id"] # compound -> structured array; one field
Editing an existing file in place ('r+', through FileEditor), with
h5py's keys, broadcasting and numeric conversion; each edit is on disk when
the statement returns:
with clawhdf5.File("data.h5", "r+") as f:
f["group/temperatures"][100:200, ::4] = 0.0
f["series"].resize(5000, axis=0) # chunked datasets, within maxshape
f["series"][4000:] = np.ones(1000)
f["group"].attrs["calibrated"] = True
'r+' cannot create or delete datasets and groups, or delete attributes
(NotImplementedError, nothing written). New files ('w') take numeric
arrays (float64, float32, int64, int32, uint8):
with clawhdf5.File("new.h5", "w") as f:
f.create_dataset("x", data=np.arange(1000.0), chunks=(100,), compression="gzip")
f.create_group("meta").attrs["version"] = np.int64(2)
A URL opens a remote file read-only, by range requests (http:// in the
default build; https:// and s3:///gs:///az:// with
--features https / s3 / gcs / azure):
with clawhdf5.File("http://data.example.org/run42.h5") as f:
first = f["group/temperatures"][0]
f = clawhdf5.File.open_url("http://data.example.org/run42.h5", block_size=256 * 1024,
headers={"Authorization": "Bearer ..."})
print(f.remote_stats)
Types, keys and limits: crates/clawhdf5-py/README.md.
3. NetCDF-4
// clawhdf5-netcdf4 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
use clawhdf5_netcdf4::NetCDF4File;
let nc = NetCDF4File::open("climate.nc")?;
let mut temp = nc.variable("temperature")?;
let values = temp.read_f64()?; // CF scale_factor/add_offset/_FillValue applied
println!("{:?} {:?}", temp.shape()?, temp.cf_attributes()?.units);
dimensions(), variables(), global_attrs() and group(..) walk the
rest of the file; hdf5_file() gives the underlying clawhdf5::File.
4. h5rs
cargo install --path crates/clawhdf5-tools # --features remote for URLs
h5rs ls -r example.h5
h5rs dump example.h5 # DDL like h5dump; --json for hdf5-json
h5rs stat example.h5
h5rs diff a.h5 b.h5
h5rs check --data example.h5 # structure + checksums + every dataset decoded
See crates/clawhdf5-tools/README.md.
5. Agent memory
[dependencies]
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions};
// A new store: 384-dim embeddings (float16 on disk and an int8 HNSW index by default).
let mut memory = HDF5Memory::create(MemoryConfig::new("agent.h5".into(), "my-agent", 384))?;
memory.save(MemoryEntry {
chunk: "User prefers dark mode and vim keybindings.".into(),
embedding: embed("User prefers dark mode and vim keybindings."), // your embedder
source_channel: "chat".into(),
timestamp: now,
session_id: "session-001".into(),
tags: "preference".into(),
})?;
// Hybrid search: HNSW vector + BM25 keyword, fused 0.4 / 0.6 (the measured default).
let query = embed("what editor does the user like?");
for r in memory.search(&query, "editor preferences", &SearchOptions::new(5)) {
println!("[{:.3}] {}", r.score, r.chunk);
}
memory.flush_wal()?; // checkpoint the WAL into agent.h5
embed is yours: clawhdf5 stores embeddings, it does not compute them.
Each agent gets its own store; a store has a single writer, and
HDF5Memory::open_read_only gives other processes a lock-free view.
Source filters, re-ranking, signed checkpoints, the knowledge graph,
consolidation and the rest: agent-memory.md.
CLI
clawhdf5-cli installs a binary named clawhdf5; output is JSON.
cargo install --path crates/clawhdf5-cli
clawhdf5 --path agent.h5 create --agent-id my-agent --dim 384 --wal
echo '{"chunk":"User prefers dark mode","embedding":[0.1, ...],"source_channel":"chat","timestamp":1700000000.0,"session_id":"s1","tags":"pref"}' \
| clawhdf5 --path agent.h5 save
clawhdf5 --path agent.h5 search --embedding '[0.1, ...]' --query 'dark mode preferences' \
--top-k 5 --vector-weight 0.4 --keyword-weight 0.6
clawhdf5 --path agent.h5 stats
clawhdf5 --path agent.h5 export > memories.jsonl
clawhdf5 --path agent.h5 snapshot backup.h5
The CLI's search defaults to weights 0.7 / 0.3, not the library's
0.4 / 0.6, so pass them.
Next
- USE_CASES.md — where clawhdf5 fits
- CONFORMANCE.md, BENCHMARKS.md — the evidence
- README.md — every document