Merge branch 'docs/refresh-crates' into docs/readme-refresh

This commit is contained in:
osobh
2026-09-28 11:14:06 -05:00
28 changed files with 1138 additions and 335 deletions
+146 -177
View File
@@ -1,193 +1,162 @@
# ClawhDF5 Roadmap — Agent Memory Evolution
# clawhdf5 roadmap
> Making clawhdf5 the defacto agentic memory solution.
> Single file. Pure Rust. Zero dependencies. Trusted everywhere.
What has shipped, and what is genuinely next. Everything here is checked
against `CHANGELOG.md`, `git log` and [`docs/known-issues.md`](docs/known-issues.md);
dates are merge dates on `main`. Nothing after v2.7.0 has been released:
the work since then is on `main` under `CHANGELOG.md` "Unreleased".
_Last updated: 2026-09-28 (at `9b5803f`, PR #21)._
---
## Track 1: Knowledge Graph in HDF5
**Status:** 🟢 Phase 1 Complete
**Priority:** Critical
**Crate:** `clawhdf5-agent`
## Done
- [x] **1.1** Entity storage — entities with properties, embeddings, timestamps (created_at/updated_at)
- [x] **1.2** Relation storage — typed edges with RelationType enum (Temporal/Causal/Associative/Hierarchical/Custom), metadata, timestamps
- [x] **1.3** Entity extraction helpers — rule-based extraction (Person, Org, Location, Date, Technology, Project) with extract_and_store_entities() integration
- [x] **1.4** Entity resolution — fuzzy name matching (Levenshtein distance) via resolve_or_create()
- [x] **1.5** Graph traversal queries — BFS neighbors with depth, subgraph extraction from seeds
- [x] **1.6** Spreading activation — weighted activation propagation with configurable decay
- [x] **1.7** Graph-aware retrieval — get_entity_context() for formatted context injection
- [x] **1.8** Tests — comprehensive tests for all new features
### Releases
**Research:** Graph-Native Cognitive Memory (2026), Graph-based Agent Memory survey (2026), SYNAPSE (2025)
| Version | Date | Headline |
|---|---|---|
| v2.0.0 | 2026-03-19 | rustyhdf5 (11 crates) and edgehdf5 (4 crates) unified into one workspace as `clawhdf5-*` |
| v2.1.0 | 2026-06-03 | HNSW backs the agent's vector search by default; live, mutable HNSW index |
| v2.2.0 – v2.7.0 | 2026-09-18 – 2026-09-20 | bounded decompression and read-path bounds checks, single-writer store locking, WAL v4, HNSW recall fix (0.31 -> 0.98 recall@10 at 100K), fusion weights tuned on LongMemEval, int8 index, Extensible Array read fix and chunk-index checksums |
Details per release: [`CHANGELOG.md`](CHANGELOG.md).
### Since v2.7.0 (unreleased, on `main`)
| PR | Merged | What |
|---|---|---|
| #3 | 2026-09-23 | pure-Rust deflate (zlib-rs) by default, no C in the core crates' default build (checked in CI), MSRV 1.92 |
| #4 | 2026-09-25 | files open in h5py again (every `f32` and every empty dataset clawhdf5 wrote was unreadable by libhdf5); float16 embedding storage |
| #5 | 2026-09-25 | `HDF5Memory::search` with `SearchOptions` (source filters, re-ranking, confidence); float16 on by default |
| #6 | 2026-09-25 | `clawhdf5-migrate` writes real agent stores; knowledge-graph fix; dated benchmark re-run |
| #7 | 2026-09-25 | consolidation benchmark completed (cheaper novelty scoring) |
| #8 | 2026-09-25 | Ed25519-signed checkpoints (`HDF5Memory::verify`) |
| #9, #10 | 2026-09-25 | OpenClaw and ZeroClaw integration claims withdrawn — neither ever integrated clawhdf5 |
| #11 | 2026-09-26 | silent wrong data and libhdf5 interop bugs found by the HDF5 audit fixed |
| #12 | 2026-09-26 | reproducible conformance sweep over eight public corpora, nightly CI job ([`CONFORMANCE.md`](CONFORMANCE.md)) |
| #13 | 2026-09-26 | reads HDF5 1.6-era layouts, user blocks, virtual datasets, dense attributes, very large groups |
| #14 | 2026-09-26 | `h5rs` tools (`ls`, `dump`, `stat`, `diff`, `check`), the browser reader (`clawhdf5-wasm`), libhdf5's header checks, plugin filters (LZF, bitshuffle, bzip2, Blosc), concurrency benchmark |
| #15 | 2026-09-26 | fast contiguous and concurrent reads, variable-length data, nested groups and links in the writer, Python bindings |
| #16 | 2026-09-26 | chunked full reads faster than an h5py process pool, writer B-trees of any size, Blosc2 (read), 599/697 conformance |
| #17 | 2026-09-26 | range reads M0/M1 (indexed name lookups, the `Storage` trait), ZFP (read), in-place editing (`FileEditor`) |
| #18 | 2026-09-27 | range reads M2/M3 (`File::open_storage`; `clawhdf5-remote`: HTTP(S), S3, GCS, Azure), in-place editing of every chunk index, shrinking, dense attributes |
| #19 | 2026-09-27 | remote files in the browser (`openUrl`, M4), SWMR reader (`File::open_swmr`, M5), Python remote reads and `'r+'` editing |
| #20 | 2026-09-28 | benchmarks re-measured: LongMemEval with real MiniLM embeddings, local reads on an idle machine |
| #21 | 2026-09-28 | remote files open in a few requests (group lookups down the B-tree, `Storage::hint`), `ObjectHeader::parse` back to its earlier speed, the last conformance mismatches resolved: 602/697 ok, 0 our-error, 0 mismatch |
### Range reads (design: [`docs/design/range-reads.md`](docs/design/range-reads.md))
- [x] M0 — indexed name lookups (#17)
- [x] M1 — metadata parsed through the `Storage` trait (#17)
- [x] M2 — raw data through `Storage`, `File::open_storage` (#18)
- [x] M3 — `clawhdf5-remote`: HTTP(S) range requests and object stores through a block cache; `h5rs` URLs (#18); Python URLs (#19)
- [x] M4 — `openUrl` in the browser, restartable "NeedBytes" cache (#19; fewer round trips in #21)
- [x] M5 — reading files a SWMR writer is appending to ([`docs/design/swmr.md`](docs/design/swmr.md), #19)
### Agent memory (`clawhdf5-agent`)
Shipped before and during the v2 releases, and kept current since:
knowledge graph with entity extraction and resolution; three-tier
consolidation with decay; hybrid retrieval (HNSW + BM25, weighted or RRF
fusion, re-ranking, confidence rejection, query expansion); temporal index
and session DAG; per-save provenance ledger and write-anomaly detection;
multi-modal embeddings; WAL with chained CRC32; single-writer locking;
signed checkpoints. Retrieval is measured, not claimed: see
[`BENCHMARKS.md`](BENCHMARKS.md) ("LongMemEval Results" reports retrieval
recall, not QA accuracy; earlier headline numbers that compared different
granularities were retracted there).
---
## Track 2: Memory Consolidation Engine
**Status:** 🟢 Phase 1 Complete
**Priority:** Critical
**Crate:** `clawhdf5-agent`
## Next
- [x] **2.1** Importance scoring — surprise (novelty), correction boost, length scoring with configurable weights
- [x] **2.2** Three-tier memory model — Working → Episodic → Semantic with bounded capacities
- [x] **2.3** Time-decay with reactivation — exponential decay with configurable half-life, access resets timestamp
- [x] **2.4** Bounded memory with graceful degradation — evict lowest-decay entries when over capacity
- [x] **2.5** Consolidation cycles — promote/evict across tiers based on importance and access thresholds
- [x] **2.6** Memory statistics — ConsolidationStats with per-tier counts, eviction/promotion tracking
- [x] **2.7** Tests — comprehensive tests for all features
Not scheduled; listed roughly by how much they unblock. None has a date.
**Research:** CraniMem (2026), D-MEM (2026), AI Hippocampus survey (2026)
### Distribution
- [ ] **Publish the crates to crates.io.** Nothing is published; the READMEs
say to depend on git. Before publishing: no `publish` settings exist
(only `clawhdf5-wasm` has `publish = false`), and several `Cargo.toml`
descriptions still name the old `rustyhdf5`/`edgehdf5` (`clawhdf5-accel`,
`-derive`, `-gpu`, `-io`, `-netcdf4`, `-android`).
- [ ] **Publish Python wheels to PyPI.** `crates/clawhdf5-py` builds with
maturin and is tested in CI, but no wheel is published. The default wheel
reads plain `http://` only; `https`/`s3`/`gcs`/`azure` wheels compile C
(ring, aws-lc-rs).
- [ ] **The Node.js package** (`packages/clawhdf5-node` over
`clawhdf5-napi`) has never worked and is not in CI: fix it and add CI, or
remove it ([known issue](docs/known-issues.md)).
### HDF5 features
- [ ] **SWMR writing.** The reader is done (M5); writing a file while
libhdf5 readers follow it is not. Also not covered: remote SWMR (a remote
file is pinned at open), `MmapFile`/`LazyFile` SWMR reads, refreshing
groups or attributes.
- [ ] **MPI collective I/O.** `clawhdf5-io`'s `MpiVol` (`mpi-io`) is
root-read + broadcast and gather-to-root writes, not collective MPI-IO
(`MPI_File_read_at_all`/`write_at_all`).
- [ ] **Paged-metadata single-request reads.** Files written with paged
aggregation (`H5Pset_file_space_strategy(PAGE)`, `h5repack -S PAGE`)
keep their metadata in a few pages; range reads could fetch those in one
request and use the file's page size as the block size. Today the block
size is fixed (1 MiB) and only the first block is read ahead
(range-reads design, option (c) as a policy).
- [ ] **Blosc2 and ZFP encoders.** Both filters are read-only; the other
plugin filters (LZF, bitshuffle, bzip2, Blosc 1) read and write.
- [ ] **External links and external raw data** are explicit errors, not
followed.
- [ ] **Virtual datasets:** the "first missing" view and printf gaps other
than 0, source-to-virtual type conversion other than a byte swap, nested
virtual sources, source files outside the virtual file's directory.
- [ ] **Datatypes:** x87 long double and binary128 are refused.
- [ ] **Writer:** one attribute or link message over 65 515 bytes in dense
storage is an error (huge fractal-heap objects); no option to write
files HDF5 1.8 can read.
- [ ] **`FileEditor`:** new chunks in implicit indexes, variable-length and
reference data, filters it cannot encode (scale-offset, N-Bit, SZIP),
some dense-attribute heap layouts, creating or deleting objects and
attributes (also from Python `'r+'`), and no journal (a crash mid-edit
can leave the file inconsistent). Freed space is reused only within one
editor.
- [ ] **Selection reads** decode the whole dataset when the selection's
bounding box covers more than half of it (a strided `ds[::100]`), and
for compact/virtual datasets or a non-default fill value: correct, but
more work than needed.
- [ ] **Readers:** `LazyFile` and `MmapFile` still need the whole file;
the zero-copy methods need the file in memory.
### Remote and browser
- [ ] Run the `s3`/`gcs`/`azure` backends against real buckets (only built
and URL-parsing-tested so far).
- [ ] `h5rs` options for request headers and cache settings.
- [ ] Browser limits in [`docs/known-issues.md`](docs/known-issues.md)
("`clawhdf5-wasm` (browser) limits"): files of 4 GiB or more (wasm32),
compound/reference/opaque datasets, round trips per index level. The
package doubled in size with `openUrl`
([size table](examples/wasm-viewer/README.md#size)); dropping the
function-name section would take a third off the raw size (13% gzipped).
### Quality
- [ ] Scheduled fuzz campaigns: the cargo-fuzz targets
([`crates/clawhdf5-format/fuzz`](crates/clawhdf5-format/fuzz/README.md),
and the agent's WAL target) run only by hand or with
`CLAWHDF5_FUZZ_SECONDS`.
---
## Track 3: Hybrid Retrieval Pipeline
**Status:** 🟢 Phase 1 Complete
**Priority:** High
**Crate:** `clawhdf5-agent`
## Withdrawn
- [x] **3.1** Reciprocal Rank Fusion (RRF) — rrf_hybrid_search() with k=60 constant
- [x] **3.2** Multi-factor re-ranking — temporal decay, source authority hierarchy, activation scores (reranker.rs)
- [x] **3.3** Low-confidence rejection — min_score threshold, gap filtering, max_results (confidence.rs)
- [x] **3.4** Query expansion — synonyms, acronyms, temporal rewrites, morphological variants, knowledge graph aliases + expanded_search() with RRF merge
- [x] **3.5** Result explanation — ReRankResult with full score breakdown per factor
- [x] **3.6** Configurable pipeline — ReRankConfig + ConfidenceConfig with tunable weights/thresholds
- [x] **3.7** Tests + MemX-comparable benchmarks — 5 integration tests (Hit@1≥90%, search<500ms@100K, BM25<200ms@100K, hybrid<50ms@10K, compact<200ms@10K)
- **OpenClaw integration** (withdrawn 2026-09-25, PR #9). clawhdf5 was
never an OpenClaw memory plugin; the documented
`memory.backend = "clawhdf5"` was never valid. The Rust `ClawhdfBackend`
remains as a library API. [`docs/openclaw.md`](docs/openclaw.md) records
what a real plugin would need.
- **ZeroClaw integration** (withdrawn 2026-09-25, PR #10). ZeroClaw has no
clawhdf5 backend, and `clawhdf5-migrate`'s SQLite layout is not
ZeroClaw's schema.
**Research:** MemX (2026), SwiftMem (2026)
---
## Track 4: Temporal Reasoning
**Status:** 🟢 Phase 1 Complete
**Priority:** High
**Crate:** `clawhdf5-agent`
- [x] **4.1** Temporal index — sorted timestamp index with binary search, insert/remove
- [x] **4.2** Time-range queries — range_query, before, after, latest, earliest
- [x] **4.3** Session DAG — parent/child linking, chain walking, time-range overlap queries
- [x] **4.4** Temporal re-ranking — query hint enum (Latest/Earliest/Around/Between/None) with boost scoring
- [x] **4.5** Temporal entity tracking — EntityTimeline with state change history + point-in-time reconstruction
- [x] **4.6** Tests — comprehensive tests for all features
**Research:** MemX temporal gaps (≤43.6% Hit@5), MemoryArena multi-session tasks (2026)
---
## Track 5: Memory Security & Provenance
**Status:** 🟢 Phase 1 Complete
**Priority:** Medium-High
**Crate:** `clawhdf5-agent`
- [x] **5.1** Source attribution — MemoryProvenance with source, creator, session, FNV-1a content hash
- [x] **5.2** Write anomaly detection — rate limiting, 15 injection patterns, source distribution analysis
- [x] **5.3** Source isolation — per-MemorySource sub-stores preventing cross-contamination
- [x] **5.4** Memory integrity verification — content hash comparison via verify_integrity()
- [x] **5.5** Poisoning resistance — pattern detection for prompt injection attempts
- [x] **5.6** Tests — comprehensive tests including adversarial patterns
**Research:** MemoryGraft (2025), SSGM Framework (2026)
---
## Track 6: Multi-Modal Memory
**Status:** 🟢 Phase 1 Complete
**Priority:** Medium
**Crate:** `clawhdf5-agent`
- [x] **6.1** Image embedding storage — ModalEmbedding with model provenance (CLIP, SigLIP, etc.)
- [x] **6.2** Audio fingerprints — Audio modality with embedding storage
- [x] **6.3** Multi-modal search — search_by_modality (filtered) + search_cross_modal (all embeddings)
- [x] **6.4** Observation records — raw perception vs interpretation with confidence scoring
- [x] **6.5** Media reference storage — MediaRef with Path/Url/Inline, MIME types, FNV-1a checksums
- [x] **6.6** Tests — 35 comprehensive tests
**Research:** Neuro-Symbolic Memory (2026), RAGdb multi-modal RAG (2025)
---
## Track 7: OpenClaw Integration — withdrawn (2026-09-25)
**Status:** ⚪ Withdrawn (the items below were library work; no OpenClaw integration shipped)
**Priority:** Critical (for adoption)
**Crates:** `clawhdf5-agent`, `clawhdf5-napi`
- [x] **7.1** Memory backend trait — MemoryBackend with search/get/write/ingest/export/stats
- [x] **7.2** Hybrid retrieval pipeline — ClawhdfBackend wires RRF → reranker → confidence rejection
- [x] **7.3** Markdown import/export — MarkdownParser + MarkdownExporter with line tracking + metadata
- [x] **7.4** `search()` — backed by the full hybrid retrieval pipeline (a Rust method; no OpenClaw tool was ever registered)
- [x] **7.5** `get()` — read back by path, with a line slice (not an OpenClaw tool either)
- [x] **7.6** Compaction integration — run_compaction() (decay + compact + WAL flush), run_consolidation() (hippocampal engine), tick_session(), flush_wal()
- [ ] **7.7** ~~Config surface — `memory.backend = "clawhdf5"`~~ — never valid OpenClaw config; docs removed
- [ ] **7.8** ~~Documentation + migration guide~~ — removed: they described an integration that never worked
**Node.js bridge:** `clawhdf5-napi` (napi-rs) and a TypeScript wrapper in `packages/clawhdf5-node` exist but are unpublished, untested in CI and known to be broken (docs/known-issues.md).
---
> **Withdrawn.** None of this track produced a working OpenClaw integration: no
> plugin was built, the documented `memory.backend = "clawhdf5"` config was never
> valid in any OpenClaw release, and the Node package was never published. The
> Rust `ClawhdfBackend` remains as a library API. Not pursued for now; see
> [docs/openclaw.md](docs/openclaw.md) for what a plugin would need today.
## Track 8: Benchmarking & Validation
**Status:** 🟢 Complete
**Priority:** High
**Crates:** `clawhdf5-agent`, `clawhdf5-bench`
- [x] **8.1** MemoryArena benchmark — 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547
- [x] **8.2** LongMemEval benchmark — 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md § LongMemEval Results](BENCHMARKS.md#longmemeval-results)
- [x] **8.3** Latency benchmarks — vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal
- [x] **8.4** Memory footprint — 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion
- [x] **8.5** Consolidation efficiency — 8.8x search speedup, 90% noise eviction, zero quality loss
- [x] **8.6** Cross-platform benchmarks — x86 measured, ARM estimated, cross_platform.sh script
- [x] **8.7** Published results in BENCHMARKS.md with ephemeral tier Redis comparison (70-140x faster)
---
## Implementation Order
**Phase 1:** ~~Tracks 1, 2, 3 — core memory intelligence~~ 🟢 Complete
**Phase 2:** ~~Track 4 (temporal) + Track 5 (security)~~ 🟢 Complete
**Phase 3:** ~~Track 6 (multi-modal)~~ 🟢 Complete; Track 7 (OpenClaw integration) withdrawn
**Phase 4:** ~~Track 8 (benchmarking + validation)~~ 🟢 Complete
All 8 tracks delivered. 1,650+ tests passing, zero clippy warnings.
---
## What's Next
Verified against current repo state on 2026-08-05 (see also `docs/superpowers/plans/` for the filter-codec/format-write/MPI-IO work, now shipped):
- [ ] TypeScript bridge not wired into CI — `packages/clawhdf5-node/` already has a complete, working napi-rs package (package.json, tsconfig, hand-written TS wrapper matching all 21 `#[napi]` items, Jest test suite, README); it isn't published to npm and has no committed lockfile
- [ ] Publish crates to crates.io — no `publish` config anywhere in the workspace yet
- [ ] Python wheel distribution via maturin — `crates/clawhdf5-py/pyproject.toml` exists (maturin-buildable locally) but wheels aren't published anywhere
- [ ] `chunked_read.rs`/`data_read.rs` full bounds-check audit + scheduled fuzz campaigns (the new `fuzz_dataset_read` target covers the two files' main entry points; a full manual audit of every indexing site is still open) — see Tier 4 below
- [ ] WAL per-entry checksum landed as CRC32 (see below); a stronger per-entry format (explicit length prefix, avoiding the read-then-verify restructuring) could still be revisited if profiling shows it matters
- [ ] HNSW build parallelism is still narrow (only `prune_connections`); the correctness-sensitive outer insert loop needs its own dedicated design pass before parallelizing
### Recently closed out (2026-08-05, Tier 3–4 hardening pass)
- [x] Academic benchmark cross-validation — LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
- [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer
- [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories
- [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open
- [x] `chunked_read.rs`/`data_read.rs`/`local_heap.rs` bounds-check audit: added `ensure_len` overflow guards, a recursion-depth guard against cyclic B-trees, and a fix for an unguarded compound-datatype byte-offset overrun. Added a new `fuzz_dataset_read` cargo-fuzz target exercising the contiguous/chunked/compact read paths — it found and we fixed 3 real crash bugs (integer-overflow panics) within the first few runs
- [x] `clawhdf5-ann`: optional `parallel` feature (rayon) for HNSW's `prune_connections` neighbor-distance computation
- [x] `[workspace.dependencies]` added for `tempfile`/`criterion`/`half`/`serde`, fixing a real version skew on `half` (2 vs 2.7)
### Recently closed out (2026-08-05 hardening pass)
- [x] CI/CD pipeline — `.gitea/workflows/ci.yml` now runs `scripts/ci-test.sh` (fmt, clippy, tests, no_std check) on push/PR to `main`
- [x] Fixed no_std build breakage in `clawhdf5-format` (missing alloc imports, `AtomicU64` unsupported on thumbv7em, `f64::powi` requiring std/libm)
- [x] Fixed version skew: `clawhdf5-py` (pyproject.toml) and `packages/clawhdf5-node` (package.json) were both behind the actual crate version
### Recently closed out (2026-08-03 cleanup pass)
- [x] Removed `clawhdf5-types` — it was an empty 1-line stub crate; shared type definitions already live in `clawhdf5-format`, so CLAUDE.md and the workspace manifest were corrected instead of filling it in
- [x] Superblock v4 (page-buffer mode) read/write — the only unimplemented task from `docs/superpowers/plans/2026-06-29-format-write-extensions.md`; now done (`Superblock::parse_v4`/`serialize`, `FileWriter::with_page_size`)
- [x] Reconciled the three `docs/superpowers/plans/*.md` docs against actual shipped code — they were pre-work plans for `d6c4d4f` (2026-06-30), committed to git late; checkboxes now reflect reality
---
_Last updated: 2026-08-05_
The old track-by-track tracker this file used to be (agent-memory
Tracks 1–8, mid-2026) is in git history (`git log -- ROADMAP.md`).
+27
View File
@@ -6,9 +6,36 @@ h5py/libhdf5, compares the two readings object by object, and writes
```sh
CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh # ~30 s once the corpus is cached
conformance/run.sh --no-fetch # use the cached corpus as is
conformance/run.sh --update-baseline # after an intended change in results
```
Latest result (tank, 2026-09-27, `conformance/run.sh --no-fetch`): 602 of
697 files ok, 0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read, and
no panic, hang, crash or out-of-memory. The report with every file is
[`CONFORMANCE.md`](../CONFORMANCE.md).
## Classes
`compare.py` puts each file in one class:
| class | meaning |
|---|---|
| **ok** | clawhdf5 and h5py read the same objects with the same values |
| **our-error** | h5py reads something clawhdf5 refuses |
| **mismatch** | both read it, with different values or structure |
| **h5py-cannot-read** | h5py (libhdf5) cannot read the file; not compared |
| **ref-bug** | h5py reads an object clawhdf5 refuses, but only through a libhdf5 over-read: `ref_bugs.py` re-reads it in six processes with different heaps (import order, `MALLOC_PERTURB_`) and its values change. The file is ref-bug only while that is confirmed in the same run; if the values become stable it counts as our-error again |
| **panic / hang / crash / oom** | a clawhdf5 failure under the timeout and address-space limit; the gate fails on any |
Where h5py itself returns wrong values through a known h5py bug (the
big-endian variable-length bug: elements returned with the file's bytes
under a little-endian dtype), `ref.py` checks that the installed h5py has
the bug, corrects the values before hashing and marks them `ref_fix`, so
those objects are still compared. The evidence for the three current
ref-bug files is under "Conformance: the last non-ok files" in
[`docs/known-issues.md`](../docs/known-issues.md).
Needs Rust, `git`, `h5dump` (Debian/Ubuntu `hdf5-tools`), `libaec` (for the
probe's `szip` feature; `libaec-dev`), and a Python with the packages in
`requirements.txt`. The first run downloads about 450 MB of sparse checkouts.
+51 -14
View File
@@ -1,24 +1,61 @@
# clawhdf5-accel
[![crates.io](https://img.shields.io/crates/v/clawhdf5-accel.svg)](https://crates.io/crates/clawhdf5-accel)
[![docs.rs](https://docs.rs/clawhdf5-accel/badge.svg)](https://docs.rs/clawhdf5-accel)
CPU SIMD kernels for vector search: dot products, cosine similarity, L2
distance, norms and int8 dot products, dispatched at run time to the best
backend the CPU has, with a portable scalar fallback for every operation.
[`clawhdf5-ann`](../clawhdf5-ann/README.md) and
[`clawhdf5-agent`](../clawhdf5-agent/README.md) use it in their distance
loops; it has nothing to do with HDF5 file I/O.
SIMD-accelerated operations for clawhdf5.
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-accel = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## API
```rust
use clawhdf5_accel::{cosine_similarity, detect_backend, dot_i8, dot_product, l2_distance};
let a = [1.0f32, 2.0, 3.0, 4.0];
let b = [4.0f32, 3.0, 2.0, 1.0];
assert_eq!(dot_product(&a, &b), 20.0);
let _cos = cosine_similarity(&a, &b);
let _l2 = l2_distance(&a, &b);
assert_eq!(dot_i8(&[1, -2, 3], &[4, 5, -6]), -24);
println!("{:?}", detect_backend()); // e.g. Avx2 on x86-64, Neon on aarch64
```
Also `vector_norm`, `batch_norms`, `batch_cosine`, `batch_cosine_prenorm`,
`f16_to_f32_batch`, `checksum_fletcher32` and `align_to_cache_line`.
## Backends
`detect_backend()` picks once per process: `Avx512` (with the `avx512`
feature), `Avx2` (AVX2 + FMA), `Neon` (every aarch64 CPU), or `Scalar`.
`Sse4` and `WasmSimd128` are reported when detected but run the scalar
kernels.
`dot_i8`, used by the agent's quantised (int8) HNSW index, runs on
AVX2 and on NEON — with the `SDOT` instruction (through inline assembly,
since the intrinsic is unstable) on cores that have dotprod, such as the
Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8
index answers 1.63x the queries per second of the f32 one on x86-64
(AVX2) and 1.18x on a Raspberry Pi 5 ([`BENCHMARKS.md`](../../BENCHMARKS.md)).
The aarch64 code is compiled out on x86, so only the `test-arm64` CI job
builds and tests it.
## Features
- AVX2 and NEON SIMD acceleration
- AVX-512 support (`avx512` feature)
- Float16 conversion (`float16` feature)
- CRC32 checksum acceleration
| Feature | Default | What | Builds C |
|---|---|---|---|
| `avx512` | no | AVX-512F kernels | no |
| `float16` | no | `f16_to_f32_batch` through the `half` crate (a software conversion otherwise) | no |
## Usage
```rust
use clawhdf5_accel::checksum::crc32_simd;
let crc = crc32_simd(&data);
```
The half-precision conversion used for stored embeddings is
`clawhdf5_format::float16`, not this crate's.
## License
+112 -16
View File
@@ -1,28 +1,124 @@
# clawhdf5-agent
[![crates.io](https://img.shields.io/crates/v/clawhdf5-agent.svg)](https://crates.io/crates/clawhdf5-agent)
[![docs.rs](https://img.shields.io/docsrs/clawhdf5-agent)](https://docs.rs/clawhdf5-agent)
Persistent memory for AI agents in a single HDF5 file: text chunks with
embeddings and metadata, hybrid search (HNSW vector search + BM25 keyword
search, fused), sessions, a knowledge graph, a write-ahead log for crash
safety, and optionally Ed25519-signed checkpoints. Stores open in h5py like
any other HDF5 file. Built on [`clawhdf5`](../clawhdf5/README.md),
[`clawhdf5-ann`](../clawhdf5-ann/README.md) and
[`clawhdf5-accel`](../clawhdf5-accel/README.md).
HDF5-backed persistent memory store for on-device AI agents.
It is a library: no agent framework integrates it (OpenClaw and ZeroClaw
integration claims were withdrawn on 2026-09-25; see
[`docs/openclaw.md`](../../docs/openclaw.md)). The command-line front end
is [`clawhdf5-cli`](../clawhdf5-cli/README.md).
Built on [clawhdf5](https://crates.io/crates/clawhdf5), clawhdf5-agent provides a vector-searchable memory backend optimized for edge AI workloads. Store embeddings, text chunks, and metadata in a single HDF5 file with SIMD-accelerated similarity search.
## Features
- Persistent vector store in HDF5 format
- Cosine similarity and L2 distance search
- SIMD-accelerated via clawhdf5-accel (AVX2, NEON)
- Optional GPU acceleration via clawhdf5-gpu
- Memory-mapped access for large stores
- f16 storage support for compact embeddings
## Usage
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-agent = "2.1.0"
clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust,no_run
use std::path::PathBuf;
use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions};
let config = MemoryConfig::new(PathBuf::from("agent.h5"), "my-agent", 384);
let mut mem = HDF5Memory::create(config)?;
mem.save(MemoryEntry {
chunk: "The deploy key rotates every Monday.".into(),
embedding: vec![0.01; 384], // from your embedding model
source_channel: "chat".into(),
timestamp: 1_790_000_000.0,
session_id: "s1".into(),
tags: "ops".into(),
})?;
let query = vec![0.01f32; 384];
let hits = mem.search(&query, "deploy key", &SearchOptions::new(5).with_sources(["chat"]));
for h in &hits {
println!("{:.3} {}", h.score, h.chunk);
}
mem.flush_wal()?; // checkpoint now; otherwise one is made after 500 WAL entries (wal_max_entries)
# Ok::<(), clawhdf5_agent::MemoryError>(())
```
## What is in it
- **`HDF5Memory`** — `create`, `open` (single writer: an exclusive lock on
`<store>.h5.lock`, a second opener gets `MemoryError::Locked`),
`open_read_only` (no lock, never writes). Through the `AgentMemory`
trait: `save`, `save_batch`, `delete`, `compact`, `count`, `snapshot`,
sessions; also `save_or_update`, `delete_batch`, `flush_wal`.
- **Search** — `search(query_embedding, text, &SearchOptions)`: optional
source-channel filter applied before ranking, vector + BM25 fusion
(weighted or RRF), Hebbian activation scaling, optional re-ranking
(`reranker::ReRankConfig`) and confidence rejection
(`confidence::ConfidenceConfig`). `hybrid_search` and
`hybrid_search_with` are thin wrappers. The vector stage uses the HNSW
index (`hnsw` feature); its graph is saved to `<store>.h5.ann` at each
checkpoint and reloaded on open (rebuilt if stale or damaged).
- **Storage settings** (`MemoryConfig`, persisted with the store):
`float16` embeddings (on by default for new stores; 48% smaller file at
100K records, same retrieval on LongMemEval), `quantized_index` (int8
copy of the vectors in the index, on by default; re-scored against the
exact embeddings), `compression` (off by default), HNSW `m`/`ef`
parameters, WAL settings (`wal_enabled`, on by default; `wal_max_entries`,
500: the WAL is checkpointed into the `.h5` once it holds more).
- **WAL** (`wal`) — every write is appended to `<store>.h5.wal` with a
chained CRC32 per entry, so a corrupted, reordered or spliced entry stops
replay. Recovers from a process crash at any point, including between a
checkpoint and the WAL truncate. WAL appends are not fsynced: saves since
the last checkpoint can be lost on power failure. An unreadable WAL is
quarantined to `<store>.h5.wal.corrupt-<ts>`.
- **Signed checkpoints** (`signing`) — `set_signing_key` signs a manifest
(SHA-256 Merkle tree over records, plus settings, sessions and graph) at
every checkpoint; `HDF5Memory::verify(path, &public_key)` checks it and
locates edits. WAL entries after the checkpoint are not covered.
- **Knowledge graph** (`knowledge`, `entity_extract`) — `add_entity`,
`add_entity_alias`, `add_relation`, `extract_and_store_entities`,
traversal and spreading activation.
- **Also:** sessions (`session`), temporal index (`temporal`),
consolidation tiers (`consolidation`), an in-memory TTL tier
(`ephemeral`), multi-modal embeddings (`multimodal`), `AGENTS.md`
generation (`agents_md`), query expansion, and a session-scoped
provenance ledger and write-anomaly detector on every save
(`take_anomaly_alerts`; alerts never block a save, and the source is
inferred from `source_channel`, not authenticated).
- `openclaw::ClawhdfBackend` is `search` with re-ranking and confidence
on, plus Markdown import/export. The module name is historical: it is not
an OpenClaw plugin.
## Features
| Feature | Default | What | Builds C |
|---|---|---|---|
| `hnsw` | yes | HNSW vector index (`clawhdf5-ann`); without it the vector stage is an exact linear cosine scan | no |
| `parallel` | yes | build the HNSW index on a rayon pool (same graph either way) | no |
| `float16` | yes | f16 helpers in `vector_search` (`half`). Stores' `MemoryConfig::float16` works without it. | no |
| `fast-math` | no | `matrixmultiply` batch distances in `strategy` | no |
| `accelerate` | no | Apple Accelerate BLAS in `strategy` (macOS) | links a system framework |
| `openblas` | no | OpenBLAS in `strategy` | yes (`openblas-src`) |
| `gpu` | no | `gpu_search` through [`clawhdf5-gpu`](../clawhdf5-gpu/README.md) (wgpu), used by `strategy`, not by `HDF5Memory::search` | no, but needs GPU drivers |
| `zstd` | no | Zstd instead of deflate when `MemoryConfig::compression` is on | yes (libzstd) |
| `async` | no | `async_memory` wrapper on tokio | no |
`--no-default-features --features float16` forces the exact linear scan.
## Measurements and limits
- Search recall and latency, file size, LongMemEval and MemoryArena
retrieval numbers: [`BENCHMARKS.md`](../../BENCHMARKS.md), measured with
the `clawhdf5-bench` binaries (`search_harness`, `longmemeval_bench`,
`footprint_bench`, ...).
- Known issues and their history: [`docs/known-issues.md`](../../docs/known-issues.md).
- Migrating a SQLite memory database:
[`clawhdf5-migrate`](../clawhdf5-migrate/README.md).
## License
MIT
+44
View File
@@ -0,0 +1,44 @@
# clawhdf5-android
A C ABI over [`clawhdf5-agent`](../clawhdf5-agent/README.md) for Android
apps: a `cdylib` exporting `extern "C"` functions (`edgehdf5_*`, a name
kept from the project's earlier "edgehdf5" days) that manage an
`HDF5Memory` through an opaque handle.
The functions are plain C symbols, not JNI-mangled `Java_...` entry points:
a Kotlin/Java app calls them through a thin JNI shim or JNA of its own. No
such shim, Gradle project or AAR is in this repository, and the crate is
not built for an Android target in CI (only its host-side unit tests run
with the workspace).
## Functions
| Function | What |
|---|---|
| `edgehdf5_create(path, agent_id, embedding_dim)` / `edgehdf5_open(path)` | a handle, or null on failure |
| `edgehdf5_close(handle)` | drop the store; what is not yet checkpointed stays in its WAL, as with any `HDF5Memory` |
| `edgehdf5_save(handle, ...)` | save one entry; the embedding length is checked against the store's dimension before the pointer is read |
| `edgehdf5_delete`, `edgehdf5_count`, `edgehdf5_count_active` | |
| `edgehdf5_hybrid_search(handle, query, len, text, vector_weight, keyword_weight, max_results, out_indices, out_scores, out_chunks)` | results into caller-provided arrays; returns the number written |
| `edgehdf5_add_session`, `edgehdf5_get_session_summary` | sessions |
| `edgehdf5_add_entity`, `edgehdf5_add_relation` | knowledge graph |
| `edgehdf5_free_string` | free a string this library returned |
Every function is `unsafe`: the caller guarantees valid, NUL-terminated
strings and correctly sized buffers (see each function's `# Safety`
section), and serialises access to a handle; separate handles are
independent.
## Build
```bash
cargo build --release -p clawhdf5-android # host build; for a device, add --target aarch64-linux-android with the NDK's linker configured
```
It depends on `clawhdf5-agent` with **default features off**, so there is
no HNSW index (the vector stage is an exact linear scan) and no rayon
pool. No C is compiled.
## License
MIT
+55 -10
View File
@@ -1,25 +1,70 @@
# clawhdf5-ann
[![crates.io](https://img.shields.io/crates/v/clawhdf5-ann.svg)](https://crates.io/crates/clawhdf5-ann)
[![docs.rs](https://docs.rs/clawhdf5-ann/badge.svg)](https://docs.rs/clawhdf5-ann)
An HNSW (Hierarchical Navigable Small World) approximate nearest-neighbour
index in pure Rust, with cosine or L2 distance, optional int8 storage of
the vectors, deletions, and persistence as an HDF5 file. It is the vector
stage of [`clawhdf5-agent`](../clawhdf5-agent/README.md)'s search (the
agent's `hnsw` feature, on by default); distances run on
[`clawhdf5-accel`](../clawhdf5-accel/README.md)'s SIMD kernels.
HNSW approximate nearest neighbor index stored as HDF5.
Neighbours are chosen with the HNSW paper's diversity heuristic, not plain
closest-M (which capped recall on clustered data at 0.31 recall@10 at 100K
vectors).
## Features
Not on crates.io yet; depend on it from git:
- Build and query HNSW indexes persisted in HDF5 format
- Pure Rust, no C dependencies
- Efficient similarity search for high-dimensional vectors
```toml
[dependencies]
clawhdf5-ann = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust
use clawhdf5_ann::HnswIndex;
use clawhdf5_ann::{DistanceMetric, HnswIndex, Storage};
let index = HnswIndex::from_hdf5("vectors.h5").unwrap();
let neighbors = index.search(&query, 10);
let vectors: Vec<Vec<f32>> = (0..500)
.map(|i| (0..16).map(|j| ((i * 31 + j * 7) % 97) as f32 / 97.0).collect())
.collect();
// m = 16 connections per node, ef_construction = 200
let mut index = HnswIndex::build_with(&vectors, 16, 200, DistanceMetric::Cosine, Storage::Int8);
let hits = index.search(&vectors[42], 10, 64); // (id, distance), closest first; ef >= k
assert!(hits[0].1 < 1e-3); // vector 42 itself (or an identical one)
let id = index.insert(vec![0.5; 16]);
index.mark_deleted(id);
// Persist as HDF5 (a self-contained file: graph and vectors) and load it back
let bytes = index.to_hdf5_bytes().unwrap();
let loaded = HnswIndex::load_from_hdf5(&bytes).unwrap();
assert_eq!(loaded.len(), index.len());
```
- `HnswIndex::build` (L2), `build_with_metric`, `build_with` (metric and
storage); `new`/`new_with` plus `insert` for an index built
incrementally.
- `Storage::Int8` keeps each vector as `i8`, a quarter of the memory; it
applies to `Cosine` only (an L2 index keeps `Float32`). Distances are then
approximate, so a caller that needs exact ranking re-scores the
candidates, as the agent does.
- `mark_deleted`, `is_deleted`, `deleted_count`, `active_len`, `compact`
(returns the old-to-new id map).
- `save_to_hdf5(&mut writer)` / `to_hdf5_bytes` / `load_from_hdf5` store
the whole index; `graph_to_bytes` / `from_graph_bytes` store only the
graph (with a CRC32) for a caller that keeps the vectors elsewhere — the
agent's `<store>.h5.ann` sidecar.
## Features
| Feature | Default | What | Builds C |
|---|---|---|---|
| `parallel` | no | build the graph on a rayon pool; the graph is identical with or without it | no |
Recall and speed against exact search, for the index alone and in the
agent: [`BENCHMARKS.md`](../../BENCHMARKS.md), measured with
`cargo run --release -p clawhdf5-bench --bin search_harness`.
## License
MIT
+49
View File
@@ -0,0 +1,49 @@
# clawhdf5-bench
The measurement harnesses behind [`BENCHMARKS.md`](../../BENCHMARKS.md):
HDF5 read and write speed (against libhdf5 and h5py where noted) and the
agent store's search, footprint and retrieval quality. Not meant for
publishing; nothing else in the workspace depends on it. Run everything with
`--release`, and quote numbers with the machine, date and command, as
`BENCHMARKS.md` does.
## Binaries
| Binary | Measures |
|---|---|
| `read_harness` | full reads vs hyperslab selections of a chunked 2-D dataset (compressed and not) and a contiguous one: does a selection cost scale with the selection or the dataset? (`-- --large` for 512 MB) |
| `concurrent_read` | decoded read throughput vs threads on one open `File`; `scripts/concurrent_read_h5py.py` runs the same workload with h5py (threads and processes) and `scripts/compare_concurrent_read.py` tabulates both |
| `search_harness` | HNSW recall@10 vs exact search, QPS and latency per `ef`, and end-to-end `HDF5Memory` ingest/checkpoint/open/search at 1K–100K (`--full`); studies: `--float16-study`, `--options-study`, `--signing-study`, `--ann-only --uniform` |
| `longmemeval_bench` | LongMemEval retrieval recall (turn and session Hit@k, MRR) — **retrieval, not QA accuracy**. Oracle or full `longmemeval_s` haystack; `--features embeddings` (or `embeddings-cuda`) embeds with MiniLM, otherwise the vector stage is inert and the run is BM25-only |
| `memory_arena` | a deterministic multi-session retrieval benchmark (BM25-only) |
| `footprint_bench` | file size and bytes per record at 100–100K records, float16 or `--f32`, WAL on/off, compressed or not |
| `consolidation_efficiency` | retrieval before and after consolidation on signal + noise records |
| `ephemeral_perf` | the in-memory ephemeral tier's set/get latency |
| `mpi_io_bench` | `clawhdf5-io`'s `MpiVol` (root-read + broadcast, not collective I/O); needs `--features mpi-io` and `mpirun` |
```bash
cargo run --release -p clawhdf5-bench --bin search_harness -- --full
cargo run --release -p clawhdf5-bench --bin read_harness
```
## Criterion benches and example
- `cargo bench -p clawhdf5-bench` runs `h5bench_write`, `h5bench_read` and
`h5bench_meta` (h5bench-style sequential, chunked, strided and metadata
workloads). `--features libhdf5-compare` adds the same workloads through
libhdf5 (the `hdf5-metno` crate; needs a system libhdf5 1.14).
- `examples/worldmodel_sampling.rs`: shuffled per-frame reads of a
`(N, H, W, C)` `uint8` dataset, clawhdf5 against h5py on the same file.
## Features
| Feature | What | Builds C |
|---|---|---|
| `libhdf5-compare` | libhdf5 variants of the Criterion benches | links the system libhdf5 |
| `mpi-io` | `mpi_io_bench` | yes (`mpi-sys`; needs an MPI installation) |
| `embeddings` | MiniLM embeddings for `longmemeval_bench` (candle) | yes (a `cc` build dependency in the candle/tokenizers tree) |
| `embeddings-cuda` | the same on a CUDA GPU (minutes instead of hours on the full haystack) | yes (CUDA) |
## License
MIT
+49
View File
@@ -0,0 +1,49 @@
# clawhdf5-cli
The `clawhdf5` command: create, fill, search and inspect a
[`clawhdf5-agent`](../clawhdf5-agent/README.md) memory store from the
shell. Output is JSON. (For general HDF5 files use `h5rs` from
[`clawhdf5-tools`](../clawhdf5-tools/README.md).)
```bash
cargo install --path crates/clawhdf5-cli # installs `clawhdf5`; not on crates.io yet
# or: cargo run -p clawhdf5-cli -- --help
```
No C is compiled.
## Commands
The store is `--path FILE` (or `CLAWHDF5_PATH`) before the subcommand.
| Command | What |
|---|---|
| `create [--agent-id ID] [--dim N] [--wal] [--f32] [--f32-index]` | a new store (dimension 384 by default); float16 embeddings and an int8 index copy unless `--f32` / `--f32-index`. The WAL is off unless `--wal` (the library's default is on), so each save is checkpointed at once |
| `save [--json '{...}']` | save one entry, from `--json` or stdin: `{"chunk", "embedding", "source_channel", "timestamp", "session_id", "tags"}` |
| `search --embedding '[...]' [--query TEXT] [-k N] [--vector-weight W] [--keyword-weight W]` | hybrid search (defaults 5 results, weights 0.7 / 0.3) |
| `recall INDEX` | one entry by index |
| `stats` | counts and configuration |
| `flush-wal` | checkpoint the WAL into the `.h5` |
| `agents-md [--output FILE]` | generate an `AGENTS.md` from the store |
| `export` | every entry as JSON lines |
| `snapshot DEST` | a copy of the store's `.h5` file |
| `keygen --out FILE` | a new Ed25519 signing key (64 hex characters, created owner-only on Unix) |
| `verify --public-key HEX_OR_FILE` | check a signed store; exit status 2 if it does not verify |
`recall`, `stats`, `agents-md` and `export` open the store read-only
(no lock, nothing written), so they work while another process has it
open. `save`, `search` (which records activation boosts) and `flush-wal`
open it for writing and take the store's lock. With
`--signing-key FILE` (or `CLAWHDF5_SIGNING_KEY`) every checkpoint a command
makes is signed; a signed store refuses to checkpoint without the key.
```bash
clawhdf5 --path mem.h5 create --agent-id demo --dim 3
echo '{"chunk":"hello","embedding":[0.1,0.2,0.3],"source_channel":"cli","timestamp":0,"session_id":"s1","tags":""}' \
| clawhdf5 --path mem.h5 save
clawhdf5 --path mem.h5 search --embedding '[0.1,0.2,0.3]' --query hello -k 3
```
## License
MIT
+34 -12
View File
@@ -1,28 +1,50 @@
# clawhdf5-derive
[![crates.io](https://img.shields.io/crates/v/clawhdf5-derive.svg)](https://crates.io/crates/clawhdf5-derive)
[![docs.rs](https://docs.rs/clawhdf5-derive/badge.svg)](https://docs.rs/clawhdf5-derive)
`#[derive(H5Type)]`: maps a Rust struct with named fields to an HDF5
compound datatype. The derive generates three inherent methods:
Derive macros for clawhdf5 HDF5 traits.
- `hdf5_datatype() -> clawhdf5_format::datatype::Datatype` — the
`Datatype::Compound` (members in field order, packed, little-endian);
- `to_bytes(&self) -> Vec<u8>` — one element in that layout;
- `from_bytes(&[u8]) -> Self` — the reverse (panics if the slice is shorter
than the compound).
## Features
Supported field types: `f32`, `f64`, `i8`–`i64`, `u8`–`u64`, `bool`
(stored as `u8`) and fixed-size arrays `[T; N]` of those numeric types.
Tuple structs, enums and nested structs are refused at compile time.
- `#[derive(HDF5Type)]` for automatic HDF5 datatype mapping
- Struct-to-compound-type derivation
The generated code names `clawhdf5_format`, so the crate using the derive
must depend on [`clawhdf5-format`](../clawhdf5-format/README.md) too. Not
on crates.io yet:
## Usage
```toml
[dependencies]
clawhdf5-derive = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Example
```rust
use clawhdf5_derive::HDF5Type;
use clawhdf5_derive::H5Type;
use clawhdf5_format::datatype::Datatype;
#[derive(HDF5Type)]
#[derive(H5Type, Debug, PartialEq)]
struct Point {
x: f64,
y: f64,
z: f64,
id: u32,
pos: [f64; 3],
valid: bool,
}
let p = Point { id: 7, pos: [1.0, 2.0, 3.0], valid: true };
let bytes = p.to_bytes();
assert_eq!(bytes.len(), 4 + 24 + 1);
assert_eq!(Point::from_bytes(&bytes), p);
assert!(matches!(Point::hdf5_datatype(), Datatype::Compound { size: 29, .. }));
```
Tests: `crates/clawhdf5-format/tests/derive_tests.rs`.
## License
MIT
+39 -10
View File
@@ -1,27 +1,56 @@
# clawhdf5-filters
[![crates.io](https://img.shields.io/crates/v/clawhdf5-filters.svg)](https://crates.io/crates/clawhdf5-filters)
[![docs.rs](https://docs.rs/clawhdf5-filters/badge.svg)](https://docs.rs/clawhdf5-filters)
Standalone deflate (zlib) compression and decompression with a choice of
backend: pure-Rust zlib-rs (default), zlib-ng, Apple's Compression
framework, or miniz_oxide.
Filter and compression pipeline for clawhdf5.
This crate holds **deflate backends only**. The HDF5 filter pipeline, the
filter registry and every other codec (shuffle, Fletcher-32, N-Bit,
scale-offset, LZ4, Zstd, SZIP, pcodec, LZF, bitshuffle, bzip2, Blosc,
Blosc2, ZFP) live in [`clawhdf5-format`](../clawhdf5-format/README.md),
which calls flate2 itself and selects its deflate backend with its own
features. No library crate of the workspace depends on this one (the
`clawhdf5` facade uses it only in tests).
## Features
Not on crates.io yet; depend on it from git:
- DEFLATE compression/decompression
- Pure-Rust deflate via zlib-rs (default, `zlib-rs` feature)
- zlib-ng instead, if you want it (`fast-deflate` feature; C, needs cmake)
- Apple Compression framework support (`apple-compression` feature)
```toml
[dependencies]
clawhdf5-filters = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
## API
```rust
use clawhdf5_filters::{deflate_compress, deflate_decompress};
use clawhdf5_filters::{deflate_backend, deflate_compress, deflate_decompress};
let data: Vec<u8> = (0..10_000u32).map(|i| (i % 251) as u8).collect();
let compressed = deflate_compress(&data, 6).unwrap();
// The second argument bounds the output: the expected decompressed size.
let decompressed = deflate_decompress(&compressed, data.len()).unwrap();
assert_eq!(decompressed, data);
println!("backend: {}", deflate_backend()); // "zlib-rs" by default
```
Also `deflate_compress_miniz`/`deflate_decompress_miniz` (always
miniz_oxide) and `fast_deflate::{compress, decompress, active_backend}`.
## Features
Backend priority: `apple-compression` (macOS only) > zlib-ng > zlib-rs >
miniz_oxide (with none enabled).
| Feature | Default | Backend | Builds C |
|---|---|---|---|
| `zlib-rs` | yes | zlib-rs through flate2, with `runtime_detection` (needed for its SIMD) | no |
| `fast-deflate` | no | zlib-ng through flate2 | yes (cmake) |
| `system-zlib` | no | the system zlib through flate2 | yes (`libz-sys`) |
| `apple-compression` | no | Apple Compression framework, macOS only (ignored elsewhere) | no (links a system framework) |
zlib-rs matches zlib-ng on HDF5 reads and writes and produces
byte-identical output: see "Deflate backend" in
[`BENCHMARKS.md`](../../BENCHMARKS.md).
## License
MIT
+95 -16
View File
@@ -1,27 +1,106 @@
# clawhdf5-format
[![crates.io](https://img.shields.io/crates/v/clawhdf5-format.svg)](https://crates.io/crates/clawhdf5-format)
[![docs.rs](https://docs.rs/clawhdf5-format/badge.svg)](https://docs.rs/clawhdf5-format)
The HDF5 file format in pure Rust: parsers and writers for every on-disk
structure, the filter pipeline and its codecs, and the shared type
definitions the other crates use. Most users want the
[`clawhdf5`](../clawhdf5/README.md) facade, which wraps this crate in an
h5py-like API; use this one directly for low-level access or in `no_std`
code.
Pure-Rust HDF5 binary format parsing and writing — no C dependencies.
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## What is in it
- **Parsing:** superblock v0–v3 (`superblock`, with the superblock
extension and metadata cache images, `superblock_ext`), object headers v1
and v2 (`object_header`), every header message the readers use
(`datatype`, `dataspace`, `data_layout` v1–v4 including virtual datasets,
`fill_value`, `attribute`, `link_message`, `shared_message`, ...), groups
old and new (`group_v1` symbol tables with local heaps, `group_v2` with
fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single
chunk, implicit, fixed array, extensible array, v2 B-tree).
- **Reading data:** `data_read` (contiguous, compact, chunked),
`partial_read` and `selection` (hyperslabs and points), `vl_data`
(variable-length strings and sequences through the global heap),
`chunk_cache`.
- **Storage:** the `storage::Storage` trait (`read_at`, `read_ranges`,
`len`, `hint`) that every read path goes through, so a file can be read
from memory, a file handle or a remote backend
([`clawhdf5-remote`](../clawhdf5-remote/README.md)).
- **Writing:** `file_writer::FileWriter` and the builders in
`type_builders` (datasets, groups, attributes, compound and enum types,
links, virtual datasets, creation-order tracking); chunk indexes and
dense-storage B-trees of any size (`chunked_write`, `btree_v2_write`,
`ea_writer`). Output is read by h5py and h5dump.
- **Filters:** `filter_pipeline` and `filter_registry` (look up by ID; other
IDs can be registered at run time with `register_filter`). Built in:
deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4,
Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle,
bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only).
- **Shared pieces:** `float16` (the one IEEE half-precision conversion the
workspace uses), `provenance` (SHA-256 dataset hashes), `checksum`
(Jenkins lookup3 for v2+ structures).
## Example
```rust
use clawhdf5_format::file_writer::{AttrValue, FileWriter};
use clawhdf5_format::{group_v2, object_header, signature, superblock};
// Write a file to memory
let mut fw = FileWriter::new();
fw.create_dataset("data")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_shape(&[3])
.set_attr("unit", AttrValue::String("m/s".into()));
let bytes = fw.finish().unwrap();
// Parse it back: superblock -> path -> object header
let (_user_block, file) = signature::split_user_block(&bytes).unwrap();
let sb = superblock::Superblock::parse(file, 0).unwrap();
let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap();
let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size)
.unwrap();
assert!(!hdr.messages.is_empty());
```
## Features
- Zero-copy superblock, object header, and B-tree parsing
- Chunked dataset read/write with filter pipelines
- `no_std` support (disable `std` feature)
- Optional parallel reads via Rayon
- SHA-256 provenance tracking
| Feature | Default | What | Builds C |
|---|---|---|---|
| `std` | yes | standard library; without it the crate is `no_std` + `alloc` (CI builds it for `thumbv7em-none-eabihf`) | no |
| `checksum` | yes | verify Jenkins lookup3 checksums | no |
| `deflate` | yes | deflate through flate2 | no |
| `zlib-rs` | yes | flate2's pure-Rust zlib-rs backend, with `runtime_detection` (without it zlib-rs loses SIMD and inflates 3.5x slower) | no |
| `system-zlib-decompress` | yes | no effect (nothing reads it; kept so existing feature lists still build) | no |
| `provenance` | yes | SHA-256 provenance hashes | no |
| `lzf` | yes | LZF (32000) | no |
| `parallel` | no | rayon-parallel chunk decoding | no |
| `fast-checksum` | no | hardware CRC32 through `crc32fast` | no |
| `lz4` | no | LZ4 (32004) | no |
| `pcodec` | no | pcodec | no |
| `bitshuffle`, `bzip2`, `blosc` | no | 32008, 307, 32001, read and write | no |
| `blosc2`, `zfp` | no | 32026, 32013, read only | no |
| `plugin-filters` | no | all six plugin filters above | no |
| `lookup-stats` | no | counters for name-lookup benchmarks | no |
| `zstd` | no | Zstandard (32015) | yes (libzstd) |
| `szip` | no | SZIP (4) decoding | links the system libaec (`libaec-dev`) |
| `fast-deflate` | no | zlib-ng | yes (cmake) |
| `system-zlib` | no | the system zlib | yes (`libz-sys`) |
| `blake3_hash` | no | `provenance::blake3_hash` | yes (`cc`) |
## Usage
## Robustness
```rust
use clawhdf5_format::Superblock;
let data = std::fs::read("data.h5").unwrap();
let sb = Superblock::from_bytes(&data).unwrap();
println!("HDF5 version {}.{}", sb.version_major(), sb.version_minor());
```
Every parser is meant to return an error, never panic, on hostile input:
nine cargo-fuzz targets live in [`fuzz/`](fuzz/README.md), the conformance
sweep includes the HDF Group's CVE corpus
([`CONFORMANCE.md`](../../CONFORMANCE.md)), and header checks follow
libhdf5's. Open gaps are in [`docs/known-issues.md`](../../docs/known-issues.md).
## License
+12 -4
View File
@@ -51,10 +51,18 @@ done
## CI
These targets are **not** run in CI (`.gitea/workflows/ci.yml`) — cargo-fuzz
requires nightly and each meaningful run takes minutes, which doesn't fit a
per-PR gate. Run them manually on a schedule (e.g. before a release, or after
touching parser code) instead.
These targets are **not** run by the CI workflows (`.gitea/workflows/ci.yml`)
— cargo-fuzz requires nightly and each meaningful run takes minutes, which
doesn't fit a per-PR gate. Run them by hand before a release or after
touching parser code. `scripts/ci-test.sh` has an opt-in smoke run: with
`CLAWHDF5_FUZZ_SECONDS=N` it runs every target of this crate and of
`crates/clawhdf5-agent/fuzz` (the WAL parser) for N seconds each.
Other robustness checks that do run: the nightly conformance sweep reads
the HDF Group's CVE reproducers and fails on any panic, hang, crash or
out-of-memory ([`conformance/README.md`](../../../conformance/README.md)),
and `scripts/h5rs-fuzz.sh` runs every `h5rs` subcommand over them, optionally
on byte-flipped copies.
## Reproducing Crashes
+47 -10
View File
@@ -1,25 +1,62 @@
# clawhdf5-gpu
[![crates.io](https://img.shields.io/crates/v/clawhdf5-gpu.svg)](https://crates.io/crates/clawhdf5-gpu)
[![docs.rs](https://docs.rs/clawhdf5-gpu/badge.svg)](https://docs.rs/clawhdf5-gpu)
GPU vector distance computation through [wgpu](https://wgpu.rs) and
hand-written WGSL compute shaders: upload a set of vectors once, then run
cosine or L2 top-k searches, dot products, distance matrices and norms
against them on Vulkan, Metal, DirectX 12 or OpenGL.
GPU-accelerated vector operations for clawhdf5 using wgpu compute shaders.
This crate does **not** read or write HDF5: dataset I/O in clawhdf5 is
CPU-only. It is a vector-search accelerator used optionally by
[`clawhdf5-agent`](../clawhdf5-agent/README.md) (its `gpu` feature exposes
`gpu_search::GpuSearchBackend` and a GPU arm of `strategy::search_with_metrics`;
`HDF5Memory::search` itself uses the HNSW index on the CPU).
## Features
Not on crates.io yet; depend on it from git:
- GPU-accelerated distance computations (L2, cosine)
- wgpu-based compute shaders for cross-platform GPU support
- Float16 support via `half` crate
```toml
[dependencies]
clawhdf5-gpu = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust
```rust,no_run
use clawhdf5_gpu::GpuAccelerator;
let accel = GpuAccelerator::new().unwrap();
let distances = accel.l2_distances(&query, &vectors).unwrap();
// Fall back to a CPU path when there is no usable GPU.
let mut gpu = match GpuAccelerator::new() {
Ok(g) => g,
Err(_) => return,
};
let dim = 128;
let vectors = vec![0.5f32; 1000 * dim]; // 1000 vectors, row-major
gpu.upload_vectors(&vectors, dim).unwrap();
let norms = gpu.compute_norms_gpu(&vectors, dim).unwrap();
gpu.upload_norms(&norms).unwrap();
let query = vec![1.0f32; dim];
let top10 = gpu.cosine_search(&query, 10).unwrap(); // (index, similarity), best first
let near10 = gpu.l2_search(&query, 10).unwrap(); // (index, distance), nearest first
```
`GpuAccelerator` also has `is_available`, `device_info`,
`batch_cosine_search`, `batch_dot_product`, `distance_matrix`,
`compute_norms`, and `f16_to_f32_batch`/`f32_to_f16_batch`. Vector sets
larger than the device's largest storage buffer binding are split into
chunks and the results merged. A GPU→CPU readback waits at most 30 s and
then fails with `GpuError::BufferMap` instead of hanging.
## Features
| Feature | Default | What |
|---|---|---|
| `gpu-wgpu` | yes | the wgpu implementation. Without it `GpuAccelerator::new()` returns `GpuError::NotCompiled` and `is_available()` is `false`. |
No C is compiled, but wgpu talks to the system's graphics drivers at run
time; the crate is exempt from CI's "no C in the default build" check for
that reason.
## License
MIT
+45 -15
View File
@@ -1,24 +1,54 @@
# clawhdf5-io
[![crates.io](https://img.shields.io/crates/v/clawhdf5-io.svg)](https://crates.io/crates/clawhdf5-io)
[![docs.rs](https://docs.rs/clawhdf5-io/badge.svg)](https://docs.rs/clawhdf5-io)
I/O building blocks under [`clawhdf5`](../clawhdf5/README.md): the
`HDF5Read`/`HDF5ReadWrite` traits with in-memory, borrowed, file and
memory-mapped readers, plus several experimental modules (async reads, an
HSDS client, a VOL-style trait, sub-filing, prefetch, and an MPI connector).
The facade uses it for memory-mapped reads (`MmapReader`, and the private
copy-on-write mapping that applies a metadata cache image).
I/O abstraction layer for clawhdf5.
Remote files are **not** read through this crate: HTTP(S) and object
stores go through `clawhdf5_format::storage::Storage` and
[`clawhdf5-remote`](../clawhdf5-remote/README.md).
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-io = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5", features = ["mmap"] }
```
## Main items
| Item | What |
|---|---|
| `HDF5Read`, `HDF5ReadWrite` | byte-level read/write traits; `MemoryReader`, `BorrowedReader`, `FileReader`, `FileWriter` implement them |
| `MmapReader`, `MmapReadWrite` (`mmap`) | memory-mapped files through `memmap2`; `HDF5Read::private_copy` gives a copy-on-write view |
| `prefetch::PrefetchReader`, `sweep::SweepDetector` | read-ahead (`madvise(MADV_WILLNEED)` on mappings) and chunk-sweep prediction |
| `ParallelConfig` | lane partitioning for parallel chunk decoding |
| `vol::VirtualObjectLayer`, `vol::NativeVol` | a backend-agnostic object-layer trait (modelled on libhdf5's VOL) |
| `async_read` (`async`) | tokio-based `AsyncHDF5Read` and `AsyncHDF5File` |
| `hsds::HsdsClient` (`hsds`) | a REST client for an HSDS server |
| `subfiling` | splitting one logical file across several physical files |
| `mpi_vol::MpiVol` (`mpi-io`) | an MPI connector: see below |
### MPI (`mpi-io`)
`MpiVol` is **not collective MPI-IO**. Reads are root-read + broadcast
(rank 0 reads the file with `std::fs::read`, parses the dataset and
broadcasts the bytes); writes gather every rank's shard to rank 0, which
writes the merged dataset. It does not call `MPI_File_read_at_all` or any
other MPI-IO routine. Collective I/O is on the [roadmap](../../ROADMAP.md).
`clawhdf5-bench`'s `mpi_io_bench` binary exercises it.
## Features
- Memory-mapped file access (`mmap` feature)
- Async I/O via Tokio (`async` feature)
- HSDS remote access (`hsds` feature)
- Prefetching and sweep optimizations
## Usage
```rust
use clawhdf5_io::MmapReader;
let reader = MmapReader::open("data.h5").unwrap();
```
| Feature | Default | What | Builds C |
|---|---|---|---|
| `mmap` | no (the `clawhdf5` facade turns it on) | `MmapReader`, `MmapReadWrite` | no |
| `async` | no | `async_read` (tokio) | no |
| `hsds` | no | `hsds` (reqwest, and `async`) | yes: reqwest's default TLS is native-tls (OpenSSL on Linux) |
| `mpi-io` | no | a real `MpiVol` (without it `MpiVol::new_world` returns an error) | yes: `mpi-sys` needs an MPI installation and libclang |
## License
+7 -5
View File
@@ -1,11 +1,8 @@
# clawhdf5-migrate
[![crates.io](https://img.shields.io/crates/v/clawhdf5-migrate.svg)](https://crates.io/crates/clawhdf5-migrate)
[![docs.rs](https://img.shields.io/docsrs/clawhdf5-migrate)](https://docs.rs/clawhdf5-migrate)
CLI tool to migrate a SQLite agent-memory database in the `memory_chunks` / `sessions` / `entities` / `relations` layout (table and
column names are configurable) to a
[clawhdf5-agent](https://crates.io/crates/clawhdf5-agent) store. This is **not**
[clawhdf5-agent](../clawhdf5-agent/README.md) store. This is **not**
ZeroClaw's schema — ZeroClaw keeps memories in a single `memories` table and
does not use clawhdf5.
@@ -15,10 +12,15 @@ the knowledge graph (entities and relations) are carried over.
## Installation
Not on crates.io yet; install from a checkout:
```bash
cargo install clawhdf5-migrate
cargo install --path crates/clawhdf5-migrate
```
It builds C: `rusqlite` is built with its `bundled` feature, which compiles
SQLite (so no system libsqlite is needed, but a C compiler is).
## Usage
```bash
+38
View File
@@ -0,0 +1,38 @@
# clawhdf5-napi
> **Status: does not work end to end.** The TypeScript package built on
> this crate (`packages/clawhdf5-node`) has never run successfully, is
> unpublished, and is not built or tested in CI. See "The Node.js package
> does not work" in [`docs/known-issues.md`](../../docs/known-issues.md).
> Fix it and add CI, or remove it, before depending on it.
A Node.js native addon ([napi-rs](https://napi.rs), N-API 9) exposing
[`clawhdf5-agent`](../clawhdf5-agent/README.md) as a `ClawhdfMemory`
class. It wraps `clawhdf5_agent::openclaw::ClawhdfBackend` (the agent's
`search` with re-ranking and confidence on) and the consolidation engine.
It was written for an OpenClaw integration that is not being pursued
([`docs/openclaw.md`](../../docs/openclaw.md)).
## What the addon exposes
`ClawhdfMemory.create(path, dim)`, `.open(path)`, `.openOrCreate(path,
dim)`, and on an instance: `search`, `get`, `write`, `ingestMarkdown`,
`exportMarkdown`, `save`, `saveBatch`, `stats`, `compact`, `tickSession`,
`flushWal`, `walPendingCount`, `runConsolidation`, and the ephemeral tier
(`enableEphemeral`, `ephemeralSet`/`Get`/`Delete`, `ephemeralStats`,
`promoteEphemeral`). napi-rs converts names and `#[napi(object)]` fields to
camelCase.
## Build
```bash
cargo build --release -p clawhdf5-napi # the Rust cdylib
# the .node package: npm install -g @napi-rs/cli; cd packages/clawhdf5-node; napi build --platform --release
```
It links against Node's N-API through `napi-sys` (a `-sys` crate), so it is
exempt from CI's "no C in the default build" check.
## License
MIT
+39 -11
View File
@@ -1,25 +1,53 @@
# clawhdf5-netcdf4
[![crates.io](https://img.shields.io/crates/v/clawhdf5-netcdf4.svg)](https://crates.io/crates/clawhdf5-netcdf4)
[![docs.rs](https://docs.rs/clawhdf5-netcdf4/badge.svg)](https://docs.rs/clawhdf5-netcdf4)
Read NetCDF-4 files in pure Rust. NetCDF-4 files are HDF5 files with
conventions for dimensions, coordinate variables and attributes; this crate
reads them through the [`clawhdf5`](../clawhdf5/README.md) facade, with no
libnetcdf or libhdf5. Read-only: NetCDF-3 (classic) files are not HDF5 and
are not supported.
NetCDF-4 read support built on clawhdf5 — pure Rust, no C dependencies.
Not on crates.io yet; depend on it from git:
## Features
- Read NetCDF-4 / HDF5-backed `.nc` files
- Dimension, variable, and CF convention support
- Climate and scientific data access
```toml
[dependencies]
clawhdf5-netcdf4 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Usage
```rust
```rust,no_run
use clawhdf5_netcdf4::NetCDF4File;
let nc = NetCDF4File::open("climate.nc").unwrap();
let temp = nc.variable("temperature").unwrap();
let nc = NetCDF4File::open("climate.nc")?;
for dim in nc.dimensions()? {
println!("{}: {} (unlimited = {})", dim.name, dim.size, dim.is_unlimited);
}
let mut temp = nc.variable("temperature")?;
let dims: Vec<&str> = temp.dimensions().iter().map(|d| d.name.as_str()).collect();
println!("{:?} over {:?}", temp.shape()?, dims);
let cf = temp.cf_attributes()?;
println!("units: {:?}", cf.units);
// scale_factor/add_offset applied; _FillValue and missing_value become NaN
let values: Vec<f64> = temp.read_f64()?;
# Ok::<(), clawhdf5_netcdf4::Error>(())
```
## API
| Item | What |
|---|---|
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable |
No cargo features. Tests compare against files written by netCDF4-python
(`tests/interop_tests.rs`; the CI job requires them with
`CLAWHDF5_REQUIRE_INTEROP=1`). What the HDF5 reader underneath cannot
read is listed in [`docs/known-issues.md`](../../docs/known-issues.md).
## License
MIT
+2 -3
View File
@@ -1,8 +1,5 @@
# clawhdf5-py
[![crates.io](https://img.shields.io/crates/v/clawhdf5-py.svg)](https://crates.io/crates/clawhdf5-py)
[![docs.rs](https://docs.rs/clawhdf5-py/badge.svg)](https://docs.rs/clawhdf5-py)
Python bindings for clawhdf5 — a pure-Rust HDF5 library. The package is
`clawhdf5` (`import clawhdf5`); it needs numpy and no libhdf5.
@@ -61,6 +58,8 @@ with clawhdf5.File("data.h5", "r") as f:
`clawhdf5.InternalError`, a `RuntimeError`.
- Attributes return what h5py returns; `clawhdf5.Empty` stands for a null
dataspace (h5py's `Empty`).
- Also as in h5py: `File.mode` (`'r'`, or `'r+'` for a writable file), `File.flush()` (a no-op:
edits are already synced), `Dataset.chunks`.
## Remote files
+31
View File
@@ -114,3 +114,34 @@ let s = storage.stats(); // requests, bytes_fetched, hits, misses, cached_bytes,
directory with range support (the server the tests use), and
`cargo run -p clawhdf5-remote --example read_url -- URL [DATASET]` lists a
file and prints what it cost.
## Other front ends
- `h5rs` (built with `--features remote`, or `remote-https`) takes URLs as
FILE arguments: [`clawhdf5-tools`](../clawhdf5-tools/README.md).
- Python: `clawhdf5.File("http://…")` and `File.open_url(url, ...)` go
through this crate: [`clawhdf5-py`](../clawhdf5-py/README.md).
- The browser does **not** use this crate (its cache fetches by blocking);
`clawhdf5-wasm`'s `openUrl` has its own restartable cache:
[`examples/wasm-viewer`](../../examples/wasm-viewer/README.md).
## Limits
Files a SWMR writer is still appending to cannot be followed remotely
(the file is pinned at open, so growth is `RemoteError::FileChanged`); the
block size is fixed rather than taken from a paged file's page size; the
cloud backends are built and unit-tested but have not been run against a
real bucket. The full list is under "Remote files (`clawhdf5-remote`)
limits" in [`docs/known-issues.md`](../../docs/known-issues.md); the design
is milestone M3 of [`docs/design/range-reads.md`](../../docs/design/range-reads.md).
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5-remote = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## License
MIT
+12 -1
View File
@@ -306,4 +306,15 @@ CLAWHDF5_PYTHON=.venv/bin/python CLAWHDF5_REQUIRE_INTEROP=1 cargo test -p clawhd
The interop tests write their files with h5py and compare with h5ls, h5stat,
h5dump and h5diff; each skips when what it needs is missing unless
`CLAWHDF5_REQUIRE_INTEROP=1`.
`CLAWHDF5_REQUIRE_INTEROP=1`. `tests/remote.rs` runs every subcommand on
URLs against a local range server.
This crate also holds the interop tests of the library's in-place editor
(`clawhdf5::FileEditor`), since they use `h5rs check` and h5dump on every
edited file and compare index and heap structures with what libhdf5 makes
of the same edits:
```bash
CLAWHDF5_PYTHON=.venv/bin/python CLAWHDF5_REQUIRE_INTEROP=1 \
cargo test -p clawhdf5-tools --test edit_interop --test edit_coverage_interop
```
+51
View File
@@ -0,0 +1,51 @@
# clawhdf5-wasm
clawhdf5's HDF5 and NetCDF-4 reader compiled to WebAssembly with
wasm-bindgen, for the browser (and Node). Read-only. Two ways in:
- `open(bytes)` — a file already in memory (a dropped file, a fetched
blob);
- `openUrl(url, opts)` — a file on a web server, read by HTTP range
requests as each call needs its bytes, without downloading it
(range-read milestone M4, [`docs/design/range-reads.md`](../../docs/design/range-reads.md)).
Both give `list`, `info`, `attrs`, `read` and `readHyperslab`; the remote
file's methods return promises, and `stats()` counts requests and bytes.
The JavaScript API, options, limits, package size and tests are documented
with the demo page, [`examples/wasm-viewer/README.md`](../../examples/wasm-viewer/README.md).
## Layout
- `src/core.rs` — the reader over any `clawhdf5_format::storage::Storage`
(`Reader::open_storage`), plain Rust and tested natively.
- `src/lazy.rs` — the restartable "NeedBytes" cache behind `openUrl`: a
call runs as a pass over the blocks fetched so far; a pass that misses is
abandoned, the missing (and hinted) blocks are fetched, and the pass is
run again. No block is evicted while a call runs.
- `js/remote.js` — the HTTP side: `fetch` with `Range`, checking every
answer (a `206` with exactly the bytes asked for, same ETag/Last-Modified
and length) so a call fails rather than return another file's bytes.
- `src/lib.rs` — the wasm-bindgen exports.
## Build and test
```bash
rustup target add wasm32-unknown-unknown
cargo install wasm-bindgen-cli --version 0.2.129 # must equal the crate's wasm-bindgen
bash examples/wasm-viewer/build.sh # -> examples/wasm-viewer/pkg/
cargo test -p clawhdf5-wasm # native: h5py_interop, lazy, vl_strings
bash examples/wasm-viewer/test/run.sh # Node + headless Chromium (not in CI)
```
`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus cargo test -p
clawhdf5-wasm --test lazy` compares every corpus file read lazily with the
same file read from bytes.
Built without `mmap` and `parallel` and without the Zstd and SZIP filters
(they link C): such datasets fail with `unsupported filter`. No C is
compiled; `publish = false` (it is distributed as the package
`build.sh` makes).
## License
MIT
+89 -15
View File
@@ -1,27 +1,101 @@
# clawhdf5
[![crates.io](https://img.shields.io/crates/v/clawhdf5.svg)](https://crates.io/crates/clawhdf5)
[![docs.rs](https://docs.rs/clawhdf5/badge.svg)](https://docs.rs/clawhdf5)
The main crate: a pure-Rust HDF5 reader, writer and in-place editor, with no
libhdf5 and, by default, no C code. It wraps
[`clawhdf5-format`](../clawhdf5-format/README.md) (the binary format) and
[`clawhdf5-io`](../clawhdf5-io/README.md) (memory-mapped reads) in an
h5py-like API.
Pure-Rust HDF5 reader/writer — no C dependencies.
Not on crates.io yet; depend on it from git:
```toml
[dependencies]
clawhdf5 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
```
## Main types
| Type | What it does |
|---|---|
| `File` | Opens a file (`open`, `open_buffered`, `from_bytes`, `open_storage` for any `Storage`), walks groups (`root`, `group`, `dataset`), lists `datasets`/`groups`/`attrs`. `File` is `Send + Sync`: several threads can read one open file. |
| `Dataset` | `shape`, `dtype`, `max_dimensions`, `attrs`; reads `read_f64`/`read_f32`/`read_i32`/`read_i64`/`read_u64`, strings (`read_string`, `read_string_bytes`), variable-length data (`read_vlen`), hyperslabs and point selections (`read_selection`, `read_f64_selection`, ...), zero-copy views of contiguous data (`read_f64_zerocopy`, ...), `verify_provenance`. |
| `FileBuilder` | Writes a new file: datasets of every numeric type, strings, compounds (`CompoundTypeBuilder`), enums, chunked and compressed layouts (deflate, shuffle, Fletcher-32, LZF, and with features LZ4, Zstd, bitshuffle, bzip2, Blosc, pcodec), nested groups, soft/hard/external links, virtual datasets, attribute creation order. Files open in h5py and h5dump. |
| `FileEditor` | Changes an existing file in place without rewriting it: `write_values`/`write_selection`/`write_all`, `resize` of chunked datasets (every chunk index), `set_attr` (compact and dense storage). Anything it cannot do safely is `Error::Unsupported` before any write. |
| `MmapFile`, `LazyFile` | Alternative readers: memory-mapped, and one that reads lazily and caches. |
| `File::open_swmr` | Reads a file a libhdf5 SWMR writer is still appending to (`Dataset::refresh`, bounded retries), as h5py's `swmr=True` reader does. |
## Examples
```rust,no_run
use clawhdf5::{AttrValue, File, FileBuilder, FileEditor, Selection};
// Write
let mut b = FileBuilder::new();
b.create_dataset("sensors/temperature")
.with_f64_data(&[20.5, 21.0, 21.5, 22.0])
.with_shape(&[4])
.with_maxshape(&[u64::MAX]) // unlimited, so it can grow
.with_chunks(&[2])
.with_deflate(4);
b.set_attr("version", AttrValue::I64(1));
b.write("data.h5")?;
// Read
let file = File::open("data.h5")?;
let ds = file.dataset("sensors/temperature")?;
assert_eq!(ds.shape()?, vec![4]);
let values = ds.read_f64()?;
// Edit in place: grow the dataset and fill the new tail
let mut ed = FileEditor::open("data.h5")?;
ed.resize("sensors/temperature", &[6])?;
let tail = Selection::Hyperslab {
start: vec![4],
stride: vec![1],
count: vec![2],
block: vec![1],
};
ed.write_values("sensors/temperature", &tail, &[22.5f64, 23.0])?;
# Ok::<(), clawhdf5::Error>(())
```
Remote files (HTTP range requests, S3/GCS/Azure) are read through
`File::open_storage`; [`clawhdf5-remote`](../clawhdf5-remote/README.md)
provides the storage and its block cache.
## Features
- Read and write HDF5 files entirely in Rust
- Memory-mapped I/O for large files (`mmap` feature, enabled by default)
- Parallel chunk reads via Rayon (`parallel` feature)
- Lazy dataset access for minimal memory usage
- h5py-compatible file output
| Feature | Default | What | Builds C |
|---|---|---|---|
| `mmap` | yes | memory-mapped reads (`File::open` maps the file; `MmapFile`) | no |
| `provenance` | yes | SHA-256 `_provenance_sha256` attributes (`DatasetBuilder::with_provenance`, `Dataset::verify_provenance`) | no |
| `lzf` | yes | LZF filter (32000), h5py's `compression="lzf"` | no |
| `parallel` | no | chunk decoding on a rayon pool | no |
| `lz4` | no | LZ4 filter (32004) | no |
| `pcodec` | no | pcodec filter | no |
| `bitshuffle`, `bzip2`, `blosc` | no | plugin filters 32008, 307, 32001 (read and write) | no (bzip2 uses the pure-Rust `libbz2-rs-sys`) |
| `blosc2`, `zfp` | no | plugin filters 32026 and 32013, **read only** | no |
| `plugin-filters` | no | `lzf`, `bitshuffle`, `bzip2`, `blosc`, `blosc2`, `zfp` | no |
| `zstd` | no | Zstandard filter (32015) | yes (libzstd) |
| `fast-deflate` | no | zlib-ng instead of the pure-Rust zlib-rs | yes (cmake) |
| `blake3_hash` | no | `provenance::blake3_hash` helpers | yes (`cc`, for blake3's SIMD code) |
| `apple-compression` | no | currently has no effect in this crate (it is not forwarded) | — |
## Usage
SZIP decoding is a `clawhdf5-format` feature (`szip`, links the system
libaec); the facade does not forward it.
```rust
use clawhdf5::File;
## Limits and further reading
let file = File::open("data.h5").unwrap();
let dataset = file.dataset("/group/data").unwrap();
let values: Vec<f64> = dataset.read_1d().unwrap();
```
- What is known not to work, and what was wrong in earlier releases:
[`docs/known-issues.md`](../../docs/known-issues.md) (editor limits, range
reads, external links and external raw data, which are explicit errors).
- Read coverage against libhdf5/h5py on eight public corpora:
[`CONFORMANCE.md`](../../CONFORMANCE.md).
- Read and write speed against libhdf5 and h5py:
[`BENCHMARKS.md`](../../BENCHMARKS.md).
- Range reads and SWMR design: [`docs/design/range-reads.md`](../../docs/design/range-reads.md),
[`docs/design/swmr.md`](../../docs/design/swmr.md).
- Changes: [`CHANGELOG.md`](../../CHANGELOG.md).
## License
@@ -1,3 +1,5 @@
> **Historical (archived 2026-09-28):** a log of an automated improvement loop's PRs from April–May 2026, on the earlier `quantumclaw/clawhdf5` PR numbering (not today's). Superseded by [`CHANGELOG.md`](../../CHANGELOG.md) and `git log`.
# Improvement Log -- clawhdf5
| Date | Loop | PR | Changes | Status |
@@ -1,3 +1,5 @@
> **Historical (archived 2026-09-28):** one automated scan's notes (2026-05-04), describing changes long since merged. Superseded by [`CHANGELOG.md`](../../CHANGELOG.md) and `git log`.
# Improvement Scan -- clawhdf5
**Date:** 2026-05-04
@@ -1,3 +1,5 @@
> **Historical (archived 2026-09-28):** a pre-work plan, implemented in `d6c4d4f` (2026-06-30). Superseded by the code (`crates/clawhdf5-format`), [`CHANGELOG.md`](../../../CHANGELOG.md) and [`ROADMAP.md`](../../../ROADMAP.md); not an open task list.
# Filter Codecs Implementation Plan
> **Status (2026-08-03):** Implemented — shipped in commit `d6c4d4f` (2026-06-30), with FFI/constant fixes in `cb0b0e9`/`e91f7fc`. This doc was authored 2026-06-29 as the pre-work plan and committed to the repo retroactively on 2026-08-03; checkboxes below have been marked complete to match. Treat this as a historical record, not an open task list.
@@ -1,3 +1,5 @@
> **Historical (archived 2026-09-28):** a pre-work plan, implemented in `d6c4d4f` (2026-06-30) and on 2026-08-03. Superseded by the code (`crates/clawhdf5-format`), [`CHANGELOG.md`](../../../CHANGELOG.md) and [`ROADMAP.md`](../../../ROADMAP.md); not an open task list.
# Format Write Extensions Implementation Plan
> **Status (2026-08-03):** Implemented. Tasks 1–3 (external links, VDS mapping serialization, VDS `FileWriter` API) shipped in commit `d6c4d4f` (2026-06-30). Tasks 4–5 (superblock v4 read/write) were not part of that commit and were completed separately as part of this cleanup pass (2026-08-03) — see `Superblock::parse_v4`/`serialize` and `FileWriter::with_page_size` in `crates/clawhdf5-format`. This doc was authored 2026-06-29 as the pre-work plan and committed to the repo retroactively; checkboxes below have been marked complete to match current state. Treat this as a historical record, not an open task list.
@@ -1,3 +1,5 @@
> **Historical (archived 2026-09-28):** a pre-work plan. Its goal of *collective* MPI-IO is not what shipped: `MpiVol` (`d6c4d4f`) is root-read + broadcast and gather-to-root writes (see [`crates/clawhdf5-io/README.md`](../../../crates/clawhdf5-io/README.md)); collective I/O is an open item in [`ROADMAP.md`](../../../ROADMAP.md).
# MPI-IO VOL Backend Implementation Plan
> **Status (2026-08-03):** Implemented — shipped in commit `d6c4d4f` (2026-06-30), with FFI/constant fixes in `cb0b0e9`/`e91f7fc`. This doc was authored 2026-06-29 as the pre-work plan and committed to the repo retroactively on 2026-08-03; checkboxes below have been marked complete to match. Treat this as a historical record, not an open task list.
+54 -16
View File
@@ -79,6 +79,30 @@ for a deep one) and one batch of requests for its chunks. Every answer is
checked — a `206` with exactly the bytes asked for, from the same file
(ETag or Last-Modified, and length) — or the call fails.
What that costs, counted on tank on 2026-09-27 (`CHANGELOG.md`, "Remote
files in the browser: fewer round trips"), on an h5py file of 3000
datasets of 16384 `f32` in one group (198 MB, h5py 3.16 / HDF5 2.0), as
passes / requests / bytes fetched, with
`CLAWHDF5_WASM_LIST_FILE=<file> CLAWHDF5_WASM_READ=/d1500 cargo test
--release -p clawhdf5-wasm --test lazy listing_cost_of_a_given_file --
--nocapture`:
| file (`libver`), block size | `list('/')` | open + read one dataset |
|---|---|---|
| earliest, 1 MiB | 4 / 68 / 192.5 MB | 6 / 5 / 5.2 MB |
| earliest, 64 KiB | 5 / 530 / 35.3 MB | 8 / 7 / 0.52 MB |
| latest, 1 MiB | 5 / 86 / 196.5 MB | 7 / 7 / 6.7 MB |
| latest, 64 KiB | 6 / 454 / 30.5 MB | 8 / 8 / 0.58 MB |
Listing reads every child's object header, and h5py spreads those through
the file, so a listing of a group this large fetches most of it at 1 MiB
blocks; a smaller `blockSize` fetches far less at the cost of more
requests. Reading one dataset does not list the group. In the test suite's
200 MB file (`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`),
listing the root, reading two small datasets, a group's attributes, the
large dataset's shape and a 10-value window of it took 5 requests and
6 MiB.
`data` is the typed array of the stored width (`Float64Array`,
`Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`,
`BigUint64Array`), or an array of strings for fixed- and variable-length
@@ -155,25 +179,39 @@ browser).
## Size
Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14,
`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is
larger now and the table has not been re-measured: the reader has grown
since, and `openUrl` (2026-09-27) made the facade's range-read path
reachable from JavaScript and added promise glue and `remote.js`.
Measured 2026-09-28 on tank at `9b5803f` (rustc 1.98.1, wasm-bindgen
0.2.129, gzip 1.14): `bash examples/wasm-viewer/build.sh`, then `wc -c` and
`gzip -9 -n -c FILE | wc -c` of each file in `pkg/`. The opt-level `z` and
`3` rows are the same build with `CARGO_PROFILE_WASM_RELEASE_OPT_LEVEL=z`
(or `3`) and the same `wasm-bindgen --target web` step.
| | raw | gzip -9 |
|---|---:|---:|
| `pkg/clawhdf5_wasm_bg.wasm` (profile `wasm-release`, opt-level `s`) | 627,501 B | 191,639 B |
| `pkg/clawhdf5_wasm.js` (wasm-bindgen glue) | 21,826 B | 4,487 B |
| same wasm at opt-level `z` | 693,068 B | 192,550 B |
| same wasm at opt-level `3` | 544,035 B | 198,803 B |
| `pkg/clawhdf5_wasm_bg.wasm` (profile `wasm-release`, opt-level `s`) | 1,384,607 B | 378,485 B |
| `pkg/clawhdf5_wasm.js` (wasm-bindgen glue) | 40,711 B | 8,181 B |
| `pkg/snippets/.../js/remote.js` (the HTTP side of `openUrl`) | 9,326 B | 3,448 B |
| same wasm at opt-level `z` | 1,533,772 B | 374,765 B |
| same wasm at opt-level `3` | 1,184,889 B | 394,569 B |
| h5wasm 0.10.3: wasm embedded in `dist/esm/hdf5_util.js` | 3,544,184 B | 907,096 B |
| h5wasm 0.10.3: `dist/esm/hdf5_util.js` as shipped | 4,150,134 B | 986,699 B |
h5wasm figures: `npm pack [email protected]` (npm reports
`dist.unpackedSize` 14,731,385 B for the whole package), wasm extracted from
the `binaryDecode` literal in `hdf5_util.js`. h5wasm is the whole of libhdf5
(writing, every datatype, plugins), so this compares download size, not
equal functionality. No `wasm-opt` pass was applied (binaryen is not
installed on tank). opt-level `s` is used because it is the smallest
compressed.
The previous measurement (2026-09-26, before `openUrl`) was 627,501 B /
191,639 B gzipped for the wasm and 21,826 B / 4,487 B for the glue. The
package roughly doubled since. `openUrl` made the facade's `Storage` read
path reachable from JavaScript (it was compiled out before) and added the
lazy cache and the promise glue (`CHANGELOG.md`, M4); the growth has not
been broken down per change. Of the
wasm's 1,384,607 bytes, 476,062 are the `name` custom section (function
names, which wasm-bindgen keeps; `wasm-bindgen --remove-name-section` or a
`wasm-opt` pass would drop them); code is 816,244 and data 83,431. Without
the name section the wasm is 908,541 B, 329,854 B gzipped (section removed
with a script, not a supported build option yet). At
opt-level `z` the gzipped wasm is now 1% smaller than at `s`, which the
profile still uses.
h5wasm figures (2026-09-26, unchanged): `npm pack [email protected]` (npm
reports `dist.unpackedSize` 14,731,385 B for the whole package), wasm
extracted from the `binaryDecode` literal in `hdf5_util.js`. h5wasm is the
whole of libhdf5 (writing, every datatype, plugins), so this compares
download size, not equal functionality. No `wasm-opt` pass was applied
(binaryen is not installed on tank).