diff --git a/ROADMAP.md b/ROADMAP.md index 319a27e..6ebf48e 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -1,193 +1,162 @@ -# ClawhDF5 Roadmap β€” Agent Memory Evolution +# clawhdf5 roadmap -> Making clawhdf5 the defacto agentic memory solution. -> Single file. Pure Rust. Zero dependencies. Trusted everywhere. +What has shipped, and what is genuinely next. Everything here is checked +against `CHANGELOG.md`, `git log` and [`docs/known-issues.md`](docs/known-issues.md); +dates are merge dates on `main`. Nothing after v2.7.0 has been released: +the work since then is on `main` under `CHANGELOG.md` "Unreleased". + +_Last updated: 2026-09-28 (at `9b5803f`, PR #21)._ --- -## Track 1: Knowledge Graph in HDF5 -**Status:** 🟒 Phase 1 Complete -**Priority:** Critical -**Crate:** `clawhdf5-agent` +## Done -- [x] **1.1** Entity storage β€” entities with properties, embeddings, timestamps (created_at/updated_at) -- [x] **1.2** Relation storage β€” typed edges with RelationType enum (Temporal/Causal/Associative/Hierarchical/Custom), metadata, timestamps -- [x] **1.3** Entity extraction helpers β€” rule-based extraction (Person, Org, Location, Date, Technology, Project) with extract_and_store_entities() integration -- [x] **1.4** Entity resolution β€” fuzzy name matching (Levenshtein distance) via resolve_or_create() -- [x] **1.5** Graph traversal queries β€” BFS neighbors with depth, subgraph extraction from seeds -- [x] **1.6** Spreading activation β€” weighted activation propagation with configurable decay -- [x] **1.7** Graph-aware retrieval β€” get_entity_context() for formatted context injection -- [x] **1.8** Tests β€” comprehensive tests for all new features +### Releases -**Research:** Graph-Native Cognitive Memory (2026), Graph-based Agent Memory survey (2026), SYNAPSE (2025) +| Version | Date | Headline | +|---|---|---| +| v2.0.0 | 2026-03-19 | rustyhdf5 (11 crates) and edgehdf5 (4 crates) unified into one workspace as `clawhdf5-*` | +| v2.1.0 | 2026-06-03 | HNSW backs the agent's vector search by default; live, mutable HNSW index | +| v2.2.0 – v2.7.0 | 2026-09-18 – 2026-09-20 | bounded decompression and read-path bounds checks, single-writer store locking, WAL v4, HNSW recall fix (0.31 -> 0.98 recall@10 at 100K), fusion weights tuned on LongMemEval, int8 index, Extensible Array read fix and chunk-index checksums | + +Details per release: [`CHANGELOG.md`](CHANGELOG.md). + +### Since v2.7.0 (unreleased, on `main`) + +| PR | Merged | What | +|---|---|---| +| #3 | 2026-09-23 | pure-Rust deflate (zlib-rs) by default, no C in the core crates' default build (checked in CI), MSRV 1.92 | +| #4 | 2026-09-25 | files open in h5py again (every `f32` and every empty dataset clawhdf5 wrote was unreadable by libhdf5); float16 embedding storage | +| #5 | 2026-09-25 | `HDF5Memory::search` with `SearchOptions` (source filters, re-ranking, confidence); float16 on by default | +| #6 | 2026-09-25 | `clawhdf5-migrate` writes real agent stores; knowledge-graph fix; dated benchmark re-run | +| #7 | 2026-09-25 | consolidation benchmark completed (cheaper novelty scoring) | +| #8 | 2026-09-25 | Ed25519-signed checkpoints (`HDF5Memory::verify`) | +| #9, #10 | 2026-09-25 | OpenClaw and ZeroClaw integration claims withdrawn β€” neither ever integrated clawhdf5 | +| #11 | 2026-09-26 | silent wrong data and libhdf5 interop bugs found by the HDF5 audit fixed | +| #12 | 2026-09-26 | reproducible conformance sweep over eight public corpora, nightly CI job ([`CONFORMANCE.md`](CONFORMANCE.md)) | +| #13 | 2026-09-26 | reads HDF5 1.6-era layouts, user blocks, virtual datasets, dense attributes, very large groups | +| #14 | 2026-09-26 | `h5rs` tools (`ls`, `dump`, `stat`, `diff`, `check`), the browser reader (`clawhdf5-wasm`), libhdf5's header checks, plugin filters (LZF, bitshuffle, bzip2, Blosc), concurrency benchmark | +| #15 | 2026-09-26 | fast contiguous and concurrent reads, variable-length data, nested groups and links in the writer, Python bindings | +| #16 | 2026-09-26 | chunked full reads faster than an h5py process pool, writer B-trees of any size, Blosc2 (read), 599/697 conformance | +| #17 | 2026-09-26 | range reads M0/M1 (indexed name lookups, the `Storage` trait), ZFP (read), in-place editing (`FileEditor`) | +| #18 | 2026-09-27 | range reads M2/M3 (`File::open_storage`; `clawhdf5-remote`: HTTP(S), S3, GCS, Azure), in-place editing of every chunk index, shrinking, dense attributes | +| #19 | 2026-09-27 | remote files in the browser (`openUrl`, M4), SWMR reader (`File::open_swmr`, M5), Python remote reads and `'r+'` editing | +| #20 | 2026-09-28 | benchmarks re-measured: LongMemEval with real MiniLM embeddings, local reads on an idle machine | +| #21 | 2026-09-28 | remote files open in a few requests (group lookups down the B-tree, `Storage::hint`), `ObjectHeader::parse` back to its earlier speed, the last conformance mismatches resolved: 602/697 ok, 0 our-error, 0 mismatch | + +### Range reads (design: [`docs/design/range-reads.md`](docs/design/range-reads.md)) + +- [x] M0 β€” indexed name lookups (#17) +- [x] M1 β€” metadata parsed through the `Storage` trait (#17) +- [x] M2 β€” raw data through `Storage`, `File::open_storage` (#18) +- [x] M3 β€” `clawhdf5-remote`: HTTP(S) range requests and object stores through a block cache; `h5rs` URLs (#18); Python URLs (#19) +- [x] M4 β€” `openUrl` in the browser, restartable "NeedBytes" cache (#19; fewer round trips in #21) +- [x] M5 β€” reading files a SWMR writer is appending to ([`docs/design/swmr.md`](docs/design/swmr.md), #19) + +### Agent memory (`clawhdf5-agent`) + +Shipped before and during the v2 releases, and kept current since: +knowledge graph with entity extraction and resolution; three-tier +consolidation with decay; hybrid retrieval (HNSW + BM25, weighted or RRF +fusion, re-ranking, confidence rejection, query expansion); temporal index +and session DAG; per-save provenance ledger and write-anomaly detection; +multi-modal embeddings; WAL with chained CRC32; single-writer locking; +signed checkpoints. Retrieval is measured, not claimed: see +[`BENCHMARKS.md`](BENCHMARKS.md) ("LongMemEval Results" reports retrieval +recall, not QA accuracy; earlier headline numbers that compared different +granularities were retracted there). --- -## Track 2: Memory Consolidation Engine -**Status:** 🟒 Phase 1 Complete -**Priority:** Critical -**Crate:** `clawhdf5-agent` +## Next -- [x] **2.1** Importance scoring β€” surprise (novelty), correction boost, length scoring with configurable weights -- [x] **2.2** Three-tier memory model β€” Working β†’ Episodic β†’ Semantic with bounded capacities -- [x] **2.3** Time-decay with reactivation β€” exponential decay with configurable half-life, access resets timestamp -- [x] **2.4** Bounded memory with graceful degradation β€” evict lowest-decay entries when over capacity -- [x] **2.5** Consolidation cycles β€” promote/evict across tiers based on importance and access thresholds -- [x] **2.6** Memory statistics β€” ConsolidationStats with per-tier counts, eviction/promotion tracking -- [x] **2.7** Tests β€” comprehensive tests for all features +Not scheduled; listed roughly by how much they unblock. None has a date. -**Research:** CraniMem (2026), D-MEM (2026), AI Hippocampus survey (2026) +### Distribution + +- [ ] **Publish the crates to crates.io.** Nothing is published; the READMEs + say to depend on git. Before publishing: no `publish` settings exist + (only `clawhdf5-wasm` has `publish = false`), and several `Cargo.toml` + descriptions still name the old `rustyhdf5`/`edgehdf5` (`clawhdf5-accel`, + `-derive`, `-gpu`, `-io`, `-netcdf4`, `-android`). +- [ ] **Publish Python wheels to PyPI.** `crates/clawhdf5-py` builds with + maturin and is tested in CI, but no wheel is published. The default wheel + reads plain `http://` only; `https`/`s3`/`gcs`/`azure` wheels compile C + (ring, aws-lc-rs). +- [ ] **The Node.js package** (`packages/clawhdf5-node` over + `clawhdf5-napi`) has never worked and is not in CI: fix it and add CI, or + remove it ([known issue](docs/known-issues.md)). + +### HDF5 features + +- [ ] **SWMR writing.** The reader is done (M5); writing a file while + libhdf5 readers follow it is not. Also not covered: remote SWMR (a remote + file is pinned at open), `MmapFile`/`LazyFile` SWMR reads, refreshing + groups or attributes. +- [ ] **MPI collective I/O.** `clawhdf5-io`'s `MpiVol` (`mpi-io`) is + root-read + broadcast and gather-to-root writes, not collective MPI-IO + (`MPI_File_read_at_all`/`write_at_all`). +- [ ] **Paged-metadata single-request reads.** Files written with paged + aggregation (`H5Pset_file_space_strategy(PAGE)`, `h5repack -S PAGE`) + keep their metadata in a few pages; range reads could fetch those in one + request and use the file's page size as the block size. Today the block + size is fixed (1 MiB) and only the first block is read ahead + (range-reads design, option (c) as a policy). +- [ ] **Blosc2 and ZFP encoders.** Both filters are read-only; the other + plugin filters (LZF, bitshuffle, bzip2, Blosc 1) read and write. +- [ ] **External links and external raw data** are explicit errors, not + followed. +- [ ] **Virtual datasets:** the "first missing" view and printf gaps other + than 0, source-to-virtual type conversion other than a byte swap, nested + virtual sources, source files outside the virtual file's directory. +- [ ] **Datatypes:** x87 long double and binary128 are refused. +- [ ] **Writer:** one attribute or link message over 65 515 bytes in dense + storage is an error (huge fractal-heap objects); no option to write + files HDF5 1.8 can read. +- [ ] **`FileEditor`:** new chunks in implicit indexes, variable-length and + reference data, filters it cannot encode (scale-offset, N-Bit, SZIP), + some dense-attribute heap layouts, creating or deleting objects and + attributes (also from Python `'r+'`), and no journal (a crash mid-edit + can leave the file inconsistent). Freed space is reused only within one + editor. +- [ ] **Selection reads** decode the whole dataset when the selection's + bounding box covers more than half of it (a strided `ds[::100]`), and + for compact/virtual datasets or a non-default fill value: correct, but + more work than needed. +- [ ] **Readers:** `LazyFile` and `MmapFile` still need the whole file; + the zero-copy methods need the file in memory. + +### Remote and browser + +- [ ] Run the `s3`/`gcs`/`azure` backends against real buckets (only built + and URL-parsing-tested so far). +- [ ] `h5rs` options for request headers and cache settings. +- [ ] Browser limits in [`docs/known-issues.md`](docs/known-issues.md) + ("`clawhdf5-wasm` (browser) limits"): files of 4 GiB or more (wasm32), + compound/reference/opaque datasets, round trips per index level. The + package doubled in size with `openUrl` + ([size table](examples/wasm-viewer/README.md#size)); dropping the + function-name section would take a third off the raw size (13% gzipped). + +### Quality + +- [ ] Scheduled fuzz campaigns: the cargo-fuzz targets + ([`crates/clawhdf5-format/fuzz`](crates/clawhdf5-format/fuzz/README.md), + and the agent's WAL target) run only by hand or with + `CLAWHDF5_FUZZ_SECONDS`. --- -## Track 3: Hybrid Retrieval Pipeline -**Status:** 🟒 Phase 1 Complete -**Priority:** High -**Crate:** `clawhdf5-agent` +## Withdrawn -- [x] **3.1** Reciprocal Rank Fusion (RRF) β€” rrf_hybrid_search() with k=60 constant -- [x] **3.2** Multi-factor re-ranking β€” temporal decay, source authority hierarchy, activation scores (reranker.rs) -- [x] **3.3** Low-confidence rejection β€” min_score threshold, gap filtering, max_results (confidence.rs) -- [x] **3.4** Query expansion β€” synonyms, acronyms, temporal rewrites, morphological variants, knowledge graph aliases + expanded_search() with RRF merge -- [x] **3.5** Result explanation β€” ReRankResult with full score breakdown per factor -- [x] **3.6** Configurable pipeline β€” ReRankConfig + ConfidenceConfig with tunable weights/thresholds -- [x] **3.7** Tests + MemX-comparable benchmarks β€” 5 integration tests (Hit@1β‰₯90%, search<500ms@100K, BM25<200ms@100K, hybrid<50ms@10K, compact<200ms@10K) +- **OpenClaw integration** (withdrawn 2026-09-25, PR #9). clawhdf5 was + never an OpenClaw memory plugin; the documented + `memory.backend = "clawhdf5"` was never valid. The Rust `ClawhdfBackend` + remains as a library API. [`docs/openclaw.md`](docs/openclaw.md) records + what a real plugin would need. +- **ZeroClaw integration** (withdrawn 2026-09-25, PR #10). ZeroClaw has no + clawhdf5 backend, and `clawhdf5-migrate`'s SQLite layout is not + ZeroClaw's schema. -**Research:** MemX (2026), SwiftMem (2026) - ---- - -## Track 4: Temporal Reasoning -**Status:** 🟒 Phase 1 Complete -**Priority:** High -**Crate:** `clawhdf5-agent` - -- [x] **4.1** Temporal index β€” sorted timestamp index with binary search, insert/remove -- [x] **4.2** Time-range queries β€” range_query, before, after, latest, earliest -- [x] **4.3** Session DAG β€” parent/child linking, chain walking, time-range overlap queries -- [x] **4.4** Temporal re-ranking β€” query hint enum (Latest/Earliest/Around/Between/None) with boost scoring -- [x] **4.5** Temporal entity tracking β€” EntityTimeline with state change history + point-in-time reconstruction -- [x] **4.6** Tests β€” comprehensive tests for all features - -**Research:** MemX temporal gaps (≀43.6% Hit@5), MemoryArena multi-session tasks (2026) - ---- - -## Track 5: Memory Security & Provenance -**Status:** 🟒 Phase 1 Complete -**Priority:** Medium-High -**Crate:** `clawhdf5-agent` - -- [x] **5.1** Source attribution β€” MemoryProvenance with source, creator, session, FNV-1a content hash -- [x] **5.2** Write anomaly detection β€” rate limiting, 15 injection patterns, source distribution analysis -- [x] **5.3** Source isolation β€” per-MemorySource sub-stores preventing cross-contamination -- [x] **5.4** Memory integrity verification β€” content hash comparison via verify_integrity() -- [x] **5.5** Poisoning resistance β€” pattern detection for prompt injection attempts -- [x] **5.6** Tests β€” comprehensive tests including adversarial patterns - -**Research:** MemoryGraft (2025), SSGM Framework (2026) - ---- - -## Track 6: Multi-Modal Memory -**Status:** 🟒 Phase 1 Complete -**Priority:** Medium -**Crate:** `clawhdf5-agent` - -- [x] **6.1** Image embedding storage β€” ModalEmbedding with model provenance (CLIP, SigLIP, etc.) -- [x] **6.2** Audio fingerprints β€” Audio modality with embedding storage -- [x] **6.3** Multi-modal search β€” search_by_modality (filtered) + search_cross_modal (all embeddings) -- [x] **6.4** Observation records β€” raw perception vs interpretation with confidence scoring -- [x] **6.5** Media reference storage β€” MediaRef with Path/Url/Inline, MIME types, FNV-1a checksums -- [x] **6.6** Tests β€” 35 comprehensive tests - -**Research:** Neuro-Symbolic Memory (2026), RAGdb multi-modal RAG (2025) - ---- - -## Track 7: OpenClaw Integration β€” withdrawn (2026-09-25) -**Status:** βšͺ Withdrawn (the items below were library work; no OpenClaw integration shipped) -**Priority:** Critical (for adoption) -**Crates:** `clawhdf5-agent`, `clawhdf5-napi` - -- [x] **7.1** Memory backend trait β€” MemoryBackend with search/get/write/ingest/export/stats -- [x] **7.2** Hybrid retrieval pipeline β€” ClawhdfBackend wires RRF β†’ reranker β†’ confidence rejection -- [x] **7.3** Markdown import/export β€” MarkdownParser + MarkdownExporter with line tracking + metadata -- [x] **7.4** `search()` β€” backed by the full hybrid retrieval pipeline (a Rust method; no OpenClaw tool was ever registered) -- [x] **7.5** `get()` β€” read back by path, with a line slice (not an OpenClaw tool either) -- [x] **7.6** Compaction integration β€” run_compaction() (decay + compact + WAL flush), run_consolidation() (hippocampal engine), tick_session(), flush_wal() -- [ ] **7.7** ~~Config surface β€” `memory.backend = "clawhdf5"`~~ β€” never valid OpenClaw config; docs removed -- [ ] **7.8** ~~Documentation + migration guide~~ β€” removed: they described an integration that never worked - -**Node.js bridge:** `clawhdf5-napi` (napi-rs) and a TypeScript wrapper in `packages/clawhdf5-node` exist but are unpublished, untested in CI and known to be broken (docs/known-issues.md). - ---- - -> **Withdrawn.** None of this track produced a working OpenClaw integration: no -> plugin was built, the documented `memory.backend = "clawhdf5"` config was never -> valid in any OpenClaw release, and the Node package was never published. The -> Rust `ClawhdfBackend` remains as a library API. Not pursued for now; see -> [docs/openclaw.md](docs/openclaw.md) for what a plugin would need today. - -## Track 8: Benchmarking & Validation -**Status:** 🟒 Complete -**Priority:** High -**Crates:** `clawhdf5-agent`, `clawhdf5-bench` - -- [x] **8.1** MemoryArena benchmark β€” 35 queries, 50 sessions, Hit@10=91.4%, MRR=0.547 -- [x] **8.2** LongMemEval benchmark β€” 500 questions, retrieval recall (not QA accuracy). Full `longmemeval_s` haystack, hybrid 0.4/0.6 with MiniLM embeddings: turn Hit@5 81.4%, MRR 0.643; session Hit@5 96.8% (re-run 2026-09-27 on tank). Oracle variant: BM25-only turn Hit@5 84.4%, MRR 0.660; hybrid 86.8%. The session Hit@1 of 100% first recorded here was degenerate on the oracle variant, and the "beats MemX 51.6%" claim compared a different granularity. Both are retracted; see [BENCHMARKS.md Β§ LongMemEval Results](BENCHMARKS.md#longmemeval-results) -- [x] **8.3** Latency benchmarks β€” vector search at 1K/10K/100K, hybrid/RRF, graph traversal, consolidation, temporal -- [x] **8.4** Memory footprint β€” 1.7 KB/record uncompressed, 282 B compressed (6.2x ratio), 100K+ rec/s ingestion -- [x] **8.5** Consolidation efficiency β€” 8.8x search speedup, 90% noise eviction, zero quality loss -- [x] **8.6** Cross-platform benchmarks β€” x86 measured, ARM estimated, cross_platform.sh script -- [x] **8.7** Published results in BENCHMARKS.md with ephemeral tier Redis comparison (70-140x faster) - ---- - -## Implementation Order - -**Phase 1:** ~~Tracks 1, 2, 3 β€” core memory intelligence~~ 🟒 Complete -**Phase 2:** ~~Track 4 (temporal) + Track 5 (security)~~ 🟒 Complete -**Phase 3:** ~~Track 6 (multi-modal)~~ 🟒 Complete; Track 7 (OpenClaw integration) withdrawn -**Phase 4:** ~~Track 8 (benchmarking + validation)~~ 🟒 Complete - -All 8 tracks delivered. 1,650+ tests passing, zero clippy warnings. - ---- - -## What's Next - -Verified against current repo state on 2026-08-05 (see also `docs/superpowers/plans/` for the filter-codec/format-write/MPI-IO work, now shipped): - -- [ ] TypeScript bridge not wired into CI β€” `packages/clawhdf5-node/` already has a complete, working napi-rs package (package.json, tsconfig, hand-written TS wrapper matching all 21 `#[napi]` items, Jest test suite, README); it isn't published to npm and has no committed lockfile -- [ ] Publish crates to crates.io β€” no `publish` config anywhere in the workspace yet -- [ ] Python wheel distribution via maturin β€” `crates/clawhdf5-py/pyproject.toml` exists (maturin-buildable locally) but wheels aren't published anywhere -- [ ] `chunked_read.rs`/`data_read.rs` full bounds-check audit + scheduled fuzz campaigns (the new `fuzz_dataset_read` target covers the two files' main entry points; a full manual audit of every indexing site is still open) β€” see Tier 4 below -- [ ] WAL per-entry checksum landed as CRC32 (see below); a stronger per-entry format (explicit length prefix, avoiding the read-then-verify restructuring) could still be revisited if profiling shows it matters -- [ ] HNSW build parallelism is still narrow (only `prune_connections`); the correctness-sensitive outer insert loop needs its own dedicated design pass before parallelizing - -### Recently closed out (2026-08-05, Tier 3–4 hardening pass) - -- [x] Academic benchmark cross-validation β€” LongMemEval reproduced on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% on the oracle variant (the comparison with MemX's 51.6% made here was later retracted, since MemX measures fact-level granularity over a far larger corpus); recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md Β§ Independent Validation: tank β€” LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05) -- [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer -- [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 β†’ 0.29, clearing two RUSTSEC advisories -- [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open -- [x] `chunked_read.rs`/`data_read.rs`/`local_heap.rs` bounds-check audit: added `ensure_len` overflow guards, a recursion-depth guard against cyclic B-trees, and a fix for an unguarded compound-datatype byte-offset overrun. Added a new `fuzz_dataset_read` cargo-fuzz target exercising the contiguous/chunked/compact read paths β€” it found and we fixed 3 real crash bugs (integer-overflow panics) within the first few runs -- [x] `clawhdf5-ann`: optional `parallel` feature (rayon) for HNSW's `prune_connections` neighbor-distance computation -- [x] `[workspace.dependencies]` added for `tempfile`/`criterion`/`half`/`serde`, fixing a real version skew on `half` (2 vs 2.7) - -### Recently closed out (2026-08-05 hardening pass) - -- [x] CI/CD pipeline β€” `.gitea/workflows/ci.yml` now runs `scripts/ci-test.sh` (fmt, clippy, tests, no_std check) on push/PR to `main` -- [x] Fixed no_std build breakage in `clawhdf5-format` (missing alloc imports, `AtomicU64` unsupported on thumbv7em, `f64::powi` requiring std/libm) -- [x] Fixed version skew: `clawhdf5-py` (pyproject.toml) and `packages/clawhdf5-node` (package.json) were both behind the actual crate version - -### Recently closed out (2026-08-03 cleanup pass) - -- [x] Removed `clawhdf5-types` β€” it was an empty 1-line stub crate; shared type definitions already live in `clawhdf5-format`, so CLAUDE.md and the workspace manifest were corrected instead of filling it in -- [x] Superblock v4 (page-buffer mode) read/write β€” the only unimplemented task from `docs/superpowers/plans/2026-06-29-format-write-extensions.md`; now done (`Superblock::parse_v4`/`serialize`, `FileWriter::with_page_size`) -- [x] Reconciled the three `docs/superpowers/plans/*.md` docs against actual shipped code β€” they were pre-work plans for `d6c4d4f` (2026-06-30), committed to git late; checkboxes now reflect reality - ---- - -_Last updated: 2026-08-05_ +The old track-by-track tracker this file used to be (agent-memory +Tracks 1–8, mid-2026) is in git history (`git log -- ROADMAP.md`). diff --git a/conformance/README.md b/conformance/README.md index 919d1fe..b4b682c 100644 --- a/conformance/README.md +++ b/conformance/README.md @@ -6,9 +6,36 @@ h5py/libhdf5, compares the two readings object by object, and writes ```sh CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh # ~30 s once the corpus is cached +conformance/run.sh --no-fetch # use the cached corpus as is conformance/run.sh --update-baseline # after an intended change in results ``` +Latest result (tank, 2026-09-27, `conformance/run.sh --no-fetch`): 602 of +697 files ok, 0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read, and +no panic, hang, crash or out-of-memory. The report with every file is +[`CONFORMANCE.md`](../CONFORMANCE.md). + +## Classes + +`compare.py` puts each file in one class: + +| class | meaning | +|---|---| +| **ok** | clawhdf5 and h5py read the same objects with the same values | +| **our-error** | h5py reads something clawhdf5 refuses | +| **mismatch** | both read it, with different values or structure | +| **h5py-cannot-read** | h5py (libhdf5) cannot read the file; not compared | +| **ref-bug** | h5py reads an object clawhdf5 refuses, but only through a libhdf5 over-read: `ref_bugs.py` re-reads it in six processes with different heaps (import order, `MALLOC_PERTURB_`) and its values change. The file is ref-bug only while that is confirmed in the same run; if the values become stable it counts as our-error again | +| **panic / hang / crash / oom** | a clawhdf5 failure under the timeout and address-space limit; the gate fails on any | + +Where h5py itself returns wrong values through a known h5py bug (the +big-endian variable-length bug: elements returned with the file's bytes +under a little-endian dtype), `ref.py` checks that the installed h5py has +the bug, corrects the values before hashing and marks them `ref_fix`, so +those objects are still compared. The evidence for the three current +ref-bug files is under "Conformance: the last non-ok files" in +[`docs/known-issues.md`](../docs/known-issues.md). + Needs Rust, `git`, `h5dump` (Debian/Ubuntu `hdf5-tools`), `libaec` (for the probe's `szip` feature; `libaec-dev`), and a Python with the packages in `requirements.txt`. The first run downloads about 450 MB of sparse checkouts. diff --git a/crates/clawhdf5-accel/README.md b/crates/clawhdf5-accel/README.md index d7c0b2f..26bacfc 100644 --- a/crates/clawhdf5-accel/README.md +++ b/crates/clawhdf5-accel/README.md @@ -1,24 +1,61 @@ # clawhdf5-accel -[![crates.io](https://img.shields.io/crates/v/clawhdf5-accel.svg)](https://crates.io/crates/clawhdf5-accel) -[![docs.rs](https://docs.rs/clawhdf5-accel/badge.svg)](https://docs.rs/clawhdf5-accel) +CPU SIMD kernels for vector search: dot products, cosine similarity, L2 +distance, norms and int8 dot products, dispatched at run time to the best +backend the CPU has, with a portable scalar fallback for every operation. +[`clawhdf5-ann`](../clawhdf5-ann/README.md) and +[`clawhdf5-agent`](../clawhdf5-agent/README.md) use it in their distance +loops; it has nothing to do with HDF5 file I/O. -SIMD-accelerated operations for clawhdf5. +Not on crates.io yet; depend on it from git: + +```toml +[dependencies] +clawhdf5-accel = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` + +## API + +```rust +use clawhdf5_accel::{cosine_similarity, detect_backend, dot_i8, dot_product, l2_distance}; + +let a = [1.0f32, 2.0, 3.0, 4.0]; +let b = [4.0f32, 3.0, 2.0, 1.0]; +assert_eq!(dot_product(&a, &b), 20.0); +let _cos = cosine_similarity(&a, &b); +let _l2 = l2_distance(&a, &b); +assert_eq!(dot_i8(&[1, -2, 3], &[4, 5, -6]), -24); +println!("{:?}", detect_backend()); // e.g. Avx2 on x86-64, Neon on aarch64 +``` + +Also `vector_norm`, `batch_norms`, `batch_cosine`, `batch_cosine_prenorm`, +`f16_to_f32_batch`, `checksum_fletcher32` and `align_to_cache_line`. + +## Backends + +`detect_backend()` picks once per process: `Avx512` (with the `avx512` +feature), `Avx2` (AVX2 + FMA), `Neon` (every aarch64 CPU), or `Scalar`. +`Sse4` and `WasmSimd128` are reported when detected but run the scalar +kernels. +`dot_i8`, used by the agent's quantised (int8) HNSW index, runs on +AVX2 and on NEON β€” with the `SDOT` instruction (through inline assembly, +since the intrinsic is unstable) on cores that have dotprod, such as the +Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8 +index answers 1.63x the queries per second of the f32 one on x86-64 +(AVX2) and 1.18x on a Raspberry Pi 5 ([`BENCHMARKS.md`](../../BENCHMARKS.md)). + +The aarch64 code is compiled out on x86, so only the `test-arm64` CI job +builds and tests it. ## Features -- AVX2 and NEON SIMD acceleration -- AVX-512 support (`avx512` feature) -- Float16 conversion (`float16` feature) -- CRC32 checksum acceleration +| Feature | Default | What | Builds C | +|---|---|---|---| +| `avx512` | no | AVX-512F kernels | no | +| `float16` | no | `f16_to_f32_batch` through the `half` crate (a software conversion otherwise) | no | -## Usage - -```rust -use clawhdf5_accel::checksum::crc32_simd; - -let crc = crc32_simd(&data); -``` +The half-precision conversion used for stored embeddings is +`clawhdf5_format::float16`, not this crate's. ## License diff --git a/crates/clawhdf5-agent/README.md b/crates/clawhdf5-agent/README.md index d0e5925..8eeafe7 100644 --- a/crates/clawhdf5-agent/README.md +++ b/crates/clawhdf5-agent/README.md @@ -1,28 +1,124 @@ # clawhdf5-agent -[![crates.io](https://img.shields.io/crates/v/clawhdf5-agent.svg)](https://crates.io/crates/clawhdf5-agent) -[![docs.rs](https://img.shields.io/docsrs/clawhdf5-agent)](https://docs.rs/clawhdf5-agent) +Persistent memory for AI agents in a single HDF5 file: text chunks with +embeddings and metadata, hybrid search (HNSW vector search + BM25 keyword +search, fused), sessions, a knowledge graph, a write-ahead log for crash +safety, and optionally Ed25519-signed checkpoints. Stores open in h5py like +any other HDF5 file. Built on [`clawhdf5`](../clawhdf5/README.md), +[`clawhdf5-ann`](../clawhdf5-ann/README.md) and +[`clawhdf5-accel`](../clawhdf5-accel/README.md). -HDF5-backed persistent memory store for on-device AI agents. +It is a library: no agent framework integrates it (OpenClaw and ZeroClaw +integration claims were withdrawn on 2026-09-25; see +[`docs/openclaw.md`](../../docs/openclaw.md)). The command-line front end +is [`clawhdf5-cli`](../clawhdf5-cli/README.md). -Built on [clawhdf5](https://crates.io/crates/clawhdf5), clawhdf5-agent provides a vector-searchable memory backend optimized for edge AI workloads. Store embeddings, text chunks, and metadata in a single HDF5 file with SIMD-accelerated similarity search. - -## Features - -- Persistent vector store in HDF5 format -- Cosine similarity and L2 distance search -- SIMD-accelerated via clawhdf5-accel (AVX2, NEON) -- Optional GPU acceleration via clawhdf5-gpu -- Memory-mapped access for large stores -- f16 storage support for compact embeddings - -## Usage +Not on crates.io yet; depend on it from git: ```toml [dependencies] -clawhdf5-agent = "2.1.0" +clawhdf5-agent = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } ``` +## Usage + +```rust,no_run +use std::path::PathBuf; +use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry, SearchOptions}; + +let config = MemoryConfig::new(PathBuf::from("agent.h5"), "my-agent", 384); +let mut mem = HDF5Memory::create(config)?; + +mem.save(MemoryEntry { + chunk: "The deploy key rotates every Monday.".into(), + embedding: vec![0.01; 384], // from your embedding model + source_channel: "chat".into(), + timestamp: 1_790_000_000.0, + session_id: "s1".into(), + tags: "ops".into(), +})?; + +let query = vec![0.01f32; 384]; +let hits = mem.search(&query, "deploy key", &SearchOptions::new(5).with_sources(["chat"])); +for h in &hits { + println!("{:.3} {}", h.score, h.chunk); +} +mem.flush_wal()?; // checkpoint now; otherwise one is made after 500 WAL entries (wal_max_entries) +# Ok::<(), clawhdf5_agent::MemoryError>(()) +``` + +## What is in it + +- **`HDF5Memory`** β€” `create`, `open` (single writer: an exclusive lock on + `.h5.lock`, a second opener gets `MemoryError::Locked`), + `open_read_only` (no lock, never writes). Through the `AgentMemory` + trait: `save`, `save_batch`, `delete`, `compact`, `count`, `snapshot`, + sessions; also `save_or_update`, `delete_batch`, `flush_wal`. +- **Search** β€” `search(query_embedding, text, &SearchOptions)`: optional + source-channel filter applied before ranking, vector + BM25 fusion + (weighted or RRF), Hebbian activation scaling, optional re-ranking + (`reranker::ReRankConfig`) and confidence rejection + (`confidence::ConfidenceConfig`). `hybrid_search` and + `hybrid_search_with` are thin wrappers. The vector stage uses the HNSW + index (`hnsw` feature); its graph is saved to `.h5.ann` at each + checkpoint and reloaded on open (rebuilt if stale or damaged). +- **Storage settings** (`MemoryConfig`, persisted with the store): + `float16` embeddings (on by default for new stores; 48% smaller file at + 100K records, same retrieval on LongMemEval), `quantized_index` (int8 + copy of the vectors in the index, on by default; re-scored against the + exact embeddings), `compression` (off by default), HNSW `m`/`ef` + parameters, WAL settings (`wal_enabled`, on by default; `wal_max_entries`, + 500: the WAL is checkpointed into the `.h5` once it holds more). +- **WAL** (`wal`) β€” every write is appended to `.h5.wal` with a + chained CRC32 per entry, so a corrupted, reordered or spliced entry stops + replay. Recovers from a process crash at any point, including between a + checkpoint and the WAL truncate. WAL appends are not fsynced: saves since + the last checkpoint can be lost on power failure. An unreadable WAL is + quarantined to `.h5.wal.corrupt-`. +- **Signed checkpoints** (`signing`) β€” `set_signing_key` signs a manifest + (SHA-256 Merkle tree over records, plus settings, sessions and graph) at + every checkpoint; `HDF5Memory::verify(path, &public_key)` checks it and + locates edits. WAL entries after the checkpoint are not covered. +- **Knowledge graph** (`knowledge`, `entity_extract`) β€” `add_entity`, + `add_entity_alias`, `add_relation`, `extract_and_store_entities`, + traversal and spreading activation. +- **Also:** sessions (`session`), temporal index (`temporal`), + consolidation tiers (`consolidation`), an in-memory TTL tier + (`ephemeral`), multi-modal embeddings (`multimodal`), `AGENTS.md` + generation (`agents_md`), query expansion, and a session-scoped + provenance ledger and write-anomaly detector on every save + (`take_anomaly_alerts`; alerts never block a save, and the source is + inferred from `source_channel`, not authenticated). +- `openclaw::ClawhdfBackend` is `search` with re-ranking and confidence + on, plus Markdown import/export. The module name is historical: it is not + an OpenClaw plugin. + +## Features + +| Feature | Default | What | Builds C | +|---|---|---|---| +| `hnsw` | yes | HNSW vector index (`clawhdf5-ann`); without it the vector stage is an exact linear cosine scan | no | +| `parallel` | yes | build the HNSW index on a rayon pool (same graph either way) | no | +| `float16` | yes | f16 helpers in `vector_search` (`half`). Stores' `MemoryConfig::float16` works without it. | no | +| `fast-math` | no | `matrixmultiply` batch distances in `strategy` | no | +| `accelerate` | no | Apple Accelerate BLAS in `strategy` (macOS) | links a system framework | +| `openblas` | no | OpenBLAS in `strategy` | yes (`openblas-src`) | +| `gpu` | no | `gpu_search` through [`clawhdf5-gpu`](../clawhdf5-gpu/README.md) (wgpu), used by `strategy`, not by `HDF5Memory::search` | no, but needs GPU drivers | +| `zstd` | no | Zstd instead of deflate when `MemoryConfig::compression` is on | yes (libzstd) | +| `async` | no | `async_memory` wrapper on tokio | no | + +`--no-default-features --features float16` forces the exact linear scan. + +## Measurements and limits + +- Search recall and latency, file size, LongMemEval and MemoryArena + retrieval numbers: [`BENCHMARKS.md`](../../BENCHMARKS.md), measured with + the `clawhdf5-bench` binaries (`search_harness`, `longmemeval_bench`, + `footprint_bench`, ...). +- Known issues and their history: [`docs/known-issues.md`](../../docs/known-issues.md). +- Migrating a SQLite memory database: + [`clawhdf5-migrate`](../clawhdf5-migrate/README.md). + ## License MIT diff --git a/crates/clawhdf5-android/README.md b/crates/clawhdf5-android/README.md new file mode 100644 index 0000000..838a088 --- /dev/null +++ b/crates/clawhdf5-android/README.md @@ -0,0 +1,44 @@ +# clawhdf5-android + +A C ABI over [`clawhdf5-agent`](../clawhdf5-agent/README.md) for Android +apps: a `cdylib` exporting `extern "C"` functions (`edgehdf5_*`, a name +kept from the project's earlier "edgehdf5" days) that manage an +`HDF5Memory` through an opaque handle. + +The functions are plain C symbols, not JNI-mangled `Java_...` entry points: +a Kotlin/Java app calls them through a thin JNI shim or JNA of its own. No +such shim, Gradle project or AAR is in this repository, and the crate is +not built for an Android target in CI (only its host-side unit tests run +with the workspace). + +## Functions + +| Function | What | +|---|---| +| `edgehdf5_create(path, agent_id, embedding_dim)` / `edgehdf5_open(path)` | a handle, or null on failure | +| `edgehdf5_close(handle)` | drop the store; what is not yet checkpointed stays in its WAL, as with any `HDF5Memory` | +| `edgehdf5_save(handle, ...)` | save one entry; the embedding length is checked against the store's dimension before the pointer is read | +| `edgehdf5_delete`, `edgehdf5_count`, `edgehdf5_count_active` | | +| `edgehdf5_hybrid_search(handle, query, len, text, vector_weight, keyword_weight, max_results, out_indices, out_scores, out_chunks)` | results into caller-provided arrays; returns the number written | +| `edgehdf5_add_session`, `edgehdf5_get_session_summary` | sessions | +| `edgehdf5_add_entity`, `edgehdf5_add_relation` | knowledge graph | +| `edgehdf5_free_string` | free a string this library returned | + +Every function is `unsafe`: the caller guarantees valid, NUL-terminated +strings and correctly sized buffers (see each function's `# Safety` +section), and serialises access to a handle; separate handles are +independent. + +## Build + +```bash +cargo build --release -p clawhdf5-android # host build; for a device, add --target aarch64-linux-android with the NDK's linker configured +``` + +It depends on `clawhdf5-agent` with **default features off**, so there is +no HNSW index (the vector stage is an exact linear scan) and no rayon +pool. No C is compiled. + +## License + +MIT diff --git a/crates/clawhdf5-ann/README.md b/crates/clawhdf5-ann/README.md index dd7a0af..233379f 100644 --- a/crates/clawhdf5-ann/README.md +++ b/crates/clawhdf5-ann/README.md @@ -1,25 +1,70 @@ # clawhdf5-ann -[![crates.io](https://img.shields.io/crates/v/clawhdf5-ann.svg)](https://crates.io/crates/clawhdf5-ann) -[![docs.rs](https://docs.rs/clawhdf5-ann/badge.svg)](https://docs.rs/clawhdf5-ann) +An HNSW (Hierarchical Navigable Small World) approximate nearest-neighbour +index in pure Rust, with cosine or L2 distance, optional int8 storage of +the vectors, deletions, and persistence as an HDF5 file. It is the vector +stage of [`clawhdf5-agent`](../clawhdf5-agent/README.md)'s search (the +agent's `hnsw` feature, on by default); distances run on +[`clawhdf5-accel`](../clawhdf5-accel/README.md)'s SIMD kernels. -HNSW approximate nearest neighbor index stored as HDF5. +Neighbours are chosen with the HNSW paper's diversity heuristic, not plain +closest-M (which capped recall on clustered data at 0.31 recall@10 at 100K +vectors). -## Features +Not on crates.io yet; depend on it from git: -- Build and query HNSW indexes persisted in HDF5 format -- Pure Rust, no C dependencies -- Efficient similarity search for high-dimensional vectors +```toml +[dependencies] +clawhdf5-ann = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` ## Usage ```rust -use clawhdf5_ann::HnswIndex; +use clawhdf5_ann::{DistanceMetric, HnswIndex, Storage}; -let index = HnswIndex::from_hdf5("vectors.h5").unwrap(); -let neighbors = index.search(&query, 10); +let vectors: Vec> = (0..500) + .map(|i| (0..16).map(|j| ((i * 31 + j * 7) % 97) as f32 / 97.0).collect()) + .collect(); + +// m = 16 connections per node, ef_construction = 200 +let mut index = HnswIndex::build_with(&vectors, 16, 200, DistanceMetric::Cosine, Storage::Int8); +let hits = index.search(&vectors[42], 10, 64); // (id, distance), closest first; ef >= k +assert!(hits[0].1 < 1e-3); // vector 42 itself (or an identical one) + +let id = index.insert(vec![0.5; 16]); +index.mark_deleted(id); + +// Persist as HDF5 (a self-contained file: graph and vectors) and load it back +let bytes = index.to_hdf5_bytes().unwrap(); +let loaded = HnswIndex::load_from_hdf5(&bytes).unwrap(); +assert_eq!(loaded.len(), index.len()); ``` +- `HnswIndex::build` (L2), `build_with_metric`, `build_with` (metric and + storage); `new`/`new_with` plus `insert` for an index built + incrementally. +- `Storage::Int8` keeps each vector as `i8`, a quarter of the memory; it + applies to `Cosine` only (an L2 index keeps `Float32`). Distances are then + approximate, so a caller that needs exact ranking re-scores the + candidates, as the agent does. +- `mark_deleted`, `is_deleted`, `deleted_count`, `active_len`, `compact` + (returns the old-to-new id map). +- `save_to_hdf5(&mut writer)` / `to_hdf5_bytes` / `load_from_hdf5` store + the whole index; `graph_to_bytes` / `from_graph_bytes` store only the + graph (with a CRC32) for a caller that keeps the vectors elsewhere β€” the + agent's `.h5.ann` sidecar. + +## Features + +| Feature | Default | What | Builds C | +|---|---|---|---| +| `parallel` | no | build the graph on a rayon pool; the graph is identical with or without it | no | + +Recall and speed against exact search, for the index alone and in the +agent: [`BENCHMARKS.md`](../../BENCHMARKS.md), measured with +`cargo run --release -p clawhdf5-bench --bin search_harness`. + ## License MIT diff --git a/crates/clawhdf5-bench/README.md b/crates/clawhdf5-bench/README.md new file mode 100644 index 0000000..afa214f --- /dev/null +++ b/crates/clawhdf5-bench/README.md @@ -0,0 +1,49 @@ +# clawhdf5-bench + +The measurement harnesses behind [`BENCHMARKS.md`](../../BENCHMARKS.md): +HDF5 read and write speed (against libhdf5 and h5py where noted) and the +agent store's search, footprint and retrieval quality. Not meant for +publishing; nothing else in the workspace depends on it. Run everything with +`--release`, and quote numbers with the machine, date and command, as +`BENCHMARKS.md` does. + +## Binaries + +| Binary | Measures | +|---|---| +| `read_harness` | full reads vs hyperslab selections of a chunked 2-D dataset (compressed and not) and a contiguous one: does a selection cost scale with the selection or the dataset? (`-- --large` for 512 MB) | +| `concurrent_read` | decoded read throughput vs threads on one open `File`; `scripts/concurrent_read_h5py.py` runs the same workload with h5py (threads and processes) and `scripts/compare_concurrent_read.py` tabulates both | +| `search_harness` | HNSW recall@10 vs exact search, QPS and latency per `ef`, and end-to-end `HDF5Memory` ingest/checkpoint/open/search at 1K–100K (`--full`); studies: `--float16-study`, `--options-study`, `--signing-study`, `--ann-only --uniform` | +| `longmemeval_bench` | LongMemEval retrieval recall (turn and session Hit@k, MRR) β€” **retrieval, not QA accuracy**. Oracle or full `longmemeval_s` haystack; `--features embeddings` (or `embeddings-cuda`) embeds with MiniLM, otherwise the vector stage is inert and the run is BM25-only | +| `memory_arena` | a deterministic multi-session retrieval benchmark (BM25-only) | +| `footprint_bench` | file size and bytes per record at 100–100K records, float16 or `--f32`, WAL on/off, compressed or not | +| `consolidation_efficiency` | retrieval before and after consolidation on signal + noise records | +| `ephemeral_perf` | the in-memory ephemeral tier's set/get latency | +| `mpi_io_bench` | `clawhdf5-io`'s `MpiVol` (root-read + broadcast, not collective I/O); needs `--features mpi-io` and `mpirun` | + +```bash +cargo run --release -p clawhdf5-bench --bin search_harness -- --full +cargo run --release -p clawhdf5-bench --bin read_harness +``` + +## Criterion benches and example + +- `cargo bench -p clawhdf5-bench` runs `h5bench_write`, `h5bench_read` and + `h5bench_meta` (h5bench-style sequential, chunked, strided and metadata + workloads). `--features libhdf5-compare` adds the same workloads through + libhdf5 (the `hdf5-metno` crate; needs a system libhdf5 1.14). +- `examples/worldmodel_sampling.rs`: shuffled per-frame reads of a + `(N, H, W, C)` `uint8` dataset, clawhdf5 against h5py on the same file. + +## Features + +| Feature | What | Builds C | +|---|---|---| +| `libhdf5-compare` | libhdf5 variants of the Criterion benches | links the system libhdf5 | +| `mpi-io` | `mpi_io_bench` | yes (`mpi-sys`; needs an MPI installation) | +| `embeddings` | MiniLM embeddings for `longmemeval_bench` (candle) | yes (a `cc` build dependency in the candle/tokenizers tree) | +| `embeddings-cuda` | the same on a CUDA GPU (minutes instead of hours on the full haystack) | yes (CUDA) | + +## License + +MIT diff --git a/crates/clawhdf5-cli/README.md b/crates/clawhdf5-cli/README.md new file mode 100644 index 0000000..4495e39 --- /dev/null +++ b/crates/clawhdf5-cli/README.md @@ -0,0 +1,49 @@ +# clawhdf5-cli + +The `clawhdf5` command: create, fill, search and inspect a +[`clawhdf5-agent`](../clawhdf5-agent/README.md) memory store from the +shell. Output is JSON. (For general HDF5 files use `h5rs` from +[`clawhdf5-tools`](../clawhdf5-tools/README.md).) + +```bash +cargo install --path crates/clawhdf5-cli # installs `clawhdf5`; not on crates.io yet +# or: cargo run -p clawhdf5-cli -- --help +``` + +No C is compiled. + +## Commands + +The store is `--path FILE` (or `CLAWHDF5_PATH`) before the subcommand. + +| Command | What | +|---|---| +| `create [--agent-id ID] [--dim N] [--wal] [--f32] [--f32-index]` | a new store (dimension 384 by default); float16 embeddings and an int8 index copy unless `--f32` / `--f32-index`. The WAL is off unless `--wal` (the library's default is on), so each save is checkpointed at once | +| `save [--json '{...}']` | save one entry, from `--json` or stdin: `{"chunk", "embedding", "source_channel", "timestamp", "session_id", "tags"}` | +| `search --embedding '[...]' [--query TEXT] [-k N] [--vector-weight W] [--keyword-weight W]` | hybrid search (defaults 5 results, weights 0.7 / 0.3) | +| `recall INDEX` | one entry by index | +| `stats` | counts and configuration | +| `flush-wal` | checkpoint the WAL into the `.h5` | +| `agents-md [--output FILE]` | generate an `AGENTS.md` from the store | +| `export` | every entry as JSON lines | +| `snapshot DEST` | a copy of the store's `.h5` file | +| `keygen --out FILE` | a new Ed25519 signing key (64 hex characters, created owner-only on Unix) | +| `verify --public-key HEX_OR_FILE` | check a signed store; exit status 2 if it does not verify | + +`recall`, `stats`, `agents-md` and `export` open the store read-only +(no lock, nothing written), so they work while another process has it +open. `save`, `search` (which records activation boosts) and `flush-wal` +open it for writing and take the store's lock. With +`--signing-key FILE` (or `CLAWHDF5_SIGNING_KEY`) every checkpoint a command +makes is signed; a signed store refuses to checkpoint without the key. + +```bash +clawhdf5 --path mem.h5 create --agent-id demo --dim 3 +echo '{"chunk":"hello","embedding":[0.1,0.2,0.3],"source_channel":"cli","timestamp":0,"session_id":"s1","tags":""}' \ + | clawhdf5 --path mem.h5 save +clawhdf5 --path mem.h5 search --embedding '[0.1,0.2,0.3]' --query hello -k 3 +``` + +## License + +MIT diff --git a/crates/clawhdf5-derive/README.md b/crates/clawhdf5-derive/README.md index 8c0160c..87ed241 100644 --- a/crates/clawhdf5-derive/README.md +++ b/crates/clawhdf5-derive/README.md @@ -1,28 +1,50 @@ # clawhdf5-derive -[![crates.io](https://img.shields.io/crates/v/clawhdf5-derive.svg)](https://crates.io/crates/clawhdf5-derive) -[![docs.rs](https://docs.rs/clawhdf5-derive/badge.svg)](https://docs.rs/clawhdf5-derive) +`#[derive(H5Type)]`: maps a Rust struct with named fields to an HDF5 +compound datatype. The derive generates three inherent methods: -Derive macros for clawhdf5 HDF5 traits. +- `hdf5_datatype() -> clawhdf5_format::datatype::Datatype` β€” the + `Datatype::Compound` (members in field order, packed, little-endian); +- `to_bytes(&self) -> Vec` β€” one element in that layout; +- `from_bytes(&[u8]) -> Self` β€” the reverse (panics if the slice is shorter + than the compound). -## Features +Supported field types: `f32`, `f64`, `i8`–`i64`, `u8`–`u64`, `bool` +(stored as `u8`) and fixed-size arrays `[T; N]` of those numeric types. +Tuple structs, enums and nested structs are refused at compile time. -- `#[derive(HDF5Type)]` for automatic HDF5 datatype mapping -- Struct-to-compound-type derivation +The generated code names `clawhdf5_format`, so the crate using the derive +must depend on [`clawhdf5-format`](../clawhdf5-format/README.md) too. Not +on crates.io yet: -## Usage +```toml +[dependencies] +clawhdf5-derive = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` + +## Example ```rust -use clawhdf5_derive::HDF5Type; +use clawhdf5_derive::H5Type; +use clawhdf5_format::datatype::Datatype; -#[derive(HDF5Type)] +#[derive(H5Type, Debug, PartialEq)] struct Point { - x: f64, - y: f64, - z: f64, + id: u32, + pos: [f64; 3], + valid: bool, } + +let p = Point { id: 7, pos: [1.0, 2.0, 3.0], valid: true }; +let bytes = p.to_bytes(); +assert_eq!(bytes.len(), 4 + 24 + 1); +assert_eq!(Point::from_bytes(&bytes), p); +assert!(matches!(Point::hdf5_datatype(), Datatype::Compound { size: 29, .. })); ``` +Tests: `crates/clawhdf5-format/tests/derive_tests.rs`. + ## License MIT diff --git a/crates/clawhdf5-filters/README.md b/crates/clawhdf5-filters/README.md index 45fcf30..78421e5 100644 --- a/crates/clawhdf5-filters/README.md +++ b/crates/clawhdf5-filters/README.md @@ -1,27 +1,56 @@ # clawhdf5-filters -[![crates.io](https://img.shields.io/crates/v/clawhdf5-filters.svg)](https://crates.io/crates/clawhdf5-filters) -[![docs.rs](https://docs.rs/clawhdf5-filters/badge.svg)](https://docs.rs/clawhdf5-filters) +Standalone deflate (zlib) compression and decompression with a choice of +backend: pure-Rust zlib-rs (default), zlib-ng, Apple's Compression +framework, or miniz_oxide. -Filter and compression pipeline for clawhdf5. +This crate holds **deflate backends only**. The HDF5 filter pipeline, the +filter registry and every other codec (shuffle, Fletcher-32, N-Bit, +scale-offset, LZ4, Zstd, SZIP, pcodec, LZF, bitshuffle, bzip2, Blosc, +Blosc2, ZFP) live in [`clawhdf5-format`](../clawhdf5-format/README.md), +which calls flate2 itself and selects its deflate backend with its own +features. No library crate of the workspace depends on this one (the +`clawhdf5` facade uses it only in tests). -## Features +Not on crates.io yet; depend on it from git: -- DEFLATE compression/decompression -- Pure-Rust deflate via zlib-rs (default, `zlib-rs` feature) -- zlib-ng instead, if you want it (`fast-deflate` feature; C, needs cmake) -- Apple Compression framework support (`apple-compression` feature) +```toml +[dependencies] +clawhdf5-filters = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` -## Usage +## API ```rust -use clawhdf5_filters::{deflate_compress, deflate_decompress}; +use clawhdf5_filters::{deflate_backend, deflate_compress, deflate_decompress}; +let data: Vec = (0..10_000u32).map(|i| (i % 251) as u8).collect(); let compressed = deflate_compress(&data, 6).unwrap(); // The second argument bounds the output: the expected decompressed size. let decompressed = deflate_decompress(&compressed, data.len()).unwrap(); +assert_eq!(decompressed, data); +println!("backend: {}", deflate_backend()); // "zlib-rs" by default ``` +Also `deflate_compress_miniz`/`deflate_decompress_miniz` (always +miniz_oxide) and `fast_deflate::{compress, decompress, active_backend}`. + +## Features + +Backend priority: `apple-compression` (macOS only) > zlib-ng > zlib-rs > +miniz_oxide (with none enabled). + +| Feature | Default | Backend | Builds C | +|---|---|---|---| +| `zlib-rs` | yes | zlib-rs through flate2, with `runtime_detection` (needed for its SIMD) | no | +| `fast-deflate` | no | zlib-ng through flate2 | yes (cmake) | +| `system-zlib` | no | the system zlib through flate2 | yes (`libz-sys`) | +| `apple-compression` | no | Apple Compression framework, macOS only (ignored elsewhere) | no (links a system framework) | + +zlib-rs matches zlib-ng on HDF5 reads and writes and produces +byte-identical output: see "Deflate backend" in +[`BENCHMARKS.md`](../../BENCHMARKS.md). + ## License MIT diff --git a/crates/clawhdf5-format/README.md b/crates/clawhdf5-format/README.md index eb3a056..3910da6 100644 --- a/crates/clawhdf5-format/README.md +++ b/crates/clawhdf5-format/README.md @@ -1,27 +1,106 @@ # clawhdf5-format -[![crates.io](https://img.shields.io/crates/v/clawhdf5-format.svg)](https://crates.io/crates/clawhdf5-format) -[![docs.rs](https://docs.rs/clawhdf5-format/badge.svg)](https://docs.rs/clawhdf5-format) +The HDF5 file format in pure Rust: parsers and writers for every on-disk +structure, the filter pipeline and its codecs, and the shared type +definitions the other crates use. Most users want the +[`clawhdf5`](../clawhdf5/README.md) facade, which wraps this crate in an +h5py-like API; use this one directly for low-level access or in `no_std` +code. -Pure-Rust HDF5 binary format parsing and writing β€” no C dependencies. +Not on crates.io yet; depend on it from git: + +```toml +[dependencies] +clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` + +## What is in it + +- **Parsing:** superblock v0–v3 (`superblock`, with the superblock + extension and metadata cache images, `superblock_ext`), object headers v1 + and v2 (`object_header`), every header message the readers use + (`datatype`, `dataspace`, `data_layout` v1–v4 including virtual datasets, + `fill_value`, `attribute`, `link_message`, `shared_message`, ...), groups + old and new (`group_v1` symbol tables with local heaps, `group_v2` with + fractal heaps and v2 B-trees), and every chunk index (v1 B-tree, single + chunk, implicit, fixed array, extensible array, v2 B-tree). +- **Reading data:** `data_read` (contiguous, compact, chunked), + `partial_read` and `selection` (hyperslabs and points), `vl_data` + (variable-length strings and sequences through the global heap), + `chunk_cache`. +- **Storage:** the `storage::Storage` trait (`read_at`, `read_ranges`, + `len`, `hint`) that every read path goes through, so a file can be read + from memory, a file handle or a remote backend + ([`clawhdf5-remote`](../clawhdf5-remote/README.md)). +- **Writing:** `file_writer::FileWriter` and the builders in + `type_builders` (datasets, groups, attributes, compound and enum types, + links, virtual datasets, creation-order tracking); chunk indexes and + dense-storage B-trees of any size (`chunked_write`, `btree_v2_write`, + `ea_writer`). Output is read by h5py and h5dump. +- **Filters:** `filter_pipeline` and `filter_registry` (look up by ID; other + IDs can be registered at run time with `register_filter`). Built in: + deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4, + Zstd, SZIP (decode), pcodec, and the plugin filters LZF, bitshuffle, + bzip2, Blosc 1 (read and write), Blosc2 and ZFP (read only). +- **Shared pieces:** `float16` (the one IEEE half-precision conversion the + workspace uses), `provenance` (SHA-256 dataset hashes), `checksum` + (Jenkins lookup3 for v2+ structures). + +## Example + +```rust +use clawhdf5_format::file_writer::{AttrValue, FileWriter}; +use clawhdf5_format::{group_v2, object_header, signature, superblock}; + +// Write a file to memory +let mut fw = FileWriter::new(); +fw.create_dataset("data") + .with_f64_data(&[1.0, 2.0, 3.0]) + .with_shape(&[3]) + .set_attr("unit", AttrValue::String("m/s".into())); +let bytes = fw.finish().unwrap(); + +// Parse it back: superblock -> path -> object header +let (_user_block, file) = signature::split_user_block(&bytes).unwrap(); +let sb = superblock::Superblock::parse(file, 0).unwrap(); +let addr = group_v2::resolve_path_any(file, &sb, "data").unwrap(); +let hdr = object_header::ObjectHeader::parse(file, addr as usize, sb.offset_size, sb.length_size) + .unwrap(); +assert!(!hdr.messages.is_empty()); +``` ## Features -- Zero-copy superblock, object header, and B-tree parsing -- Chunked dataset read/write with filter pipelines -- `no_std` support (disable `std` feature) -- Optional parallel reads via Rayon -- SHA-256 provenance tracking +| Feature | Default | What | Builds C | +|---|---|---|---| +| `std` | yes | standard library; without it the crate is `no_std` + `alloc` (CI builds it for `thumbv7em-none-eabihf`) | no | +| `checksum` | yes | verify Jenkins lookup3 checksums | no | +| `deflate` | yes | deflate through flate2 | no | +| `zlib-rs` | yes | flate2's pure-Rust zlib-rs backend, with `runtime_detection` (without it zlib-rs loses SIMD and inflates 3.5x slower) | no | +| `system-zlib-decompress` | yes | no effect (nothing reads it; kept so existing feature lists still build) | no | +| `provenance` | yes | SHA-256 provenance hashes | no | +| `lzf` | yes | LZF (32000) | no | +| `parallel` | no | rayon-parallel chunk decoding | no | +| `fast-checksum` | no | hardware CRC32 through `crc32fast` | no | +| `lz4` | no | LZ4 (32004) | no | +| `pcodec` | no | pcodec | no | +| `bitshuffle`, `bzip2`, `blosc` | no | 32008, 307, 32001, read and write | no | +| `blosc2`, `zfp` | no | 32026, 32013, read only | no | +| `plugin-filters` | no | all six plugin filters above | no | +| `lookup-stats` | no | counters for name-lookup benchmarks | no | +| `zstd` | no | Zstandard (32015) | yes (libzstd) | +| `szip` | no | SZIP (4) decoding | links the system libaec (`libaec-dev`) | +| `fast-deflate` | no | zlib-ng | yes (cmake) | +| `system-zlib` | no | the system zlib | yes (`libz-sys`) | +| `blake3_hash` | no | `provenance::blake3_hash` | yes (`cc`) | -## Usage +## Robustness -```rust -use clawhdf5_format::Superblock; - -let data = std::fs::read("data.h5").unwrap(); -let sb = Superblock::from_bytes(&data).unwrap(); -println!("HDF5 version {}.{}", sb.version_major(), sb.version_minor()); -``` +Every parser is meant to return an error, never panic, on hostile input: +nine cargo-fuzz targets live in [`fuzz/`](fuzz/README.md), the conformance +sweep includes the HDF Group's CVE corpus +([`CONFORMANCE.md`](../../CONFORMANCE.md)), and header checks follow +libhdf5's. Open gaps are in [`docs/known-issues.md`](../../docs/known-issues.md). ## License diff --git a/crates/clawhdf5-format/fuzz/README.md b/crates/clawhdf5-format/fuzz/README.md index 1343bf6..c80e82d 100644 --- a/crates/clawhdf5-format/fuzz/README.md +++ b/crates/clawhdf5-format/fuzz/README.md @@ -51,10 +51,18 @@ done ## CI -These targets are **not** run in CI (`.gitea/workflows/ci.yml`) β€” cargo-fuzz -requires nightly and each meaningful run takes minutes, which doesn't fit a -per-PR gate. Run them manually on a schedule (e.g. before a release, or after -touching parser code) instead. +These targets are **not** run by the CI workflows (`.gitea/workflows/ci.yml`) +β€” cargo-fuzz requires nightly and each meaningful run takes minutes, which +doesn't fit a per-PR gate. Run them by hand before a release or after +touching parser code. `scripts/ci-test.sh` has an opt-in smoke run: with +`CLAWHDF5_FUZZ_SECONDS=N` it runs every target of this crate and of +`crates/clawhdf5-agent/fuzz` (the WAL parser) for N seconds each. + +Other robustness checks that do run: the nightly conformance sweep reads +the HDF Group's CVE reproducers and fails on any panic, hang, crash or +out-of-memory ([`conformance/README.md`](../../../conformance/README.md)), +and `scripts/h5rs-fuzz.sh` runs every `h5rs` subcommand over them, optionally +on byte-flipped copies. ## Reproducing Crashes diff --git a/crates/clawhdf5-gpu/README.md b/crates/clawhdf5-gpu/README.md index e758e1a..cc6710b 100644 --- a/crates/clawhdf5-gpu/README.md +++ b/crates/clawhdf5-gpu/README.md @@ -1,25 +1,62 @@ # clawhdf5-gpu -[![crates.io](https://img.shields.io/crates/v/clawhdf5-gpu.svg)](https://crates.io/crates/clawhdf5-gpu) -[![docs.rs](https://docs.rs/clawhdf5-gpu/badge.svg)](https://docs.rs/clawhdf5-gpu) +GPU vector distance computation through [wgpu](https://wgpu.rs) and +hand-written WGSL compute shaders: upload a set of vectors once, then run +cosine or L2 top-k searches, dot products, distance matrices and norms +against them on Vulkan, Metal, DirectX 12 or OpenGL. -GPU-accelerated vector operations for clawhdf5 using wgpu compute shaders. +This crate does **not** read or write HDF5: dataset I/O in clawhdf5 is +CPU-only. It is a vector-search accelerator used optionally by +[`clawhdf5-agent`](../clawhdf5-agent/README.md) (its `gpu` feature exposes +`gpu_search::GpuSearchBackend` and a GPU arm of `strategy::search_with_metrics`; +`HDF5Memory::search` itself uses the HNSW index on the CPU). -## Features +Not on crates.io yet; depend on it from git: -- GPU-accelerated distance computations (L2, cosine) -- wgpu-based compute shaders for cross-platform GPU support -- Float16 support via `half` crate +```toml +[dependencies] +clawhdf5-gpu = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` ## Usage -```rust +```rust,no_run use clawhdf5_gpu::GpuAccelerator; -let accel = GpuAccelerator::new().unwrap(); -let distances = accel.l2_distances(&query, &vectors).unwrap(); +// Fall back to a CPU path when there is no usable GPU. +let mut gpu = match GpuAccelerator::new() { + Ok(g) => g, + Err(_) => return, +}; + +let dim = 128; +let vectors = vec![0.5f32; 1000 * dim]; // 1000 vectors, row-major +gpu.upload_vectors(&vectors, dim).unwrap(); +let norms = gpu.compute_norms_gpu(&vectors, dim).unwrap(); +gpu.upload_norms(&norms).unwrap(); + +let query = vec![1.0f32; dim]; +let top10 = gpu.cosine_search(&query, 10).unwrap(); // (index, similarity), best first +let near10 = gpu.l2_search(&query, 10).unwrap(); // (index, distance), nearest first ``` +`GpuAccelerator` also has `is_available`, `device_info`, +`batch_cosine_search`, `batch_dot_product`, `distance_matrix`, +`compute_norms`, and `f16_to_f32_batch`/`f32_to_f16_batch`. Vector sets +larger than the device's largest storage buffer binding are split into +chunks and the results merged. A GPUβ†’CPU readback waits at most 30 s and +then fails with `GpuError::BufferMap` instead of hanging. + +## Features + +| Feature | Default | What | +|---|---|---| +| `gpu-wgpu` | yes | the wgpu implementation. Without it `GpuAccelerator::new()` returns `GpuError::NotCompiled` and `is_available()` is `false`. | + +No C is compiled, but wgpu talks to the system's graphics drivers at run +time; the crate is exempt from CI's "no C in the default build" check for +that reason. + ## License MIT diff --git a/crates/clawhdf5-io/README.md b/crates/clawhdf5-io/README.md index fb34684..a372453 100644 --- a/crates/clawhdf5-io/README.md +++ b/crates/clawhdf5-io/README.md @@ -1,24 +1,54 @@ # clawhdf5-io -[![crates.io](https://img.shields.io/crates/v/clawhdf5-io.svg)](https://crates.io/crates/clawhdf5-io) -[![docs.rs](https://docs.rs/clawhdf5-io/badge.svg)](https://docs.rs/clawhdf5-io) +I/O building blocks under [`clawhdf5`](../clawhdf5/README.md): the +`HDF5Read`/`HDF5ReadWrite` traits with in-memory, borrowed, file and +memory-mapped readers, plus several experimental modules (async reads, an +HSDS client, a VOL-style trait, sub-filing, prefetch, and an MPI connector). +The facade uses it for memory-mapped reads (`MmapReader`, and the private +copy-on-write mapping that applies a metadata cache image). -I/O abstraction layer for clawhdf5. +Remote files are **not** read through this crate: HTTP(S) and object +stores go through `clawhdf5_format::storage::Storage` and +[`clawhdf5-remote`](../clawhdf5-remote/README.md). + +Not on crates.io yet; depend on it from git: + +```toml +[dependencies] +clawhdf5-io = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5", features = ["mmap"] } +``` + +## Main items + +| Item | What | +|---|---| +| `HDF5Read`, `HDF5ReadWrite` | byte-level read/write traits; `MemoryReader`, `BorrowedReader`, `FileReader`, `FileWriter` implement them | +| `MmapReader`, `MmapReadWrite` (`mmap`) | memory-mapped files through `memmap2`; `HDF5Read::private_copy` gives a copy-on-write view | +| `prefetch::PrefetchReader`, `sweep::SweepDetector` | read-ahead (`madvise(MADV_WILLNEED)` on mappings) and chunk-sweep prediction | +| `ParallelConfig` | lane partitioning for parallel chunk decoding | +| `vol::VirtualObjectLayer`, `vol::NativeVol` | a backend-agnostic object-layer trait (modelled on libhdf5's VOL) | +| `async_read` (`async`) | tokio-based `AsyncHDF5Read` and `AsyncHDF5File` | +| `hsds::HsdsClient` (`hsds`) | a REST client for an HSDS server | +| `subfiling` | splitting one logical file across several physical files | +| `mpi_vol::MpiVol` (`mpi-io`) | an MPI connector: see below | + +### MPI (`mpi-io`) + +`MpiVol` is **not collective MPI-IO**. Reads are root-read + broadcast +(rank 0 reads the file with `std::fs::read`, parses the dataset and +broadcasts the bytes); writes gather every rank's shard to rank 0, which +writes the merged dataset. It does not call `MPI_File_read_at_all` or any +other MPI-IO routine. Collective I/O is on the [roadmap](../../ROADMAP.md). +`clawhdf5-bench`'s `mpi_io_bench` binary exercises it. ## Features -- Memory-mapped file access (`mmap` feature) -- Async I/O via Tokio (`async` feature) -- HSDS remote access (`hsds` feature) -- Prefetching and sweep optimizations - -## Usage - -```rust -use clawhdf5_io::MmapReader; - -let reader = MmapReader::open("data.h5").unwrap(); -``` +| Feature | Default | What | Builds C | +|---|---|---|---| +| `mmap` | no (the `clawhdf5` facade turns it on) | `MmapReader`, `MmapReadWrite` | no | +| `async` | no | `async_read` (tokio) | no | +| `hsds` | no | `hsds` (reqwest, and `async`) | yes: reqwest's default TLS is native-tls (OpenSSL on Linux) | +| `mpi-io` | no | a real `MpiVol` (without it `MpiVol::new_world` returns an error) | yes: `mpi-sys` needs an MPI installation and libclang | ## License diff --git a/crates/clawhdf5-migrate/README.md b/crates/clawhdf5-migrate/README.md index 092bc94..9e00f38 100644 --- a/crates/clawhdf5-migrate/README.md +++ b/crates/clawhdf5-migrate/README.md @@ -1,11 +1,8 @@ # clawhdf5-migrate -[![crates.io](https://img.shields.io/crates/v/clawhdf5-migrate.svg)](https://crates.io/crates/clawhdf5-migrate) -[![docs.rs](https://img.shields.io/docsrs/clawhdf5-migrate)](https://docs.rs/clawhdf5-migrate) - CLI tool to migrate a SQLite agent-memory database in the `memory_chunks` / `sessions` / `entities` / `relations` layout (table and column names are configurable) to a -[clawhdf5-agent](https://crates.io/crates/clawhdf5-agent) store. This is **not** +[clawhdf5-agent](../clawhdf5-agent/README.md) store. This is **not** ZeroClaw's schema β€” ZeroClaw keeps memories in a single `memories` table and does not use clawhdf5. @@ -15,10 +12,15 @@ the knowledge graph (entities and relations) are carried over. ## Installation +Not on crates.io yet; install from a checkout: + ```bash -cargo install clawhdf5-migrate +cargo install --path crates/clawhdf5-migrate ``` +It builds C: `rusqlite` is built with its `bundled` feature, which compiles +SQLite (so no system libsqlite is needed, but a C compiler is). + ## Usage ```bash diff --git a/crates/clawhdf5-napi/README.md b/crates/clawhdf5-napi/README.md new file mode 100644 index 0000000..f102069 --- /dev/null +++ b/crates/clawhdf5-napi/README.md @@ -0,0 +1,38 @@ +# clawhdf5-napi + +> **Status: does not work end to end.** The TypeScript package built on +> this crate (`packages/clawhdf5-node`) has never run successfully, is +> unpublished, and is not built or tested in CI. See "The Node.js package +> does not work" in [`docs/known-issues.md`](../../docs/known-issues.md). +> Fix it and add CI, or remove it, before depending on it. + +A Node.js native addon ([napi-rs](https://napi.rs), N-API 9) exposing +[`clawhdf5-agent`](../clawhdf5-agent/README.md) as a `ClawhdfMemory` +class. It wraps `clawhdf5_agent::openclaw::ClawhdfBackend` (the agent's +`search` with re-ranking and confidence on) and the consolidation engine. +It was written for an OpenClaw integration that is not being pursued +([`docs/openclaw.md`](../../docs/openclaw.md)). + +## What the addon exposes + +`ClawhdfMemory.create(path, dim)`, `.open(path)`, `.openOrCreate(path, +dim)`, and on an instance: `search`, `get`, `write`, `ingestMarkdown`, +`exportMarkdown`, `save`, `saveBatch`, `stats`, `compact`, `tickSession`, +`flushWal`, `walPendingCount`, `runConsolidation`, and the ephemeral tier +(`enableEphemeral`, `ephemeralSet`/`Get`/`Delete`, `ephemeralStats`, +`promoteEphemeral`). napi-rs converts names and `#[napi(object)]` fields to +camelCase. + +## Build + +```bash +cargo build --release -p clawhdf5-napi # the Rust cdylib +# the .node package: npm install -g @napi-rs/cli; cd packages/clawhdf5-node; napi build --platform --release +``` + +It links against Node's N-API through `napi-sys` (a `-sys` crate), so it is +exempt from CI's "no C in the default build" check. + +## License + +MIT diff --git a/crates/clawhdf5-netcdf4/README.md b/crates/clawhdf5-netcdf4/README.md index ddc3452..9e82b97 100644 --- a/crates/clawhdf5-netcdf4/README.md +++ b/crates/clawhdf5-netcdf4/README.md @@ -1,25 +1,53 @@ # clawhdf5-netcdf4 -[![crates.io](https://img.shields.io/crates/v/clawhdf5-netcdf4.svg)](https://crates.io/crates/clawhdf5-netcdf4) -[![docs.rs](https://docs.rs/clawhdf5-netcdf4/badge.svg)](https://docs.rs/clawhdf5-netcdf4) +Read NetCDF-4 files in pure Rust. NetCDF-4 files are HDF5 files with +conventions for dimensions, coordinate variables and attributes; this crate +reads them through the [`clawhdf5`](../clawhdf5/README.md) facade, with no +libnetcdf or libhdf5. Read-only: NetCDF-3 (classic) files are not HDF5 and +are not supported. -NetCDF-4 read support built on clawhdf5 β€” pure Rust, no C dependencies. +Not on crates.io yet; depend on it from git: -## Features - -- Read NetCDF-4 / HDF5-backed `.nc` files -- Dimension, variable, and CF convention support -- Climate and scientific data access +```toml +[dependencies] +clawhdf5-netcdf4 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` ## Usage -```rust +```rust,no_run use clawhdf5_netcdf4::NetCDF4File; -let nc = NetCDF4File::open("climate.nc").unwrap(); -let temp = nc.variable("temperature").unwrap(); +let nc = NetCDF4File::open("climate.nc")?; +for dim in nc.dimensions()? { + println!("{}: {} (unlimited = {})", dim.name, dim.size, dim.is_unlimited); +} +let mut temp = nc.variable("temperature")?; +let dims: Vec<&str> = temp.dimensions().iter().map(|d| d.name.as_str()).collect(); +println!("{:?} over {:?}", temp.shape()?, dims); +let cf = temp.cf_attributes()?; +println!("units: {:?}", cf.units); +// scale_factor/add_offset applied; _FillValue and missing_value become NaN +let values: Vec = temp.read_f64()?; +# Ok::<(), clawhdf5_netcdf4::Error>(()) ``` +## API + +| Item | What | +|---|---| +| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | +| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) | +| `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` | +| `Dimension` | `name`, `size`, `is_unlimited` | +| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` | +| `NcType` | the NetCDF type of a variable | + +No cargo features. Tests compare against files written by netCDF4-python +(`tests/interop_tests.rs`; the CI job requires them with +`CLAWHDF5_REQUIRE_INTEROP=1`). What the HDF5 reader underneath cannot +read is listed in [`docs/known-issues.md`](../../docs/known-issues.md). + ## License MIT diff --git a/crates/clawhdf5-py/README.md b/crates/clawhdf5-py/README.md index 06eebec..73f5053 100644 --- a/crates/clawhdf5-py/README.md +++ b/crates/clawhdf5-py/README.md @@ -1,8 +1,5 @@ # clawhdf5-py -[![crates.io](https://img.shields.io/crates/v/clawhdf5-py.svg)](https://crates.io/crates/clawhdf5-py) -[![docs.rs](https://docs.rs/clawhdf5-py/badge.svg)](https://docs.rs/clawhdf5-py) - Python bindings for clawhdf5 β€” a pure-Rust HDF5 library. The package is `clawhdf5` (`import clawhdf5`); it needs numpy and no libhdf5. @@ -61,6 +58,8 @@ with clawhdf5.File("data.h5", "r") as f: `clawhdf5.InternalError`, a `RuntimeError`. - Attributes return what h5py returns; `clawhdf5.Empty` stands for a null dataspace (h5py's `Empty`). +- Also as in h5py: `File.mode` (`'r'`, or `'r+'` for a writable file), `File.flush()` (a no-op: + edits are already synced), `Dataset.chunks`. ## Remote files diff --git a/crates/clawhdf5-remote/README.md b/crates/clawhdf5-remote/README.md index cd663b4..876f2a2 100644 --- a/crates/clawhdf5-remote/README.md +++ b/crates/clawhdf5-remote/README.md @@ -114,3 +114,34 @@ let s = storage.stats(); // requests, bytes_fetched, hits, misses, cached_bytes, directory with range support (the server the tests use), and `cargo run -p clawhdf5-remote --example read_url -- URL [DATASET]` lists a file and prints what it cost. + +## Other front ends + +- `h5rs` (built with `--features remote`, or `remote-https`) takes URLs as + FILE arguments: [`clawhdf5-tools`](../clawhdf5-tools/README.md). +- Python: `clawhdf5.File("http://…")` and `File.open_url(url, ...)` go + through this crate: [`clawhdf5-py`](../clawhdf5-py/README.md). +- The browser does **not** use this crate (its cache fetches by blocking); + `clawhdf5-wasm`'s `openUrl` has its own restartable cache: + [`examples/wasm-viewer`](../../examples/wasm-viewer/README.md). + +## Limits + +Files a SWMR writer is still appending to cannot be followed remotely +(the file is pinned at open, so growth is `RemoteError::FileChanged`); the +block size is fixed rather than taken from a paged file's page size; the +cloud backends are built and unit-tested but have not been run against a +real bucket. The full list is under "Remote files (`clawhdf5-remote`) +limits" in [`docs/known-issues.md`](../../docs/known-issues.md); the design +is milestone M3 of [`docs/design/range-reads.md`](../../docs/design/range-reads.md). + +Not on crates.io yet; depend on it from git: + +```toml +[dependencies] +clawhdf5-remote = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` + +## License + +MIT diff --git a/crates/clawhdf5-tools/README.md b/crates/clawhdf5-tools/README.md index 24ecb9c..5c1cd99 100644 --- a/crates/clawhdf5-tools/README.md +++ b/crates/clawhdf5-tools/README.md @@ -306,4 +306,15 @@ CLAWHDF5_PYTHON=.venv/bin/python CLAWHDF5_REQUIRE_INTEROP=1 cargo test -p clawhd The interop tests write their files with h5py and compare with h5ls, h5stat, h5dump and h5diff; each skips when what it needs is missing unless -`CLAWHDF5_REQUIRE_INTEROP=1`. +`CLAWHDF5_REQUIRE_INTEROP=1`. `tests/remote.rs` runs every subcommand on +URLs against a local range server. + +This crate also holds the interop tests of the library's in-place editor +(`clawhdf5::FileEditor`), since they use `h5rs check` and h5dump on every +edited file and compare index and heap structures with what libhdf5 makes +of the same edits: + +```bash +CLAWHDF5_PYTHON=.venv/bin/python CLAWHDF5_REQUIRE_INTEROP=1 \ + cargo test -p clawhdf5-tools --test edit_interop --test edit_coverage_interop +``` diff --git a/crates/clawhdf5-wasm/README.md b/crates/clawhdf5-wasm/README.md new file mode 100644 index 0000000..80c27a2 --- /dev/null +++ b/crates/clawhdf5-wasm/README.md @@ -0,0 +1,51 @@ +# clawhdf5-wasm + +clawhdf5's HDF5 and NetCDF-4 reader compiled to WebAssembly with +wasm-bindgen, for the browser (and Node). Read-only. Two ways in: + +- `open(bytes)` β€” a file already in memory (a dropped file, a fetched + blob); +- `openUrl(url, opts)` β€” a file on a web server, read by HTTP range + requests as each call needs its bytes, without downloading it + (range-read milestone M4, [`docs/design/range-reads.md`](../../docs/design/range-reads.md)). + +Both give `list`, `info`, `attrs`, `read` and `readHyperslab`; the remote +file's methods return promises, and `stats()` counts requests and bytes. +The JavaScript API, options, limits, package size and tests are documented +with the demo page, [`examples/wasm-viewer/README.md`](../../examples/wasm-viewer/README.md). + +## Layout + +- `src/core.rs` β€” the reader over any `clawhdf5_format::storage::Storage` + (`Reader::open_storage`), plain Rust and tested natively. +- `src/lazy.rs` β€” the restartable "NeedBytes" cache behind `openUrl`: a + call runs as a pass over the blocks fetched so far; a pass that misses is + abandoned, the missing (and hinted) blocks are fetched, and the pass is + run again. No block is evicted while a call runs. +- `js/remote.js` β€” the HTTP side: `fetch` with `Range`, checking every + answer (a `206` with exactly the bytes asked for, same ETag/Last-Modified + and length) so a call fails rather than return another file's bytes. +- `src/lib.rs` β€” the wasm-bindgen exports. + +## Build and test + +```bash +rustup target add wasm32-unknown-unknown +cargo install wasm-bindgen-cli --version 0.2.129 # must equal the crate's wasm-bindgen +bash examples/wasm-viewer/build.sh # -> examples/wasm-viewer/pkg/ +cargo test -p clawhdf5-wasm # native: h5py_interop, lazy, vl_strings +bash examples/wasm-viewer/test/run.sh # Node + headless Chromium (not in CI) +``` + +`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus cargo test -p +clawhdf5-wasm --test lazy` compares every corpus file read lazily with the +same file read from bytes. + +Built without `mmap` and `parallel` and without the Zstd and SZIP filters +(they link C): such datasets fail with `unsupported filter`. No C is +compiled; `publish = false` (it is distributed as the package +`build.sh` makes). + +## License + +MIT diff --git a/crates/clawhdf5/README.md b/crates/clawhdf5/README.md index 707715b..6ba17e5 100644 --- a/crates/clawhdf5/README.md +++ b/crates/clawhdf5/README.md @@ -1,27 +1,101 @@ # clawhdf5 -[![crates.io](https://img.shields.io/crates/v/clawhdf5.svg)](https://crates.io/crates/clawhdf5) -[![docs.rs](https://docs.rs/clawhdf5/badge.svg)](https://docs.rs/clawhdf5) +The main crate: a pure-Rust HDF5 reader, writer and in-place editor, with no +libhdf5 and, by default, no C code. It wraps +[`clawhdf5-format`](../clawhdf5-format/README.md) (the binary format) and +[`clawhdf5-io`](../clawhdf5-io/README.md) (memory-mapped reads) in an +h5py-like API. -Pure-Rust HDF5 reader/writer β€” no C dependencies. +Not on crates.io yet; depend on it from git: + +```toml +[dependencies] +clawhdf5 = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" } +``` + +## Main types + +| Type | What it does | +|---|---| +| `File` | Opens a file (`open`, `open_buffered`, `from_bytes`, `open_storage` for any `Storage`), walks groups (`root`, `group`, `dataset`), lists `datasets`/`groups`/`attrs`. `File` is `Send + Sync`: several threads can read one open file. | +| `Dataset` | `shape`, `dtype`, `max_dimensions`, `attrs`; reads `read_f64`/`read_f32`/`read_i32`/`read_i64`/`read_u64`, strings (`read_string`, `read_string_bytes`), variable-length data (`read_vlen`), hyperslabs and point selections (`read_selection`, `read_f64_selection`, ...), zero-copy views of contiguous data (`read_f64_zerocopy`, ...), `verify_provenance`. | +| `FileBuilder` | Writes a new file: datasets of every numeric type, strings, compounds (`CompoundTypeBuilder`), enums, chunked and compressed layouts (deflate, shuffle, Fletcher-32, LZF, and with features LZ4, Zstd, bitshuffle, bzip2, Blosc, pcodec), nested groups, soft/hard/external links, virtual datasets, attribute creation order. Files open in h5py and h5dump. | +| `FileEditor` | Changes an existing file in place without rewriting it: `write_values`/`write_selection`/`write_all`, `resize` of chunked datasets (every chunk index), `set_attr` (compact and dense storage). Anything it cannot do safely is `Error::Unsupported` before any write. | +| `MmapFile`, `LazyFile` | Alternative readers: memory-mapped, and one that reads lazily and caches. | +| `File::open_swmr` | Reads a file a libhdf5 SWMR writer is still appending to (`Dataset::refresh`, bounded retries), as h5py's `swmr=True` reader does. | + +## Examples + +```rust,no_run +use clawhdf5::{AttrValue, File, FileBuilder, FileEditor, Selection}; + +// Write +let mut b = FileBuilder::new(); +b.create_dataset("sensors/temperature") + .with_f64_data(&[20.5, 21.0, 21.5, 22.0]) + .with_shape(&[4]) + .with_maxshape(&[u64::MAX]) // unlimited, so it can grow + .with_chunks(&[2]) + .with_deflate(4); +b.set_attr("version", AttrValue::I64(1)); +b.write("data.h5")?; + +// Read +let file = File::open("data.h5")?; +let ds = file.dataset("sensors/temperature")?; +assert_eq!(ds.shape()?, vec![4]); +let values = ds.read_f64()?; + +// Edit in place: grow the dataset and fill the new tail +let mut ed = FileEditor::open("data.h5")?; +ed.resize("sensors/temperature", &[6])?; +let tail = Selection::Hyperslab { + start: vec![4], + stride: vec![1], + count: vec![2], + block: vec![1], +}; +ed.write_values("sensors/temperature", &tail, &[22.5f64, 23.0])?; +# Ok::<(), clawhdf5::Error>(()) +``` + +Remote files (HTTP range requests, S3/GCS/Azure) are read through +`File::open_storage`; [`clawhdf5-remote`](../clawhdf5-remote/README.md) +provides the storage and its block cache. ## Features -- Read and write HDF5 files entirely in Rust -- Memory-mapped I/O for large files (`mmap` feature, enabled by default) -- Parallel chunk reads via Rayon (`parallel` feature) -- Lazy dataset access for minimal memory usage -- h5py-compatible file output +| Feature | Default | What | Builds C | +|---|---|---|---| +| `mmap` | yes | memory-mapped reads (`File::open` maps the file; `MmapFile`) | no | +| `provenance` | yes | SHA-256 `_provenance_sha256` attributes (`DatasetBuilder::with_provenance`, `Dataset::verify_provenance`) | no | +| `lzf` | yes | LZF filter (32000), h5py's `compression="lzf"` | no | +| `parallel` | no | chunk decoding on a rayon pool | no | +| `lz4` | no | LZ4 filter (32004) | no | +| `pcodec` | no | pcodec filter | no | +| `bitshuffle`, `bzip2`, `blosc` | no | plugin filters 32008, 307, 32001 (read and write) | no (bzip2 uses the pure-Rust `libbz2-rs-sys`) | +| `blosc2`, `zfp` | no | plugin filters 32026 and 32013, **read only** | no | +| `plugin-filters` | no | `lzf`, `bitshuffle`, `bzip2`, `blosc`, `blosc2`, `zfp` | no | +| `zstd` | no | Zstandard filter (32015) | yes (libzstd) | +| `fast-deflate` | no | zlib-ng instead of the pure-Rust zlib-rs | yes (cmake) | +| `blake3_hash` | no | `provenance::blake3_hash` helpers | yes (`cc`, for blake3's SIMD code) | +| `apple-compression` | no | currently has no effect in this crate (it is not forwarded) | β€” | -## Usage +SZIP decoding is a `clawhdf5-format` feature (`szip`, links the system +libaec); the facade does not forward it. -```rust -use clawhdf5::File; +## Limits and further reading -let file = File::open("data.h5").unwrap(); -let dataset = file.dataset("/group/data").unwrap(); -let values: Vec = dataset.read_1d().unwrap(); -``` +- What is known not to work, and what was wrong in earlier releases: + [`docs/known-issues.md`](../../docs/known-issues.md) (editor limits, range + reads, external links and external raw data, which are explicit errors). +- Read coverage against libhdf5/h5py on eight public corpora: + [`CONFORMANCE.md`](../../CONFORMANCE.md). +- Read and write speed against libhdf5 and h5py: + [`BENCHMARKS.md`](../../BENCHMARKS.md). +- Range reads and SWMR design: [`docs/design/range-reads.md`](../../docs/design/range-reads.md), + [`docs/design/swmr.md`](../../docs/design/swmr.md). +- Changes: [`CHANGELOG.md`](../../CHANGELOG.md). ## License diff --git a/IMPROVEMENT_LOG.md b/docs/archive/IMPROVEMENT_LOG.md similarity index 71% rename from IMPROVEMENT_LOG.md rename to docs/archive/IMPROVEMENT_LOG.md index 5c43d8b..af45f00 100644 --- a/IMPROVEMENT_LOG.md +++ b/docs/archive/IMPROVEMENT_LOG.md @@ -1,3 +1,5 @@ +> **Historical (archived 2026-09-28):** a log of an automated improvement loop's PRs from April–May 2026, on the earlier `quantumclaw/clawhdf5` PR numbering (not today's). Superseded by [`CHANGELOG.md`](../../CHANGELOG.md) and `git log`. + # Improvement Log -- clawhdf5 | Date | Loop | PR | Changes | Status | diff --git a/IMPROVEMENT_SCAN.md b/docs/archive/IMPROVEMENT_SCAN.md similarity index 82% rename from IMPROVEMENT_SCAN.md rename to docs/archive/IMPROVEMENT_SCAN.md index a4fdcf2..e0eccd1 100644 --- a/IMPROVEMENT_SCAN.md +++ b/docs/archive/IMPROVEMENT_SCAN.md @@ -1,3 +1,5 @@ +> **Historical (archived 2026-09-28):** one automated scan's notes (2026-05-04), describing changes long since merged. Superseded by [`CHANGELOG.md`](../../CHANGELOG.md) and `git log`. + # Improvement Scan -- clawhdf5 **Date:** 2026-05-04 diff --git a/docs/superpowers/plans/2026-06-29-filter-codecs.md b/docs/archive/plans/2026-06-29-filter-codecs.md similarity index 98% rename from docs/superpowers/plans/2026-06-29-filter-codecs.md rename to docs/archive/plans/2026-06-29-filter-codecs.md index 9f44186..8a78d49 100644 --- a/docs/superpowers/plans/2026-06-29-filter-codecs.md +++ b/docs/archive/plans/2026-06-29-filter-codecs.md @@ -1,3 +1,5 @@ +> **Historical (archived 2026-09-28):** a pre-work plan, implemented in `d6c4d4f` (2026-06-30). Superseded by the code (`crates/clawhdf5-format`), [`CHANGELOG.md`](../../../CHANGELOG.md) and [`ROADMAP.md`](../../../ROADMAP.md); not an open task list. + # Filter Codecs Implementation Plan > **Status (2026-08-03):** Implemented β€” shipped in commit `d6c4d4f` (2026-06-30), with FFI/constant fixes in `cb0b0e9`/`e91f7fc`. This doc was authored 2026-06-29 as the pre-work plan and committed to the repo retroactively on 2026-08-03; checkboxes below have been marked complete to match. Treat this as a historical record, not an open task list. diff --git a/docs/superpowers/plans/2026-06-29-format-write-extensions.md b/docs/archive/plans/2026-06-29-format-write-extensions.md similarity index 99% rename from docs/superpowers/plans/2026-06-29-format-write-extensions.md rename to docs/archive/plans/2026-06-29-format-write-extensions.md index 71f0380..6f8cdca 100644 --- a/docs/superpowers/plans/2026-06-29-format-write-extensions.md +++ b/docs/archive/plans/2026-06-29-format-write-extensions.md @@ -1,3 +1,5 @@ +> **Historical (archived 2026-09-28):** a pre-work plan, implemented in `d6c4d4f` (2026-06-30) and on 2026-08-03. Superseded by the code (`crates/clawhdf5-format`), [`CHANGELOG.md`](../../../CHANGELOG.md) and [`ROADMAP.md`](../../../ROADMAP.md); not an open task list. + # Format Write Extensions Implementation Plan > **Status (2026-08-03):** Implemented. Tasks 1–3 (external links, VDS mapping serialization, VDS `FileWriter` API) shipped in commit `d6c4d4f` (2026-06-30). Tasks 4–5 (superblock v4 read/write) were not part of that commit and were completed separately as part of this cleanup pass (2026-08-03) β€” see `Superblock::parse_v4`/`serialize` and `FileWriter::with_page_size` in `crates/clawhdf5-format`. This doc was authored 2026-06-29 as the pre-work plan and committed to the repo retroactively; checkboxes below have been marked complete to match current state. Treat this as a historical record, not an open task list. diff --git a/docs/superpowers/plans/2026-06-29-mpi-io-vol-backend.md b/docs/archive/plans/2026-06-29-mpi-io-vol-backend.md similarity index 98% rename from docs/superpowers/plans/2026-06-29-mpi-io-vol-backend.md rename to docs/archive/plans/2026-06-29-mpi-io-vol-backend.md index b772ba1..2570335 100644 --- a/docs/superpowers/plans/2026-06-29-mpi-io-vol-backend.md +++ b/docs/archive/plans/2026-06-29-mpi-io-vol-backend.md @@ -1,3 +1,5 @@ +> **Historical (archived 2026-09-28):** a pre-work plan. Its goal of *collective* MPI-IO is not what shipped: `MpiVol` (`d6c4d4f`) is root-read + broadcast and gather-to-root writes (see [`crates/clawhdf5-io/README.md`](../../../crates/clawhdf5-io/README.md)); collective I/O is an open item in [`ROADMAP.md`](../../../ROADMAP.md). + # MPI-IO VOL Backend Implementation Plan > **Status (2026-08-03):** Implemented β€” shipped in commit `d6c4d4f` (2026-06-30), with FFI/constant fixes in `cb0b0e9`/`e91f7fc`. This doc was authored 2026-06-29 as the pre-work plan and committed to the repo retroactively on 2026-08-03; checkboxes below have been marked complete to match. Treat this as a historical record, not an open task list. diff --git a/examples/wasm-viewer/README.md b/examples/wasm-viewer/README.md index 7eb0849..c697357 100644 --- a/examples/wasm-viewer/README.md +++ b/examples/wasm-viewer/README.md @@ -79,6 +79,30 @@ for a deep one) and one batch of requests for its chunks. Every answer is checked β€” a `206` with exactly the bytes asked for, from the same file (ETag or Last-Modified, and length) β€” or the call fails. +What that costs, counted on tank on 2026-09-27 (`CHANGELOG.md`, "Remote +files in the browser: fewer round trips"), on an h5py file of 3000 +datasets of 16384 `f32` in one group (198 MB, h5py 3.16 / HDF5 2.0), as +passes / requests / bytes fetched, with +`CLAWHDF5_WASM_LIST_FILE= CLAWHDF5_WASM_READ=/d1500 cargo test +--release -p clawhdf5-wasm --test lazy listing_cost_of_a_given_file -- +--nocapture`: + +| file (`libver`), block size | `list('/')` | open + read one dataset | +|---|---|---| +| earliest, 1 MiB | 4 / 68 / 192.5 MB | 6 / 5 / 5.2 MB | +| earliest, 64 KiB | 5 / 530 / 35.3 MB | 8 / 7 / 0.52 MB | +| latest, 1 MiB | 5 / 86 / 196.5 MB | 7 / 7 / 6.7 MB | +| latest, 64 KiB | 6 / 454 / 30.5 MB | 8 / 8 / 0.58 MB | + +Listing reads every child's object header, and h5py spreads those through +the file, so a listing of a group this large fetches most of it at 1 MiB +blocks; a smaller `blockSize` fetches far less at the cost of more +requests. Reading one dataset does not list the group. In the test suite's +200 MB file (`WASM_BIG_MB=200 bash examples/wasm-viewer/test/run.sh`), +listing the root, reading two small datasets, a group's attributes, the +large dataset's shape and a 10-value window of it took 5 requests and +6 MiB. + `data` is the typed array of the stored width (`Float64Array`, `Float32Array` also for `f16`, `Int8Array` ... `BigInt64Array`, `BigUint64Array`), or an array of strings for fixed- and variable-length @@ -155,25 +179,39 @@ browser). ## Size -Measured 2026-09-26 on tank (rustc 1.98.1, wasm-bindgen 0.2.129, gzip 1.14, -`gzip -9 -n`), after `bash examples/wasm-viewer/build.sh`. The package is -larger now and the table has not been re-measured: the reader has grown -since, and `openUrl` (2026-09-27) made the facade's range-read path -reachable from JavaScript and added promise glue and `remote.js`. +Measured 2026-09-28 on tank at `9b5803f` (rustc 1.98.1, wasm-bindgen +0.2.129, gzip 1.14): `bash examples/wasm-viewer/build.sh`, then `wc -c` and +`gzip -9 -n -c FILE | wc -c` of each file in `pkg/`. The opt-level `z` and +`3` rows are the same build with `CARGO_PROFILE_WASM_RELEASE_OPT_LEVEL=z` +(or `3`) and the same `wasm-bindgen --target web` step. | | raw | gzip -9 | |---|---:|---:| -| `pkg/clawhdf5_wasm_bg.wasm` (profile `wasm-release`, opt-level `s`) | 627,501 B | 191,639 B | -| `pkg/clawhdf5_wasm.js` (wasm-bindgen glue) | 21,826 B | 4,487 B | -| same wasm at opt-level `z` | 693,068 B | 192,550 B | -| same wasm at opt-level `3` | 544,035 B | 198,803 B | +| `pkg/clawhdf5_wasm_bg.wasm` (profile `wasm-release`, opt-level `s`) | 1,384,607 B | 378,485 B | +| `pkg/clawhdf5_wasm.js` (wasm-bindgen glue) | 40,711 B | 8,181 B | +| `pkg/snippets/.../js/remote.js` (the HTTP side of `openUrl`) | 9,326 B | 3,448 B | +| same wasm at opt-level `z` | 1,533,772 B | 374,765 B | +| same wasm at opt-level `3` | 1,184,889 B | 394,569 B | | h5wasm 0.10.3: wasm embedded in `dist/esm/hdf5_util.js` | 3,544,184 B | 907,096 B | | h5wasm 0.10.3: `dist/esm/hdf5_util.js` as shipped | 4,150,134 B | 986,699 B | -h5wasm figures: `npm pack h5wasm@0.10.3` (npm reports -`dist.unpackedSize` 14,731,385 B for the whole package), wasm extracted from -the `binaryDecode` literal in `hdf5_util.js`. h5wasm is the whole of libhdf5 -(writing, every datatype, plugins), so this compares download size, not -equal functionality. No `wasm-opt` pass was applied (binaryen is not -installed on tank). opt-level `s` is used because it is the smallest -compressed. +The previous measurement (2026-09-26, before `openUrl`) was 627,501 B / +191,639 B gzipped for the wasm and 21,826 B / 4,487 B for the glue. The +package roughly doubled since. `openUrl` made the facade's `Storage` read +path reachable from JavaScript (it was compiled out before) and added the +lazy cache and the promise glue (`CHANGELOG.md`, M4); the growth has not +been broken down per change. Of the +wasm's 1,384,607 bytes, 476,062 are the `name` custom section (function +names, which wasm-bindgen keeps; `wasm-bindgen --remove-name-section` or a +`wasm-opt` pass would drop them); code is 816,244 and data 83,431. Without +the name section the wasm is 908,541 B, 329,854 B gzipped (section removed +with a script, not a supported build option yet). At +opt-level `z` the gzipped wasm is now 1% smaller than at `s`, which the +profile still uses. + +h5wasm figures (2026-09-26, unchanged): `npm pack h5wasm@0.10.3` (npm +reports `dist.unpackedSize` 14,731,385 B for the whole package), wasm +extracted from the `binaryDecode` literal in `hdf5_util.js`. h5wasm is the +whole of libhdf5 (writing, every datatype, plugins), so this compares +download size, not equal functionality. No `wasm-opt` pass was applied +(binaryen is not installed on tank).