Files
clawhdf5/ROADMAP.md
T
osobhandClaude Opus 5.5 c27a478e44 docs: fact-check the refreshed documentation against its sources
Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:23:51 -05:00

9.2 KiB
Raw Blame History

clawhdf5 roadmap

What has shipped, and what is genuinely next. Everything here is checked against CHANGELOG.md, git log and docs/known-issues.md; dates are merge dates on main. Nothing after v2.7.0 has been released: the work since then is on main under CHANGELOG.md "Unreleased".

Last updated: 2026-09-28 (at 9b5803f, PR #21).


Done

Releases

Version Date Headline
v2.0.0 2026-03-19 rustyhdf5 (11 crates) and edgehdf5 (4 crates) unified into one workspace as clawhdf5-*
v2.1.0 2026-06-03 HNSW backs the agent's vector search by default; live, mutable HNSW index
v2.2.0 – v2.7.0 2026-09-18 – 2026-09-20 bounded decompression and read-path bounds checks, single-writer store locking, WAL v4, HNSW recall fix (0.31 -> 0.98 recall@10 at 100K), fusion weights tuned on LongMemEval, int8 index, Extensible Array read fix and chunk-index checksums

Details per release: CHANGELOG.md.

Since v2.7.0 (unreleased, on main)

PR Merged What
#3 2026-09-23 pure-Rust deflate (zlib-rs) by default, no C in the core crates' default build (checked in CI), MSRV 1.92
#4 2026-09-25 files open in h5py again (every f32 and every empty dataset clawhdf5 wrote was unreadable by libhdf5); float16 embedding storage
#5 2026-09-25 HDF5Memory::search with SearchOptions (source filters, re-ranking, confidence); float16 on by default
#6 2026-09-25 clawhdf5-migrate writes real agent stores; knowledge-graph fix; dated benchmark re-run
#7 2026-09-25 consolidation benchmark completed (cheaper novelty scoring)
#8 2026-09-25 Ed25519-signed checkpoints (HDF5Memory::verify)
#9, #10 2026-09-25 OpenClaw and ZeroClaw integration claims withdrawn — neither ever integrated clawhdf5
#11 2026-09-26 silent wrong data and libhdf5 interop bugs found by the HDF5 audit fixed
#12 2026-09-26 reproducible conformance sweep over eight public corpora, nightly CI job (CONFORMANCE.md)
#13 2026-09-26 reads HDF5 1.6-era layouts, user blocks, virtual datasets, dense attributes, very large groups
#14 2026-09-26 h5rs tools (ls, dump, stat, diff, check), the browser reader (clawhdf5-wasm), libhdf5's header checks, plugin filters (LZF, bitshuffle, bzip2, Blosc), concurrency benchmark
#15 2026-09-26 fast contiguous and concurrent reads, variable-length data, nested groups and links in the writer, Python bindings
#16 2026-09-26 chunked full reads faster than an h5py process pool, writer B-trees of any size, Blosc2 (read), 599/697 conformance
#17 2026-09-26 range reads M0/M1 (indexed name lookups, the Storage trait), ZFP (read), in-place editing (FileEditor)
#18 2026-09-27 range reads M2/M3 (File::open_storage; clawhdf5-remote: HTTP(S), S3, GCS, Azure), in-place editing of every chunk index, shrinking, dense attributes
#19 2026-09-27 remote files in the browser (openUrl, M4), SWMR reader (File::open_swmr, M5), Python remote reads and 'r+' editing
#20 2026-09-28 benchmarks re-measured: LongMemEval with real MiniLM embeddings, local reads on an idle machine
#21 2026-09-28 remote files open in a few requests (group lookups down the B-tree, Storage::hint), ObjectHeader::parse back to its earlier speed, the last conformance mismatches resolved: 602/697 ok, 0 mismatch (the run of 2026-09-28 in CONFORMANCE.md still counts 1 our-error, a corrupt N-Bit file libhdf5's own tests refuse)

Range reads (design: docs/design/range-reads.md)

  • M0 — indexed name lookups (#17)
  • M1 — metadata parsed through the Storage trait (#17)
  • M2 — raw data through Storage, File::open_storage (#18)
  • M3 — clawhdf5-remote: HTTP(S) range requests and object stores through a block cache; h5rs URLs (#18); Python URLs (#19)
  • M4 — openUrl in the browser, restartable "NeedBytes" cache (#19; fewer round trips in #21)
  • M5 — reading files a SWMR writer is appending to (docs/design/swmr.md, #19)

Agent memory (clawhdf5-agent)

Shipped before and during the v2 releases, and kept current since: knowledge graph with entity extraction and resolution; three-tier consolidation with decay; hybrid retrieval (HNSW + BM25, weighted or RRF fusion, re-ranking, confidence rejection, query expansion); temporal index and session DAG; per-save provenance ledger and write-anomaly detection; multi-modal embeddings; WAL with chained CRC32; single-writer locking; signed checkpoints. Retrieval is measured, not claimed: see BENCHMARKS.md ("LongMemEval Results" reports retrieval recall, not QA accuracy; earlier headline numbers that compared different granularities were retracted there).


Next

Not scheduled; listed roughly by how much they unblock. None has a date.

Distribution

  • Publish the crates to crates.io. Nothing is published; the READMEs say to depend on git. Before publishing: no publish settings exist (only clawhdf5-wasm has publish = false).
  • Publish Python wheels to PyPI. crates/clawhdf5-py builds with maturin and is tested in CI, but no wheel is published. The default wheel reads plain http:// only; https/s3/gcs/azure wheels compile C (ring, aws-lc-rs).
  • The Node.js package (packages/clawhdf5-node over clawhdf5-napi) has never worked and is not in CI: fix it and add CI, or remove it (known issue).

HDF5 features

  • SWMR writing. The reader is done (M5); writing a file while libhdf5 readers follow it is not. Also not covered: remote SWMR (a remote file is pinned at open), MmapFile/LazyFile SWMR reads, refreshing groups or attributes.
  • MPI collective I/O. clawhdf5-io's MpiVol (mpi-io) is root-read + broadcast and gather-to-root writes, not collective MPI-IO (MPI_File_read_at_all/write_at_all).
  • Paged-metadata single-request reads. Files written with paged aggregation (H5Pset_file_space_strategy(PAGE), h5repack -S PAGE) keep their metadata in a few pages; range reads could fetch those in one request and use the file's page size as the block size. Today the block size is fixed (1 MiB) and only the first block is read ahead (range-reads design, option (c) as a policy).
  • Blosc2 and ZFP encoders. Both filters are read-only; the other plugin filters (LZF, bitshuffle, bzip2, Blosc 1) read and write.
  • External links and external raw data are explicit errors, not followed.
  • Virtual datasets: the "first missing" view and printf gaps other than 0, source-to-virtual type conversion other than a byte swap, nested virtual sources, source files outside the virtual file's directory.
  • Datatypes: x87 long double and binary128 are refused.
  • Writer: one attribute or link message over 65 515 bytes in dense storage is an error (huge fractal-heap objects); no option to write files HDF5 1.8 can read.
  • FileEditor: new chunks in implicit indexes, variable-length and reference data, filters it cannot encode (scale-offset, N-Bit, SZIP), some dense-attribute heap layouts, creating or deleting objects and attributes (also from Python 'r+'), and no journal (a crash mid-edit can leave the file inconsistent). Freed space is reused only within one editor.
  • Selection reads decode the whole dataset when the selection's bounding box covers more than half of it (a strided ds[::100]), and for compact/virtual datasets or a non-default fill value: correct, but more work than needed.
  • Readers: LazyFile and MmapFile still need the whole file; the zero-copy methods need the file in memory.

Remote and browser

  • Run the s3/gcs/azure backends against real buckets (only built and URL-parsing-tested so far).
  • h5rs options for request headers and cache settings.
  • Browser limits in docs/known-issues.md ("clawhdf5-wasm (browser) limits"): files of 4 GiB or more (wasm32), compound/reference/opaque datasets, round trips per index level. The package doubled in size with openUrl (size table); dropping the function-name section would take a third off the raw size (13% gzipped).

Quality

  • Scheduled fuzz campaigns: the cargo-fuzz targets (crates/clawhdf5-format/fuzz, and the agent's WAL target) run only by hand or with CLAWHDF5_FUZZ_SECONDS.

Withdrawn

  • OpenClaw integration (withdrawn 2026-09-25, PR #9). clawhdf5 was never an OpenClaw memory plugin; the documented memory.backend = "clawhdf5" was never valid. The Rust ClawhdfBackend remains as a library API. docs/openclaw.md records what a real plugin would need.
  • ZeroClaw integration (withdrawn 2026-09-25, PR #10). ZeroClaw has no clawhdf5 backend, and clawhdf5-migrate's SQLite layout is not ZeroClaw's schema.

The old track-by-track tracker this file used to be (agent-memory Tracks 1–8, mid-2026) is in git history (git log -- ROADMAP.md).