Numbers, API names, feature defaults and PR references checked against CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history. - int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives them (2026-09-19/20, machine not recorded, not re-run) instead of none; the Pi 5 1.18x carries 2026-09-21. - BENCHMARKS headline: the libhdf5 chunked-write figure is the newest measurement (35x, 2026-09-23), not 45.3x (2026-08-03). - Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug) in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to the bad_nbit_parms_walk.h5 flip. - README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/ time datasets too; zlib-rs byte-identity scoped to what was measured; macOS default links the system libz for inflate. - Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4 unlimited-dimension size warning. - agent-memory.md: string-dataset compression threshold, agents-md prints Markdown, float16 file sizes linked to their study. - known-issues.md: contiguous selection reads, 1.21x vs h5py threads. - docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify, dated figures, fast-math is not BLAS. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
161 lines
9.2 KiB
Markdown
161 lines
9.2 KiB
Markdown
# clawhdf5 roadmap
|
||
|
||
What has shipped, and what is genuinely next. Everything here is checked
|
||
against `CHANGELOG.md`, `git log` and [`docs/known-issues.md`](docs/known-issues.md);
|
||
dates are merge dates on `main`. Nothing after v2.7.0 has been released:
|
||
the work since then is on `main` under `CHANGELOG.md` "Unreleased".
|
||
|
||
_Last updated: 2026-09-28 (at `9b5803f`, PR #21)._
|
||
|
||
---
|
||
|
||
## Done
|
||
|
||
### Releases
|
||
|
||
| Version | Date | Headline |
|
||
|---|---|---|
|
||
| v2.0.0 | 2026-03-19 | rustyhdf5 (11 crates) and edgehdf5 (4 crates) unified into one workspace as `clawhdf5-*` |
|
||
| v2.1.0 | 2026-06-03 | HNSW backs the agent's vector search by default; live, mutable HNSW index |
|
||
| v2.2.0 – v2.7.0 | 2026-09-18 – 2026-09-20 | bounded decompression and read-path bounds checks, single-writer store locking, WAL v4, HNSW recall fix (0.31 -> 0.98 recall@10 at 100K), fusion weights tuned on LongMemEval, int8 index, Extensible Array read fix and chunk-index checksums |
|
||
|
||
Details per release: [`CHANGELOG.md`](CHANGELOG.md).
|
||
|
||
### Since v2.7.0 (unreleased, on `main`)
|
||
|
||
| PR | Merged | What |
|
||
|---|---|---|
|
||
| #3 | 2026-09-23 | pure-Rust deflate (zlib-rs) by default, no C in the core crates' default build (checked in CI), MSRV 1.92 |
|
||
| #4 | 2026-09-25 | files open in h5py again (every `f32` and every empty dataset clawhdf5 wrote was unreadable by libhdf5); float16 embedding storage |
|
||
| #5 | 2026-09-25 | `HDF5Memory::search` with `SearchOptions` (source filters, re-ranking, confidence); float16 on by default |
|
||
| #6 | 2026-09-25 | `clawhdf5-migrate` writes real agent stores; knowledge-graph fix; dated benchmark re-run |
|
||
| #7 | 2026-09-25 | consolidation benchmark completed (cheaper novelty scoring) |
|
||
| #8 | 2026-09-25 | Ed25519-signed checkpoints (`HDF5Memory::verify`) |
|
||
| #9, #10 | 2026-09-25 | OpenClaw and ZeroClaw integration claims withdrawn — neither ever integrated clawhdf5 |
|
||
| #11 | 2026-09-26 | silent wrong data and libhdf5 interop bugs found by the HDF5 audit fixed |
|
||
| #12 | 2026-09-26 | reproducible conformance sweep over eight public corpora, nightly CI job ([`CONFORMANCE.md`](CONFORMANCE.md)) |
|
||
| #13 | 2026-09-26 | reads HDF5 1.6-era layouts, user blocks, virtual datasets, dense attributes, very large groups |
|
||
| #14 | 2026-09-26 | `h5rs` tools (`ls`, `dump`, `stat`, `diff`, `check`), the browser reader (`clawhdf5-wasm`), libhdf5's header checks, plugin filters (LZF, bitshuffle, bzip2, Blosc), concurrency benchmark |
|
||
| #15 | 2026-09-26 | fast contiguous and concurrent reads, variable-length data, nested groups and links in the writer, Python bindings |
|
||
| #16 | 2026-09-26 | chunked full reads faster than an h5py process pool, writer B-trees of any size, Blosc2 (read), 599/697 conformance |
|
||
| #17 | 2026-09-26 | range reads M0/M1 (indexed name lookups, the `Storage` trait), ZFP (read), in-place editing (`FileEditor`) |
|
||
| #18 | 2026-09-27 | range reads M2/M3 (`File::open_storage`; `clawhdf5-remote`: HTTP(S), S3, GCS, Azure), in-place editing of every chunk index, shrinking, dense attributes |
|
||
| #19 | 2026-09-27 | remote files in the browser (`openUrl`, M4), SWMR reader (`File::open_swmr`, M5), Python remote reads and `'r+'` editing |
|
||
| #20 | 2026-09-28 | benchmarks re-measured: LongMemEval with real MiniLM embeddings, local reads on an idle machine |
|
||
| #21 | 2026-09-28 | remote files open in a few requests (group lookups down the B-tree, `Storage::hint`), `ObjectHeader::parse` back to its earlier speed, the last conformance mismatches resolved: 602/697 ok, 0 mismatch (the run of 2026-09-28 in [`CONFORMANCE.md`](CONFORMANCE.md) still counts 1 our-error, a corrupt N-Bit file libhdf5's own tests refuse) |
|
||
|
||
### Range reads (design: [`docs/design/range-reads.md`](docs/design/range-reads.md))
|
||
|
||
- [x] M0 — indexed name lookups (#17)
|
||
- [x] M1 — metadata parsed through the `Storage` trait (#17)
|
||
- [x] M2 — raw data through `Storage`, `File::open_storage` (#18)
|
||
- [x] M3 — `clawhdf5-remote`: HTTP(S) range requests and object stores through a block cache; `h5rs` URLs (#18); Python URLs (#19)
|
||
- [x] M4 — `openUrl` in the browser, restartable "NeedBytes" cache (#19; fewer round trips in #21)
|
||
- [x] M5 — reading files a SWMR writer is appending to ([`docs/design/swmr.md`](docs/design/swmr.md), #19)
|
||
|
||
### Agent memory (`clawhdf5-agent`)
|
||
|
||
Shipped before and during the v2 releases, and kept current since:
|
||
knowledge graph with entity extraction and resolution; three-tier
|
||
consolidation with decay; hybrid retrieval (HNSW + BM25, weighted or RRF
|
||
fusion, re-ranking, confidence rejection, query expansion); temporal index
|
||
and session DAG; per-save provenance ledger and write-anomaly detection;
|
||
multi-modal embeddings; WAL with chained CRC32; single-writer locking;
|
||
signed checkpoints. Retrieval is measured, not claimed: see
|
||
[`BENCHMARKS.md`](BENCHMARKS.md) ("LongMemEval Results" reports retrieval
|
||
recall, not QA accuracy; earlier headline numbers that compared different
|
||
granularities were retracted there).
|
||
|
||
---
|
||
|
||
## Next
|
||
|
||
Not scheduled; listed roughly by how much they unblock. None has a date.
|
||
|
||
### Distribution
|
||
|
||
- [ ] **Publish the crates to crates.io.** Nothing is published; the READMEs
|
||
say to depend on git. Before publishing: no `publish` settings exist
|
||
(only `clawhdf5-wasm` has `publish = false`).
|
||
- [ ] **Publish Python wheels to PyPI.** `crates/clawhdf5-py` builds with
|
||
maturin and is tested in CI, but no wheel is published. The default wheel
|
||
reads plain `http://` only; `https`/`s3`/`gcs`/`azure` wheels compile C
|
||
(ring, aws-lc-rs).
|
||
- [ ] **The Node.js package** (`packages/clawhdf5-node` over
|
||
`clawhdf5-napi`) has never worked and is not in CI: fix it and add CI, or
|
||
remove it ([known issue](docs/known-issues.md)).
|
||
|
||
### HDF5 features
|
||
|
||
- [ ] **SWMR writing.** The reader is done (M5); writing a file while
|
||
libhdf5 readers follow it is not. Also not covered: remote SWMR (a remote
|
||
file is pinned at open), `MmapFile`/`LazyFile` SWMR reads, refreshing
|
||
groups or attributes.
|
||
- [ ] **MPI collective I/O.** `clawhdf5-io`'s `MpiVol` (`mpi-io`) is
|
||
root-read + broadcast and gather-to-root writes, not collective MPI-IO
|
||
(`MPI_File_read_at_all`/`write_at_all`).
|
||
- [ ] **Paged-metadata single-request reads.** Files written with paged
|
||
aggregation (`H5Pset_file_space_strategy(PAGE)`, `h5repack -S PAGE`)
|
||
keep their metadata in a few pages; range reads could fetch those in one
|
||
request and use the file's page size as the block size. Today the block
|
||
size is fixed (1 MiB) and only the first block is read ahead
|
||
(range-reads design, option (c) as a policy).
|
||
- [ ] **Blosc2 and ZFP encoders.** Both filters are read-only; the other
|
||
plugin filters (LZF, bitshuffle, bzip2, Blosc 1) read and write.
|
||
- [ ] **External links and external raw data** are explicit errors, not
|
||
followed.
|
||
- [ ] **Virtual datasets:** the "first missing" view and printf gaps other
|
||
than 0, source-to-virtual type conversion other than a byte swap, nested
|
||
virtual sources, source files outside the virtual file's directory.
|
||
- [ ] **Datatypes:** x87 long double and binary128 are refused.
|
||
- [ ] **Writer:** one attribute or link message over 65 515 bytes in dense
|
||
storage is an error (huge fractal-heap objects); no option to write
|
||
files HDF5 1.8 can read.
|
||
- [ ] **`FileEditor`:** new chunks in implicit indexes, variable-length and
|
||
reference data, filters it cannot encode (scale-offset, N-Bit, SZIP),
|
||
some dense-attribute heap layouts, creating or deleting objects and
|
||
attributes (also from Python `'r+'`), and no journal (a crash mid-edit
|
||
can leave the file inconsistent). Freed space is reused only within one
|
||
editor.
|
||
- [ ] **Selection reads** decode the whole dataset when the selection's
|
||
bounding box covers more than half of it (a strided `ds[::100]`), and
|
||
for compact/virtual datasets or a non-default fill value: correct, but
|
||
more work than needed.
|
||
- [ ] **Readers:** `LazyFile` and `MmapFile` still need the whole file;
|
||
the zero-copy methods need the file in memory.
|
||
|
||
### Remote and browser
|
||
|
||
- [ ] Run the `s3`/`gcs`/`azure` backends against real buckets (only built
|
||
and URL-parsing-tested so far).
|
||
- [ ] `h5rs` options for request headers and cache settings.
|
||
- [ ] Browser limits in [`docs/known-issues.md`](docs/known-issues.md)
|
||
("`clawhdf5-wasm` (browser) limits"): files of 4 GiB or more (wasm32),
|
||
compound/reference/opaque datasets, round trips per index level. The
|
||
package doubled in size with `openUrl`
|
||
([size table](examples/wasm-viewer/README.md#size)); dropping the
|
||
function-name section would take a third off the raw size (13% gzipped).
|
||
|
||
### Quality
|
||
|
||
- [ ] Scheduled fuzz campaigns: the cargo-fuzz targets
|
||
([`crates/clawhdf5-format/fuzz`](crates/clawhdf5-format/fuzz/README.md),
|
||
and the agent's WAL target) run only by hand or with
|
||
`CLAWHDF5_FUZZ_SECONDS`.
|
||
|
||
---
|
||
|
||
## Withdrawn
|
||
|
||
- **OpenClaw integration** (withdrawn 2026-09-25, PR #9). clawhdf5 was
|
||
never an OpenClaw memory plugin; the documented
|
||
`memory.backend = "clawhdf5"` was never valid. The Rust `ClawhdfBackend`
|
||
remains as a library API. [`docs/openclaw.md`](docs/openclaw.md) records
|
||
what a real plugin would need.
|
||
- **ZeroClaw integration** (withdrawn 2026-09-25, PR #10). ZeroClaw has no
|
||
clawhdf5 backend, and `clawhdf5-migrate`'s SQLite layout is not
|
||
ZeroClaw's schema.
|
||
|
||
The old track-by-track tracker this file used to be (agent-memory
|
||
Tracks 1–8, mid-2026) is in git history (`git log -- ROADMAP.md`).
|