Merge branch 'docs/refresh-reference' into docs/readme-refresh
This commit is contained in:
+100
-12
@@ -51,6 +51,31 @@ target: Criterion stretched it where 5 s could not hold the samples it needed
|
||||
|
||||
---
|
||||
|
||||
## Current headline numbers
|
||||
|
||||
The newest dated measurement of each headline figure, as of 2026-09-28.
|
||||
Everything below this section is the dated record behind them; sections whose
|
||||
figures a later run replaced are marked *Superseded*. Machine "tank" is an AMD
|
||||
Ryzen 7 7800X3D (8C/16T); rows marked idle were run with the 1-minute load
|
||||
average below 2.
|
||||
|
||||
| Figure | Value | Measured | Command | Details |
|
||||
|---|---|---|---|---|
|
||||
| Agent memory search, `HDF5Memory::hybrid_search` p50 | 0.49 ms at 10K, 4.69 ms at 100K records | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` | [Current: search harness](#current-search-harness-2026-09-24) |
|
||||
| LongMemEval `longmemeval_s` (full haystack), default hybrid 0.4/0.6, turn-level retrieval Hit@5 (not QA accuracy) | 81.4% | 2026-09-27, tank, search code of `7a8fae0` | `longmemeval_bench … --embeddings weights/all-minilm-l6-v2` | [Re-run with real embeddings](#re-run-with-real-embeddings-2026-09-27-tank), [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500) |
|
||||
| Loaded store memory, 100K × 384 | 399 MiB (2.72x raw) with the `f32` index; 256 MiB (1.74x) with the int8 index (int8 side not re-run since it was first measured) | 2026-09-24, tank, `5c8323c` | `search_harness -- --footprint --full [--int8]` | [Memory footprint](#memory-footprint), [Quantising the index copy](#quantising-the-index-copy-quantized_index) |
|
||||
| int8 index vs `f32` index, QPS at equal recall | 1.63x (x86-64 AVX2), 1.18x (Raspberry Pi 5, `SDOT`) | undated; not re-checked against the 2026-09-24 `f32` figure | `search_harness -- --full` | [Quantising the index copy](#quantising-the-index-copy-quantized_index), [On ARM](#on-arm-raspberry-pi-5-cortex-a76) |
|
||||
| `float16` store file size, 100K × 384 | 80.8 MiB vs 154.0 MiB `f32` (48% smaller) | 2026-09-23, tank | `search_harness -- --float16-study --full` | [float16 embedding storage](#float16-embedding-storage-memoryconfigfloat16) |
|
||||
| Full reads of chunked deflate data, 16 threads on one `File` | 4944 MB/s, 1.58x 16 h5py processes (noisy run: compare ratios, not MB/s) | 2026-09-26, tank, `c5334b1` | `concurrent_read` + `concurrent_read_h5py.py` | [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) |
|
||||
| Same, clawhdf5 only, against the build before range-read M2/M3 | 8525 MB/s vs 6258 (+36%); contiguous and metadata reads at parity | 2026-09-27, tank (idle), `7a8fae0` vs `8f59b2e` | `concurrent_read --decode-threads 1 --reps 3` | [Local metadata and data reads after range-read M2/M3](#local-metadata-and-data-reads-after-range-read-m2m3-2026-09-27-tank) |
|
||||
| `ObjectHeader::parse` (401 headers) | 23.5–23.6 µs, 1.0–2.6% below `8f59b2e` | 2026-09-27, tank (idle), `96086ad` | `cargo bench -p clawhdf5 --bench local_metadata_bench` | [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) |
|
||||
| Selection reads, 64 MB chunked + deflate `f64` | full 63.2 ms; one 64 × 64 window 0.18 ms | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin read_harness` | [Current: read harness](#current-read-harness-2026-09-24) |
|
||||
| Deflate backend, zlib-rs (default) vs zlib-ng | within 6% on every HDF5 read/write path | 2026-09-23, tank | `cargo bench -p clawhdf5-filters --bench deflate_bench` (and the two commands with it) | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng) |
|
||||
| vs libhdf5 1.14.6: chunked deflate-6 write 512×512 / 128 attributes / 64 groups | 45.3x / 10.3x / 10.6x | 2026-08-03, tank | `cargo bench -p clawhdf5-bench --features libhdf5-compare` | [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) |
|
||||
| Signed checkpoints | about 20% of a checkpoint (598 vs 495 ms at 100K) | 2026-09-25, tank | `search_harness -- --signing-study --full` | [Signed checkpoints](#signed-checkpoints) |
|
||||
|
||||
---
|
||||
|
||||
## Memory footprint
|
||||
|
||||
`cargo run --release -p clawhdf5-bench --bin search_harness -- --footprint --full`,
|
||||
@@ -64,6 +89,10 @@ change at all. Measured that way a store holding the corpus twice and one
|
||||
holding it once came out *identical* (1.00x both), which is how the first
|
||||
attempt at this measurement went.
|
||||
|
||||
> *Superseded* by the current figures below (2026-09-24): this table is the
|
||||
> record of the double-copy fix (commit 2e7e045, undated); the store measured
|
||||
> 2.72x, not 2.43x, by the time the int8 index landed.
|
||||
|
||||
| N | vectors (raw) | reopened, before | reopened, after |
|
||||
|---:|---:|---:|---:|
|
||||
| 1 000 | 1 MiB | 5 MiB (3.41x) | 4 MiB (2.39x) |
|
||||
@@ -379,6 +408,9 @@ point: does a selection cost what the *selection* costs?
|
||||
|
||||
### Baseline (v2.4.0): every selection decodes the whole dataset
|
||||
|
||||
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
|
||||
> (2026-09-24). Kept as the before picture.
|
||||
|
||||
4096 x 2048 f64 (64 MB per dataset), chunks 256 x 256, file 129 MB
|
||||
|
||||
| layout | read | selected | time ms | MB/s of selection | vs full read |
|
||||
@@ -404,6 +436,9 @@ point: does a selection cost what the *selection* costs?
|
||||
|
||||
### After: partial reads
|
||||
|
||||
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
|
||||
> (2026-09-24).
|
||||
|
||||
Only the rows of a contiguous dataset, or the chunks, that overlap the
|
||||
selection's bounding box are read/decoded. A 64 x 64 window of the compressed
|
||||
dataset: **105 -> 0.39 ms**; one row: **106 -> 2.7 ms**; one column:
|
||||
@@ -435,6 +470,9 @@ because the machine's speed drifted; compare the *vs full read* column.)
|
||||
|
||||
### After: parallel cached decode, fewer copies (full reads)
|
||||
|
||||
> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24)
|
||||
> (2026-09-24).
|
||||
|
||||
Full-read times, old and new binaries run alternately at the same moment (this
|
||||
machine's absolute speed drifts over a long session, so only same-moment
|
||||
comparisons mean anything):
|
||||
@@ -542,6 +580,10 @@ rounds; run 2 also alternated `main` `425585e`.
|
||||
|
||||
### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank)
|
||||
|
||||
> The `object_header_parse_x401` row (+4.2%) is *superseded* by
|
||||
> [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank)
|
||||
> (2026-09-27, `96086ad`); the other rows are current.
|
||||
|
||||
`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main`
|
||||
`7a8fae0` (PRs #18 and #19), each built in its own worktree and run as
|
||||
separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen
|
||||
@@ -585,8 +627,8 @@ What this shows:
|
||||
- **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header;
|
||||
the base and candidate ranges do not overlap). It is the cost of reading
|
||||
continuation chunks from a bounded queue (the fix for unbounded reads on
|
||||
crafted headers) and does not show in the listing. Kept open in
|
||||
`docs/known-issues.md`.
|
||||
crafted headers) and does not show in the listing. (Fixed later the
|
||||
same day; see the section above and `docs/known-issues.md`.)
|
||||
- **Full reads of deflate data got faster** after #18 (in-place chunk
|
||||
decoding into the typed output and per-thread scratch buffers): +1.7% on
|
||||
one thread, +36% at 16.
|
||||
@@ -639,6 +681,10 @@ saturate memory bandwidth (1.05x).
|
||||
|
||||
### Results after the read fixes (2026-09-26, tank, `408f69e`)
|
||||
|
||||
> *Superseded* by [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)
|
||||
> (2026-09-26, `c5334b1`), which closed the 16-thread gap listed at the end
|
||||
> of this section.
|
||||
|
||||
Same machine, files and commands as the first run below, re-run on an idle
|
||||
tank (load average 1.60 at the start; the 1-minute figure rose to about 5
|
||||
during the clawhdf5 runs, mostly their own threads) after two fixes:
|
||||
@@ -679,11 +725,17 @@ Read with care:
|
||||
- At 16 threads every tool dropped in this run (h5py threads on contiguous
|
||||
data from 8002 to 2285 MB/s, processes from 12846 to 6942), so the
|
||||
16-thread rows are noisier than the others.
|
||||
- Still behind: full reads of chunked data at 16 threads (0.69x-0.76x h5py
|
||||
processes). See `docs/known-issues.md`.
|
||||
- Still behind at this commit: full reads of chunked data at 16 threads
|
||||
(0.69x-0.76x h5py processes); fixed by `c5334b1` (above), recorded as
|
||||
fixed in `docs/known-issues.md`.
|
||||
|
||||
### First run, before the read fixes (2026-09-26, tank, `91644d8`)
|
||||
|
||||
> *Superseded* results: the tables and "What this shows" are the before
|
||||
> picture for [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1)
|
||||
> (2026-09-26). The workload description and the **Run** box below are
|
||||
> still how every `concurrent_read` figure in this file is produced.
|
||||
|
||||
Measured on tank (AMD Ryzen 7 7800X3D, 8 cores / 16 threads, 61 GiB, Linux
|
||||
7.0) at commit `91644d8`, load average 1.84 when the run started (the
|
||||
1-minute figure rose to 3.7 during the runs; that is mostly the benchmark's
|
||||
@@ -720,8 +772,8 @@ What this shows:
|
||||
- **clawhdf5 threads on one `File` do, for hyperslab reads of compressed
|
||||
data:** 1244 MB/s at 16 threads, 9.7x h5py threads and 0.89x h5py
|
||||
processes, without a process pool.
|
||||
- **Where clawhdf5 is behind** (open performance bugs, see
|
||||
`docs/known-issues.md`):
|
||||
- **Where clawhdf5 was behind** at `91644d8` (both since fixed; see
|
||||
`docs/known-issues.md`, "Concurrent and contiguous read performance"):
|
||||
- *Full reads of chunked datasets stop scaling at about 4 threads*
|
||||
(about 880 MB/s) while h5py processes reach 4424 MB/s. Hyperslab
|
||||
reads, which bypass the `File`'s chunk cache, keep scaling, so the
|
||||
@@ -811,6 +863,11 @@ Other flags (both harnesses): `--threads`, `--reps`, `--slab`, `--slabs`,
|
||||
|
||||
## Search harness baseline (v2.3.0)
|
||||
|
||||
> *Historical.* This baseline and the "After: …" subsections that follow
|
||||
> record each step of the search work; they are *superseded* by
|
||||
> [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24),
|
||||
> the last subsection of this part.
|
||||
|
||||
Produced by `cargo run --release -p clawhdf5-bench --bin search_harness -- --full`
|
||||
on deterministic **clustered** synthetic data (384-dim, unit-normalised; points =
|
||||
cluster centre + noise — uniform random vectors are nearly equidistant in high
|
||||
@@ -873,10 +930,11 @@ build: 9752.6 ms (10254 vectors/s) · exact scan: 40 QPS, p50 24648 µs
|
||||
| 1000 | 11 | 3.9 | 0.9 | 68.1 | 5.48 | 5.57 | 182.5 |
|
||||
| 10000 | 114 | 32.2 | 10.9 | 845.0 | 48.56 | 78.65 | 19.8 |
|
||||
| 100000 | 1486 | 713.0 | 354.5 | 10486.5 | 883.51 | 975.23 | 1.1 |
|
||||
wrote /tmp/claude-1000/-home-osobh-projects-clawhdf5/422f755e-dd25-4c35-8613-5439087e3aaa/scratchpad/baseline_full.json
|
||||
|
||||
### After: HNSW neighbour-selection heuristic
|
||||
|
||||
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
|
||||
|
||||
Same harness, same data, after replacing closest-M neighbour selection with the
|
||||
HNSW paper's diversity heuristic (Algorithm 4, keeping pruned connections) for
|
||||
both new links and back-link pruning. Recall@10 at `ef = 64`: **0.87 → 1.00**
|
||||
@@ -922,6 +980,8 @@ build: 36472.8 ms (2742 vectors/s) · exact scan: 40 QPS, p50 24644 µs
|
||||
|
||||
### After: persistent keyword index, no store rewrite per query
|
||||
|
||||
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
|
||||
|
||||
`hybrid_search` used to rebuild the BM25 index from scratch (re-tokenising every
|
||||
record) and rewrite the whole `.h5` file on **every query**. The index is now
|
||||
kept for the life of the store and updated incrementally, and activation boosts
|
||||
@@ -942,6 +1002,8 @@ index removes that.
|
||||
|
||||
### After: vector index persisted with the checkpoint
|
||||
|
||||
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
|
||||
|
||||
The HNSW graph (not the vectors, which the store already holds) is saved to
|
||||
`<store>.h5.ann` at each checkpoint and reloaded by `open()`, tied to that
|
||||
checkpoint by a generation id. The index is now built once per store (the *cold
|
||||
@@ -959,6 +1021,8 @@ index incrementally.
|
||||
|
||||
### After: unit-vector dot product, reusable visited set
|
||||
|
||||
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
|
||||
|
||||
Cosine distance recomputed both vector norms on every evaluation; the index now
|
||||
stores unit vectors and uses a plain dot product. The per-call `HashSet` of
|
||||
visited nodes became a reusable epoch-stamped array. Recall is unchanged.
|
||||
@@ -1004,6 +1068,8 @@ build: 21084.6 ms (4743 vectors/s) · exact scan: 39 QPS, p50 24739 µs
|
||||
|
||||
### After: unranked keyword scores, top-k merge (rankings unchanged)
|
||||
|
||||
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
|
||||
|
||||
A fusion study (`search_harness --fusion-study`) showed that capping the
|
||||
keyword candidate pool is **not** a safe optimisation: against the current
|
||||
full-corpus normalisation the final top-10 overlap is only 0.83-0.92 and the
|
||||
@@ -1026,6 +1092,8 @@ results.
|
||||
|
||||
### After: batched bulk build (optionally parallel); deletions handled in search
|
||||
|
||||
> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24).
|
||||
|
||||
Profiling showed **90% of a build's distance evaluations are in back-link
|
||||
pruning**. The bulk build now inserts in batches: plan each node's neighbours
|
||||
against the graph as it stood at the start of the batch, link, then prune every
|
||||
@@ -1588,6 +1656,11 @@ MRR, or a one-question change in recency, is within this variation.
|
||||
|
||||
### Full haystack — `longmemeval_s`, n=500 (the number to cite)
|
||||
|
||||
This table is **BM25-only** (zero embeddings). With real embeddings and the
|
||||
default hybrid 0.4/0.6 the same corpus gives turn Hit@5 **81.4%** (2026-09-27;
|
||||
see [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500)),
|
||||
which is the headline figure.
|
||||
|
||||
47.7 sessions and 493.5 turns per question; 4.0% of haystack sessions are evidence
|
||||
sessions, so retrieval has to actually discriminate.
|
||||
|
||||
@@ -1695,7 +1768,9 @@ over rank-1 precision.
|
||||
activation. Until now its combined score contained **no relevance term at
|
||||
all** — `RerankInput` did not carry the retrieval score — so a caller that
|
||||
re-ranked its candidates threw the retriever's ordering away and returned them
|
||||
ordered by age. The OpenClaw backend did exactly that on every search.
|
||||
ordered by age. `ClawhdfBackend` (the `openclaw` module) did exactly that on
|
||||
every search. (OpenClaw itself never integrated clawhdf5; see
|
||||
`docs/openclaw.md`.)
|
||||
|
||||
Measuring that is unambiguous. "Recency" below is the share of
|
||||
`knowledge-update` questions where the newest gold session outranked the stale
|
||||
@@ -1793,7 +1868,7 @@ worth stating plainly rather than hiding: LongMemEval questions share substantia
|
||||
vocabulary with their evidence turns, which is close to the best case for lexical
|
||||
matching, and MiniLM at 384 dimensions is a small embedding model.
|
||||
|
||||
> **Run:** `cargo run --release --bin longmemeval_bench --features embeddings -- \
|
||||
> **Run:** `cargo run --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings -- \
|
||||
> benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings weights/all-minilm-l6-v2`
|
||||
> For the GPU path use `--features embeddings-cuda`. That requires `nvcc` on
|
||||
> `PATH` at *build* time — cudarc's build script shells out to it. The toolkit
|
||||
@@ -2212,7 +2287,8 @@ The tank row was measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D), commit
|
||||
### Reproducibility
|
||||
|
||||
```bash
|
||||
rustup override set nightly
|
||||
# Any stable toolchain at or above the MSRV (1.92) works; the original
|
||||
# 2026-07-01 run used a nightly, later runs stable.
|
||||
|
||||
# Latency benchmarks (Criterion)
|
||||
cargo bench -p clawhdf5-agent
|
||||
@@ -2315,7 +2391,9 @@ libhdf5 reads from a temp file including `open` + `read` + `close` overhead.
|
||||
| clawhdf5 hyperslab (f64, 10% slice) | — | 4.09 µs / **1.8 GiB/s** | 50.1 µs / **1.5 GiB/s** |
|
||||
|
||||
libhdf5 f64 comparison excluded — clawhdf5's datatype encoding differs from libhdf5's (known
|
||||
gap), making cross-format reads unreliable for comparison.
|
||||
gap), making cross-format reads unreliable for comparison. (That gap was the float sign-bit
|
||||
bug, fixed 2026-09-23: `docs/known-issues.md`, "Every `f32` dataset we wrote was unreadable by
|
||||
h5py / libhdf5". The comparison has not been re-run since.)
|
||||
|
||||
### Chunked Read Throughput
|
||||
|
||||
@@ -2423,6 +2501,12 @@ global file mutex and flushes to disk on every attribute write or group creation
|
||||
|
||||
## vs libhdf5 Summary
|
||||
|
||||
> Measured on the original i7-12650H (clawhdf5 2026-07-01, libhdf5
|
||||
> 2026-06-30). The newest run of this table is
|
||||
> [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03)
|
||||
> (2026-08-03), which reproduces every row within ~15% except chunked
|
||||
> write (45.3x on tank).
|
||||
|
||||
| Workload | clawhdf5 | libhdf5 | Speedup |
|
||||
|----------|----------|---------|---------|
|
||||
| Sequential read, 1K f32 | 634 ns | 45.2 µs | **71×** |
|
||||
@@ -2454,7 +2538,7 @@ to the page cache. There is no algorithmic headroom above ~1.7 GiB/s on this har
|
||||
|
||||
### Caveats
|
||||
|
||||
- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap). f64 results are clawhdf5-only.
|
||||
- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap at the time; fixed 2026-09-23, see [Sequential Read Throughput](#sequential-read-throughput)). f64 results are clawhdf5-only.
|
||||
- Serial benchmarks. clawhdf5 uses Rayon for chunk compression when > 2 chunks; that parallelism is already reflected in the chunked write numbers.
|
||||
- clawhdf5 reads from `Vec<u8>` (zero-copy from mmap in production); libhdf5 reads from a temp file. This gives clawhdf5 a structural read advantage that reflects realistic API usage.
|
||||
|
||||
@@ -2651,6 +2735,10 @@ Same not-like-for-like caveat as the "Comparison to MemX" section at the top of
|
||||
file applies — MemX's figure is end-to-end, these are a single component. Ratios are
|
||||
an order-of-magnitude indication, not a benchmark result.
|
||||
|
||||
> The Ratio column below was retracted afterwards: see
|
||||
> [Comparison to MemX](#comparison-to-memx-arxiv260316171). Kept as recorded
|
||||
> on 2026-08-05; do not cite it.
|
||||
|
||||
| Metric | MemX (claimed, end-to-end) | ClawhDF5 (tank, component only) | Ratio |
|
||||
|--------|----------------------------|----------------------------------|-------|
|
||||
| 100K flat search | <90 ms | 6.60 ms | ~14x |
|
||||
|
||||
@@ -1,253 +1,266 @@
|
||||
# clawhdf5
|
||||
|
||||
## Purpose
|
||||
Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated vector search. A standalone library. Its one verified consumer is ClawBrainHub (`.brain` files); no agent framework integrates it (OpenClaw and ZeroClaw claims were withdrawn on 2026-09-25 — neither was ever true).
|
||||
Pure-Rust HDF5 implementation (read, write, in-place edit, remote and browser
|
||||
reads) plus agent memory on top of it: HNSW vector search, a WAL-backed store,
|
||||
and GPU vector distances. A standalone library. Its one verified consumer is
|
||||
ClawBrainHub (`.brain` files); no agent framework integrates it (see
|
||||
*Standing rules*).
|
||||
|
||||
## Architecture
|
||||
|
||||
Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal FFI bindings crate for the optional `szip` feature):
|
||||
Cargo workspace, 19 crates under `crates/` (plus `libaec-sys`, the FFI crate
|
||||
behind the optional `szip` feature). MSRV 1.92 (`rust-version`, checked in CI).
|
||||
|
||||
| Crate | Role |
|
||||
|-------|------|
|
||||
| `clawhdf5-format` | HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants |
|
||||
| `clawhdf5-io` | Read/write implementation |
|
||||
| `clawhdf5-filters` | Deflate backends (zlib-rs, zlib-ng, Apple Compression); the HDF5 filter pipeline, the filter registry (`clawhdf5_format::filter_registry`) and the other codecs (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec, and the pure-Rust plugin filters LZF, bitshuffle, bzip2, Blosc 1, and Blosc2 and ZFP read-only) live in `clawhdf5-format`. |
|
||||
| `clawhdf5-derive` | Proc-macro derive for HDF5-serializable structs |
|
||||
| `clawhdf5` | Main facade crate |
|
||||
| `clawhdf5-netcdf4` | NetCDF-4 compatibility layer |
|
||||
| `clawhdf5-ann` | HNSW approximate nearest-neighbor vector index |
|
||||
| `clawhdf5-agent` | Agent memory, session history, knowledge graph storage |
|
||||
| `clawhdf5-gpu` | GPU vector distance computation via wgpu (hand-written WGSL compute shaders) — not dataset I/O |
|
||||
| `clawhdf5-accel` | CPU SIMD acceleration path |
|
||||
| `clawhdf5-migrate` | SQLite → HDF5 agent-memory migration |
|
||||
| `clawhdf5-format` | The HDF5 format: parsers and writer (superblock, headers, B-trees, heaps, chunk indexes), the `Storage` trait, the filter pipeline and registry (`filter_registry`), every codec except deflate (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec; pure-Rust LZF, bitshuffle, bzip2, Blosc 1; Blosc2 and ZFP read-only), `float16`, `checksum` |
|
||||
| `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) |
|
||||
| `clawhdf5-io` | I/O adapters (buffers, mmap, prefetch) |
|
||||
| `clawhdf5-derive` | `#[derive(H5Type)]` for compound types |
|
||||
| `clawhdf5` | Facade: `File`, `FileBuilder`, `Dataset`, `FileEditor` (`src/edit/`), SWMR reading (`src/swmr.rs`) |
|
||||
| `clawhdf5-netcdf4` | NetCDF-4 read support |
|
||||
| `clawhdf5-remote` | `open_url`: HTTP(S) range requests and object stores (S3, GCS, Azure) through `BlockCache` |
|
||||
| `clawhdf5-tools` | `h5rs`: `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` |
|
||||
| `clawhdf5-py` | PyO3 bindings (h5py-like API, remote files, `'r+'` editing) |
|
||||
| `clawhdf5-wasm` | wasm-bindgen browser reader (`open(bytes)`, `openUrl(url)`); demo in `examples/wasm-viewer/` |
|
||||
| `clawhdf5-ann` | HNSW index |
|
||||
| `clawhdf5-agent` | Agent memory store (`HDF5Memory`), sessions, knowledge graph, BM25 |
|
||||
| `clawhdf5-accel` | CPU SIMD kernels (AVX2, NEON) |
|
||||
| `clawhdf5-gpu` | wgpu vector distances (WGSL) — not dataset I/O; HDF5 I/O is CPU-only |
|
||||
| `clawhdf5-migrate` | SQLite → agent store migration |
|
||||
| `clawhdf5-cli` | Agent-memory CLI |
|
||||
| `clawhdf5-napi` | Node.js addon (the `packages/clawhdf5-node` wrapper is broken; `docs/known-issues.md`) |
|
||||
| `clawhdf5-android` | Android JNI bindings |
|
||||
| `clawhdf5-cli` | Command-line interface (agent memory) |
|
||||
| `clawhdf5-tools` | `h5rs`: pure-Rust HDF5 tools — `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` (structural + checksum validator) |
|
||||
| `clawhdf5-napi` | Node.js native addon bindings |
|
||||
| `clawhdf5-py` | PyO3 Python bindings |
|
||||
| `clawhdf5-wasm` | WebAssembly (wasm-bindgen) reader for the browser; demo in `examples/wasm-viewer/` |
|
||||
| `clawhdf5-remote` | Remote files: `open_url` over HTTP(S) range requests and object stores (`object_store`: S3, GCS, Azure) through a mandatory block cache (`BlockCache`) |
|
||||
| `clawhdf5-bench` | Benchmark suite |
|
||||
| `clawhdf5-bench` | Benchmarks and harnesses (`search_harness`, `read_harness`, `concurrent_read`, `longmemeval_bench`, …) |
|
||||
|
||||
## Key Features
|
||||
- Zero-C-dependency HDF5 read/write: no libhdf5, and deflate defaults to
|
||||
pure-Rust zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake).
|
||||
`ci-test.sh` fails if a C-building crate enters the core crates' default
|
||||
tree. flate2 must keep `runtime_detection` with zlib-rs — without it zlib-rs
|
||||
loses SIMD and inflates 3.5x slower. MSRV is 1.92 (`rust-version`, checked
|
||||
in CI).
|
||||
- HNSW vector index for semantic similarity search over agent memories — the
|
||||
`clawhdf5-agent` `hnsw` feature is **on by default**, so `hybrid_search` uses
|
||||
the approximate `clawhdf5-ann` index for the vector stage (the index mirrors
|
||||
the cache and self-heals on drift). Build the agent with
|
||||
`--no-default-features --features float16` to force the exact linear cosine scan.
|
||||
The agent's `parallel` feature (also default) builds the index on a thread
|
||||
pool; the graph is identical with or without it.
|
||||
The index uses the HNSW paper's diversity heuristic for neighbour selection
|
||||
(plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its
|
||||
graph is saved to `<store>.h5.ann` at each checkpoint and reloaded by `open()`
|
||||
(tied to the checkpoint by a generation id; stale/damaged sidecars are
|
||||
ignored and the index rebuilt). `MemoryConfig::quantized_index` (**on by
|
||||
default** for new stores, persisted; stores predating the setting load as
|
||||
`false` and keep their f32 index — guarded by
|
||||
`tests/fixtures/store_v2_5_0.h5`; CLI opt-out is `create --f32-index`)
|
||||
stores the index's own copy of the embeddings as `i8`,
|
||||
which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors
|
||||
at 100K); because quantised distances are approximate and `ef` cannot
|
||||
compensate, the query path then re-scores the candidate pool against the
|
||||
exact embeddings, which holds recall at the f32 index's level. It is also
|
||||
faster at equal recall: 1.63x the QPS on x86-64 (AVX2) and 1.18x on a
|
||||
Raspberry Pi 5 (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm since
|
||||
the intrinsic is unstable; plain NEON on pre-dotprod cores). The aarch64
|
||||
code is `cfg`'d out on x86, so x86 CI never compiles or lints it — test it
|
||||
on real ARM (`rpivision02`, 10.0.2.3, is a Pi 5). `hybrid_search` keeps one incremental BM25
|
||||
index for the life of the store and never writes the store: Hebbian
|
||||
activation boosts are persisted by the next checkpoint (or on drop), not per
|
||||
query. Measure any search-path change with
|
||||
`cargo run --release -p clawhdf5-bench --bin search_harness` (baselines in
|
||||
`BENCHMARKS.md`).
|
||||
- WAL (write-ahead log) for crash-safe persistence, with a chained CRC32
|
||||
trailer per entry (each entry's CRC folds in the previous entry's CRC) so a
|
||||
corrupted, reordered, duplicated, or spliced entry stops replay cleanly
|
||||
instead of loading bad or tampered data. The pre-chaining per-entry-CRC
|
||||
format (v2) is still fully readable; the oldest no-CRC format (v1) is only
|
||||
reachable through the one-time migration path in `HDF5Memory::open`, not
|
||||
through the public `WalFile::read_entries`.
|
||||
**What the WAL guarantees:** integrity, ordering, and recovery from a
|
||||
*process* crash at any point — including between a checkpoint and the WAL
|
||||
truncate (each checkpoint records a `WalMark` in `/meta`, and `open()` skips
|
||||
the WAL prefix the `.h5` already contains, so entries are never applied
|
||||
twice). Checkpoints and snapshots are made durable as a unit (temp file
|
||||
synced, renamed, directory synced). **What it does not guarantee:**
|
||||
individual WAL appends are *not* fsynced (a deliberate latency trade-off), so
|
||||
saves made since the last checkpoint can be lost on power failure or kernel
|
||||
panic. Current header version is 4 (adds the `Update` record used by
|
||||
`save_or_update`); v3 files are read and upgraded in place.
|
||||
- A store has a **single writer**: `HDF5Memory::create`/`open` hold an exclusive
|
||||
advisory lock on `<store>.h5.lock` and a second opener gets
|
||||
`MemoryError::Locked`. Use `HDF5Memory::open_read_only` for a lock-free,
|
||||
never-writing point-in-time view (the CLI's `recall`/`stats`/`agents-md`/
|
||||
`export` do). An unreadable WAL (torn header, bad magic) is quarantined to
|
||||
`<store>.h5.wal.corrupt-<ts>` rather than blocking `open()`; a WAL with an
|
||||
unknown *newer* version still fails and is left untouched.
|
||||
- `MemoryConfig::float16` (**on by default** for new stores, persisted;
|
||||
existing stores keep their recorded `false` — guarded by the v2.5.0
|
||||
fixture in `tests/float16_store.rs`; CLI opt-out is `create --f32`) writes
|
||||
`/memory/embeddings` as IEEE half precision (48% smaller file at 100K;
|
||||
LongMemEval with real MiniLM embeddings identical to f32).
|
||||
`MemoryCache::half_precision` rounds each embedding as it enters the cache (push, update, WAL replay, and on load of a store still
|
||||
`f32` on disk), so memory and file agree bit for bit; the conversions live
|
||||
in `clawhdf5_format::float16` and must stay the single implementation.
|
||||
Values beyond ±65504 are `MemoryError::InvalidEntry`. Interop: every file
|
||||
must open in h5py — `f32` datasets and empty datasets did not until
|
||||
2026-09-23 (see `docs/known-issues.md`); the agent's `h5py_interop` test
|
||||
guards a whole store.
|
||||
- `HDF5Memory::search(query_emb, text, &SearchOptions)` is the full search
|
||||
path: optional source-channel filter (applied before ranking; exact scan of
|
||||
the allowed records whenever cheaper than `pool × M` index distance
|
||||
Reference docs: `docs/known-issues.md` (open issues table first — check it
|
||||
before calling something a bug or a feature), `BENCHMARKS.md` (headline
|
||||
numbers first), `CONFORMANCE.md` (generated), `docs/design/range-reads.md`
|
||||
and `docs/design/swmr.md`, `CHANGELOG.md` (full detail of every fix).
|
||||
|
||||
## Standing rules
|
||||
|
||||
- **No C in the default build.** No libhdf5; deflate defaults to pure-Rust
|
||||
zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake). `ci-test.sh`
|
||||
fails if a C-building crate enters the core crates' default tree. Zstd,
|
||||
SZIP, `https` (ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are opt-in. flate2
|
||||
must keep `runtime_detection` with zlib-rs — without it zlib-rs loses SIMD
|
||||
and inflates 3.5x slower.
|
||||
- **Every file we write must open in h5py/libhdf5.** Interop tests compare
|
||||
against h5py and h5dump; `f32` and empty datasets did not open until
|
||||
2026-09-23.
|
||||
- **float16 has one implementation:** `clawhdf5_format::float16`.
|
||||
- **Claims need evidence.** Performance and integration claims in docs must
|
||||
be measured, dated (with machine and command), or withdrawn. Benchmark
|
||||
numbers are dated records: never edit a measured value, add a new dated
|
||||
section and mark the old one superseded.
|
||||
- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not and
|
||||
never was an OpenClaw memory plugin; the old `memory.backend = "clawhdf5"`
|
||||
config was never valid. `docs/openclaw.md` records what a real plugin would
|
||||
need. The `openclaw` module's `ClawhdfBackend` is just `search` with
|
||||
re-rank + confidence on.
|
||||
- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream
|
||||
v0.8.5 and the `osobh/zeroclaw` fork and their history): its memory
|
||||
backends are its own; `clawhdf5-migrate`'s default SQLite layout is not
|
||||
ZeroClaw's schema. Don't reintroduce integration claims without an
|
||||
integration and a test against the real consumer.
|
||||
- **known-issues.md:** one entry per bug; when fixed, record it in
|
||||
`CHANGELOG.md` and move the entry to *Fixed (history)* with date, PR,
|
||||
affected releases and what users must do — never delete it.
|
||||
|
||||
## HDF5 library: invariants and gotchas
|
||||
|
||||
- **Remote/range reads** (`docs/design/range-reads.md`, M0-M5 merged in PRs
|
||||
#17-#21): every format-crate read path goes through `Storage`
|
||||
(`read_at`/`read_ranges`/`hint`). `File::open_storage` takes any
|
||||
`Storage`; `clawhdf5_remote::open_url` wraps HTTP (`HttpStorage`, ureq) or
|
||||
`ObjectStoreStorage` in `BlockCache` (1 MiB blocks, LRU budget, in-flight
|
||||
dedup, coalesced runs). Remote files are pinned by ETag/Last-Modified and
|
||||
length (`RemoteError::FileChanged`). Zero-copy APIs and `File::as_bytes`
|
||||
need an in-memory file. Parse through `File::storage()` and the `*_in`
|
||||
functions, not `as_bytes`, in new code (the Python bindings do).
|
||||
`ObjectStoreStorage` runs reads on its own small tokio runtime, so it
|
||||
works from any thread.
|
||||
- **SWMR** (`docs/design/swmr.md`): `File::open_swmr` reads a file a libhdf5
|
||||
SWMR writer is appending to — positioned reads, no chunk cache, bounded
|
||||
retries (100), `Dataset::refresh()`. clawhdf5 has no SWMR writer; remote
|
||||
SWMR is out of scope.
|
||||
- **Browser** (`clawhdf5-wasm`, read-only, no Zstd/SZIP): `openUrl` reads
|
||||
through the restartable "NeedBytes" cache (`src/lazy.rs`: a call is re-run
|
||||
after each wave of misses; no block is evicted while a call runs); the HTTP
|
||||
is JavaScript (`js/remote.js`).
|
||||
- **In-place editing** (`clawhdf5::FileEditor`): overwrites values, grows and
|
||||
shrinks chunked datasets (every chunk index) and sets attributes (compact
|
||||
and dense) without rewriting the file, changing indexes and heaps as
|
||||
libhdf5 does; freed space is reused within one editor. Anything it cannot
|
||||
do safely is `Error::Unsupported` before any write (limits in
|
||||
`docs/known-issues.md`). The algorithms follow libhdf5 `hdf5_1_14_6`
|
||||
(github.com/HDFGroup/hdf5). Test changes with `cargo test -p
|
||||
clawhdf5-tools --test edit_interop --test edit_coverage_interop`.
|
||||
- **Provenance:** `Dataset::verify_provenance()` (facade `provenance`
|
||||
feature, default) re-hashes a dataset against its `_provenance_sha256`
|
||||
attribute (`DatasetBuilder::with_provenance`). Opt-in per call; unkeyed
|
||||
hash — tamper-evident, not tamper-proof.
|
||||
|
||||
## Agent memory: invariants and gotchas
|
||||
|
||||
- **Search.** `HDF5Memory::search(query_emb, text, &SearchOptions)` is the
|
||||
full path: optional source-channel filter (before ranking; exact scan of
|
||||
the allowed records when cheaper than `pool × M` index distance
|
||||
evaluations, and as the fallback when the pool comes back short), fusion,
|
||||
activation scaling, optional re-ranking and confidence rejection.
|
||||
`hybrid_search`/`hybrid_search_with` are thin wrappers; `ClawhdfBackend`
|
||||
(the `openclaw` module) is `search` with re-rank + confidence on.
|
||||
- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not an
|
||||
OpenClaw memory plugin and never was — the old `memory.backend = "clawhdf5"`
|
||||
config was never valid. Don't reintroduce OpenClaw claims; `docs/openclaw.md`
|
||||
records what a real plugin would need.
|
||||
- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream
|
||||
v0.8.5 and the `osobh/zeroclaw` fork, and their full history): no
|
||||
`clawhdf5` feature or backend exists; ZeroClaw's memory backends are
|
||||
sqlite/lucid/postgres/qdrant/markdown/none behind its own `Memory` trait.
|
||||
`clawhdf5-migrate`'s default SQLite layout (`memory_chunks`, `sessions`,
|
||||
`entities`, `relations`) is not ZeroClaw's schema either (ZeroClaw's is a
|
||||
`memories` table). Don't reintroduce integration claims without an
|
||||
integration and a test against the real consumer. Measure changes with
|
||||
`search_harness --options-study`.
|
||||
- `MemoryConfig::compression` is off by default; when on, embeddings are
|
||||
deflate-compressed, or Zstd with the agent's `zstd` feature (links libzstd).
|
||||
- Signed checkpoints (`clawhdf5-agent` `signing` module): with
|
||||
`HDF5Memory::set_signing_key` every checkpoint stores an Ed25519-signed
|
||||
manifest (SHA-256 per record in a Merkle tree + settings/sessions/graph
|
||||
hashes; per-record hashes in `/integrity/record_hashes`);
|
||||
`HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover exactly
|
||||
what the file persists in the form the loader returns it (strings lose
|
||||
trailing NULs; an empty WAL mark is not written) or untouched stores stop
|
||||
verifying — `tests/signed_store.rs` round-trips awkward strings. The key is
|
||||
never persisted; a signed store refuses to checkpoint without it
|
||||
(`MemoryError::SigningKeyRequired`, and `MemoryError` is `#[non_exhaustive]`).
|
||||
`hybrid_search`/`hybrid_search_with` are thin wrappers. It keeps one
|
||||
incremental BM25 index for the life of the store and never writes the
|
||||
store: Hebbian activation boosts are persisted by the next checkpoint (or
|
||||
on drop).
|
||||
- **HNSW** (`hnsw` feature, default): the approximate `clawhdf5-ann` index
|
||||
mirrors the cache and self-heals on drift; build the agent with
|
||||
`--no-default-features --features float16` for the exact linear scan.
|
||||
`parallel` (default) builds it on a thread pool with an identical graph.
|
||||
Neighbour selection uses the HNSW paper's diversity heuristic (closest-M
|
||||
capped recall at 0.31 recall@10 at 100K on clustered data). The graph is
|
||||
saved to `<store>.h5.ann` at each checkpoint, tied to it by a generation
|
||||
id; a stale or damaged sidecar is ignored and the index rebuilt.
|
||||
- **`MemoryConfig::quantized_index`** (default on for new stores, persisted;
|
||||
older stores load as `false` — guarded by `tests/fixtures/store_v2_5_0.h5`;
|
||||
CLI `create --f32-index`): the index's copy of the embeddings is `i8`, and
|
||||
the query path re-scores candidates against the exact embeddings. The
|
||||
aarch64 kernels (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm) are
|
||||
`cfg`'d out on x86, so x86 CI never compiles them — test on real ARM
|
||||
(`rpivision02`, 10.0.2.3, a Pi 5) or rely on the `test-arm64` job.
|
||||
- **`MemoryConfig::float16`** (default on for new stores, persisted; older
|
||||
stores keep `false` — guarded in `tests/float16_store.rs`; CLI `create
|
||||
--f32`): `/memory/embeddings` is IEEE half. `MemoryCache::half_precision`
|
||||
rounds each embedding as it enters the cache (push, update, WAL replay, and
|
||||
load of a store still `f32` on disk) so memory and file agree bit for bit.
|
||||
Values beyond ±65504 are `MemoryError::InvalidEntry`. The agent's
|
||||
`h5py_interop` test guards that a whole store opens in h5py.
|
||||
- **WAL.** Chained CRC32 per entry (a corrupted, reordered, duplicated or
|
||||
spliced entry stops replay cleanly). Header version 4 (`Update` record for
|
||||
`save_or_update`); v3 is upgraded in place, v2 read, v1 only through the
|
||||
one-time migration in `HDF5Memory::open`. Each checkpoint records a
|
||||
`WalMark` in `/meta` so `open()` never applies an entry twice; checkpoints
|
||||
and snapshots are durable as a unit (temp file synced, renamed, directory
|
||||
synced). Individual WAL appends are **not** fsynced (deliberate): saves
|
||||
since the last checkpoint can be lost on power failure or kernel panic.
|
||||
- **Single writer.** `create`/`open` hold an exclusive lock on
|
||||
`<store>.h5.lock` (`MemoryError::Locked` for a second opener);
|
||||
`open_read_only` is a lock-free point-in-time view (CLI `recall`/`stats`/
|
||||
`agents-md`/`export`). An unreadable WAL is quarantined to
|
||||
`<store>.h5.wal.corrupt-<ts>`; a WAL of an unknown newer version fails and
|
||||
is left untouched.
|
||||
- **Signed checkpoints** (`signing` module): with `set_signing_key` each
|
||||
checkpoint stores an Ed25519-signed manifest (per-record SHA-256 in a
|
||||
Merkle tree plus settings/sessions/graph hashes; `/integrity/record_hashes`);
|
||||
`HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover
|
||||
exactly what the file persists in the form the loader returns it (strings
|
||||
lose trailing NULs; an empty WAL mark is not written) —
|
||||
`tests/signed_store.rs` round-trips awkward strings. The key is never
|
||||
persisted; a signed store refuses to checkpoint without it
|
||||
(`MemoryError::SigningKeyRequired`; `MemoryError` is `#[non_exhaustive]`).
|
||||
WAL entries after the checkpoint are not covered.
|
||||
- `Dataset::verify_provenance()` (clawhdf5 facade, `provenance` feature, on by
|
||||
default) recomputes a dataset's SHA-256 and compares it against the
|
||||
`_provenance_sha256` attribute written automatically on save when
|
||||
`DatasetBuilder::with_provenance` is used. It's opt-in per call, not run
|
||||
automatically on open — it decodes and hashes the whole dataset. The hash
|
||||
is unkeyed (tamper-*evident*, not tamper-*proof*): it detects accidental
|
||||
corruption, not a deliberate actor able to modify both the data and the
|
||||
stored hash.
|
||||
- `clawhdf5-agent`'s `HDF5Memory::save`/`save_batch`/`save_or_update` run every
|
||||
write through an in-memory (session-scoped, not persisted to disk)
|
||||
provenance ledger and write-anomaly detector: a content hash per record
|
||||
(`provenance.rs`) for detecting accidental mid-session corruption, plus
|
||||
rate-limit/injection-pattern/source-distribution checks (`anomaly.rs`).
|
||||
Alerts never block a save — drain them with `HDF5Memory::take_anomaly_alerts`.
|
||||
`MemorySource` for this bookkeeping is inferred from the caller-supplied
|
||||
`source_channel` string (a heuristic, not an authenticated trust boundary).
|
||||
- In-place modification: `clawhdf5::FileEditor` (`crates/clawhdf5/src/edit/`)
|
||||
overwrites values, grows and shrinks chunked datasets (every chunk index,
|
||||
version-2 B-trees included) and sets attributes (compact and dense
|
||||
storage) in existing files (h5py- or clawhdf5-written) without rewriting
|
||||
them, changing indexes and heaps as libhdf5 does (index shapes and heap
|
||||
bookkeeping are compared with libhdf5's in the tests); space an edit
|
||||
frees is reused by later edits of the same editor. Anything it cannot do
|
||||
safely is `Error::Unsupported` before any write (limits in
|
||||
`docs/known-issues.md`). Test changes with
|
||||
`cargo test -p clawhdf5-tools --test edit_interop --test
|
||||
edit_coverage_interop` (h5py, h5dump, `h5rs check`, structure comparisons
|
||||
with libhdf5; libhdf5 sources for the algorithms are at
|
||||
github.com/HDFGroup/hdf5, tag `hdf5_1_14_6`).
|
||||
- Remote files (`clawhdf5-remote`, range-read milestone M3 of
|
||||
`docs/design/range-reads.md`): `open_url("http://…")` gives a
|
||||
`clawhdf5::File` over `File::open_storage`, read through `BlockCache`
|
||||
(1 MiB blocks, LRU byte budget, per-block in-flight dedup across threads,
|
||||
runs coalesced into parallel requests). `HttpStorage` pins the file by
|
||||
ETag/Last-Modified and length (a change is `RemoteError::FileChanged`),
|
||||
refuses servers that ignore `Range` unless a full download is allowed,
|
||||
and retries transient failures. `ObjectStoreStorage` (feature
|
||||
`object-store`, pure Rust) runs each read on a small owned tokio
|
||||
runtime and waits on a channel, so it works from any thread, including
|
||||
inside `spawn_blocking` or another runtime. Default build is plain HTTP with
|
||||
no C; `https` (rustls + ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are
|
||||
opt-in. Tests run a std-only HTTP server
|
||||
(`tests/common/server.rs`, also the `range_server` example);
|
||||
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
|
||||
file over HTTP with `File::open`.
|
||||
- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only
|
||||
- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since
|
||||
they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds
|
||||
the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range
|
||||
requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a
|
||||
call is re-run after each wave of misses; no block evicted while a call
|
||||
runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh`
|
||||
builds the package (needs the `wasm-bindgen` CLI at the crate's exact
|
||||
version) and tests it under Node and headless Chromium (a Playwright
|
||||
download in `~/.cache/ms-playwright` on tank) against `test/serve.py`
|
||||
(range server with request counts, 200 MB budget file); the CI container
|
||||
has neither, so CI runs the native `h5py_interop` and `lazy` tests
|
||||
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size
|
||||
numbers in the example's README predate `openUrl`.
|
||||
- Python and Node.js bindings for cross-language use
|
||||
- NetCDF-4 compatibility for scientific data interop
|
||||
- **Write bookkeeping.** `save`/`save_batch`/`save_or_update` feed an
|
||||
in-memory, session-scoped provenance ledger and anomaly detector
|
||||
(`provenance.rs`, `anomaly.rs`); alerts never block a save
|
||||
(`take_anomaly_alerts`). `MemorySource` is inferred from the caller's
|
||||
`source_channel` string — a heuristic, not a trust boundary.
|
||||
- `MemoryConfig::compression` is off by default (deflate, or Zstd with the
|
||||
agent's `zstd` feature, which links libzstd).
|
||||
|
||||
## Workflows
|
||||
|
||||
### Build
|
||||
Put `$HOME/.cargo/bin` on `PATH`. The h5py/netCDF4 interop tests find their
|
||||
Python through `CLAWHDF5_PYTHON` (or `.venv/bin/python`); create it with
|
||||
`python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 hdf5plugin`.
|
||||
Set `CLAWHDF5_REQUIRE_INTEROP=1` to make a missing interpreter a failure.
|
||||
|
||||
```bash
|
||||
cargo build --release
|
||||
```
|
||||
|
||||
### Test
|
||||
```bash
|
||||
cargo test --workspace
|
||||
bash scripts/ci-test.sh # everything CI runs (see below)
|
||||
```
|
||||
|
||||
### CI
|
||||
`.gitea/workflows/ci.yml` has two jobs, both green as of 2026-09-22:
|
||||
- **`test`** (`ubuntu-latest`, in `rust:latest`) runs `scripts/ci-test.sh` with
|
||||
the h5py/netCDF4 interop suites required (`CLAWHDF5_REQUIRE_INTEROP=1`).
|
||||
Served by the `tank` and `architect` runners.
|
||||
- **`test-arm64`** (`linux_arm64`) lints and tests the aarch64 code — the NEON
|
||||
kernels are `cfg`'d out on x86, so this is the only place they are built.
|
||||
Served by `vision-01` (host mode) and `vision-02` (Docker), so steps must
|
||||
work in both.
|
||||
### CI (`.gitea/workflows/`)
|
||||
- **`ci.yml` `test`** (`ubuntu-latest`, `rust:latest` container; runners
|
||||
`tank`, `architect`): installs h5py/netCDF4/xarray/hdf5plugin/maturin/pytest,
|
||||
`hdf5-tools` and `cmake`, then runs `scripts/ci-test.sh` with
|
||||
`CLAWHDF5_REQUIRE_INTEROP=1`. The script runs: fmt; clippy (workspace, the
|
||||
format feature matrix, each plugin filter alone, parallel, fast-deflate,
|
||||
remote with all backends, h5rs remote); "no C in the default build";
|
||||
wasm32 build and clippy; `check-32bit-casts.sh`; the wasm package under
|
||||
Node when `node` and `wasm-bindgen` exist (not in CI); the MSRV check;
|
||||
`cargo test` (workspace plus feature variants: format matrix, parallel,
|
||||
remote/object_store, h5rs URLs, ann parallel, fast-deflate); the h5py
|
||||
interop suites (`writer_h5py_tests --include-ignored`, plugin filters,
|
||||
ZFP); the Python package (clippy, `maturin build`, pytest vs h5py);
|
||||
`cargo bench --no-run`; `check-nostd.sh`; an optional fuzz smoke run
|
||||
(`CLAWHDF5_FUZZ_SECONDS`).
|
||||
- **`ci.yml` `test-arm64`** (`linux_arm64`; `vision-01` host mode,
|
||||
`vision-02` Docker — steps must work in both): clippy and tests of
|
||||
`clawhdf5-accel`, `-ann`, `-format`; the only place the NEON kernels build.
|
||||
- **`conformance.yml`** (nightly 03:17 UTC and manual): probe unit tests,
|
||||
`conformance/test_ref.py`, then `conformance/run.sh` (gate:
|
||||
`conformance/check.py` against `baseline.json`).
|
||||
|
||||
Keep workflows free of JavaScript actions (`actions/checkout`, `actions/cache`,
|
||||
…): `rust:latest` has no `node`, and not every runner reaches GitHub, where
|
||||
they are fetched from. Check out with plain `git` instead. The `test` job
|
||||
installs `cmake` for the opt-in `fast-deflate` (zlib-ng) steps; the default
|
||||
build needs no C toolchain, so `test-arm64` does not.
|
||||
All runners are on `gitea-runner` 3.5.0, from `docker.gitea.com/act_runner`
|
||||
— `gitea/act_runner:latest` on Docker Hub is frozen at 0.6.1.
|
||||
Keep workflows free of JavaScript actions (`actions/checkout`,
|
||||
`actions/cache`, …): `rust:latest` has no `node` and not every runner reaches
|
||||
GitHub. Check out with plain `git`. Runners are `gitea-runner` 3.5.0 from
|
||||
`docker.gitea.com/act_runner` (`gitea/act_runner:latest` on Docker Hub is
|
||||
frozen at 0.6.1).
|
||||
|
||||
### CLI
|
||||
### Conformance
|
||||
```bash
|
||||
cargo run -p clawhdf5-cli -- --help
|
||||
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands
|
||||
CLAWHDF5_PYTHON=.venv/bin/python bash conformance/run.sh --no-fetch # writes CONFORMANCE.md
|
||||
```
|
||||
Reads 697 files of eight pinned corpora with clawhdf5 and h5py and compares
|
||||
them object by object (602 ok as of 2026-09-27). `CONFORMANCE.md` is
|
||||
generated — never hand-edit it (its wording lives in `conformance/report.py`).
|
||||
Use `--update-baseline` only after an intended change in results.
|
||||
`CONFORMANCE_CACHE` points at an existing corpus cache (`conformance/.cache`,
|
||||
about 450 MB). See `conformance/README.md`.
|
||||
|
||||
### HDF5 tools (`h5rs`, crate `clawhdf5-tools`)
|
||||
### HDF5 tools (`h5rs`)
|
||||
```bash
|
||||
cargo run -p clawhdf5-tools -- ls -r file.h5 # also dump [--json], stat, diff, check
|
||||
bash scripts/h5rs-fuzz.sh # every subcommand over the CVE corpus: no panic/crash/hang
|
||||
bash scripts/h5rs-check-ok-files.sh --data # check passes every fully-read conformance file
|
||||
```
|
||||
Its interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian
|
||||
`hdf5-tools`, installed in CI); `dump` must stay byte-identical to h5dump on
|
||||
the test files.
|
||||
Interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian `hdf5-tools`);
|
||||
`dump` must stay byte-identical to h5dump on the test files.
|
||||
|
||||
### Remote and browser tests
|
||||
- `clawhdf5-remote` tests run a std-only HTTP server
|
||||
(`tests/common/server.rs`, also the `range_server` example);
|
||||
`CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus
|
||||
file over HTTP with `File::open`.
|
||||
- wasm: `bash examples/wasm-viewer/test/run.sh` builds the package (needs the
|
||||
`wasm-bindgen` CLI at the crate's exact version) and tests it under Node and
|
||||
headless Chromium (Playwright's download in `~/.cache/ms-playwright` on
|
||||
tank) against `test/serve.py` (range server with request counts). CI has
|
||||
neither, so it runs the native `h5py_interop` and `lazy` tests
|
||||
(`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus).
|
||||
|
||||
### Python bindings
|
||||
```bash
|
||||
cd crates/clawhdf5-py
|
||||
maturin develop
|
||||
python -c "import clawhdf5; print(clawhdf5.__version__)"
|
||||
cd crates/clawhdf5-py && maturin develop
|
||||
python -m pytest crates/clawhdf5-py/tests # compares with h5py; editing tests want CLAWHDF5_H5RS=<path to h5rs>
|
||||
```
|
||||
|
||||
### Benchmarks
|
||||
- Search path: `cargo run --release -p clawhdf5-bench --bin search_harness`
|
||||
(`--full`, `--options-study`, `--footprint`, …); reads: `read_harness`,
|
||||
`concurrent_read`; criterion benches with `cargo bench -p <crate>`.
|
||||
- Run on an idle machine (1-minute load average below 2; wait otherwise),
|
||||
alternate base and candidate binaries for A/B comparisons, and record date,
|
||||
machine, commit and command with every number in `BENCHMARKS.md`.
|
||||
- `scripts/run-benchmarks.sh` is stale (it benchmarks `rustyhdf5-format` and
|
||||
overwrites `BENCHMARKS.md`) — do not run it.
|
||||
|
||||
### CLI
|
||||
```bash
|
||||
cargo run -p clawhdf5-cli -- --help
|
||||
# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot
|
||||
```
|
||||
|
||||
## Integration
|
||||
@@ -255,8 +268,7 @@ python -c "import clawhdf5; print(clawhdf5.__version__)"
|
||||
verified consumer: `cbh-core` reads and writes `.brain` files through the
|
||||
facade (`File`, `FileBuilder`, `AttrValue`, `Selection`), `cbh-scanner`
|
||||
uses the facade, and `cbh-cli` uses `clawhdf5_agent::bm25::BM25Index`. It
|
||||
depends on this repo by path (`../clawhdf5`), so it builds against whatever
|
||||
is checked out — changes to those APIs reach it directly. Verified
|
||||
2026-09-25 against main: builds, and its 204 tests pass.
|
||||
- OpenClaw and ZeroClaw were both described as consumers; neither integrates
|
||||
clawhdf5 (see Key Features and `docs/openclaw.md`).
|
||||
depends on this repo by path (`../clawhdf5`), so changes to those APIs
|
||||
reach it directly. Verified 2026-09-25 against main: builds, and its 204
|
||||
tests pass.
|
||||
- OpenClaw and ZeroClaw integrate nothing (see *Standing rules*).
|
||||
|
||||
+40
-36
@@ -1,40 +1,33 @@
|
||||
# Design: range reads (reading HDF5 without holding the whole file)
|
||||
|
||||
Status: proposal, 2026-09-26; the plan for Phase 3's largest architectural
|
||||
change. Progress: M0 and M1 are done, and so is M2 (branch
|
||||
`feat/p3-m2-raw-data`): every read path of the format crate works through
|
||||
`Storage`, v2 B-trees, dense groups and raw data included, and
|
||||
`File::open_storage` gives the facade's read API over any `Storage` (see
|
||||
`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch
|
||||
`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S),
|
||||
object stores) and URLs in `h5rs` (see the M3 status below). M5 (SWMR) is
|
||||
done on branch `feat/p3-m5-swmr-reader`, with its own design in
|
||||
[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is
|
||||
next. Every count in §1–§2 was
|
||||
Status (updated 2026-09-28): **implemented and merged.** Proposed
|
||||
2026-09-26 as the plan for Phase 3's largest architectural change; every
|
||||
milestone below is on `main`:
|
||||
|
||||
object stores) and URLs in `h5rs` (see the M3 status below); the Python
|
||||
bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27),
|
||||
which completes M3. M4 (wasm) is next. Every count in §1–§2 was
|
||||
| Milestone | What | Merged |
|
||||
|---|---|---|
|
||||
| M0 | indexed name lookups, checked address conversion | PR #17 (`8f59b2e`) |
|
||||
| M1 | metadata parsers over the `Storage` trait | PR #17 (`8f59b2e`) |
|
||||
| M2 | raw data over `Storage`, `File::open_storage` | PR #18 (`a4c2ace`) |
|
||||
| M3 | `clawhdf5-remote` (block cache, HTTP(S), object stores), URLs in `h5rs`; Python `clawhdf5.File(url)` | PR #18 (`a4c2ace`); Python in PR #19 (`7a8fae0`) |
|
||||
| M4 | wasm `openUrl` through the restartable `NeedBytes` mode | PR #19 (`7a8fae0`); fewer round trips in PR #21 (`9b5803f`) |
|
||||
| M5 | SWMR reader (`File::open_swmr`), design in [`swmr.md`](swmr.md) | PR #19 (`7a8fae0`) |
|
||||
|
||||
change. Progress: M1, first part (the `Storage` trait and the metadata
|
||||
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
|
||||
group B-tree v2 lookups, dense groups and the facade are not converted yet.
|
||||
Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||
Each milestone's own *Status* note in §4 records what was built and how it
|
||||
differs from the plan. What is still missing is tracked in
|
||||
[`docs/known-issues.md`](../known-issues.md) ("Range reads", "Remote
|
||||
files" and "`clawhdf5-wasm`" limits); the main gaps are a paged file's page
|
||||
size as the block size, remote SWMR, and a SWMR writer.
|
||||
|
||||
object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on
|
||||
branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser
|
||||
reader, through the restartable `NeedBytes` mode (see the M4 status
|
||||
below). M5 (SWMR) is not started.
|
||||
Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched
|
||||
converted code without changing the plan: object-header continuation chunks
|
||||
are followed without recursion (still one bounded `read_at` per chunk), and
|
||||
implicit chunk indexes are addressed over the maximum chunk grid (in
|
||||
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps
|
||||
working on the whole file in memory; it is not part of this design. Every
|
||||
count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and
|
||||
every count in a milestone's status on the date it gives, with the
|
||||
commands given next to it. No timing numbers appear here on purpose: the machine was shared
|
||||
with other build jobs when this was written.
|
||||
`chunked_read`, an M2 module). The in-place editor (`FileEditor`) is not part
|
||||
of this design. Every count in §1–§2 was taken on `tank` on 2026-09-26 at
|
||||
commit `de2a53f`, and every count in a milestone's status on the date it
|
||||
gives, with the commands given next to it. No timing numbers appear here on
|
||||
purpose: the machine was shared with other build jobs when this was written.
|
||||
|
||||
## The problem
|
||||
|
||||
@@ -398,7 +391,8 @@ Rejected. It is how one would retrofit a C library that cannot change; we can.
|
||||
Adopt **(a)**, with a block cache as a required part of every non-local
|
||||
backend, **(c)** as a cache policy, and the wasm path through the restartable
|
||||
`NeedBytes` mode. Every milestone keeps `main` green: `cargo test
|
||||
--workspace`, clippy, the conformance gate at 575/697 unchanged, and the mmap
|
||||
--workspace`, clippy, the conformance gate unchanged (575/697 when this was written; 602/697
|
||||
after PR #21), and the mmap
|
||||
fast path within benchmark noise.
|
||||
|
||||
**M0 — prerequisites (≈1 week).**
|
||||
@@ -414,7 +408,8 @@ fast path within benchmark noise.
|
||||
n children decodes its links O(n) times. Look names up through the index
|
||||
(above) and let a listing hand out its entries, so the cache has less to
|
||||
absorb.
|
||||
- *Status 2026-09-26:* done on branch `perf/p3-indexed-lookups` — link and
|
||||
- *Status 2026-09-26:* done on branch `perf/p3-indexed-lookups` (merged in
|
||||
PR #17) — link and
|
||||
attribute names through the name indexes (`group_v2::resolve_child`,
|
||||
`attribute::find_attribute_in_file`; creation-order lookups by name do
|
||||
not exist in the API, so the creation-order index is still only listed),
|
||||
@@ -438,6 +433,11 @@ fast path within benchmark noise.
|
||||
(`fn parse(data: &[u8], ..) { parse_in(data, ..) }`, generic core), so
|
||||
callers and the other crates don't move yet.
|
||||
- Replace the 5 open-ended slices and 38 `len()` checks with bounded reads.
|
||||
- *Status 2026-09-26:* done on branch `feat/p3-storage-trait` (merged in
|
||||
PR #17) for the `Storage` trait and the metadata parsers listed in
|
||||
`CHANGELOG.md` under "Range reads, milestone M1"; group B-tree v2 lookups,
|
||||
dense groups and the facade were converted in M2. Both error enums are
|
||||
`#[non_exhaustive]`; storage failures are `FormatError::Storage`.
|
||||
|
||||
**M2 — raw data over the trait (1–2 weeks).**
|
||||
- `data_read`, `chunked_read`, `parallel_read`, `partial_read`, `vds`,
|
||||
@@ -449,7 +449,8 @@ fast path within benchmark noise.
|
||||
on other backends (they already return `Option`/`Result`).
|
||||
- Facade: `File::open_storage(Box<dyn Storage + Send + Sync>)`; `File::open`
|
||||
keeps mmap and `from_bytes` keeps `Vec`, both through `impl Storage for [u8]`.
|
||||
- *Status 2026-09-26:* done on branch `feat/p3-m2-raw-data`. As planned,
|
||||
- *Status 2026-09-26:* done on branch `feat/p3-m2-raw-data` (merged in PR
|
||||
#18). As planned,
|
||||
with these choices:
|
||||
- `File::open_storage` takes an `Arc<dyn Storage + Send + Sync>` (the
|
||||
file handle is shared by its datasets and may be sent across threads).
|
||||
@@ -487,7 +488,8 @@ fast path within benchmark noise.
|
||||
§2; the page size for paged files; the first block prefetched on open) and
|
||||
a request counter exposed for tests and users.
|
||||
- Python bindings: `clawhdf5.File("s3://…")` / `https://` through it.
|
||||
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the
|
||||
- *Status 2026-09-26:* done on branch `feat/p3-m3-remote` (merged in PR
|
||||
#18), except the
|
||||
Python bindings (done 2026-09-27, below), with these choices:
|
||||
- A new crate, `clawhdf5-remote`, instead of a `remote` feature of
|
||||
`clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io`
|
||||
@@ -526,7 +528,7 @@ fast path within benchmark noise.
|
||||
requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole
|
||||
6.4 MB: 35 001 object headers spread over the file).
|
||||
- *Status 2026-09-27, Python bindings:* done on branch
|
||||
`feat/p3-python-remote-edit`. `clawhdf5.File(url)` and
|
||||
`feat/p3-python-remote-edit` (merged in PR #19). `clawhdf5.File(url)` and
|
||||
`File.open_url(url, **options)` (cache and HTTP options) go through
|
||||
`clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no
|
||||
C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own
|
||||
@@ -543,7 +545,8 @@ fast path within benchmark noise.
|
||||
Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back
|
||||
to a whole download when the server does not answer 206.
|
||||
- `examples/wasm-viewer`: open by URL.
|
||||
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned,
|
||||
- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy` (merged in PR
|
||||
#19), as planned,
|
||||
with these choices:
|
||||
- **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a
|
||||
`Storage` over the blocks fetched so far. A call (open, list, read)
|
||||
@@ -607,7 +610,7 @@ fast path within benchmark noise.
|
||||
missing blocks: listing 3000 datasets went from 185 passes to 6.
|
||||
The 32-bit risk below is covered by a Node test that reads data at
|
||||
3 GiB from a mock server and is refused a 4 GiB file.
|
||||
- Fewer round trips (2026-09-27, later): the walks descend into every
|
||||
- Fewer round trips (2026-09-27, later; merged in PR #21): the walks descend into every
|
||||
child after a failure (not only read the siblings), and parsers
|
||||
call `Storage::hint` for what they read next (node bodies, object
|
||||
header chunks, a dense group's heap blocks, a listing's child
|
||||
@@ -621,7 +624,8 @@ fast path within benchmark noise.
|
||||
**M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow;
|
||||
add `File::refresh()` that re-reads the superblock/EOF and invalidates cached
|
||||
blocks past the old end. Needs libhdf5 SWMR semantics research first.
|
||||
- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader`; design and
|
||||
- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader` (merged in
|
||||
PR #19); design and
|
||||
libhdf5 research in [`swmr.md`](swmr.md). Differences from the sketch
|
||||
above: the refresh is per dataset (`Dataset::refresh`, as libhdf5's
|
||||
`H5Drefresh`), not per file — a SWMR writer only grows datasets, and the
|
||||
|
||||
+9
-4
@@ -1,7 +1,8 @@
|
||||
# Design: reading files a SWMR writer is still appending to (range-read M5)
|
||||
|
||||
Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader`
|
||||
(see "Status" at the end). This is milestone M5 of
|
||||
Status: design 2026-09-27; the reader is implemented and merged (branch
|
||||
`feat/p3-m5-swmr-reader`, PR #19, `7a8fae0`; see "Status" at the end).
|
||||
clawhdf5 has no SWMR writer. This is milestone M5 of
|
||||
[`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a
|
||||
refresh". It covers the reader only; clawhdf5 does not write SWMR files.
|
||||
|
||||
@@ -165,7 +166,7 @@ writer cannot add them), and `MmapFile`/`LazyFile`.
|
||||
## Status
|
||||
|
||||
Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed
|
||||
above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
|
||||
above, merged to `main` in PR #19 (`7a8fae0`) (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the
|
||||
same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test
|
||||
swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release
|
||||
build): no read returned a value the writer had not written at that
|
||||
@@ -175,4 +176,8 @@ variant of the test with the chunk cache left on in live mode fails it
|
||||
(stale chunk index / edge chunk), which is why live files do not use it.
|
||||
|
||||
Also found: `File::open` of such a file had been failing since the
|
||||
end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`).
|
||||
end-of-file check of 2026-09-26 (item 1; fixed before any release, see
|
||||
[`docs/known-issues.md`](../known-issues.md#files-a-swmr-writer-had-open-could-not-be-read-past-a-stale-end-of-file)).
|
||||
|
||||
Not done (tracked in `docs/known-issues.md`, "Range reads" limits): SWMR
|
||||
writing, remote SWMR, and live reading through `MmapFile`/`LazyFile`.
|
||||
|
||||
+571
-970
File diff suppressed because it is too large
Load Diff
+3
-1
@@ -71,4 +71,6 @@ Building blocks, usable as a library today, but not an OpenClaw plugin:
|
||||
and export rewrites every heading as `##`.
|
||||
- `crates/clawhdf5-napi` and `packages/clawhdf5-node` — Node bindings and a
|
||||
TypeScript wrapper. **Not published, not built or tested in CI, and known to
|
||||
be broken**; see `docs/known-issues.md`.
|
||||
be broken**; see
|
||||
[`docs/known-issues.md`](known-issues.md#the-nodejs-package-packagesclawhdf5-node-does-not-work)
|
||||
(re-checked 2026-09-28: unchanged).
|
||||
|
||||
Reference in New Issue
Block a user