diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 621f912..7910b60 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -51,6 +51,31 @@ target: Criterion stretched it where 5 s could not hold the samples it needed --- +## Current headline numbers + +The newest dated measurement of each headline figure, as of 2026-09-28. +Everything below this section is the dated record behind them; sections whose +figures a later run replaced are marked *Superseded*. Machine "tank" is an AMD +Ryzen 7 7800X3D (8C/16T); rows marked idle were run with the 1-minute load +average below 2. + +| Figure | Value | Measured | Command | Details | +|---|---|---|---|---| +| Agent memory search, `HDF5Memory::hybrid_search` p50 | 0.49 ms at 10K, 4.69 ms at 100K records | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` | [Current: search harness](#current-search-harness-2026-09-24) | +| LongMemEval `longmemeval_s` (full haystack), default hybrid 0.4/0.6, turn-level retrieval Hit@5 (not QA accuracy) | 81.4% | 2026-09-27, tank, search code of `7a8fae0` | `longmemeval_bench … --embeddings weights/all-minilm-l6-v2` | [Re-run with real embeddings](#re-run-with-real-embeddings-2026-09-27-tank), [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500) | +| Loaded store memory, 100K × 384 | 399 MiB (2.72x raw) with the `f32` index; 256 MiB (1.74x) with the int8 index (int8 side not re-run since it was first measured) | 2026-09-24, tank, `5c8323c` | `search_harness -- --footprint --full [--int8]` | [Memory footprint](#memory-footprint), [Quantising the index copy](#quantising-the-index-copy-quantized_index) | +| int8 index vs `f32` index, QPS at equal recall | 1.63x (x86-64 AVX2), 1.18x (Raspberry Pi 5, `SDOT`) | undated; not re-checked against the 2026-09-24 `f32` figure | `search_harness -- --full` | [Quantising the index copy](#quantising-the-index-copy-quantized_index), [On ARM](#on-arm-raspberry-pi-5-cortex-a76) | +| `float16` store file size, 100K × 384 | 80.8 MiB vs 154.0 MiB `f32` (48% smaller) | 2026-09-23, tank | `search_harness -- --float16-study --full` | [float16 embedding storage](#float16-embedding-storage-memoryconfigfloat16) | +| Full reads of chunked deflate data, 16 threads on one `File` | 4944 MB/s, 1.58x 16 h5py processes (noisy run: compare ratios, not MB/s) | 2026-09-26, tank, `c5334b1` | `concurrent_read` + `concurrent_read_h5py.py` | [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) | +| Same, clawhdf5 only, against the build before range-read M2/M3 | 8525 MB/s vs 6258 (+36%); contiguous and metadata reads at parity | 2026-09-27, tank (idle), `7a8fae0` vs `8f59b2e` | `concurrent_read --decode-threads 1 --reps 3` | [Local metadata and data reads after range-read M2/M3](#local-metadata-and-data-reads-after-range-read-m2m3-2026-09-27-tank) | +| `ObjectHeader::parse` (401 headers) | 23.5–23.6 µs, 1.0–2.6% below `8f59b2e` | 2026-09-27, tank (idle), `96086ad` | `cargo bench -p clawhdf5 --bench local_metadata_bench` | [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) | +| Selection reads, 64 MB chunked + deflate `f64` | full 63.2 ms; one 64 × 64 window 0.18 ms | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin read_harness` | [Current: read harness](#current-read-harness-2026-09-24) | +| Deflate backend, zlib-rs (default) vs zlib-ng | within 6% on every HDF5 read/write path | 2026-09-23, tank | `cargo bench -p clawhdf5-filters --bench deflate_bench` (and the two commands with it) | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng) | +| vs libhdf5 1.14.6: chunked deflate-6 write 512×512 / 128 attributes / 64 groups | 45.3x / 10.3x / 10.6x | 2026-08-03, tank | `cargo bench -p clawhdf5-bench --features libhdf5-compare` | [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) | +| Signed checkpoints | about 20% of a checkpoint (598 vs 495 ms at 100K) | 2026-09-25, tank | `search_harness -- --signing-study --full` | [Signed checkpoints](#signed-checkpoints) | + +--- + ## Memory footprint `cargo run --release -p clawhdf5-bench --bin search_harness -- --footprint --full`, @@ -64,6 +89,10 @@ change at all. Measured that way a store holding the corpus twice and one holding it once came out *identical* (1.00x both), which is how the first attempt at this measurement went. +> *Superseded* by the current figures below (2026-09-24): this table is the +> record of the double-copy fix (commit 2e7e045, undated); the store measured +> 2.72x, not 2.43x, by the time the int8 index landed. + | N | vectors (raw) | reopened, before | reopened, after | |---:|---:|---:|---:| | 1 000 | 1 MiB | 5 MiB (3.41x) | 4 MiB (2.39x) | @@ -379,6 +408,9 @@ point: does a selection cost what the *selection* costs? ### Baseline (v2.4.0): every selection decodes the whole dataset +> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24) +> (2026-09-24). Kept as the before picture. + 4096 x 2048 f64 (64 MB per dataset), chunks 256 x 256, file 129 MB | layout | read | selected | time ms | MB/s of selection | vs full read | @@ -404,6 +436,9 @@ point: does a selection cost what the *selection* costs? ### After: partial reads +> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24) +> (2026-09-24). + Only the rows of a contiguous dataset, or the chunks, that overlap the selection's bounding box are read/decoded. A 64 x 64 window of the compressed dataset: **105 -> 0.39 ms**; one row: **106 -> 2.7 ms**; one column: @@ -435,6 +470,9 @@ because the machine's speed drifted; compare the *vs full read* column.) ### After: parallel cached decode, fewer copies (full reads) +> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24) +> (2026-09-24). + Full-read times, old and new binaries run alternately at the same moment (this machine's absolute speed drifts over a long session, so only same-moment comparisons mean anything): @@ -542,6 +580,10 @@ rounds; run 2 also alternated `main` `425585e`. ### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) +> The `object_header_parse_x401` row (+4.2%) is *superseded* by +> [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) +> (2026-09-27, `96086ad`); the other rows are current. + `main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main` `7a8fae0` (PRs #18 and #19), each built in its own worktree and run as separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen @@ -585,8 +627,8 @@ What this shows: - **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header; the base and candidate ranges do not overlap). It is the cost of reading continuation chunks from a bounded queue (the fix for unbounded reads on - crafted headers) and does not show in the listing. Kept open in - `docs/known-issues.md`. + crafted headers) and does not show in the listing. (Fixed later the + same day; see the section above and `docs/known-issues.md`.) - **Full reads of deflate data got faster** after #18 (in-place chunk decoding into the typed output and per-thread scratch buffers): +1.7% on one thread, +36% at 16. @@ -639,6 +681,10 @@ saturate memory bandwidth (1.05x). ### Results after the read fixes (2026-09-26, tank, `408f69e`) +> *Superseded* by [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) +> (2026-09-26, `c5334b1`), which closed the 16-thread gap listed at the end +> of this section. + Same machine, files and commands as the first run below, re-run on an idle tank (load average 1.60 at the start; the 1-minute figure rose to about 5 during the clawhdf5 runs, mostly their own threads) after two fixes: @@ -679,11 +725,17 @@ Read with care: - At 16 threads every tool dropped in this run (h5py threads on contiguous data from 8002 to 2285 MB/s, processes from 12846 to 6942), so the 16-thread rows are noisier than the others. -- Still behind: full reads of chunked data at 16 threads (0.69x-0.76x h5py - processes). See `docs/known-issues.md`. +- Still behind at this commit: full reads of chunked data at 16 threads + (0.69x-0.76x h5py processes); fixed by `c5334b1` (above), recorded as + fixed in `docs/known-issues.md`. ### First run, before the read fixes (2026-09-26, tank, `91644d8`) +> *Superseded* results: the tables and "What this shows" are the before +> picture for [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) +> (2026-09-26). The workload description and the **Run** box below are +> still how every `concurrent_read` figure in this file is produced. + Measured on tank (AMD Ryzen 7 7800X3D, 8 cores / 16 threads, 61 GiB, Linux 7.0) at commit `91644d8`, load average 1.84 when the run started (the 1-minute figure rose to 3.7 during the runs; that is mostly the benchmark's @@ -720,8 +772,8 @@ What this shows: - **clawhdf5 threads on one `File` do, for hyperslab reads of compressed data:** 1244 MB/s at 16 threads, 9.7x h5py threads and 0.89x h5py processes, without a process pool. -- **Where clawhdf5 is behind** (open performance bugs, see - `docs/known-issues.md`): +- **Where clawhdf5 was behind** at `91644d8` (both since fixed; see + `docs/known-issues.md`, "Concurrent and contiguous read performance"): - *Full reads of chunked datasets stop scaling at about 4 threads* (about 880 MB/s) while h5py processes reach 4424 MB/s. Hyperslab reads, which bypass the `File`'s chunk cache, keep scaling, so the @@ -811,6 +863,11 @@ Other flags (both harnesses): `--threads`, `--reps`, `--slab`, `--slabs`, ## Search harness baseline (v2.3.0) +> *Historical.* This baseline and the "After: …" subsections that follow +> record each step of the search work; they are *superseded* by +> [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24), +> the last subsection of this part. + Produced by `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` on deterministic **clustered** synthetic data (384-dim, unit-normalised; points = cluster centre + noise — uniform random vectors are nearly equidistant in high @@ -873,10 +930,11 @@ build: 9752.6 ms (10254 vectors/s) · exact scan: 40 QPS, p50 24648 µs | 1000 | 11 | 3.9 | 0.9 | 68.1 | 5.48 | 5.57 | 182.5 | | 10000 | 114 | 32.2 | 10.9 | 845.0 | 48.56 | 78.65 | 19.8 | | 100000 | 1486 | 713.0 | 354.5 | 10486.5 | 883.51 | 975.23 | 1.1 | -wrote /tmp/claude-1000/-home-osobh-projects-clawhdf5/422f755e-dd25-4c35-8613-5439087e3aaa/scratchpad/baseline_full.json ### After: HNSW neighbour-selection heuristic +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + Same harness, same data, after replacing closest-M neighbour selection with the HNSW paper's diversity heuristic (Algorithm 4, keeping pruned connections) for both new links and back-link pruning. Recall@10 at `ef = 64`: **0.87 → 1.00** @@ -922,6 +980,8 @@ build: 36472.8 ms (2742 vectors/s) · exact scan: 40 QPS, p50 24644 µs ### After: persistent keyword index, no store rewrite per query +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + `hybrid_search` used to rebuild the BM25 index from scratch (re-tokenising every record) and rewrite the whole `.h5` file on **every query**. The index is now kept for the life of the store and updated incrementally, and activation boosts @@ -942,6 +1002,8 @@ index removes that. ### After: vector index persisted with the checkpoint +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + The HNSW graph (not the vectors, which the store already holds) is saved to `.h5.ann` at each checkpoint and reloaded by `open()`, tied to that checkpoint by a generation id. The index is now built once per store (the *cold @@ -959,6 +1021,8 @@ index incrementally. ### After: unit-vector dot product, reusable visited set +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + Cosine distance recomputed both vector norms on every evaluation; the index now stores unit vectors and uses a plain dot product. The per-call `HashSet` of visited nodes became a reusable epoch-stamped array. Recall is unchanged. @@ -1004,6 +1068,8 @@ build: 21084.6 ms (4743 vectors/s) · exact scan: 39 QPS, p50 24739 µs ### After: unranked keyword scores, top-k merge (rankings unchanged) +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + A fusion study (`search_harness --fusion-study`) showed that capping the keyword candidate pool is **not** a safe optimisation: against the current full-corpus normalisation the final top-10 overlap is only 0.83-0.92 and the @@ -1026,6 +1092,8 @@ results. ### After: batched bulk build (optionally parallel); deletions handled in search +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + Profiling showed **90% of a build's distance evaluations are in back-link pruning**. The bulk build now inserts in batches: plan each node's neighbours against the graph as it stood at the start of the batch, link, then prune every @@ -1588,6 +1656,11 @@ MRR, or a one-question change in recency, is within this variation. ### Full haystack — `longmemeval_s`, n=500 (the number to cite) +This table is **BM25-only** (zero embeddings). With real embeddings and the +default hybrid 0.4/0.6 the same corpus gives turn Hit@5 **81.4%** (2026-09-27; +see [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500)), +which is the headline figure. + 47.7 sessions and 493.5 turns per question; 4.0% of haystack sessions are evidence sessions, so retrieval has to actually discriminate. @@ -1695,7 +1768,9 @@ over rank-1 precision. activation. Until now its combined score contained **no relevance term at all** — `RerankInput` did not carry the retrieval score — so a caller that re-ranked its candidates threw the retriever's ordering away and returned them -ordered by age. The OpenClaw backend did exactly that on every search. +ordered by age. `ClawhdfBackend` (the `openclaw` module) did exactly that on +every search. (OpenClaw itself never integrated clawhdf5; see +`docs/openclaw.md`.) Measuring that is unambiguous. "Recency" below is the share of `knowledge-update` questions where the newest gold session outranked the stale @@ -1793,7 +1868,7 @@ worth stating plainly rather than hiding: LongMemEval questions share substantia vocabulary with their evidence turns, which is close to the best case for lexical matching, and MiniLM at 384 dimensions is a small embedding model. -> **Run:** `cargo run --release --bin longmemeval_bench --features embeddings -- \ +> **Run:** `cargo run --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings -- \ > benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings weights/all-minilm-l6-v2` > For the GPU path use `--features embeddings-cuda`. That requires `nvcc` on > `PATH` at *build* time — cudarc's build script shells out to it. The toolkit @@ -2212,7 +2287,8 @@ The tank row was measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D), commit ### Reproducibility ```bash -rustup override set nightly +# Any stable toolchain at or above the MSRV (1.92) works; the original +# 2026-07-01 run used a nightly, later runs stable. # Latency benchmarks (Criterion) cargo bench -p clawhdf5-agent @@ -2315,7 +2391,9 @@ libhdf5 reads from a temp file including `open` + `read` + `close` overhead. | clawhdf5 hyperslab (f64, 10% slice) | — | 4.09 µs / **1.8 GiB/s** | 50.1 µs / **1.5 GiB/s** | libhdf5 f64 comparison excluded — clawhdf5's datatype encoding differs from libhdf5's (known -gap), making cross-format reads unreliable for comparison. +gap), making cross-format reads unreliable for comparison. (That gap was the float sign-bit +bug, fixed 2026-09-23: `docs/known-issues.md`, "Every `f32` dataset we wrote was unreadable by +h5py / libhdf5". The comparison has not been re-run since.) ### Chunked Read Throughput @@ -2423,6 +2501,12 @@ global file mutex and flushes to disk on every attribute write or group creation ## vs libhdf5 Summary +> Measured on the original i7-12650H (clawhdf5 2026-07-01, libhdf5 +> 2026-06-30). The newest run of this table is +> [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) +> (2026-08-03), which reproduces every row within ~15% except chunked +> write (45.3x on tank). + | Workload | clawhdf5 | libhdf5 | Speedup | |----------|----------|---------|---------| | Sequential read, 1K f32 | 634 ns | 45.2 µs | **71×** | @@ -2454,7 +2538,7 @@ to the page cache. There is no algorithmic headroom above ~1.7 GiB/s on this har ### Caveats -- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap). f64 results are clawhdf5-only. +- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap at the time; fixed 2026-09-23, see [Sequential Read Throughput](#sequential-read-throughput)). f64 results are clawhdf5-only. - Serial benchmarks. clawhdf5 uses Rayon for chunk compression when > 2 chunks; that parallelism is already reflected in the chunked write numbers. - clawhdf5 reads from `Vec` (zero-copy from mmap in production); libhdf5 reads from a temp file. This gives clawhdf5 a structural read advantage that reflects realistic API usage. @@ -2651,6 +2735,10 @@ Same not-like-for-like caveat as the "Comparison to MemX" section at the top of file applies — MemX's figure is end-to-end, these are a single component. Ratios are an order-of-magnitude indication, not a benchmark result. +> The Ratio column below was retracted afterwards: see +> [Comparison to MemX](#comparison-to-memx-arxiv260316171). Kept as recorded +> on 2026-08-05; do not cite it. + | Metric | MemX (claimed, end-to-end) | ClawhDF5 (tank, component only) | Ratio | |--------|----------------------------|----------------------------------|-------| | 100K flat search | <90 ms | 6.60 ms | ~14x | diff --git a/CLAUDE.md b/CLAUDE.md index 8c95cc5..cc90aed 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,253 +1,266 @@ # clawhdf5 ## Purpose -Pure-Rust HDF5 format implementation with HNSW vector search, WAL-backed persistence, agent memory storage, and GPU-accelerated vector search. A standalone library. Its one verified consumer is ClawBrainHub (`.brain` files); no agent framework integrates it (OpenClaw and ZeroClaw claims were withdrawn on 2026-09-25 — neither was ever true). +Pure-Rust HDF5 implementation (read, write, in-place edit, remote and browser +reads) plus agent memory on top of it: HNSW vector search, a WAL-backed store, +and GPU vector distances. A standalone library. Its one verified consumer is +ClawBrainHub (`.brain` files); no agent framework integrates it (see +*Standing rules*). ## Architecture -Cargo workspace with 19 crates under `crates/` (plus `libaec-sys`, an internal FFI bindings crate for the optional `szip` feature): +Cargo workspace, 19 crates under `crates/` (plus `libaec-sys`, the FFI crate +behind the optional `szip` feature). MSRV 1.92 (`rust-version`, checked in CI). | Crate | Role | |-------|------| -| `clawhdf5-format` | HDF5 binary spec parser (superblock, B-tree, heap) — also holds shared type definitions and physical constants | -| `clawhdf5-io` | Read/write implementation | -| `clawhdf5-filters` | Deflate backends (zlib-rs, zlib-ng, Apple Compression); the HDF5 filter pipeline, the filter registry (`clawhdf5_format::filter_registry`) and the other codecs (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec, and the pure-Rust plugin filters LZF, bitshuffle, bzip2, Blosc 1, and Blosc2 and ZFP read-only) live in `clawhdf5-format`. | -| `clawhdf5-derive` | Proc-macro derive for HDF5-serializable structs | -| `clawhdf5` | Main facade crate | -| `clawhdf5-netcdf4` | NetCDF-4 compatibility layer | -| `clawhdf5-ann` | HNSW approximate nearest-neighbor vector index | -| `clawhdf5-agent` | Agent memory, session history, knowledge graph storage | -| `clawhdf5-gpu` | GPU vector distance computation via wgpu (hand-written WGSL compute shaders) — not dataset I/O | -| `clawhdf5-accel` | CPU SIMD acceleration path | -| `clawhdf5-migrate` | SQLite → HDF5 agent-memory migration | +| `clawhdf5-format` | The HDF5 format: parsers and writer (superblock, headers, B-trees, heaps, chunk indexes), the `Storage` trait, the filter pipeline and registry (`filter_registry`), every codec except deflate (LZ4, Zstd, SZIP, N-Bit, scale-offset, pcodec; pure-Rust LZF, bitshuffle, bzip2, Blosc 1; Blosc2 and ZFP read-only), `float16`, `checksum` | +| `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) | +| `clawhdf5-io` | I/O adapters (buffers, mmap, prefetch) | +| `clawhdf5-derive` | `#[derive(H5Type)]` for compound types | +| `clawhdf5` | Facade: `File`, `FileBuilder`, `Dataset`, `FileEditor` (`src/edit/`), SWMR reading (`src/swmr.rs`) | +| `clawhdf5-netcdf4` | NetCDF-4 read support | +| `clawhdf5-remote` | `open_url`: HTTP(S) range requests and object stores (S3, GCS, Azure) through `BlockCache` | +| `clawhdf5-tools` | `h5rs`: `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` | +| `clawhdf5-py` | PyO3 bindings (h5py-like API, remote files, `'r+'` editing) | +| `clawhdf5-wasm` | wasm-bindgen browser reader (`open(bytes)`, `openUrl(url)`); demo in `examples/wasm-viewer/` | +| `clawhdf5-ann` | HNSW index | +| `clawhdf5-agent` | Agent memory store (`HDF5Memory`), sessions, knowledge graph, BM25 | +| `clawhdf5-accel` | CPU SIMD kernels (AVX2, NEON) | +| `clawhdf5-gpu` | wgpu vector distances (WGSL) — not dataset I/O; HDF5 I/O is CPU-only | +| `clawhdf5-migrate` | SQLite → agent store migration | +| `clawhdf5-cli` | Agent-memory CLI | +| `clawhdf5-napi` | Node.js addon (the `packages/clawhdf5-node` wrapper is broken; `docs/known-issues.md`) | | `clawhdf5-android` | Android JNI bindings | -| `clawhdf5-cli` | Command-line interface (agent memory) | -| `clawhdf5-tools` | `h5rs`: pure-Rust HDF5 tools — `ls`, `dump` (DDL / hdf5-json), `stat`, `diff`, `check` (structural + checksum validator) | -| `clawhdf5-napi` | Node.js native addon bindings | -| `clawhdf5-py` | PyO3 Python bindings | -| `clawhdf5-wasm` | WebAssembly (wasm-bindgen) reader for the browser; demo in `examples/wasm-viewer/` | -| `clawhdf5-remote` | Remote files: `open_url` over HTTP(S) range requests and object stores (`object_store`: S3, GCS, Azure) through a mandatory block cache (`BlockCache`) | -| `clawhdf5-bench` | Benchmark suite | +| `clawhdf5-bench` | Benchmarks and harnesses (`search_harness`, `read_harness`, `concurrent_read`, `longmemeval_bench`, …) | -## Key Features -- Zero-C-dependency HDF5 read/write: no libhdf5, and deflate defaults to - pure-Rust zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake). - `ci-test.sh` fails if a C-building crate enters the core crates' default - tree. flate2 must keep `runtime_detection` with zlib-rs — without it zlib-rs - loses SIMD and inflates 3.5x slower. MSRV is 1.92 (`rust-version`, checked - in CI). -- HNSW vector index for semantic similarity search over agent memories — the - `clawhdf5-agent` `hnsw` feature is **on by default**, so `hybrid_search` uses - the approximate `clawhdf5-ann` index for the vector stage (the index mirrors - the cache and self-heals on drift). Build the agent with - `--no-default-features --features float16` to force the exact linear cosine scan. - The agent's `parallel` feature (also default) builds the index on a thread - pool; the graph is identical with or without it. - The index uses the HNSW paper's diversity heuristic for neighbour selection - (plain closest-M capped recall on clustered data: 0.31 recall@10 at 100K). Its - graph is saved to `.h5.ann` at each checkpoint and reloaded by `open()` - (tied to the checkpoint by a generation id; stale/damaged sidecars are - ignored and the index rebuilt). `MemoryConfig::quantized_index` (**on by - default** for new stores, persisted; stores predating the setting load as - `false` and keep their f32 index — guarded by - `tests/fixtures/store_v2_5_0.h5`; CLI opt-out is `create --f32-index`) - stores the index's own copy of the embeddings as `i8`, - which roughly halves a loaded store's memory (2.72x -> 1.74x the raw vectors - at 100K); because quantised distances are approximate and `ef` cannot - compensate, the query path then re-scores the candidate pool against the - exact embeddings, which holds recall at the f32 index's level. It is also - faster at equal recall: 1.63x the QPS on x86-64 (AVX2) and 1.18x on a - Raspberry Pi 5 (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm since - the intrinsic is unstable; plain NEON on pre-dotprod cores). The aarch64 - code is `cfg`'d out on x86, so x86 CI never compiles or lints it — test it - on real ARM (`rpivision02`, 10.0.2.3, is a Pi 5). `hybrid_search` keeps one incremental BM25 - index for the life of the store and never writes the store: Hebbian - activation boosts are persisted by the next checkpoint (or on drop), not per - query. Measure any search-path change with - `cargo run --release -p clawhdf5-bench --bin search_harness` (baselines in - `BENCHMARKS.md`). -- WAL (write-ahead log) for crash-safe persistence, with a chained CRC32 - trailer per entry (each entry's CRC folds in the previous entry's CRC) so a - corrupted, reordered, duplicated, or spliced entry stops replay cleanly - instead of loading bad or tampered data. The pre-chaining per-entry-CRC - format (v2) is still fully readable; the oldest no-CRC format (v1) is only - reachable through the one-time migration path in `HDF5Memory::open`, not - through the public `WalFile::read_entries`. - **What the WAL guarantees:** integrity, ordering, and recovery from a - *process* crash at any point — including between a checkpoint and the WAL - truncate (each checkpoint records a `WalMark` in `/meta`, and `open()` skips - the WAL prefix the `.h5` already contains, so entries are never applied - twice). Checkpoints and snapshots are made durable as a unit (temp file - synced, renamed, directory synced). **What it does not guarantee:** - individual WAL appends are *not* fsynced (a deliberate latency trade-off), so - saves made since the last checkpoint can be lost on power failure or kernel - panic. Current header version is 4 (adds the `Update` record used by - `save_or_update`); v3 files are read and upgraded in place. -- A store has a **single writer**: `HDF5Memory::create`/`open` hold an exclusive - advisory lock on `.h5.lock` and a second opener gets - `MemoryError::Locked`. Use `HDF5Memory::open_read_only` for a lock-free, - never-writing point-in-time view (the CLI's `recall`/`stats`/`agents-md`/ - `export` do). An unreadable WAL (torn header, bad magic) is quarantined to - `.h5.wal.corrupt-` rather than blocking `open()`; a WAL with an - unknown *newer* version still fails and is left untouched. -- `MemoryConfig::float16` (**on by default** for new stores, persisted; - existing stores keep their recorded `false` — guarded by the v2.5.0 - fixture in `tests/float16_store.rs`; CLI opt-out is `create --f32`) writes - `/memory/embeddings` as IEEE half precision (48% smaller file at 100K; - LongMemEval with real MiniLM embeddings identical to f32). - `MemoryCache::half_precision` rounds each embedding as it enters the cache (push, update, WAL replay, and on load of a store still - `f32` on disk), so memory and file agree bit for bit; the conversions live - in `clawhdf5_format::float16` and must stay the single implementation. - Values beyond ±65504 are `MemoryError::InvalidEntry`. Interop: every file - must open in h5py — `f32` datasets and empty datasets did not until - 2026-09-23 (see `docs/known-issues.md`); the agent's `h5py_interop` test - guards a whole store. -- `HDF5Memory::search(query_emb, text, &SearchOptions)` is the full search - path: optional source-channel filter (applied before ranking; exact scan of - the allowed records whenever cheaper than `pool × M` index distance +Reference docs: `docs/known-issues.md` (open issues table first — check it +before calling something a bug or a feature), `BENCHMARKS.md` (headline +numbers first), `CONFORMANCE.md` (generated), `docs/design/range-reads.md` +and `docs/design/swmr.md`, `CHANGELOG.md` (full detail of every fix). + +## Standing rules + +- **No C in the default build.** No libhdf5; deflate defaults to pure-Rust + zlib-rs (`fast-deflate` opts into zlib-ng, which needs cmake). `ci-test.sh` + fails if a C-building crate enters the core crates' default tree. Zstd, + SZIP, `https` (ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are opt-in. flate2 + must keep `runtime_detection` with zlib-rs — without it zlib-rs loses SIMD + and inflates 3.5x slower. +- **Every file we write must open in h5py/libhdf5.** Interop tests compare + against h5py and h5dump; `f32` and empty datasets did not open until + 2026-09-23. +- **float16 has one implementation:** `clawhdf5_format::float16`. +- **Claims need evidence.** Performance and integration claims in docs must + be measured, dated (with machine and command), or withdrawn. Benchmark + numbers are dated records: never edit a measured value, add a new dated + section and mark the old one superseded. +- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not and + never was an OpenClaw memory plugin; the old `memory.backend = "clawhdf5"` + config was never valid. `docs/openclaw.md` records what a real plugin would + need. The `openclaw` module's `ClawhdfBackend` is just `search` with + re-rank + confidence on. +- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream + v0.8.5 and the `osobh/zeroclaw` fork and their history): its memory + backends are its own; `clawhdf5-migrate`'s default SQLite layout is not + ZeroClaw's schema. Don't reintroduce integration claims without an + integration and a test against the real consumer. +- **known-issues.md:** one entry per bug; when fixed, record it in + `CHANGELOG.md` and move the entry to *Fixed (history)* with date, PR, + affected releases and what users must do — never delete it. + +## HDF5 library: invariants and gotchas + +- **Remote/range reads** (`docs/design/range-reads.md`, M0-M5 merged in PRs + #17-#21): every format-crate read path goes through `Storage` + (`read_at`/`read_ranges`/`hint`). `File::open_storage` takes any + `Storage`; `clawhdf5_remote::open_url` wraps HTTP (`HttpStorage`, ureq) or + `ObjectStoreStorage` in `BlockCache` (1 MiB blocks, LRU budget, in-flight + dedup, coalesced runs). Remote files are pinned by ETag/Last-Modified and + length (`RemoteError::FileChanged`). Zero-copy APIs and `File::as_bytes` + need an in-memory file. Parse through `File::storage()` and the `*_in` + functions, not `as_bytes`, in new code (the Python bindings do). + `ObjectStoreStorage` runs reads on its own small tokio runtime, so it + works from any thread. +- **SWMR** (`docs/design/swmr.md`): `File::open_swmr` reads a file a libhdf5 + SWMR writer is appending to — positioned reads, no chunk cache, bounded + retries (100), `Dataset::refresh()`. clawhdf5 has no SWMR writer; remote + SWMR is out of scope. +- **Browser** (`clawhdf5-wasm`, read-only, no Zstd/SZIP): `openUrl` reads + through the restartable "NeedBytes" cache (`src/lazy.rs`: a call is re-run + after each wave of misses; no block is evicted while a call runs); the HTTP + is JavaScript (`js/remote.js`). +- **In-place editing** (`clawhdf5::FileEditor`): overwrites values, grows and + shrinks chunked datasets (every chunk index) and sets attributes (compact + and dense) without rewriting the file, changing indexes and heaps as + libhdf5 does; freed space is reused within one editor. Anything it cannot + do safely is `Error::Unsupported` before any write (limits in + `docs/known-issues.md`). The algorithms follow libhdf5 `hdf5_1_14_6` + (github.com/HDFGroup/hdf5). Test changes with `cargo test -p + clawhdf5-tools --test edit_interop --test edit_coverage_interop`. +- **Provenance:** `Dataset::verify_provenance()` (facade `provenance` + feature, default) re-hashes a dataset against its `_provenance_sha256` + attribute (`DatasetBuilder::with_provenance`). Opt-in per call; unkeyed + hash — tamper-evident, not tamper-proof. + +## Agent memory: invariants and gotchas + +- **Search.** `HDF5Memory::search(query_emb, text, &SearchOptions)` is the + full path: optional source-channel filter (before ranking; exact scan of + the allowed records when cheaper than `pool × M` index distance evaluations, and as the fallback when the pool comes back short), fusion, activation scaling, optional re-ranking and confidence rejection. - `hybrid_search`/`hybrid_search_with` are thin wrappers; `ClawhdfBackend` - (the `openclaw` module) is `search` with re-rank + confidence on. -- **OpenClaw is not supported** (decided 2026-09-25): clawhdf5 is not an - OpenClaw memory plugin and never was — the old `memory.backend = "clawhdf5"` - config was never valid. Don't reintroduce OpenClaw claims; `docs/openclaw.md` - records what a real plugin would need. -- **ZeroClaw does not use clawhdf5** (checked 2026-09-25 against upstream - v0.8.5 and the `osobh/zeroclaw` fork, and their full history): no - `clawhdf5` feature or backend exists; ZeroClaw's memory backends are - sqlite/lucid/postgres/qdrant/markdown/none behind its own `Memory` trait. - `clawhdf5-migrate`'s default SQLite layout (`memory_chunks`, `sessions`, - `entities`, `relations`) is not ZeroClaw's schema either (ZeroClaw's is a - `memories` table). Don't reintroduce integration claims without an - integration and a test against the real consumer. Measure changes with - `search_harness --options-study`. -- `MemoryConfig::compression` is off by default; when on, embeddings are - deflate-compressed, or Zstd with the agent's `zstd` feature (links libzstd). -- Signed checkpoints (`clawhdf5-agent` `signing` module): with - `HDF5Memory::set_signing_key` every checkpoint stores an Ed25519-signed - manifest (SHA-256 per record in a Merkle tree + settings/sessions/graph - hashes; per-record hashes in `/integrity/record_hashes`); - `HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover exactly - what the file persists in the form the loader returns it (strings lose - trailing NULs; an empty WAL mark is not written) or untouched stores stop - verifying — `tests/signed_store.rs` round-trips awkward strings. The key is - never persisted; a signed store refuses to checkpoint without it - (`MemoryError::SigningKeyRequired`, and `MemoryError` is `#[non_exhaustive]`). + `hybrid_search`/`hybrid_search_with` are thin wrappers. It keeps one + incremental BM25 index for the life of the store and never writes the + store: Hebbian activation boosts are persisted by the next checkpoint (or + on drop). +- **HNSW** (`hnsw` feature, default): the approximate `clawhdf5-ann` index + mirrors the cache and self-heals on drift; build the agent with + `--no-default-features --features float16` for the exact linear scan. + `parallel` (default) builds it on a thread pool with an identical graph. + Neighbour selection uses the HNSW paper's diversity heuristic (closest-M + capped recall at 0.31 recall@10 at 100K on clustered data). The graph is + saved to `.h5.ann` at each checkpoint, tied to it by a generation + id; a stale or damaged sidecar is ignored and the index rebuilt. +- **`MemoryConfig::quantized_index`** (default on for new stores, persisted; + older stores load as `false` — guarded by `tests/fixtures/store_v2_5_0.h5`; + CLI `create --f32-index`): the index's copy of the embeddings is `i8`, and + the query path re-scores candidates against the exact embeddings. The + aarch64 kernels (`clawhdf5_accel::dot_i8`, NEON `SDOT` via inline asm) are + `cfg`'d out on x86, so x86 CI never compiles them — test on real ARM + (`rpivision02`, 10.0.2.3, a Pi 5) or rely on the `test-arm64` job. +- **`MemoryConfig::float16`** (default on for new stores, persisted; older + stores keep `false` — guarded in `tests/float16_store.rs`; CLI `create + --f32`): `/memory/embeddings` is IEEE half. `MemoryCache::half_precision` + rounds each embedding as it enters the cache (push, update, WAL replay, and + load of a store still `f32` on disk) so memory and file agree bit for bit. + Values beyond ±65504 are `MemoryError::InvalidEntry`. The agent's + `h5py_interop` test guards that a whole store opens in h5py. +- **WAL.** Chained CRC32 per entry (a corrupted, reordered, duplicated or + spliced entry stops replay cleanly). Header version 4 (`Update` record for + `save_or_update`); v3 is upgraded in place, v2 read, v1 only through the + one-time migration in `HDF5Memory::open`. Each checkpoint records a + `WalMark` in `/meta` so `open()` never applies an entry twice; checkpoints + and snapshots are durable as a unit (temp file synced, renamed, directory + synced). Individual WAL appends are **not** fsynced (deliberate): saves + since the last checkpoint can be lost on power failure or kernel panic. +- **Single writer.** `create`/`open` hold an exclusive lock on + `.h5.lock` (`MemoryError::Locked` for a second opener); + `open_read_only` is a lock-free point-in-time view (CLI `recall`/`stats`/ + `agents-md`/`export`). An unreadable WAL is quarantined to + `.h5.wal.corrupt-`; a WAL of an unknown newer version fails and + is left untouched. +- **Signed checkpoints** (`signing` module): with `set_signing_key` each + checkpoint stores an Ed25519-signed manifest (per-record SHA-256 in a + Merkle tree plus settings/sessions/graph hashes; `/integrity/record_hashes`); + `HDF5Memory::verify(path, &pk)` locates edits. The hashes must cover + exactly what the file persists in the form the loader returns it (strings + lose trailing NULs; an empty WAL mark is not written) — + `tests/signed_store.rs` round-trips awkward strings. The key is never + persisted; a signed store refuses to checkpoint without it + (`MemoryError::SigningKeyRequired`; `MemoryError` is `#[non_exhaustive]`). WAL entries after the checkpoint are not covered. -- `Dataset::verify_provenance()` (clawhdf5 facade, `provenance` feature, on by - default) recomputes a dataset's SHA-256 and compares it against the - `_provenance_sha256` attribute written automatically on save when - `DatasetBuilder::with_provenance` is used. It's opt-in per call, not run - automatically on open — it decodes and hashes the whole dataset. The hash - is unkeyed (tamper-*evident*, not tamper-*proof*): it detects accidental - corruption, not a deliberate actor able to modify both the data and the - stored hash. -- `clawhdf5-agent`'s `HDF5Memory::save`/`save_batch`/`save_or_update` run every - write through an in-memory (session-scoped, not persisted to disk) - provenance ledger and write-anomaly detector: a content hash per record - (`provenance.rs`) for detecting accidental mid-session corruption, plus - rate-limit/injection-pattern/source-distribution checks (`anomaly.rs`). - Alerts never block a save — drain them with `HDF5Memory::take_anomaly_alerts`. - `MemorySource` for this bookkeeping is inferred from the caller-supplied - `source_channel` string (a heuristic, not an authenticated trust boundary). -- In-place modification: `clawhdf5::FileEditor` (`crates/clawhdf5/src/edit/`) - overwrites values, grows and shrinks chunked datasets (every chunk index, - version-2 B-trees included) and sets attributes (compact and dense - storage) in existing files (h5py- or clawhdf5-written) without rewriting - them, changing indexes and heaps as libhdf5 does (index shapes and heap - bookkeeping are compared with libhdf5's in the tests); space an edit - frees is reused by later edits of the same editor. Anything it cannot do - safely is `Error::Unsupported` before any write (limits in - `docs/known-issues.md`). Test changes with - `cargo test -p clawhdf5-tools --test edit_interop --test - edit_coverage_interop` (h5py, h5dump, `h5rs check`, structure comparisons - with libhdf5; libhdf5 sources for the algorithms are at - github.com/HDFGroup/hdf5, tag `hdf5_1_14_6`). -- Remote files (`clawhdf5-remote`, range-read milestone M3 of - `docs/design/range-reads.md`): `open_url("http://…")` gives a - `clawhdf5::File` over `File::open_storage`, read through `BlockCache` - (1 MiB blocks, LRU byte budget, per-block in-flight dedup across threads, - runs coalesced into parallel requests). `HttpStorage` pins the file by - ETag/Last-Modified and length (a change is `RemoteError::FileChanged`), - refuses servers that ignore `Range` unless a full download is allowed, - and retries transient failures. `ObjectStoreStorage` (feature - `object-store`, pure Rust) runs each read on a small owned tokio - runtime and waits on a channel, so it works from any thread, including - inside `spawn_blocking` or another runtime. Default build is plain HTTP with - no C; `https` (rustls + ring) and `s3`/`gcs`/`azure` (aws-lc-rs) are - opt-in. Tests run a std-only HTTP server - (`tests/common/server.rs`, also the `range_server` example); - `CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus - file over HTTP with `File::open`. -- GPU-accelerated vector distance computation (`clawhdf5-gpu`, wgpu); HDF5 I/O itself is CPU-only -- Browser: `clawhdf5-wasm` (wasm-bindgen, read-only; no Zstd/SZIP since - they link C) and the `examples/wasm-viewer/` page. `open(bytes)` holds - the file in memory; `openUrl(url)` (range-read M4) reads it by HTTP range - requests through the restartable "NeedBytes" cache (`src/lazy.rs`: a - call is re-run after each wave of misses; no block evicted while a call - runs), the HTTP in `js/remote.js`. `examples/wasm-viewer/test/run.sh` - builds the package (needs the `wasm-bindgen` CLI at the crate's exact - version) and tests it under Node and headless Chromium (a Playwright - download in `~/.cache/ms-playwright` on tank) against `test/serve.py` - (range server with request counts, 200 MB budget file); the CI container - has neither, so CI runs the native `h5py_interop` and `lazy` tests - (`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). Size - numbers in the example's README predate `openUrl`. -- Python and Node.js bindings for cross-language use -- NetCDF-4 compatibility for scientific data interop +- **Write bookkeeping.** `save`/`save_batch`/`save_or_update` feed an + in-memory, session-scoped provenance ledger and anomaly detector + (`provenance.rs`, `anomaly.rs`); alerts never block a save + (`take_anomaly_alerts`). `MemorySource` is inferred from the caller's + `source_channel` string — a heuristic, not a trust boundary. +- `MemoryConfig::compression` is off by default (deflate, or Zstd with the + agent's `zstd` feature, which links libzstd). ## Workflows -### Build +Put `$HOME/.cargo/bin` on `PATH`. The h5py/netCDF4 interop tests find their +Python through `CLAWHDF5_PYTHON` (or `.venv/bin/python`); create it with +`python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 hdf5plugin`. +Set `CLAWHDF5_REQUIRE_INTEROP=1` to make a missing interpreter a failure. + ```bash cargo build --release -``` - -### Test -```bash cargo test --workspace +bash scripts/ci-test.sh # everything CI runs (see below) ``` -### CI -`.gitea/workflows/ci.yml` has two jobs, both green as of 2026-09-22: -- **`test`** (`ubuntu-latest`, in `rust:latest`) runs `scripts/ci-test.sh` with - the h5py/netCDF4 interop suites required (`CLAWHDF5_REQUIRE_INTEROP=1`). - Served by the `tank` and `architect` runners. -- **`test-arm64`** (`linux_arm64`) lints and tests the aarch64 code — the NEON - kernels are `cfg`'d out on x86, so this is the only place they are built. - Served by `vision-01` (host mode) and `vision-02` (Docker), so steps must - work in both. +### CI (`.gitea/workflows/`) +- **`ci.yml` `test`** (`ubuntu-latest`, `rust:latest` container; runners + `tank`, `architect`): installs h5py/netCDF4/xarray/hdf5plugin/maturin/pytest, + `hdf5-tools` and `cmake`, then runs `scripts/ci-test.sh` with + `CLAWHDF5_REQUIRE_INTEROP=1`. The script runs: fmt; clippy (workspace, the + format feature matrix, each plugin filter alone, parallel, fast-deflate, + remote with all backends, h5rs remote); "no C in the default build"; + wasm32 build and clippy; `check-32bit-casts.sh`; the wasm package under + Node when `node` and `wasm-bindgen` exist (not in CI); the MSRV check; + `cargo test` (workspace plus feature variants: format matrix, parallel, + remote/object_store, h5rs URLs, ann parallel, fast-deflate); the h5py + interop suites (`writer_h5py_tests --include-ignored`, plugin filters, + ZFP); the Python package (clippy, `maturin build`, pytest vs h5py); + `cargo bench --no-run`; `check-nostd.sh`; an optional fuzz smoke run + (`CLAWHDF5_FUZZ_SECONDS`). +- **`ci.yml` `test-arm64`** (`linux_arm64`; `vision-01` host mode, + `vision-02` Docker — steps must work in both): clippy and tests of + `clawhdf5-accel`, `-ann`, `-format`; the only place the NEON kernels build. +- **`conformance.yml`** (nightly 03:17 UTC and manual): probe unit tests, + `conformance/test_ref.py`, then `conformance/run.sh` (gate: + `conformance/check.py` against `baseline.json`). -Keep workflows free of JavaScript actions (`actions/checkout`, `actions/cache`, -…): `rust:latest` has no `node`, and not every runner reaches GitHub, where -they are fetched from. Check out with plain `git` instead. The `test` job -installs `cmake` for the opt-in `fast-deflate` (zlib-ng) steps; the default -build needs no C toolchain, so `test-arm64` does not. -All runners are on `gitea-runner` 3.5.0, from `docker.gitea.com/act_runner` -— `gitea/act_runner:latest` on Docker Hub is frozen at 0.6.1. +Keep workflows free of JavaScript actions (`actions/checkout`, +`actions/cache`, …): `rust:latest` has no `node` and not every runner reaches +GitHub. Check out with plain `git`. Runners are `gitea-runner` 3.5.0 from +`docker.gitea.com/act_runner` (`gitea/act_runner:latest` on Docker Hub is +frozen at 0.6.1). -### CLI +### Conformance ```bash -cargo run -p clawhdf5-cli -- --help -# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot subcommands +CLAWHDF5_PYTHON=.venv/bin/python bash conformance/run.sh --no-fetch # writes CONFORMANCE.md ``` +Reads 697 files of eight pinned corpora with clawhdf5 and h5py and compares +them object by object (602 ok as of 2026-09-27). `CONFORMANCE.md` is +generated — never hand-edit it (its wording lives in `conformance/report.py`). +Use `--update-baseline` only after an intended change in results. +`CONFORMANCE_CACHE` points at an existing corpus cache (`conformance/.cache`, +about 450 MB). See `conformance/README.md`. -### HDF5 tools (`h5rs`, crate `clawhdf5-tools`) +### HDF5 tools (`h5rs`) ```bash cargo run -p clawhdf5-tools -- ls -r file.h5 # also dump [--json], stat, diff, check bash scripts/h5rs-fuzz.sh # every subcommand over the CVE corpus: no panic/crash/hang bash scripts/h5rs-check-ok-files.sh --data # check passes every fully-read conformance file ``` -Its interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian -`hdf5-tools`, installed in CI); `dump` must stay byte-identical to h5dump on -the test files. +Interop tests compare against h5ls/h5stat/h5dump/h5diff (Debian `hdf5-tools`); +`dump` must stay byte-identical to h5dump on the test files. + +### Remote and browser tests +- `clawhdf5-remote` tests run a std-only HTTP server + (`tests/common/server.rs`, also the `range_server` example); + `CLAWHDF5_REMOTE_CORPUS=conformance/.cache/corpus` compares every corpus + file over HTTP with `File::open`. +- wasm: `bash examples/wasm-viewer/test/run.sh` builds the package (needs the + `wasm-bindgen` CLI at the crate's exact version) and tests it under Node and + headless Chromium (Playwright's download in `~/.cache/ms-playwright` on + tank) against `test/serve.py` (range server with request counts). CI has + neither, so it runs the native `h5py_interop` and `lazy` tests + (`CLAWHDF5_WASM_CORPUS=conformance/.cache/corpus` for the corpus). ### Python bindings ```bash -cd crates/clawhdf5-py -maturin develop -python -c "import clawhdf5; print(clawhdf5.__version__)" +cd crates/clawhdf5-py && maturin develop +python -m pytest crates/clawhdf5-py/tests # compares with h5py; editing tests want CLAWHDF5_H5RS= +``` + +### Benchmarks +- Search path: `cargo run --release -p clawhdf5-bench --bin search_harness` + (`--full`, `--options-study`, `--footprint`, …); reads: `read_harness`, + `concurrent_read`; criterion benches with `cargo bench -p `. +- Run on an idle machine (1-minute load average below 2; wait otherwise), + alternate base and candidate binaries for A/B comparisons, and record date, + machine, commit and command with every number in `BENCHMARKS.md`. +- `scripts/run-benchmarks.sh` is stale (it benchmarks `rustyhdf5-format` and + overwrites `BENCHMARKS.md`) — do not run it. + +### CLI +```bash +cargo run -p clawhdf5-cli -- --help +# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot ``` ## Integration @@ -255,8 +268,7 @@ python -c "import clawhdf5; print(clawhdf5.__version__)" verified consumer: `cbh-core` reads and writes `.brain` files through the facade (`File`, `FileBuilder`, `AttrValue`, `Selection`), `cbh-scanner` uses the facade, and `cbh-cli` uses `clawhdf5_agent::bm25::BM25Index`. It - depends on this repo by path (`../clawhdf5`), so it builds against whatever - is checked out — changes to those APIs reach it directly. Verified - 2026-09-25 against main: builds, and its 204 tests pass. -- OpenClaw and ZeroClaw were both described as consumers; neither integrates - clawhdf5 (see Key Features and `docs/openclaw.md`). + depends on this repo by path (`../clawhdf5`), so changes to those APIs + reach it directly. Verified 2026-09-25 against main: builds, and its 204 + tests pass. +- OpenClaw and ZeroClaw integrate nothing (see *Standing rules*). diff --git a/docs/design/range-reads.md b/docs/design/range-reads.md index 9dea022..18ebec4 100644 --- a/docs/design/range-reads.md +++ b/docs/design/range-reads.md @@ -1,40 +1,33 @@ # Design: range reads (reading HDF5 without holding the whole file) -Status: proposal, 2026-09-26; the plan for Phase 3's largest architectural -change. Progress: M0 and M1 are done, and so is M2 (branch -`feat/p3-m2-raw-data`): every read path of the format crate works through -`Storage`, v2 B-trees, dense groups and raw data included, and -`File::open_storage` gives the facade's read API over any `Storage` (see -`CHANGELOG.md`, "Range reads, milestone M2"). M3 is done on branch -`feat/p3-m3-remote`: the `clawhdf5-remote` crate (block cache, HTTP(S), -object stores) and URLs in `h5rs` (see the M3 status below). M5 (SWMR) is -done on branch `feat/p3-m5-swmr-reader`, with its own design in -[`swmr.md`](swmr.md) (see the M5 status below). M4 (wasm) is -next. Every count in §1–§2 was +Status (updated 2026-09-28): **implemented and merged.** Proposed +2026-09-26 as the plan for Phase 3's largest architectural change; every +milestone below is on `main`: -object stores) and URLs in `h5rs` (see the M3 status below); the Python -bindings followed on branch `feat/p3-python-remote-edit` (2026-09-27), -which completes M3. M4 (wasm) is next. Every count in §1–§2 was +| Milestone | What | Merged | +|---|---|---| +| M0 | indexed name lookups, checked address conversion | PR #17 (`8f59b2e`) | +| M1 | metadata parsers over the `Storage` trait | PR #17 (`8f59b2e`) | +| M2 | raw data over `Storage`, `File::open_storage` | PR #18 (`a4c2ace`) | +| M3 | `clawhdf5-remote` (block cache, HTTP(S), object stores), URLs in `h5rs`; Python `clawhdf5.File(url)` | PR #18 (`a4c2ace`); Python in PR #19 (`7a8fae0`) | +| M4 | wasm `openUrl` through the restartable `NeedBytes` mode | PR #19 (`7a8fae0`); fewer round trips in PR #21 (`9b5803f`) | +| M5 | SWMR reader (`File::open_swmr`), design in [`swmr.md`](swmr.md) | PR #19 (`7a8fae0`) | -change. Progress: M1, first part (the `Storage` trait and the metadata -parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done; -group B-tree v2 lookups, dense groups and the facade are not converted yet. -Later the same day (branch `feat/p3-editor-coverage`) two reader fixes touched +Each milestone's own *Status* note in §4 records what was built and how it +differs from the plan. What is still missing is tracked in +[`docs/known-issues.md`](../known-issues.md) ("Range reads", "Remote +files" and "`clawhdf5-wasm`" limits); the main gaps are a paged file's page +size as the block size, remote SWMR, and a SWMR writer. -object stores) and URLs in `h5rs` (see the M3 status below). M4 is done on -branch `feat/p3-m4-wasm-lazy` (2026-09-27): `openUrl` in the browser -reader, through the restartable `NeedBytes` mode (see the M4 status -below). M5 (SWMR) is not started. Also on 2026-09-26 (branch `feat/p3-editor-coverage`) two reader fixes touched converted code without changing the plan: object-header continuation chunks are followed without recursion (still one bounded `read_at` per chunk), and implicit chunk indexes are addressed over the maximum chunk grid (in -`chunked_read`, an M2 module). The in-place editor (`FileEditor`) keeps -working on the whole file in memory; it is not part of this design. Every -count in §1–§2 was taken on `tank` on 2026-09-26 at commit `de2a53f`, and -every count in a milestone's status on the date it gives, with the -commands given next to it. No timing numbers appear here on purpose: the machine was shared -with other build jobs when this was written. +`chunked_read`, an M2 module). The in-place editor (`FileEditor`) is not part +of this design. Every count in §1–§2 was taken on `tank` on 2026-09-26 at +commit `de2a53f`, and every count in a milestone's status on the date it +gives, with the commands given next to it. No timing numbers appear here on +purpose: the machine was shared with other build jobs when this was written. ## The problem @@ -398,7 +391,8 @@ Rejected. It is how one would retrofit a C library that cannot change; we can. Adopt **(a)**, with a block cache as a required part of every non-local backend, **(c)** as a cache policy, and the wasm path through the restartable `NeedBytes` mode. Every milestone keeps `main` green: `cargo test ---workspace`, clippy, the conformance gate at 575/697 unchanged, and the mmap +--workspace`, clippy, the conformance gate unchanged (575/697 when this was written; 602/697 +after PR #21), and the mmap fast path within benchmark noise. **M0 — prerequisites (≈1 week).** @@ -414,7 +408,8 @@ fast path within benchmark noise. n children decodes its links O(n) times. Look names up through the index (above) and let a listing hand out its entries, so the cache has less to absorb. -- *Status 2026-09-26:* done on branch `perf/p3-indexed-lookups` — link and +- *Status 2026-09-26:* done on branch `perf/p3-indexed-lookups` (merged in + PR #17) — link and attribute names through the name indexes (`group_v2::resolve_child`, `attribute::find_attribute_in_file`; creation-order lookups by name do not exist in the API, so the creation-order index is still only listed), @@ -438,6 +433,11 @@ fast path within benchmark noise. (`fn parse(data: &[u8], ..) { parse_in(data, ..) }`, generic core), so callers and the other crates don't move yet. - Replace the 5 open-ended slices and 38 `len()` checks with bounded reads. +- *Status 2026-09-26:* done on branch `feat/p3-storage-trait` (merged in + PR #17) for the `Storage` trait and the metadata parsers listed in + `CHANGELOG.md` under "Range reads, milestone M1"; group B-tree v2 lookups, + dense groups and the facade were converted in M2. Both error enums are + `#[non_exhaustive]`; storage failures are `FormatError::Storage`. **M2 — raw data over the trait (1–2 weeks).** - `data_read`, `chunked_read`, `parallel_read`, `partial_read`, `vds`, @@ -449,7 +449,8 @@ fast path within benchmark noise. on other backends (they already return `Option`/`Result`). - Facade: `File::open_storage(Box)`; `File::open` keeps mmap and `from_bytes` keeps `Vec`, both through `impl Storage for [u8]`. -- *Status 2026-09-26:* done on branch `feat/p3-m2-raw-data`. As planned, +- *Status 2026-09-26:* done on branch `feat/p3-m2-raw-data` (merged in PR + #18). As planned, with these choices: - `File::open_storage` takes an `Arc` (the file handle is shared by its datasets and may be sent across threads). @@ -487,7 +488,8 @@ fast path within benchmark noise. §2; the page size for paged files; the first block prefetched on open) and a request counter exposed for tests and users. - Python bindings: `clawhdf5.File("s3://…")` / `https://` through it. -- *Status 2026-09-26:* done on branch `feat/p3-m3-remote`, except the +- *Status 2026-09-26:* done on branch `feat/p3-m3-remote` (merged in PR + #18), except the Python bindings (done 2026-09-27, below), with these choices: - A new crate, `clawhdf5-remote`, instead of a `remote` feature of `clawhdf5-io`: `open_url` returns a `clawhdf5::File`, and `clawhdf5-io` @@ -526,7 +528,7 @@ fast path within benchmark noise. requests (§2 predicted 2 blocks of 1 MiB), B in 1, C in 7 (its whole 6.4 MB: 35 001 object headers spread over the file). - *Status 2026-09-27, Python bindings:* done on branch - `feat/p3-python-remote-edit`. `clawhdf5.File(url)` and + `feat/p3-python-remote-edit` (merged in PR #19). `clawhdf5.File(url)` and `File.open_url(url, **options)` (cache and HTTP options) go through `clawhdf5_remote::storage_for_url`; the default wheel is plain HTTP (no C), `https`/`s3`/`gcs`/`azure` are build features. The bindings' own @@ -543,7 +545,8 @@ fast path within benchmark noise. Worker, no synchronous XHR — the thing h5wasm's lazy files need). Falls back to a whole download when the server does not answer 206. - `examples/wasm-viewer`: open by URL. -- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy`, as planned, +- *Status 2026-09-27:* done on branch `feat/p3-m4-wasm-lazy` (merged in PR + #19), as planned, with these choices: - **NeedBytes, not a Worker.** `clawhdf5_wasm::lazy::LazyStorage` is a `Storage` over the blocks fetched so far. A call (open, list, read) @@ -607,7 +610,7 @@ fast path within benchmark noise. missing blocks: listing 3000 datasets went from 185 passes to 6. The 32-bit risk below is covered by a Node test that reads data at 3 GiB from a mock server and is refused a 4 GiB file. - - Fewer round trips (2026-09-27, later): the walks descend into every + - Fewer round trips (2026-09-27, later; merged in PR #21): the walks descend into every child after a failure (not only read the siblings), and parsers call `Storage::hint` for what they read next (node bodies, object header chunks, a dense group's heap blocks, a listing's child @@ -621,7 +624,8 @@ fast path within benchmark noise. **M5 — SWMR and growth (later, separate design).** `Storage::len()` may grow; add `File::refresh()` that re-reads the superblock/EOF and invalidates cached blocks past the old end. Needs libhdf5 SWMR semantics research first. -- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader`; design and +- *Status 2026-09-27:* done on branch `feat/p3-m5-swmr-reader` (merged in + PR #19); design and libhdf5 research in [`swmr.md`](swmr.md). Differences from the sketch above: the refresh is per dataset (`Dataset::refresh`, as libhdf5's `H5Drefresh`), not per file — a SWMR writer only grows datasets, and the diff --git a/docs/design/swmr.md b/docs/design/swmr.md index dba542a..4b89845 100644 --- a/docs/design/swmr.md +++ b/docs/design/swmr.md @@ -1,7 +1,8 @@ # Design: reading files a SWMR writer is still appending to (range-read M5) -Status: design 2026-09-27, implemented on branch `feat/p3-m5-swmr-reader` -(see "Status" at the end). This is milestone M5 of +Status: design 2026-09-27; the reader is implemented and merged (branch +`feat/p3-m5-swmr-reader`, PR #19, `7a8fae0`; see "Status" at the end). +clawhdf5 has no SWMR writer. This is milestone M5 of [`range-reads.md`](range-reads.md): "`Storage::len()` may grow; add a refresh". It covers the reader only; clawhdf5 does not write SWMR files. @@ -165,7 +166,7 @@ writer cannot add them), and `MmapFile`/`LazyFile`. ## Status Implemented 2026-09-27 on branch `feat/p3-m5-swmr-reader` as designed -above (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the +above, merged to `main` in PR #19 (`7a8fae0`) (`CHANGELOG.md`, "Range reads, milestone M5"). Observed on tank the same day (h5py 3.16 / HDF5 2.0, `cargo test -p clawhdf5 --test swmr_interop`, and once with `CLAWHDF5_SWMR_STEPS=20000` in a release build): no read returned a value the writer had not written at that @@ -175,4 +176,8 @@ variant of the test with the chunk cache left on in live mode fails it (stale chunk index / edge chunk), which is why live files do not use it. Also found: `File::open` of such a file had been failing since the -end-of-file check of 2026-09-26 (item 1; `docs/known-issues.md`). +end-of-file check of 2026-09-26 (item 1; fixed before any release, see +[`docs/known-issues.md`](../known-issues.md#files-a-swmr-writer-had-open-could-not-be-read-past-a-stale-end-of-file)). + +Not done (tracked in `docs/known-issues.md`, "Range reads" limits): SWMR +writing, remote SWMR, and live reading through `MmapFile`/`LazyFile`. diff --git a/docs/known-issues.md b/docs/known-issues.md index 0b25635..43f945c 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -1,189 +1,39 @@ # Known Issues -Bugs found during development or downstream use, tracked here because this -repository's issue tracker is disabled. One entry per bug; when an entry is -fixed, record the fix in `CHANGELOG.md` and update its status here rather than -deleting it. +Bugs and limits found during development or downstream use, tracked here +because this repository's issue tracker is disabled. One entry per bug; when +an entry is fixed, record the fix in `CHANGELOG.md` and move it to +[Fixed (history)](#fixed-history) with its date, the releases it affected and +what users must do, rather than deleting it. + +Releases referred to below: v2.1.0 (2026-06-03) to v2.7.0 (2026-09-20). +Everything fixed after v2.7.0 is on `main` and unreleased. + +## Open issues + +Checked against `main` at `9b5803f` on 2026-09-28. + +| Issue | Kind | Since | +|---|---|---| +| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | +| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | +| [Selection reads that decode more than the selection](#selection-reads-that-decode-more-than-the-selection) | speed only | 2026-09-26 | +| [HDF5 features still unsupported](#hdf5-features-still-unsupported) | errors, never wrong data | 2026-09-25 audit | +| [External links and external raw data are not followed](#external-links-and-external-raw-data-are-not-followed) | error, by design for now | 2026-09-19 | +| [Range reads (`File::open_storage`) limits](#range-reads-fileopen_storage-limits) | cost; zero-copy APIs need an in-memory file | 2026-09-26 | +| [Remote files (`clawhdf5-remote`) limits](#remote-files-clawhdf5-remote-limits) | untested backends, fixed block size, timeouts | 2026-09-26 | +| [`clawhdf5-wasm` (browser) limits](#clawhdf5-wasm-browser-limits) | memory, round trips, unsupported types | 2026-09-26 | +| [Conformance: `bad_nbit_parms_walk.h5` flips between ref-bug and our-error](#conformance-bad_nbit_parms_walkh5-flips-between-ref-bug-and-our-error) | report noise (ok count unaffected) | 2026-09-28 | +| [The Node.js package does not work](#the-nodejs-package-packagesclawhdf5-node-does-not-work) | broken, unpublished, not in CI | 2026-09-25 | --- -## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27) - -**Status:** fixed 2026-09-27 (`96086ad`, branch -`perf/header-parse-and-last-files`): the remaining cost was the call to the -version-1 message loop, kept out of line by `#[inline(never)]`; inlined, -`object_header_parse_x401` is 1.0% and 2.6% below `8f59b2e` in two idle -A/B runs (`BENCHMARKS.md`, "`ObjectHeader::parse` back at 8f59b2e's -speed"), with every chunk-queue check unchanged. - -Was: open (speed only; values are correct). Measured on an idle tank -(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`) -against `main` `7a8fae0`, separate binaries alternating, 3 rounds -(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"): -`object_header_parse_x401` 23.86 → 24.86 µs median (+4.2%; ranges -23.79–23.96 vs 24.69–24.98), about 2.5 ns per header. The parser has read -continuation chunks from a bounded queue since `a69c5be` (bounding what a -crafted header can make it read); `4313917` removed its per-header -allocations but not all of the cost. Listing the 400-group file through the -facade, which parses the same headers, is 1.7% faster, so no user-visible -path is slower. - -An earlier run the same day under load also listed single-thread contiguous -hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with -overlapping ranges (noise), so that item is withdrawn. - -## Conformance: the last non-ok files (checked 2026-09-27) - -**Status:** classified with evidence — none is a clawhdf5 bug. The -conformance report had 3 our-errors and 2 mismatches left, each documented -as "not ours" but still counted against us. Re-checked on tank -(h5py 3.16 / HDF5 2.0.0, h5dump 1.14.6): - -- **2 mismatches, h5py's big-endian VL bug:** - `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64` and - `hdf5/tools/test/testfiles/tcomplex_be.h5` - `/VariableLengthDatasetFloatComplex`. h5py returns the elements with the - file's big-endian bytes under a little-endian dtype (a `vlen('>f4')` - holding `[1.0, 2.0]` reads back as `[4.6e-41, 9.0e-44]`). h5dump 1.14.6 - prints `(1, 2), (3, 4, 5), (42)` for `vlen_uint64`, which is what we read; - it cannot print `tcomplex_be.h5` (complex types are HDF5 2.0). The - reference side (`conformance/ref.py`) now checks that the installed h5py - has the bug and relabels such elements with the file's byte order, so the - values are compared rather than excused: both objects are identical to - ours, and both files are now *ok*. -- **3 our-errors, objects HDF5 2.0 reads only by over-reading memory:** - `cve-2025-2308.h5` `/Scale_offset_long_long_data_le` (first chunk records - `minbits` 11; its 12 values need 17 bytes of codes and the 26-byte chunk - holds 5 after its 21-byte header), `cve-2025-44904.h5` - `/Scale_offset_float_data_le` (unfiltered chunks stored as 38 and 37 bytes - for 48-byte chunks; 1.14.6's `H5D__chunk_lock` reads the stored bytes - into a buffer of that size and uses it as the whole chunk) and - `bad_nbit_parms_walk.h5` `/Nbit_int_data_le` (N-Bit parameters - `(7, 0, 40, 1, 4, 0, 20)`: an integer needs 8, and the decoder takes the - bit offset from `cd_values[7]`, past the list). **Could we match - libhdf5?** No: its output is not determined by the file. h5py's values for - the first two change from one run to the next (three plain runs gave three - different results), and all three change with `MALLOC_PERTURB_` or with - whether numpy is imported before h5py (`bad_nbit_parms_walk` then fails - with "filter returned failure during read" or reads all zeros); h5dump - 1.14.6 prints yet other values for the over-read parts (`1280, 0, 0` where - h5py gave e.g. `1283, 749, 1713`). libhdf5's develop branch refuses all - three. We keep - refusing them. `conformance/ref_bugs.py` repeats the check in every - conformance run (six reads per object in differently set-up processes) and - such a file is classified *ref-bug* only while its values keep changing. - -Result (`conformance/run.sh --no-fetch`, tank, 2026-09-27): 602 of 697 ok -(was 600), 0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read. - -## Files a SWMR writer had open could not be read past a stale end of file - -**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any -release: reads have been bounded by the recorded end of file since -`7d7a7e7` (2026-09-26), which no release contains. - -A libhdf5 writer in SWMR mode (h5py `f.swmr_mode = True`) sets the -superblock's SWMR-write flag and does not keep its end-of-file address up to -date: a copy h5py made of its own file mid-write records 715 in a -6 030-byte file (h5py 3.16 / HDF5 2.0, tank). `Superblock::data_end` only -ignored the recorded end when it lay *past* the end of the file, so every -reader bounded such a file at 715 bytes: it listed, but every chunked read -failed ("unexpected EOF: need 787 bytes, have 715") and `h5rs check` -reported the chunk indexes past the end of the file. Never wrong data. - -**Fix:** for a v3 superblock with the SWMR-write flag, the data ends at the -end of the file, as libhdf5's SWMR reader reads it. **Test:** -`crates/clawhdf5/tests/swmr_interop.rs` (the mid-write copy, -`tests/fixtures/swmr_mid_write.h5`, through every open path and against -h5py's SWMR reader). A file still being written is read with -`File::open_swmr` (see `docs/design/swmr.md`); `File::open` maps the file -at its length at open and is not meant for files that change while open. - -## Shrinking a chunked dataset with no recorded maximum scrambled it - -**Status:** fixed 2026-09-27, before any release (`FileEditor::resize` -shipped on main in PR #18, a4c2ace). Files the writer produced before the -fix still lack the maximum; see *Existing files*. - -clawhdf5's writer stored no maximum dimensions for a chunked dataset -created without a `maxshape` (libhdf5 always stores one, equal to the -dimensions when none is given). With none recorded, the maximum is the -current dimensions (`H5S_extent_get_dims`), and a Fixed Array chunk index -places chunks by the maximum. `FileEditor::resize` changed only the -current dimensions, so a shrink moved every existing chunk and every -reader returned wrong values; after a shrink the dataset could not grow -back. Found by the review of the Python editing work. - -**Fix:** before the first resize of such a dataset the editor records the -maximum libhdf5 would have written (the dimensions the index was built -with); the writer now records it for every chunked dataset. **Test:** -`crates/clawhdf5/tests/edit_resize_interop.rs`. **Existing files:** -datasets written before the fix have no recorded maximum. The fixed -editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`) -does not** — it scrambles them the same way, and lets them grow past -their Fixed Array. Resize them once with a fixed `FileEditor` (a resize -to the same shape changes nothing; shrink and grow back) before letting -libhdf5 resize them. A dataset already shrunk by the unfixed -editor holds misplaced chunks; rewrite it from a good copy. - -## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768 - -**Status:** fixed 2026-09-26, after v2.7.0. **Every release (v2.1.0 to -v2.7.0) is affected**, in both directions. - -Our Fletcher-32 reduced its two running sums with `% 65535`; libhdf5's -`H5_checksum_fletcher32` (H5checksum.c) folds them with -`(s & 0xffff) + (s >> 16)`. Both are arithmetic mod 65535, but where a sum -is a non-zero multiple of 65535 the fold leaves 0xffff and the modulo 0, so -the checksums differ — for random data about one chunk in 32768 (each of -the two sums hits it with probability about 1/65535). Found by the review -of the editor work: a random-edit fuzzer with gzip + Fletcher-32 hit it on -13 of about 100 seeds. - -- Chunks we wrote (`FileBuilder`/`FileWriter` `with_fletcher32`, and the - unreleased `FileEditor`) with such a sum are refused by h5py and libhdf5: - "filter returned failure during read". h5py writing `[1, 0xfffe]` as - big-endian `u2` stores checksum `0x0001ffff`; we computed `0x00010000`. -- Chunks libhdf5 wrote with such a sum were refused by every reader here - with `Fletcher32Mismatch`; the data itself was never wrong. - -**Fix:** `clawhdf5_format::checksum::fletcher32`, a port of -`H5_checksum_fletcher32`, used by the filter for writing and verifying. It -also accepts a checksum whose 16-bit halves are byte-swapped, as libhdf5 -does for files from 1.6.2 and earlier, and the `% 65535` form clawhdf5 -v2.7.0 and earlier wrote (the two differ only in a half that is 0xffff). -**Test:** -`crates/clawhdf5/tests/fletcher32_interop.rs` (libhdf5's own function -through ctypes on every 1- and 2-byte input plus 40 000 random and -fold-heavy inputs; h5py reads fold-case chunks from `FileBuilder` and -`FileEditor`; we read h5py's). **Existing data:** a Fletcher-32 dataset -written by v2.7.0 or earlier may hold chunks libhdf5 cannot read; a fixed -build reads them. Rewrite such datasets with a fixed build (read, then -write them again) before handing the file to libhdf5 or h5py. - -## LZF/Blosc chunks written with a stale filter mask - -**Status:** fixed 2026-09-26, before any release (the LZF and Blosc writers -were added the same day; v2.7.0 and earlier write neither). - -`FileBuilder` stored every chunk of an LZF or Blosc dataset through the -filter with filter mask 0. libhdf5 counts LZF and Blosc output no smaller -than the chunk as a failure of the (optional) filter and stores the chunk -raw with the filter's mask bit set. When a chunk's LZF stream was exactly -the chunk's size, the first libhdf5 rewrite of it stored raw data at the -same size and left our mask 0 in the index, so h5py could no longer read -the dataset. `FileEditor` had the same bug, fixed earlier the same day. -Both now use `clawhdf5_format::filters::compress_chunk_masked`, and every -chunk index the writer builds records the real mask (see `CHANGELOG.md`). -Files written before the fix read correctly; rewrite them before letting -libhdf5 modify them. - ## In-place modification (`FileEditor`) limits -**Status:** open (documented 2026-09-26, updated the same day when -version-2 B-tree chunk indexes, shrinking, dense attributes and space -reuse were added). `clawhdf5::FileEditor` refuses, with -`Error::Unsupported` and without writing anything: +**Status:** open (documented 2026-09-26, updated when version-2 B-tree +chunk indexes, shrinking, dense attributes and space reuse were added). +`clawhdf5::FileEditor` refuses, with `Error::Unsupported` and without +writing anything: - new chunks in an **implicit** index (it has all of its chunks from the start; they are written in place, and allocated/filled on growth under early allocation as libhdf5 does); @@ -201,11 +51,7 @@ reuse were added). `clawhdf5::FileEditor` refuses, with random-edit harness (120 runs of 150 random edits, `earliest`/`v110`/ `latest`, about 5600 `set_attr` calls of 8 bytes to 6 KiB): 2.2% of `set_attr` calls are refused, every one the last-attribute-in-a-block - replacement; before blocks could be skipped (an attribute needing a heap - block larger than the next one — any attribute of about 1 KiB or more - once a heap has started, or at the move to dense storage), 24% were, - since an object whose move to dense storage was refused kept refusing - every new attribute; + replacement (24% before heap blocks could be skipped); - version-1 object headers asked for an attribute larger than a header message (they have no dense storage); - partial edge chunks stored unfiltered (`H5Pset_chunk_opts`), external @@ -219,18 +65,19 @@ heap's replaced blocks) is reused by later edits of the same `FileEditor`; what is left when it is dropped is leaked, as libhdf5 leaks it without a persistent free-space manager (`h5repack` reclaims it). A chunk that is the last thing in the file grows in place, which covers the usual append. -Measured 2026-09-26 on tank with `cargo test -p clawhdf5-tools --test -edit_interop -- --ignored --nocapture measure_append_waste` (one editor for -the whole workload; file sizes are deterministic): 1000 appends of 100 `f8` -values to a 1-D dataset with 1024-element chunks give 810 504 bytes -unfiltered, as libhdf5's file (`h5repack` of either: 810 360), and 306 780 -bytes with gzip (307 210 before reuse; libhdf5's file: 306 058; `h5repack` -of the editor's file: 306 104, of libhdf5's: 305 954); 2000 appends of 10 -values with 4096-element gzip chunks give 79 829 bytes (119 684 before -reuse) against libhdf5's 50 292 (`h5repack` of the editor's file: 49 930, -of libhdf5's: 50 188): the chunk being appended to -is followed by new index blocks and moves each time it grows, and the -space it leaves is too small for its next, larger version. +Measured 2026-09-26 on tank (`cargo test -p clawhdf5-tools --test +edit_interop -- --ignored --nocapture measure_append_waste`, one editor for +the whole workload; sizes are deterministic): + +| Workload | `FileEditor` | libhdf5 | `h5repack` of ours | +|---|---|---|---| +| 1000 appends of 100 `f8`, 1024-element chunks, unfiltered | 810 504 B | 810 504 B | 810 360 B | +| same, gzip | 306 780 B | 306 058 B | 306 104 B | +| 2000 appends of 10 values, 4096-element gzip chunks | 79 829 B | 50 292 B | 49 930 B | + +In the last case the chunk being appended to is followed by new index +blocks and moves each time it grows, and the space it leaves is too small +for its next, larger version. **No journal.** A crash while an edit patches existing structures can leave the file inconsistent; see the `FileEditor` documentation. @@ -288,711 +135,94 @@ before anything is written. On top of them: ## Selection reads that decode more than the selection -**Status:** open (documented 2026-09-26). `Dataset::read_selection` (and so -the Python `ds[...]`) materialises only the selection's bounding box when -that box covers at most half the dataset (`partial_read`). It decodes the -whole dataset and extracts the selection instead when: +**Status:** open (documented 2026-09-26; re-checked 2026-09-28 in +`Dataset::read_selection`, `crates/clawhdf5/src/reader.rs`). A selection +read (and so the Python `ds[...]`) materialises only the selection's +bounding box when that box covers at most half the dataset +(`partial_read`). It decodes the whole dataset and extracts the selection +instead when: - the bounding box covers more than half the dataset — which a strided selection across a chunked dataset (`ds[::100]`) always does, although it may touch few chunks; - the dataset is compact or virtual, or has no storage; - it is chunked with a non-default fill value (the box path does not fill unallocated chunks, so the fill-aware full read is used). + Values are correct in every case; this is cost only. Selections other than -`Selection::All` also bypass the file's chunk cache. The bounding-box -heuristic's other cost is measured under "Concurrent and contiguous read -performance" below. +`Selection::All` also bypass the file's chunk cache. -## Concurrent and contiguous read performance (measured 2026-09-26) +## HDF5 features still unsupported -**Status:** fixed (2026-09-26). Re-measured on tank at `c5334b1`: full -reads of chunked deflate data at 16 threads run at 4944 MB/s against 3135 -for 16 h5py processes (1.58x; 0.69x-0.76x before), and contiguous reads -are 1.29x h5py on one thread (`BENCHMARKS.md`, "Results after in-place -chunk decoding"). The history below is kept for reference. Measured on -tank with `concurrent_read` against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`, -"Concurrent reads"): -- **Partly fixed 2026-09-26.** Full reads of chunked datasets from several threads - through one `File` stop scaling at about 4 threads (880 MB/s on deflate - data vs 4424 MB/s for 16 h5py processes). Hyperslab reads, which skip the - chunk cache, scale to 1244 MB/s, so the `File`'s shared chunk cache is the - suspect. *Cause:* not the cache. Those numbers were taken with - `--decode-threads 1`, a one-thread rayon pool, and every full read handed - its chunks to that pool, so all reader threads queued behind its single - worker (per-thread CPU time: one thread did all the decoding, the 16 - readers almost none). Hyperslab reads touch one chunk each and never used - the pool. Reads now decode on the calling thread when the pool has one - thread (`tests/single_thread_decode_pool.rs`); re-measured at - `408f69e`, 8 threads went from 887 to 2943 MB/s (h5py processes: 3042). - **Still open:** at 16 threads full chunked reads reach 2142-2341 MB/s, - 0.69x-0.76x 16 h5py processes (3083 MB/s in the same run), with the - default pool as with a one-thread one; with a small pool (2-4 threads) - readers outside it still wait on its workers. Datasets larger than the - cache's budget were already read without inserting into it, and skipping - its lookups entirely gained only a few percent at 16 threads. Remaining - per-read overhead: each full `read_f32` of a chunked dataset faults in - about three times its size in fresh pages (the output, the `f32` copy of - it, and a new buffer per decoded chunk). - **Fixed 2026-09-26** (both causes; see `CHANGELOG.md`, "Chunked full - reads"): chunks are decoded into buffers each thread reuses and copied - straight into the output, and the typed readers decode into the `Vec` - they return, so a full read no longer faults in a buffer per chunk or a - second copy of its output; and the reading - thread decodes its own chunks with pool workers helping when free, so no - reader waits on a small or busy pool - (`crates/clawhdf5/tests/busy_decode_pool.rs`). The 16-thread comparison - with h5py processes has not been re-measured yet (tank was busy with - other work); this item stays open until it is. -- Contiguous datasets read 4x slower than h5py on one thread (2.5 vs - 9.8 GB/s full, 0.12x for 256 x 256 hyperslabs). - **Fixed 2026-09-26** (re-measured on tank at `408f69e`: 13665 MB/s - full and 31991 MB/s for 256 x 256 hyperslabs on one thread, 1.44x and - 6.3x h5py; `BENCHMARKS.md`): full - reads were dominated by 4 KiB page faults on the fresh output buffer, - which is now backed by transparent huge pages as numpy's is; hyperslab - reads copied the selection three times, element by element, and now copy - each contiguous run once, straight from the file into the output (see - `CHANGELOG.md`). The chunked-read scaling item above is still open. -Values are correct; this is speed only. +**Status:** open. What remains of the gaps the +[2026-09-25 audit](#silent-wrong-data-found-by-the-2026-09-25-hdf5-audit) +found (the rest are [fixed](#gaps-found-by-the-2026-09-25-hdf5-audit-fixed-parts)), +plus later residuals. Each is an error or a documented difference, never +wrong data. -## Scale-offset data read back wrong values - -**Status:** fixed 2026-09-26, after v2.7.0. **Every release that decoded -the scale-offset filter (v2.2.0 to v2.7.0) is affected**, on ordinary -files h5py writes, with no error. - -Found by the review of the 2026-09-26 conformance work: h5py wrote 1480 -scale-offset datasets (every integer type `i1` .. `u8`, `f4` and `f8`, -little- and big-endian, with no fill value and with the type's minimum, -maximum or another value as fill, random, constant, all-fill and extreme -data, `scaleoffset` 0, 1, 3, full width less one and full width for -integers, decimal scale factors 0, 1, 2, 4 and 7 for floats). v2.7.0's -decoder read 332 of them differently from h5py: **151 returned wrong -values with no error**, 181 failed to read. - -| Case | Datasets | v2.7.0 | -|---|---|---| -| integer, `scaleoffset=0` (libhdf5 picks the bits), data spanning most of the type's range | 82 | wrong values | -| integer, `scaleoffset` = full width: `i4`/`u4` (10), `i8`/`u8` (41) | 51 | wrong values | -| integer, `scaleoffset` = full width, the other datasets (every `i1`..`u2` one, most `i4`/`u4`, some `i8`/`u8`) | 181 | "truncated minval" / "implausible minbits" | -| `f4` D-scale, factor 4 or 7, values up to about 10^6 | 18 | wrong values | - -In every case libhdf5 stored a chunk at full width (`minbits` equal to the -type's width): the chunk then holds the elements as they are, and they -were decoded as offsets from `minval`. Two more differences were found on -crafted files and fixed with them: the packed codes start at byte 21 -whatever size the chunk records for `minval` (`cve-2025-44905` -`/Scale_offset_short_data_be`), and a chunk with `minbits` 0 and a fill -value is all fill values (it read as `minval`). - -**Fix:** `clawhdf5_format::filters` decodes a scale-offset chunk as -`H5Z__filter_scaleoffset` does. **Test:** the whole matrix is -`crates/clawhdf5/tests/scaleoffset_interop.rs`, generated by h5py at test -time and compared dataset by dataset; on v2.7.0's decoder it reports the -332. **Existing data:** the files were always right; only reads were -wrong, so re-reading with a fixed build gives the correct values. - -## Silent wrong data found by the 2026-09-25 HDF5 audit - -**Status:** fixed after v2.7.0 (2026-09-25). **Every release up -to and including v2.7.0 is affected.** - -An audit on tank checked clawhdf5 against libhdf5 in three ways: -- a sweep of 686 public files: the libhdf5 test files, the HDF Group's - `cve_hdf5` reproducers, and the pyfive, netcdf-c, netcdf4-python, h5wasm, - h5py and xarray corpora; -- 567 read cases generated with h5py 3.16 / HDF5 2.0; -- 96 write cases checked with h5py builds linking HDF5 1.10, 1.12, 1.14 and - 2.0, plus h5dump 1.14.6. - -It found these cases where a value came back wrong **without an error**: - -| Area | What happened | Who is affected | -|---|---|---| -| Chunk index (read) | Fixed/Extensible Array indexes laid out by the current shape, not the max shape: chunks returned from the wrong place | any file with a max shape larger than its shape and `libver='latest'` (h5py `maxshape=(10, None)`, `(20, 10)`) | -| Chunk index (write) | Extensible Array chunks from index 244 on never indexed (read as 0); unlimited dimension not first: data scrambled | files we wrote with one unlimited dimension and > 244 chunks, or e.g. `maxshape=(20, None)` | -| 4-byte offsets | unfiltered chunked datasets read as zeros | files created with `sizeof_addr = 4` | -| Filter mask | any skipped filter skipped the whole pipeline | files with partially filtered chunks (optional filters, direct chunk writes) | -| Numeric reads | float read as integer returned the bit pattern; narrowing integer reads kept the low bits; bfloat16 decoded as IEEE half | `read_i32`/`read_i64`/`read_u64` callers on float or wider data; HDF5 2.0 bf16 data | -| SZIP | garbage or zeros | every libhdf5-written SZIP dataset | -| Scale-offset | float values 1 ULP off | libhdf5 D-scale float data | -| Shared fill value | read as zero fill | fill values stored as shared messages | -| VL sequences | `read_vl_bytes` truncated non-byte base types | VL int/float sequences | -| Chunk cache | two threads reading two chunked datasets through one `File` could get each other's chunks | multi-threaded readers, including Python with the GIL released | - -The audit also found files we wrote that libhdf5 **refuses**, now fixed: -- Fixed Array datasets with more than 1 024 chunks. -- Header messages over 64 KiB (large attributes). -- Reference, Opaque, BitField and Time datatypes. -- Files written with `with_page_size`. -- Several unlimited dimensions. -- A finite max shape larger than the shape. -- An empty-string attribute, which broke every attribute on its object. -- `FillTime` codes, which were rotated. - -Our LZ4 and Zstd output could not be read by libhdf5's registered plugins, and -our pcodec filter used Granular BitRound's ID. The details are in -`CHANGELOG.md` under Correctness and Interop. - -Before the fix, 419 of the 686 files read correctly and 43 differed from h5py. -After it, 448 read correctly and 23 differ. Of those 23: -- 17 are N-Bit float files. The probe compares raw file-type bytes; the typed - reader returns libhdf5's values (`nbit_custom_float_decodes_like_libhdf5`). -- 2 are an h5py bug: VL data with a big-endian base type comes back - byte-swapped in h5py, and h5dump agrees with us. -- The rest are object or attribute listing differences. - -**Update 2026-09-25:** the sweep is now in the repo (`conformance/run.sh`, -corpora pinned by commit) and its current numbers are in `CONFORMANCE.md`, -regenerated nightly by `.gitea/workflows/conformance.yml`. The probe now -compares N-Bit floats as the values libhdf5 converts them to, so the N-Bit -files above count as identical. Its file list is defined by -`conformance/list_files.py` (697 files: netCDF classic files are left out, -and 11 HDF5 files the ad-hoc sweep missed are in). On 42b81d9: 467 identical, -123 our-error, 15 mismatch (2 are the h5py bug above), 92 that libhdf5 cannot -read, and no panics, hangs or crashes. - -There were no panics, hangs or crashes before or after, including on all 147 -CVE and fuzzer files. On some of those files, h5dump 1.14.6 and h5py/HDF5 2.0 -segfault or abort. - -## Gaps found by the 2026-09-25 HDF5 audit (open) - -**Status:** open. These fail with an error; none returns wrong data (the VDS -fill-value item that did is fixed). - -- **Layout message versions 1 and 2** (HDF5 1.6-era files): 84 of the 686 - sweep files, `InvalidLayoutVersion`. This is the largest single gap. - **Fixed 2026-09-25:** versions 1 and 2 are parsed (compact, contiguous, - chunked via the v1 B-tree). -- **Compound datatype version 1 array members** (found with the layout - fix; pre-1.4 files such as `tarrold.h5`): **wrong data** — the legacy - per-member dimensions were skipped, so an array member read as one scalar. - **Fixed 2026-09-25.** -- **Virtual datasets:** - - ~~**Wrong data:** unmapped regions read as 0 instead of the fill value.~~ - Fixed 2026-09-25: unmapped elements and missing sources read as the - virtual dataset's fill value. - - ~~`%b` printf-style source names are not expanded.~~ Fixed 2026-09-25: - printf-style and unlimited mappings are read, and the extent is - recomputed from the sources as libhdf5 does. Still open: the - "first missing" view and a printf gap other than 0 (libhdf5 access - properties we always read at their defaults), source-to-virtual type - conversion other than a byte swap, nested virtual sources, and source - files outside the virtual file's directory (refused with an error). - - ~~Hyperslab selection versions 1 and 2 are refused.~~ Fixed 2026-09-25: - versions 1-3 and irregular hyperslabs are decoded. - - ~~The version-1 mapping list written with a 2.0 low bound (flags byte, - shared names) was misparsed.~~ Found and fixed 2026-09-25. -- **Files with a user block:** the base address is not applied. - **Fixed 2026-09-25:** every reader views the file from the superblock on - (`twithub.h5`, `twithub513.h5`, `h5clear_fsm_persist_user_*.h5`; the - `twithub` files still stop at the user-defined link type below). -- **Old-style shared messages (version 1)** read the wrong address. - **Fixed 2026-09-25:** the address follows the length-sized link-name - offset of the embedded symbol table entry (`tcompound.h5`, `tcompound2.h5`). -- **Groups and links:** - - Groups with a user-defined link type (e.g. 187) cannot be listed. - **Fixed 2026-09-25:** user-defined links are skipped; the rest of the - group lists. - - Dense groups with more than about 22 000 links cannot be listed. - **Fixed 2026-09-25:** two bugs — fractal-heap child indirect blocks had - the wrong row count, and v2 B-tree internal nodes at depth 3+ were read - with the wrong pointer widths. - - Soft links are left out of `datasets()`. **Fixed 2026-09-25:** soft - links are listed as their targets; dangling ones are left out. - - **Wrong data (found while fixing user blocks):** an old-style group whose - local-heap free list points outside the heap listed garbage names where - libhdf5 refuses the heap. **Fixed 2026-09-25** - (`InvalidLocalHeapFreeList`, checked when a name is first read, as - libhdf5 does). -- **Dense attributes:** a large attribute stored as a fractal-heap "huge" - object makes every attribute on the object fail. This affects real NetCDF - files (`issue671.nc`). **Fixed 2026-09-25:** huge and tiny heap objects, - and filtered heaps, are read; and an attribute that still cannot be read - is left out of `attrs()` (reported by `attrs_with_errors()`) instead of - failing the others. -- **Other readers:** - - VL-string datasets are not readable through `File`. **Fixed - 2026-09-26:** `read_string` reads them (also `read_string_bytes`, - `read_string_selection`, and on `MmapFile`/`LazyFile`), with h5py's - values: strings end at a NUL, null elements (heap address 0) are `""`, - and an element at the undefined heap address is an error as in libhdf5 - (it read as `""` until 2026-09-26); VL sequences of - numbers read with `read_vlen::()`, and VL values inside compounds or - `AttrValue::Raw` attributes decode with `File::decode_strings` / - `File::decode_vlen` (`crates/clawhdf5/tests/vl_data_interop.rs`). - - Variable-length values inside a compound (and VL-string attributes) in - a file with 4-byte offsets (`sizeof_addr = 4`) fail with - `GlobalHeapObjectNotFound` or come back as `Raw`: these paths assume - the 16-byte element of an 8-byte-offset file. The datatype itself reads - (it was refused as "member overlaps with previous member" until - 2026-09-26). **Fixed 2026-09-26:** a VL type's element size is the one - its datatype message stores (12 with 4-byte offsets), and the global - heap is read with libhdf5's header padding - (`crates/clawhdf5/tests/vl_offset4_interop.rs`). - - Metadata cache images are not supported. **Fixed 2026-09-26:** the - image is applied at open, as libhdf5 loads it over the file's metadata - (`clawhdf5_format::superblock_ext`); `h5clear_mdc_image.h5` reads - (`crates/clawhdf5/tests/metadata_cache_image.rs`), without copying the - file (a private copy-on-write mapping takes the image's entries; - `tests/cache_image_memory.rs`). A file whose image libhdf5 cannot load - (`cve-2025-6269-*`, `cve-2025-6516`) opens, as in libhdf5, and every - object lookup fails with the image's error. Differences from libhdf5 - that remain: libhdf5 fails only the first metadata read and then reads - the file's own (possibly stale) metadata, where we keep failing; an - image entry that runs past the end of file is refused (libhdf5 checks - only its start); a flush-dependency parent flag is checked against the - child count as libhdf5's debug build checks it (HDF5 2.0 release - builds refuse every entry that has children, even in images they - wrote); the superblock extension's driver-info and shared-message table - messages are not decoded at open. - - x87 long double and binary128 are refused. - - N-Bit on 64-bit scale-offset data and some N-Bit parameter layouts fail. - **Not our bug (checked 2026-09-26):** both are corrupt files HDF5 2.0 - reads only by reading past a buffer. `cve-2025-2308`'s - `/Scale_offset_long_long_data_le` has scale-offset codes that run past - the end of the chunk (libhdf5's develop branch refuses it, "Buffer too - short"); `bad_nbit_parms_walk.h5` has an N-Bit parameter list one value - short (libhdf5's own `test_filter_bad_params` in `test/dsets.c` now - requires that read to fail). We refuse both; `CONFORMANCE.md` lists them - under *Known not-our-bug* (since 2026-09-27: the *ref-bug* class, with - the evidence re-checked every run — see *Conformance: the last non-ok - files* above). Scale-offset did decode three cases - differently from libhdf5 (codes after a `minval` of recorded size other - than 8, `minbits` 0 with a fill value, full-width `minbits`): fixed - 2026-09-26, and the full-width case was silent wrong data on ordinary - h5py files (see *Scale-offset data read back wrong values* above). -- **Filters:** blosc, blosc2, bitshuffle, bzip2, LZF and zfp are not - implemented. **Fixed 2026-09-26** for LZF (default-on `lzf` feature), - bitshuffle, bzip2 and Blosc 1 (`bitshuffle`, `bzip2`, `blosc`, or - `plugin-filters` for all), read and write, pure Rust; h5ex_d_lzf, - h5ex_d_bshuf, h5ex_d_bzip2 and h5ex_d_blosc now read (conformance 573 of - 697 ok). **Fixed 2026-09-26** for Blosc2 (32026, `blosc2` feature, also - in `plugin-filters`), read only: hdf5plugin's frames and B2ND arrays, every - codec and filter it offers; h5ex_d_blosc2 now reads (conformance 576 of - 697 ok, tank, `conformance/run.sh --no-fetch`). Blosc2 frames using - dictionaries, lazy chunks, variable-length blocks, user-defined codecs or - registered filters (e.g. bytedelta) are refused with an error. - **Fixed 2026-09-26** for ZFP (32013, `zfp` feature, also in - `plugin-filters`), read only: every H5Z-ZFP mode and type, bit-exact - against h5py + hdf5plugin 7.1 (`crates/clawhdf5/tests/zfp_interop.rs`); - h5ex_d_zfp now reads (conformance 600 of 697 ok, tank, - `conformance/run.sh --no-fetch`). **Still open:** clawhdf5 cannot write - Blosc2 or ZFP. Other filters (32023, Granular BitRound, too, since - 2026-09-26 even with the `pcodec` feature) can be plugged in with +- **Virtual datasets:** the "first missing" view and a printf gap other + than 0 (libhdf5 access properties we always read at their defaults), + source-to-virtual type conversion other than a byte swap, nested virtual + sources, and source files outside the virtual file's directory (refused + with an error). +- **Datatypes:** x87 long double and binary128 are refused. Revised + (version-4) region and attribute references are recognised but not + decoded, and external references are an error (object references + decode). A multi-dimensional numeric attribute is returned as a flat + array (its shape is not reported; `AttrValue::Raw` carries the shape). +- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in + that: libhdf5 fails only the first metadata read of an image it cannot + load and then reads the file's own (possibly stale) metadata, where we + keep failing; an image entry that runs past the end of file is refused + (libhdf5 checks only its start); a flush-dependency parent flag is + checked against the child count as libhdf5's debug build checks it (HDF5 + 2.0 release builds refuse every entry that has children, even in images + they wrote); the superblock extension's driver-info and shared-message + table messages are not decoded at open. +- **Filters:** Blosc2 and ZFP are read-only (clawhdf5 cannot write them). + Blosc2 frames using dictionaries, lazy chunks, variable-length blocks, + user-defined codecs or registered filters (e.g. bytedelta) are refused. + Other filters (32023 Granular BitRound among them, even with the + `pcodec` feature) can be plugged in with `filter_registry::register_filter`. -- **Wrong data: a chunk whose filters decode to fewer bytes than the chunk - read with zeros for the missing bytes** (any filter; found reviewing the plugin - filters). **Fixed 2026-09-26:** it is an error naming the chunk. A corrupt - chunk must never read as zeros. Unfiltered chunks are read at their stored - size and are not checked this way. **Fixed 2026-09-26** for unfiltered - chunks too: in a dataset without filters a chunk the index records at - other than the chunk's size is refused, as libhdf5's develop branch - refuses it (`cve-2025-44904`, where HDF5 2.0 fills the rest from its - buffer). -- **Crash:** a hostile Blosc chunk (frame size below its header) panicked in - builds with overflow checks. **Fixed 2026-09-26**; the new decoders are - fuzzed in the unit tests. -- ~~**Header checks:** on 12 CVE datasets libhdf5 rejects a corrupt header and - we read data anyway. We need stricter header checks.~~ **Fixed - 2026-09-26** (counted again: 18 objects on the CVE corpus that libhdf5 - refuses; some read as wrong data, e.g. a zero chunk dimension read as all - fill values): object headers, datatypes, chunk dimensions and chunk-index - offsets are checked as libhdf5 checks them, and truncated files are - refused. 17 of the 18 now - fail as in libhdf5 (conformance on tank, `conformance/run.sh --no-fetch`, - 2026-09-26: 571 of 697 ok). Still read where libhdf5 refuses: - - ~~`cve-2024-32624.h5` `/Dset_OBJREF`: a dataspace whose storage size - overflows 64 bits. `File::dataset` and `shape()` succeed (libhdf5 - refuses at open); reading the values fails.~~ **Fixed 2026-09-26:** - `File::dataset` (and `MmapFile`, `LazyFile`) refuse it at open - (`FormatError::InvalidDatasetStorage`), as they do contiguous storage - past the end of the file. - - ~~`cve-2020-10810.h5`, `cve-2020-10812.h5` (whole files libhdf5 cannot - open, not among the 18): libhdf5 decodes the superblock extension's File - Space Info and metadata-cache-image messages at open and refuses these - files; we do not decode those messages at open.~~ **Fixed 2026-09-26:** - the superblock extension is decoded at open with libhdf5's checks, and - both files are refused. - - Deliberately not refused, because clawhdf5 up to v2.7.0 wrote them: a - float sign bit position outside the type, and a size-0 string type. - - Not refused because current libhdf5 reads it though HDF5 2.0.0 - (h5py 3.16) refuses it: a v4 chunked layout whose dimensions are - encoded in more bytes than they need (HDFGroup/hdf5@e124c36, - 2026-06-05, relaxed that check; clawhdf5 wrote such layouts until - 2026-09-26). - - Not refused because HDF5 2.0 (h5py 3.16) reads them though newer - libhdf5 refuses them: bit-field offset/precision outside the type, an - unknown variable-length kind, an array type whose stored size is not - its element count times its base size. - - (`cve-2024-32616` `/group1/dset3` and `cve-2025-2309`'s `Comp_OBJREF` - attribute are h5py/numpy type-mapping failures, not libhdf5 refusals.) - - `h5rs check` validates with the library's parsers, so it inherits what - they accept: of the 150 CVE and fuzzer files, `check --data` passes 15, - and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26; 28 and 21 - before these checks, 16 and 9 before a VL type's stored element size - was checked, which flags `cve-2024-32608`). -- **Writer:** - - ~~Nested groups beyond one level: path-like names are now refused, not - created.~~ **Fixed 2026-09-26:** groups nest to any depth (path names - create intermediate groups, as h5py does), with soft, extra hard and - external links at any depth and optional creation-order tracking; - h5py, h5dump and `h5rs check --data` read them - (`crates/clawhdf5/tests/writer_groups_interop.rs`, - `crates/clawhdf5-tools/tests/h5rs_interop.rs`). ~~Still missing: - attribute creation order is not tracked.~~ **Fixed 2026-09-26:** - `track_order` (file default, `GroupBuilder`, and the new - `DatasetBuilder::track_order`) tracks and indexes attribute creation - order as h5py's `track_order=True` does; h5py lists the attributes in - the order they were set, inline and dense (20 000 on one dataset), and - keeps numbering them in "r+" mode (tank, - `cargo test -p clawhdf5 --test writer_groups_interop - track_order_lists_attributes_in_creation_order`). libhdf5 numbers at - most 65 535 attributes on such an object, so more is an error. - - ~~A group with more than 65 535 links, or an object with more than - 65 535 dense attributes, is an error (the index is one B-tree leaf).~~ - **Fixed 2026-09-26:** the dense indexes are v2 B-trees of any depth - (libhdf5's 512-byte nodes once the records outgrow the one-leaf layout, - which smaller indexes keep byte for byte). Tested on tank with 100 000 - links in one group (short names, and 111-byte names with creation order - tracked) and 70 000 attributes on one object: h5py lists them in order - and reads the values, h5dump reads spot checks, `h5rs check` reads every - record, and h5py in "r+" mode adds and deletes thousands of links and - attributes in those trees (`cargo test -p clawhdf5 --test - deep_btree_interop`; `cargo test -p clawhdf5-tools --test h5rs_interop - check_files_with_deep_btrees`). The name indexes are ordered by hash and - then name, as libhdf5 needs for names whose hashes collide. - - ~~Dense link or attribute storage past 512 KiB of messages was written - unreadable (child indirect blocks of the fractal heap written as direct - blocks).~~ **Fixed 2026-09-26** (it affected 2.7.0 too): tested with - 20 000 and 65 535 links and with 8 MB of dense attributes, read by - h5py, h5dump, `h5rs check` and clawhdf5, and h5py can add links to - such groups. - - ~~libhdf5 could not add a link to a group we wrote (no Group Info - message).~~ **Fixed 2026-09-26.** - - Huge fractal heap objects: in dense storage (more than 8 attributes on - an object, or more than 8 links in a group) one attribute or link - message over 65 515 bytes is an error. - - Output that HDF5 1.8 can read. - - ~~A B-tree v2 chunk index larger than one leaf, so datasets with - several unlimited dimensions are limited to 65 535 chunks.~~ **Fixed - 2026-09-26:** more chunks get libhdf5's 2048-byte nodes with internal - nodes above the leaves. Tested on tank with 200 000 chunks (and 80 000 - deflated): h5py and clawhdf5 read every value, and h5py resizes the - dataset and writes 4 510 new chunks into the tree (same commands as - above). +- **Writer:** in dense storage (more than 8 attributes on an object, or + more than 8 links in a group) one attribute or link message over 65 515 + bytes is an error (no huge fractal-heap objects). The writer does not + produce output that HDF5 1.8 can read. +- **Checks we deliberately do not make:** + - a float sign bit position outside the type, and a size-0 string type: + clawhdf5 up to v2.7.0 wrote them; + - a v4 chunked layout whose dimensions are encoded in more bytes than + they need: current libhdf5 reads it (HDFGroup/hdf5@e124c36, + 2026-06-05) though HDF5 2.0.0 (h5py 3.16) refuses it, and clawhdf5 + wrote such layouts until 2026-09-26; + - bit-field offset/precision outside the type, an unknown + variable-length kind, an array type whose stored size is not its + element count times its base size: HDF5 2.0 reads them though newer + libhdf5 refuses them. ---- - -## Compound datatype message version 5 is not parsed (HDF5 2.0) - -**Status:** fixed on `main` in `a13ff51` (2026-06-03); **not in the v2.1.0 -tag**, which was cut five commits earlier. Ships in the next release. - -**Reported by:** M. Scot Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0. - -**Summary:** `clawhdf5-format` v2.1.0 rejects any dataset with a compound -(struct) datatype written by an HDF5 2.0 library in `libver='latest'` mode: -`InvalidDatatypeVersion { class: 6, version: 5 }`. - -**Reproduction** (h5py 3.16.0 / HDF5 2.0.0): - -```python -import h5py, numpy as np -dt = np.dtype([('x', 'f8'), ('y', 'f8'), ('id', 'i4')]) -data = np.array([(1.0, 2.0, 10), (3.0, 4.0, 20)], dtype=dt) -f = h5py.File('compound.h5', 'w', libver='latest') -f.create_dataset('particles', data=data) -f.close() -``` - -Committed as `crates/clawhdf5-format/tests/writer_h5py_tests.rs::read_h5py_generated_compound` -(`#[ignore]`d; needs `python3` with h5py on `PATH`). Run with -`cargo test -p clawhdf5-format --test writer_h5py_tests -- --include-ignored`: -v2.1.0 gives 25 passed / 1 failed; `main` passes everything. - -**Root cause:** the compound (class 6) branch of `Datatype::parse` -(`crates/clawhdf5-format/src/datatype.rs`) accepted only versions 1–4. Datatype -message versions 4 and 5 changed only the Reference and Complex classes, so a -v5-tagged compound uses the unchanged v3 member-list layout. - -**Fix:** versions 3–5 are accepted for compound (class 6) and array (class 10) -datatypes, and data layout message version 5 is accepted too (needed for every -chunked dataset written by HDF5 2.0). Byte-level regression tests: -`test_compound_v5_from_hdf5_2_0`, `test_array_v5_from_hdf5_2_0`. - -## Native complex datatype (class 11) is mis-parsed (HDF5 2.0) - -**Status:** fixed 2026-09-18. Found while validating the report above. - -**Summary:** HDF5 2.0 native complex types (`H5T_COMPLEX_IEEE_F64LE` etc.) -were parsed as if they carried a compound-style member list. The properties are -actually a single base floating-point datatype, so the parser produced a garbage -datatype, or `UnexpectedEof` when the complex type was a compound member. h5py's -default numpy-complex mapping is unaffected (it writes a `{r, i}` compound); -only files using the native type through the C API / h5py low-level API hit this. - -**Fix:** class 11 parses its base type and is surfaced as the equivalent -`{r, i}` compound. Tests: `test_complex_v5_from_hdf5_2_0`, -`test_compound_with_complex_member_from_hdf5_2_0`, -`writer_h5py_tests.rs::read_h5py_generated_native_complex`. - -## Revised reference datatype (class 7, version 4) is not parsed - -**Status:** fixed 2026-09-19 for object references; region and attribute -references are recognised but not decoded. - -**Summary:** HDF5 1.12+ `H5T_STD_REF` references use datatype message version 4 -with reference types 2-4 (object / region / attribute), which `Datatype::parse` -rejected with `InvalidReferenceType`. h5py still writes the legacy references, -so no file had been available to test against. - -**Fix:** a real file was produced by driving the libhdf5 bundled in the h5py -wheel through ctypes (`tests/fixtures/gen_std_ref.py` -> -`std_ref_hdf5_2_0.h5`). The three new types parse as -`ReferenceType::{Object2, DatasetRegion2, Attribute}`, and -`read_object_references` decodes `Object2` elements (type, flags, token size, -token = target object header address). External references (flag bit 0) and -the region/attribute payloads are errors rather than misreads. - -## `clawhdf5-gpu` `gpu_tests` can hang under the default parallel test runner - -**Status:** fixed 2026-09-19. - -**Summary:** during `cargo test --workspace` the `gpu_tests` binary sat idle for -25+ minutes. Every test created its own `wgpu::Instance` + device (requesting -adapter-maximum limits) concurrently, and readback used an unbounded -`device.poll(Wait)`. - -**Fix:** tests hold a process-wide lock while they own a device, and -`GpuAccelerator` readback waits time out after 30 s with `GpuError::BufferMap`. - -## Compound datatype versions 1 and 2 are mis-parsed (default libver files) - -**Status:** fixed 2026-09-19. Found by adding a default-libver axis to the h5py -interop tests. - -**Summary:** any compound dataset written with default libver bounds (plain -`h5py.File(path, 'w')`, datatype message version 1) failed to read, typically -with `Overflow("compound member 'x': byte_offset(0) + field_size(4136977) ...")`. -Only `libver='latest'` files (version 3+) and files written by clawhdf5 itself -worked, which is why the existing tests never caught it. - -**Root cause:** `Datatype::parse` skipped 24 bytes of legacy per-member array -fields for v1 where the format has 28 (dimensionality 1 + reserved 3 + -permutation 4 + reserved 4 + 4 dimension sizes 16), and treated v2 like v1 minus -name padding, whereas v2 keeps the 8-byte name padding and has no array fields. - -## Attributes with unsupported datatypes are silently dropped - -**Status:** fixed 2026-09-19. - -**Summary:** `Dataset::attrs()` / `Group::attrs()` returned only attributes -convertible to `AttrValue` and omitted the rest without any indication — every -Python `bool` (an HDF5 enum), complex, compound and reference attributes. Unsigned -64-bit arrays were also cast to `I64Array`, turning values above `i64::MAX` -negative. - -**Fix:** booleans decode as 0/1 integers, `AttrValue::U64Array` keeps unsigned -arrays unsigned, and `AttrValue::Raw { datatype, shape, data }` carries any other -attribute verbatim. Both new variants are writable. Still lossy: a -multi-dimensional numeric attribute is returned as a flat array (its shape is -not reported). - -## B-tree v2 chunk index (layout v4, index type 5) is not supported - -**Status:** fixed 2026-09-19. - -**Summary:** a chunked dataset with **two or more unlimited dimensions** written -with `libver='latest'` indexes its chunks with a version-2 B-tree, and reading it -failed with `unsupported chunked layout version=4, index_type=Some(5)`. - -**Fix:** record types 10 (unfiltered) and 11 (filtered) are decoded — address, -stored size, filter mask, scaled offsets — through the shared chunk-listing -function, so full reads, cached reads, partial reads and fill-value handling -all work. Covered by an h5py interop test (plain, gzip+shuffle, a 2500-chunk -tree with internal nodes, a sparse dataset with a fill value, a hyperslab). + `h5rs check` validates with the library's parsers, so it inherits what + they accept: of the 150 CVE and fuzzer files, `check --data` passes 15, + and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26). +- **VL sequences:** a file may point many elements at one large global + heap object, and a VL-sequence read then returns that object once per + element, as h5py would (memory is bounded otherwise; see + [Crafted global heaps](#crafted-global-heaps-exhaust-the-variable-length-readers-memory)). ## External links and external raw data are not followed -**Status:** open (by design for now); both are explicit errors. - -**Summary:** a path through an external link returns -`FormatError::ExternalLinkUnsupported { filename, object_path }`, and a dataset -created with `external=[...]` storage returns -`FormatError::ExternalDataFilesUnsupported`. Neither is resolved. If support is -added, file names must be confined to the opened file's directory, as the -virtual-dataset resolver now does. - ---- - -## Python interop suites skip silently when no interpreter has h5py - -**Status:** fixed on `main` in `a29c1b2` (2026-09-19). - -On a system where `python3` is a PEP 668 "externally managed" interpreter, -h5py cannot be installed into it at all, and every interop suite — the h5py -writer round-trips, the facade suite, netCDF4, and the reference files — -returned `false` from its availability probe and skipped without failing. CI -reported `SKIP` and a green run. This is the same class of gap that let the -compound-datatype v5 bug above reach a release. - -The probes now read `CLAWHDF5_PYTHON`, and `scripts/ci-test.sh` picks up -`.venv/bin/python` automatically. To restore the coverage on a fresh checkout: - -```bash -python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 -``` - -Set `CLAWHDF5_REQUIRE_INTEROP=1` in any automated runner so a missing -interpreter is a failure rather than a skip. - - ---- - -## Crafted B-tree v2 structures crash or exhaust the reader - -**Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to -and including v2.6.0 is affected.** - -B-tree v2 traversal (`clawhdf5-format`, `btree_v2::collect_btree_v2_records`) -recursed one frame per level with the depth taken from the file, and followed -child addresses without checking whether they were shared. Two consequences -for anyone reading untrusted files: - -- A node that is its own child, under a header claiming 65 535 levels, overflows - the stack and aborts the process. The file is under 100 bytes. -- Levels whose children all point at one node below make the traversal visit it - fan-out^depth times: ~30 million records from ~5 KB, and memory exhaustion one - level deeper. - -B-tree v2 backs dense attribute storage, v2 groups, shared object header -messages and chunk indexes, so opening an object that uses any of them is -enough. Both are now errors: depth is capped at 64, and traversal stops once it -has produced more records than the file could physically hold. - ---- - -## Crafted global heaps exhaust the variable-length reader's memory - -**Status:** fixed on `feat/p2-vl-strings` (2026-09-26). Not a regression of -that branch: every earlier release is affected through `read_vl_strings`. - -Reading variable-length values kept an owned copy of every object of every -global heap collection visited, for the whole read. A file whose collections -nest inside one another's object data (32 bytes apart, each element pointing -at a different one) made retained memory O(elements × file size): a 744 KB -file reached 1.58 GB. Letting every collection's object chain jump to one -shared run of tiny objects made the parse time O(elements × objects) too. -libhdf5 refuses such files. - -Now `VlResolver` caches where each object lies instead of a copy, drops its -cache past a 32 MiB budget, and refuses a collection that overlaps one it -has already read (libhdf5 gives each collection its own block, so only a -crafted file has them). `GlobalHeapCollection::parse` (and the new -`parse_index`) also refuse a collection that runs past the end of the file, -or an object that runs past the end of its collection. Guarded by -`crates/clawhdf5-format/tests/vl_heap_bounds.rs`, which measures peak heap -use with a counting allocator. Still open: a file may point many elements -at one large heap object, and a VL-*sequence* read then returns that -object once per element, as h5py would. - ---- - -## Extensible Array chunk indexes read back wrong data past the inline elements - -**Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to -and including v2.6.0 is affected.** - -A dataset created with exactly one unlimited dimension (`maxshape=(None, ...)`, -the usual append-only/resizable case) is indexed by an Extensible Array. Its -index block holds the first `idx_blk_elmts` chunk entries inline — 4 by -default — and everything after that lives in data blocks and super blocks whose -layout `clawhdf5-format` computed incorrectly. - -Consequences, by dataset size (1 chunk per element): - -| chunks | result before the fix | -|---|---| -| <= 36 | correct (inline, plus two data blocks that happened to line up) | -| 37 | 1 element wrong | -| 400 | 364 elements wrong | -| >= ~1000 | `invalid Extensible Array data block signature` | - -The dangerous case is the middle one: values were returned from the wrong -chunks rather than an error being raised. Any reader that accepted the data at -face value saw plausible but incorrect numbers. - -The root causes were the super block sizing formulas (`ndblks` and -`dblk_nelmts` each double every *other* level, a half-step apart), a missing -block-offset field in the super block, and a page-init bitmap read from the -wrong structure. All four are fixed and covered by interop tests against -HDF5 2.0 at sizes that cross each boundary, including paged data blocks. - -Files written by this crate were not affected by *this* read bug, but the -writer had its own: it indexed only the first 244 chunks, so later chunks -read back as 0 in libhdf5 and in clawhdf5. See "Silent wrong data found by -the 2026-09-25 HDF5 audit" below. - -## Every `f32` dataset we wrote was unreadable by h5py / libhdf5 - -**Status:** fixed 2026-09-23, after v2.7.0. **Every -release up to and including v2.7.0 is affected** — the encoder was already -wrong in v2.1.0. - -The floating-point datatype message carries the position of the sign bit -(bits 8–15 of its class bit field). `clawhdf5-format` wrote 63 for every -float, which is correct only for `f64`. libhdf5 validates the field, so opening -any `f32` dataset written by this crate failed: - -``` -KeyError: 'Unable to synchronously open object (sign bit position out of bounds)' -``` - -That covers every agent store (`/memory/embeddings`, `norms` and -`activation_weights` are `f32`). `clawhdf5` itself ignores the field on read, -and the interop suites only ever wrote `f64` from our side, so nothing here -noticed. - -**Fix:** the sign position is computed from the type (`bit_offset + -bit_precision - 1`: 15, 31, 63 for half, single, double). Regression tests: -`float_sign_location_is_the_top_bit_of_the_value` (byte level), -`clawhdf5_writes_f32_h5py_reads` and the agent's -`h5py_reads_every_dataset_of_an_agent_store`. - -**Existing files:** an agent store is rewritten in full at every checkpoint, so -it becomes readable by h5py at its next checkpoint with a fixed build. Other -files with `f32` datasets need to be rewritten. - -## Empty datasets we wrote were unreadable by h5py / libhdf5 - -**Status:** fixed 2026-09-23, after v2.7.0. Every -release up to and including v2.7.0 is affected. - -A dataset with no elements was written with a real file address and a storage -size of 0. libhdf5 guards contiguous storage with an overflow check -(`addr + size <= addr`) that is always true when the size is 0, so it rejected -the dataset: - -``` -KeyError: 'Unable to synchronously open object (invalid dataset size, likely file corruption)' -``` - -In practice: every agent store without sessions or a knowledge graph — the -`/sessions` and `/knowledge_graph` datasets are empty until something is added -— could not be read by h5py even once the `f32` bug above was fixed. Found by -the same agent-store interop test. - -**Fix:** an empty contiguous dataset gets the undefined address (all `0xff`), -which is what libhdf5 itself writes. +**Status:** open (by design for now; re-checked 2026-09-28); both are +explicit errors. A path through an external link returns +`FormatError::ExternalLinkUnsupported { filename, object_path }`, and a +dataset created with `external=[...]` storage returns +`FormatError::ExternalDataFilesUnsupported`. If support is added, file +names must be confined to the opened file's directory, as the +virtual-dataset resolver does. ## Range reads (`File::open_storage`) limits -**Status:** open (added 2026-09-26, milestone M2 of -`docs/design/range-reads.md`; remote backends added by M3). `File::open_storage` +**Status:** open (added 2026-09-26 with milestone M2 of +`docs/design/range-reads.md`; updated for M3-M5). `File::open_storage` reads any `clawhdf5_format::storage::Storage` through the whole read API, -every format-crate read path works through `Storage::read_at`/`read_ranges`, and `clawhdf5-remote` serves HTTP(S) and object-store files through a block cache, but: @@ -1008,7 +238,8 @@ cache, but: range; coalescing is the backend's (or the cache's) job. - A group lookup by name in a version-1 (symbol-table) group lists the whole group (dense groups use their name index). Over a range backend that is - one read per symbol-table node and name, per lookup. + one read per symbol-table node and name, per lookup. (The wasm lazy + reader walks the group's B-tree instead; see its entry.) - External virtual-dataset source files are loaded whole through the resolver (`File::set_vds_resolver`), as bytes; they are not read through a `Storage`. @@ -1016,41 +247,29 @@ cache, but: `read_*_zerocopy`) need the file in memory and answer `FormatError::ContiguousStorageRequired` otherwise; `File::as_bytes()` panics for such a file (`File::contiguous_bytes()` is the fallible form). - `LazyFile`, `MmapFile` and the wasm bindings still read a whole file - (`h5rs` and the Python bindings read through `File::storage`, and take - URLs: `h5rs` with its `remote` feature, Python with `clawhdf5.File(url)`). - - `LazyFile`, `MmapFile` and the Python bindings still read a whole file - (`h5rs` reads through `File::storage`, and takes URLs with its `remote` - feature; the wasm reader's `openUrl` reads by range requests since - 2026-09-27, its `open(bytes)` takes a whole file). -- The file's length is read once, at open: a growing file (SWMR) is not - followed (milestone M5). A remote file is pinned at open, so one that - grows is `RemoteError::FileChanged`. -- Not new, but visible through the equivalence tests: a full read through - the file's chunk cache (`read_raw_data_cached`, `read_raw_data_indexed`, - and so `Dataset::read_*`) lists a damaged dataset's chunks in hash-map - order, so which failing chunk it reports can differ from one `File` to - the next (`cve-2025-2310.h5`); the values of a dataset that reads are - not affected. **Fixed 2026-09-27:** the chunk cache keeps the chunks in - the order the index lists them, as the uncached readers do - (`several_damaged_chunks_report_the_same_chunk_every_time`). +- `LazyFile` and `MmapFile` still read a whole local file. `h5rs` and the + Python bindings read through `File::storage` and take URLs (`h5rs` with + its `remote` feature, Python with `clawhdf5.File(url)`); the wasm + reader's `openUrl` reads by range requests, its `open(bytes)` takes a + whole file. +- A growing local file is followed only when opened with + `File::open_swmr` (milestone M5, 2026-09-27; see `docs/design/swmr.md`): + `File::open` and `open_storage` read the length once, at open. SWMR + reading is not available for remote files (a remote file is pinned at + open, so one that grows is `RemoteError::FileChanged`), `MmapFile` or + `LazyFile`, and clawhdf5 has no SWMR writer. ## Remote files (`clawhdf5-remote`) limits -**Status:** open (added 2026-09-26, milestone M3 of -`docs/design/range-reads.md`). +**Status:** open (added 2026-09-26 with milestone M3 of +`docs/design/range-reads.md`). Python (`clawhdf5.File(url)`) and the +browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27. -- **The browser cannot open URLs yet**: the wasm reader's `openUrl` is - milestone M4. Python can (`clawhdf5.File(url)`, since 2026-09-27), but - the default wheel reads plain `http://` only: `https://` needs a wheel - built with `--features https` (rustls with ring, which compiles C), and - `s3://`, `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs). - The Python tests run against an in-process `http.server` only. - -- **Python cannot open URLs yet.** `clawhdf5.File` (PyO3) parses through - `File::as_bytes`, which a remote file does not have. (The browser can - since 2026-09-27: `clawhdf5-wasm`'s `openUrl`, below.) +- **Default builds read plain `http://` only.** `https://` needs the + `https` feature (rustls with ring, which compiles C) and `s3://`, + `gs://`, `az://` the `s3`, `gcs`, `azure` features (aws-lc-rs); the + default Python wheel has none of them. The Python tests run against an + in-process `http.server` only. - **The block size is fixed** (1 MiB unless `CacheConfig` says otherwise). The design's policy of using a paged file's page size as the block size is not implemented, and only the first block is read ahead. @@ -1089,7 +308,8 @@ cache, but: ## `clawhdf5-wasm` (browser) limits -**Status:** open (by design for now; added 2026-09-26, `openUrl` 2026-09-27). +**Status:** open (by design for now; added 2026-09-26, `openUrl` +2026-09-27). - `open()` holds the whole file in memory (it takes its bytes), so a multi-GB local file does not fit a browser tab. A file on a web server @@ -1098,33 +318,27 @@ cache, but: with these limits: - **Round trips:** a call runs as passes over the blocks fetched so far and is re-run after each wave of misses, so a call costs one round - trip per wave, not one for everything: a chunk index is walked a - level (or a node) per round trip, while the chunks of a read are - fetched together. Listing a group asks for every child's object - header, and every node of a level of the group's index, in one pass - (since 2026-09-27; it was one round trip per header block): 3000 - datasets of an h5py file took 6 passes at 1 MiB blocks, 9 for a - `libver="latest"` file (dense links). **Since 2026-09-27 (later):** - 4 and 5 passes (5 and 6 at 64 KiB, from 8 and 11): the index walks go - on past a missing node, and parsers hint what they read next - (`Storage::hint`: node bodies, the heap's blocks, each child's - header), which the lazy reader fetches with a pass's misses. That is - the depth of the chain (index levels, then symbol table nodes or - heap objects, then headers) plus the pass that finishes; it cannot - go lower without reading structures before their addresses are - known. Opening one dataset of a v1 group looks its name up down the - group's B-tree (it read every entry: 74 requests, 193 MB at 1 MiB - blocks for one 64 KiB dataset of the 3000; now 5 requests, 5 MB). - Each pass re-parses what the call reads (CPU, not network). With - headers spread through the file (h5py writes each next to its data) - a listing still fetches most of the file at 1 MiB blocks (192 of - 198 MB; 35 MB in 530 requests at 64 KiB); a smaller `blockSize` - fetches less. Merging nearby requests does not help such a file: the - blocks a listing needs are five or six apart at 64 KiB, so fewer requests would - mean fetching most of the file. Listing it a second time is free at - 64 KiB blocks, but at 1 MiB its metadata blocks (192 MB) exceed the - 64 MiB `cacheSize`, so they are fetched again (the earliest file: 4 - passes, 50 requests); a larger `cacheSize` keeps them. A file's paged + trip per wave: a chunk index is walked a level (or a node) per round + trip, while the chunks of a read are fetched together. Since + 2026-09-27 listing 3000 datasets takes 4 passes for an h5py file and + 5 for a `libver="latest"` file at 1 MiB blocks (5 and 6 at 64 KiB): + index walks go on past a missing node, and parsers hint what they read + next (`Storage::hint`), which the lazy reader fetches with a pass's + misses. That is the depth of the chain (index levels, then symbol + table nodes or heap objects, then headers) plus the pass that + finishes; it cannot go lower without reading structures before their + addresses are known. Opening one dataset of a v1 group looks its name + up down the group's B-tree (5 requests, 5 MB at 1 MiB blocks for one + 64 KiB dataset of the 3000). Each pass re-parses what the call reads + (CPU, not network). + - **Scattered metadata:** with headers spread through the file (h5py + writes each next to its data) a listing still fetches most of the file + at 1 MiB blocks (192 of 198 MB; 35 MB in 530 requests at 64 KiB); a + smaller `blockSize` fetches less. Merging nearby requests does not + help such a file (the blocks a listing needs are five or six apart at + 64 KiB). Listing it a second time is free at 64 KiB blocks, but at + 1 MiB its metadata blocks (192 MB) exceed the 64 MiB `cacheSize`, so + they are fetched again; a larger `cacheSize` keeps them. A file's paged aggregation (metadata in pages) is not used to fetch its metadata in one request. - **Memory:** a call keeps every block it reads until it finishes (the @@ -1139,11 +353,9 @@ cache, but: whole module (every open file on the page), which these limits keep from happening; before 2026-09-27 both did abort it. - **File size:** at most 4 GiB - 1 bytes; a larger file is refused at - open. The format code turns file offsets into `usize` to use them - (with a clean error past it), which is 32 bits on wasm32, so nothing - at 4 GiB or beyond could be read. Offsets between 2 and 4 GiB are - tested (with a mock server); files above 200 MB have not been served - for real. + open (file offsets become `usize`, 32 bits on wasm32). Offsets between + 2 and 4 GiB are tested with a mock server; files above 200 MB have not + been served for real. - **Cross-origin servers** must allow CORS for the page's origin and either expose `Content-Range` (`Access-Control-Expose-Headers`) or answer `HEAD` with `Content-Length`. The file is pinned at open by its @@ -1151,48 +363,53 @@ cache, but: either header, only a change of length is detected. - **A server without range support** (it answers `200`) costs a whole download, up to `maxDownload` (512 MiB, at most 1 GiB), or an error - with `fallback: "error"`. Every body, this one and each `206`, is read - as it arrives and cut off past its limit (the range asked for, or - `maxDownload`): a server cannot make the page buffer more. + with `fallback: "error"`. Every body is read as it arrives and cut off + past its limit: a server cannot make the page buffer more. - Fixed block size (`blockSize`, 1 MiB by default); a paged file's page size is not used. No retries: a failed request fails the call (calling again retries it; what was fetched stays cached). - Tested under Node 22 and headless Chromium (Playwright's build) against - a local server, cross-origin included (a page on 127.0.0.1 reading a - file from localhost, with and without exposed headers); not in Firefox - or Safari. - - The native corpus comparison (`tests/lazy.rs` with - `CLAWHDF5_WASM_CORPUS`) fails now and then on one CVE file, - `cve-2025-2310.h5`: two of its datasets have more than one bad chunk, - and which chunk's error is reported depends on the iteration order of - the chunk index (a `HashMap`, seeded per process), so the lazy and - the range-storage reads can name different errors. Both are errors; - not specific to `openUrl` (it predates it). **Fixed 2026-09-27:** the - chunk cache keeps the index's chunk order, so every read path names the - same (first) damaged chunk. + a local server, cross-origin included; not in Firefox or Safari. - Compound, reference, opaque, bitfield, time and VL-sequence datasets are refused with an error naming the type; attributes of those types come back - as `value: null` with their `dtype`. + as `value: null` with their `dtype`. (VL strings read, with h5py's + values, through the same `VlResolver` as `File` and `h5rs`.) - No Zstd or SZIP (both link C): such datasets fail with `unsupported filter: 32015` / `: 4`. pcodec is not enabled either. - External links and virtual-dataset sources in other files cannot be followed (no file system). -- Variable-length string datasets are read by decoding `read_selection`'s - bytes with `clawhdf5_format::vl_data` in the wasm crate; `File` itself still - cannot (see the audit gaps above). (`File` can since 2026-09-26. Since - 2026-09-26 the wasm crate resolves them with the same `VlResolver` as - `File` and `h5rs`, so all three return h5py's values.) + +## Conformance: `bad_nbit_parms_walk.h5` flips between ref-bug and our-error + +**Status:** open (found 2026-09-28). Report noise, not a clawhdf5 bug: the +ok count (602 of 697) and the gate are unaffected. + +`hdf5/test/testfiles/bad_nbit_parms_walk.h5` `/Nbit_int_data_le` is an +object clawhdf5 refuses and HDF5 2.0 reads only by over-reading memory (see +[the last non-ok files](#conformance-the-last-non-ok-files)). +`conformance/ref_bugs.py` counts it as *ref-bug* only when h5py's values +differ across its six differently set-up reads. The committed +`CONFORMANCE.md` (run of 2026-09-28 04:29 UTC, `bf5a163`) saw one distinct +result in six reads, so it reports the file as **1 our-error** and 2 +ref-bug; a rerun on tank the same day (16:06 UTC, `conformance/run.sh +--no-fetch`, docs-only changes on top of `9b5803f`) saw three distinct +results and reports 0 our-error, 3 ref-bug. Whether libhdf5's over-read +changes between processes depends on heap layout, so six reads do not +always expose it. A fix belongs in `conformance/ref_bugs.py` (more or more +varied reads for this object) or in documenting the file as a known +refusal; neither is done. ## The Node.js package (`packages/clawhdf5-node`) does not work -**Status:** open (found 2026-09-25). Unpublished; not built or tested in CI. +**Status:** open (found 2026-09-25; re-checked 2026-09-28, unchanged). +Unpublished; not built or tested in CI. The TypeScript wrapper over `crates/clawhdf5-napi` has never run successfully: - napi-rs converts `#[napi(object)]` fields to camelCase, but the wrapper reads snake_case (`r.line_range`, `s.total_records`, `s.working_count`, …), so every stats and consolidation field comes back `undefined` - (`src/index.ts:76-120`). + (`src/index.ts`). - It loads `../clawhdf5.node`, but `napi build --platform` produces `clawhdf5..node`; `main` points at `index.js` while `tsc` writes to `dist/`; `napi prepublish` expects per-platform packages that are not @@ -1205,3 +422,387 @@ The TypeScript wrapper over `crates/clawhdf5-napi` has never run successfully: It was written for an OpenClaw integration that is not being pursued (see `docs/openclaw.md`). Fix and add CI, or remove it, before anyone depends on it. + +--- + +# Fixed (history) + +Newest first. "Before any release" means no tagged release (v2.7.0 and +earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date +given. + +## `ObjectHeader::parse` 4% slower after range-read M2/M3 + +**Status:** fixed 2026-09-27 (`96086ad`, PR #21), before any release. +Speed only; values were always correct. + +Measured on an idle tank (`BENCHMARKS.md`, "Local metadata and data reads +after range-read M2/M3"): `object_header_parse_x401` 23.86 → 24.86 µs +median (+4.2%) from `8f59b2e` to `7a8fae0`, after the parser began reading +continuation chunks from a bounded queue (`a69c5be`). The remaining cost was +the call to the version-1 message loop, kept out of line by +`#[inline(never)]`; inlined, it is 1.0% and 2.6% below `8f59b2e` in two idle +A/B runs (`BENCHMARKS.md`, "`ObjectHeader::parse` back at 8f59b2e's speed"). +A reported 5.6% slowdown of contiguous hyperslab reads, measured under load, +was noise and was withdrawn. + +## Conformance: the last non-ok files + +**Status:** classified 2026-09-27 (PR #21) — none is a clawhdf5 bug. +Result (`conformance/run.sh --no-fetch`, tank): 602 of 697 ok (was 600), +0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read. + +- **2 mismatches were h5py's big-endian VL bug** + (`NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64`, + `hdf5/tools/test/testfiles/tcomplex_be.h5` + `/VariableLengthDatasetFloatComplex`): h5py returns the file's + big-endian bytes under a little-endian dtype. `conformance/ref.py` now + detects the bug in the installed h5py and relabels such elements, so the + values are compared; both files are *ok*. +- **3 our-errors are objects HDF5 2.0 reads only by over-reading memory** + (`cve-2025-2308.h5` `/Scale_offset_long_long_data_le`, + `cve-2025-44904.h5` `/Scale_offset_float_data_le`, + `bad_nbit_parms_walk.h5` `/Nbit_int_data_le`). libhdf5's output for them + is not determined by the file (it changes from run to run and with + `MALLOC_PERTURB_`), and libhdf5's develop branch refuses all three. We + keep refusing them; `conformance/ref_bugs.py` re-checks every run and + classifies such a file *ref-bug* only while its values keep changing. + +## Damaged chunked datasets reported a different failing chunk per open + +**Status:** fixed 2026-09-27 (`4ad8073`, PR #19). Errors only; the values of +readable datasets were never affected. + +A full read through the file's chunk cache listed a damaged dataset's +chunks in hash-map order, seeded per `File`, so two opens of +`cve-2025-2310.h5` could name different failing chunks, and the storage and +wasm corpus comparisons failed now and then. The cache now keeps the chunk +index's order, as the uncached readers do +(`several_damaged_chunks_report_the_same_chunk_every_time`). + +## Files a SWMR writer had open could not be read past a stale end of file + +**Status:** fixed 2026-09-27 (PR #19), before any release: reads have been +bounded by the recorded end of file only since `7d7a7e7` (2026-09-26). + +A libhdf5 SWMR writer does not keep the superblock's end-of-file address up +to date (a mid-write copy records 715 in a 6 030-byte file), so every reader +bounded such a file there: it listed, but chunked reads failed and `h5rs +check` reported chunk indexes past the end. Never wrong data. For a v3 +superblock with the SWMR-write flag the data now ends at the end of the +file, as libhdf5's SWMR reader reads it. Test: +`crates/clawhdf5/tests/swmr_interop.rs` (fixture +`tests/fixtures/swmr_mid_write.h5`). A file still being written is read +with `File::open_swmr` (`docs/design/swmr.md`). + +## Shrinking a chunked dataset with no recorded maximum scrambled it + +**Status:** fixed 2026-09-27 (PR #19), before any release +(`FileEditor::resize` shipped on main in PR #18). + +clawhdf5's writer stored no maximum dimensions for a chunked dataset +created without a `maxshape`; a Fixed Array index then places chunks by the +current dimensions, and `FileEditor::resize` changed only those, so a +shrink moved every chunk (wrong values in every reader) and the dataset +could not grow back. The editor now records the maximum libhdf5 would have +written before the first resize, and the writer records it for every +chunked dataset. Test: `crates/clawhdf5/tests/edit_resize_interop.rs`. + +**What users must do:** datasets written before the fix (v2.7.0 and +earlier, and `main` before 2026-09-27) have no recorded maximum. The fixed +editor handles them; **libhdf5 (h5py `Dataset.resize`, `H5Dset_extent`) +does not** — it scrambles them the same way and lets them grow past their +Fixed Array. Resize them once with a fixed `FileEditor` (a resize to the +same shape changes nothing; shrink and grow back) before letting libhdf5 +resize them. A dataset already shrunk by the unfixed editor holds misplaced +chunks; rewrite it from a good copy. + +## Fletcher-32 checksums disagreed with libhdf5 on about 1 chunk in 32768 + +**Status:** fixed 2026-09-26 (PR #18), after v2.7.0. **Every release +(v2.1.0 to v2.7.0) is affected**, in both directions. + +Our Fletcher-32 reduced its sums with `% 65535`; libhdf5 folds them with +`(s & 0xffff) + (s >> 16)`, which differs where a sum is a non-zero +multiple of 65535. Chunks we wrote with such a sum are refused by h5py and +libhdf5 ("filter returned failure during read"); chunks libhdf5 wrote with +one were refused here with `Fletcher32Mismatch` (the data itself was never +wrong). The fix, `clawhdf5_format::checksum::fletcher32`, ports +`H5_checksum_fletcher32` and also accepts the byte-swapped (libhdf5 ≤ 1.6.2) +and the old clawhdf5 forms. Test: `crates/clawhdf5/tests/fletcher32_interop.rs`. + +**What users must do:** a Fletcher-32 dataset written by v2.7.0 or earlier +may hold chunks libhdf5 cannot read; a fixed build reads them. Rewrite such +datasets with a fixed build before handing the file to libhdf5 or h5py. + +## LZF/Blosc chunks written with a stale filter mask + +**Status:** fixed 2026-09-26 (PR #17), before any release (v2.7.0 and +earlier write neither filter). + +`FileBuilder` (and `FileEditor`) stored every LZF or Blosc chunk with +filter mask 0, where libhdf5 stores a chunk whose output is no smaller than +the chunk raw with the filter's mask bit set; after libhdf5 rewrote such a +chunk h5py could no longer read the dataset. Both now use +`clawhdf5_format::filters::compress_chunk_masked`. **What users must do:** +files written before the fix read correctly; rewrite them before letting +libhdf5 modify them. + +## Concurrent and contiguous read performance + +**Status:** fixed 2026-09-26 (PRs #15 and #16). Speed only. + +Measured on tank against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`, "First run, +before the read fixes"): full reads of chunked data from 16 threads through +one `File` stopped scaling at about 4 threads (880 MB/s vs 4424 MB/s for 16 +h5py processes), and contiguous datasets read 4x slower than h5py on one +thread. Causes: every full read queued on a one-thread decode pool, fresh +buffers per chunk and a second copy of the output, 4 KiB page faults on the +output buffer, and element-by-element hyperslab copies. Re-measured at +`c5334b1` (`BENCHMARKS.md`, "Results after in-place chunk decoding"): +16-thread chunked deflate reads 4944 MB/s against 3135 for 16 h5py +processes (1.58x), contiguous reads 1.29x h5py on one thread. Tests: +`single_thread_decode_pool.rs`, `busy_decode_pool.rs`. + +## Scale-offset data read back wrong values + +**Status:** fixed 2026-09-26 (PR #16), after v2.7.0. **Every release that +decoded the scale-offset filter (v2.2.0 to v2.7.0) is affected**, on +ordinary files h5py writes, with no error. + +Of 1480 scale-offset datasets h5py wrote (every integer type, `f4`/`f8`, +both byte orders, many fill values and scale factors), v2.7.0 read 332 +differently from h5py: **151 returned wrong values with no error**, 181 +failed. In every case libhdf5 had stored a chunk at full width (`minbits` +equal to the type's width), which was decoded as offsets from `minval`; +two smaller differences (codes start at byte 21 whatever `minval`'s +recorded size; `minbits` 0 with a fill value is all fill) were fixed with +it. Test: `crates/clawhdf5/tests/scaleoffset_interop.rs`. + +**What users must do:** nothing to the files — they were always right; re-read +them with a fixed build. + +## Crafted global heaps exhaust the variable-length reader's memory + +**Status:** fixed 2026-09-26 (PR #15). Every earlier release is affected +through `read_vl_strings`. + +Global heap collections nested inside one another's object data made +retained memory O(elements × file size) (a 744 KB file reached 1.58 GB) and +parse time O(elements × objects). `VlResolver` now caches object locations +within a 32 MiB budget and refuses overlapping collections, and collections +or objects running past their bounds are refused. Test: +`crates/clawhdf5-format/tests/vl_heap_bounds.rs`. The remaining +one-object-many-elements case is listed under +[HDF5 features still unsupported](#hdf5-features-still-unsupported). + +## Gaps found by the 2026-09-25 HDF5 audit (fixed parts) + +**Status:** fixed 2026-09-25 and 2026-09-26 (PRs #11 to #17). Each was an +error unless marked **wrong data**; every release up to v2.7.0 has them. +What remains open is under +[HDF5 features still unsupported](#hdf5-features-still-unsupported). + +- **Layout message versions 1 and 2** (HDF5 1.6-era files; 84 of the 686 + sweep files) and **compound datatype version 1 array members** (**wrong + data**: an array member read as one scalar) — fixed 2026-09-25. +- **Virtual datasets:** unmapped regions read as 0 instead of the fill + value (**wrong data**); printf-style and unlimited mappings; hyperslab + selection versions 1-3; the version-1 mapping list with a 2.0 low bound — + fixed 2026-09-25. +- **User blocks**, **old-style shared messages**, **user-defined link + types**, **dense groups over about 22 000 links**, **soft links in + `datasets()`**, a local-heap free list outside the heap (**wrong data**: + garbage names) — fixed 2026-09-25. +- **Dense attributes stored as huge/tiny/filtered fractal-heap objects** + (real NetCDF files, `issue671.nc`); an unreadable attribute no longer + hides the others (`attrs_with_errors()`) — fixed 2026-09-25. +- **VL strings through `File`** (`read_string`, `read_vlen::()`, + `File::decode_strings`/`decode_vlen`), **VL data with 4-byte offsets**, + **metadata cache images** — fixed 2026-09-26. +- **Filters:** LZF, bitshuffle, bzip2 and Blosc 1 (read and write), + Blosc2 and ZFP (read) — fixed 2026-09-26, pure Rust. A chunk whose + filters decode to fewer bytes than the chunk read with zeros for the rest + (**wrong data**); now an error, and unfiltered chunks of the wrong stored + size are refused. A hostile Blosc chunk panicked with overflow checks. +- **Header checks:** 18 CVE objects libhdf5 refuses were read (some as + wrong data); object headers, datatypes, chunk dimensions and chunk-index + offsets are now checked as libhdf5 checks them, a dataspace whose storage + size overflows is refused at open, and the superblock extension is + decoded at open with libhdf5's checks — fixed 2026-09-26. +- **N-Bit/scale-offset on corrupt files** (`cve-2025-2308`, + `bad_nbit_parms_walk.h5`): not our bug — see + [the last non-ok files](#conformance-the-last-non-ok-files). +- **Writer:** groups nest to any depth with soft, hard and external links; + attribute creation order (`track_order`); dense link/attribute indexes of + any size (v2 B-trees of any depth); dense storage past 512 KiB was + written unreadable (it affected v2.7.0); libhdf5 could not add a link to + a group we wrote (no Group Info message); B-tree v2 chunk indexes larger + than one leaf — fixed 2026-09-26. Tests: `writer_groups_interop.rs`, + `deep_btree_interop.rs`. + +## Silent wrong data found by the 2026-09-25 HDF5 audit + +**Status:** fixed 2026-09-25 (PR #11), after v2.7.0. **Every release up to +and including v2.7.0 is affected.** + +The audit (686 public files, 567 read cases and 96 write cases against h5py +3.16 / HDF5 1.10-2.0 and h5dump 1.14.6) found these values returned wrong +**without an error**: + +| Area | What happened | Who is affected | +|---|---|---| +| Chunk index (read) | Fixed/Extensible Array indexes laid out by the current shape, not the max shape: chunks returned from the wrong place | any file with a max shape larger than its shape and `libver='latest'` (h5py `maxshape=(10, None)`, `(20, 10)`) | +| Chunk index (write) | Extensible Array chunks from index 244 on never indexed (read as 0); unlimited dimension not first: data scrambled | files we wrote with one unlimited dimension and > 244 chunks, or e.g. `maxshape=(20, None)` | +| 4-byte offsets | unfiltered chunked datasets read as zeros | files created with `sizeof_addr = 4` | +| Filter mask | any skipped filter skipped the whole pipeline | files with partially filtered chunks (optional filters, direct chunk writes) | +| Numeric reads | float read as integer returned the bit pattern; narrowing integer reads kept the low bits; bfloat16 decoded as IEEE half | `read_i32`/`read_i64`/`read_u64` callers on float or wider data; HDF5 2.0 bf16 data | +| SZIP | garbage or zeros | every libhdf5-written SZIP dataset | +| Scale-offset | float values 1 ULP off | libhdf5 D-scale float data | +| Shared fill value | read as zero fill | fill values stored as shared messages | +| VL sequences | `read_vl_bytes` truncated non-byte base types | VL int/float sequences | +| Chunk cache | two threads reading two chunked datasets through one `File` could get each other's chunks | multi-threaded readers, including Python with the GIL released | + +It also found files we wrote that libhdf5 **refuses**, fixed with it: Fixed +Array datasets with more than 1 024 chunks, header messages over 64 KiB, +Reference/Opaque/BitField/Time datatypes, files written with +`with_page_size`, several unlimited dimensions, a finite max shape larger +than the shape, an empty-string attribute (which broke every attribute on +its object), and rotated `FillTime` codes. Our LZ4 and Zstd output could not +be read by libhdf5's registered plugins, and our pcodec filter used Granular +BitRound's ID. Details: `CHANGELOG.md`, Correctness and Interop. + +**What users must do:** re-read affected files with a fixed build; rewrite +files clawhdf5 wrote in the affected cases (one unlimited dimension with more +than 244 chunks, an unlimited dimension not first, the refused cases) before +handing them to libhdf5. + +The sweep became `conformance/run.sh` (corpora pinned by commit; nightly by +`.gitea/workflows/conformance.yml`); current numbers are in +`CONFORMANCE.md`. No sweep run found a panic, hang or crash, including on +the 147 CVE and fuzzer files on some of which h5dump 1.14.6 and h5py/HDF5 +2.0 segfault or abort. + +## Every `f32` dataset we wrote was unreadable by h5py / libhdf5 + +**Status:** fixed 2026-09-23 (PR #4), after v2.7.0. **Every release up to +and including v2.7.0 is affected** (the encoder was already wrong in +v2.1.0). + +The floating-point datatype message's sign-bit position was written as 63 +for every float; libhdf5 validates it, so every `f32` dataset — including +every agent store's `/memory/embeddings`, `norms` and `activation_weights` +— failed to open ("sign bit position out of bounds"). It is now +`bit_offset + bit_precision - 1`. Tests: +`float_sign_location_is_the_top_bit_of_the_value`, +`clawhdf5_writes_f32_h5py_reads`, the agent's +`h5py_reads_every_dataset_of_an_agent_store`. + +**What users must do:** an agent store is rewritten at every checkpoint, so +it becomes readable by h5py at its next checkpoint with a fixed build. Other +files with `f32` datasets need to be rewritten. + +## Empty datasets we wrote were unreadable by h5py / libhdf5 + +**Status:** fixed 2026-09-23 (PR #4), after v2.7.0. Every release up to and +including v2.7.0 is affected. + +An empty dataset was written with a real address and size 0, which +libhdf5's overflow check rejects ("invalid dataset size, likely file +corruption") — every agent store without sessions or a knowledge graph. An +empty contiguous dataset now gets the undefined address, as libhdf5 writes. +**What users must do:** as for `f32` above (agent stores heal at their next +checkpoint; rewrite other files). + +## Extensible Array chunk indexes read back wrong data past the inline elements + +**Status:** fixed 2026-09-20, in v2.7.0. **Every release up to and including +v2.6.0 is affected.** + +A dataset with exactly one unlimited dimension is indexed by an Extensible +Array, whose data and super block layout was computed wrongly: with more +than 36 chunks values came back from the wrong chunks **with no error** (37 +chunks: 1 element wrong; 400: 364 wrong), and from about 1000 chunks the +read failed. Fixed with interop tests against HDF5 2.0 across every +boundary. **What users must do:** re-read with v2.7.0 or later. (The +writer's own 244-chunk bug is under the +[2026-09-25 audit](#silent-wrong-data-found-by-the-2026-09-25-hdf5-audit).) + +## Crafted B-tree v2 structures crash or exhaust the reader + +**Status:** fixed 2026-09-20, in v2.7.0. **Every release up to and including +v2.6.0 is affected** (for untrusted files). + +A self-referencing node under a header claiming 65 535 levels overflowed the +stack and aborted the process (under 100 bytes), and shared children made +traversal visit a node fan-out^depth times (memory exhaustion from ~5 KB). +Depth is now capped at 64 and traversal stops past the records the file +could hold. + +## Python interop suites skip silently when no interpreter has h5py + +**Status:** fixed 2026-09-19 (`a29c1b2`), in v2.6.0. + +On a PEP 668 system every h5py/netCDF4 interop suite skipped and CI stayed +green. The probes now read `CLAWHDF5_PYTHON`, `scripts/ci-test.sh` picks up +`.venv/bin/python`, and `CLAWHDF5_REQUIRE_INTEROP=1` (set in CI) makes a +missing interpreter a failure. To restore coverage on a fresh checkout: + +```bash +python3 -m venv .venv && .venv/bin/pip install h5py numpy netCDF4 +``` + +## B-tree v2 chunk index (layout v4, index type 5) is not supported + +**Status:** fixed 2026-09-19, in v2.5.0. Datasets with two or more +unlimited dimensions written with `libver='latest'` failed to read in +earlier releases; record types 10 and 11 are now decoded on every read path. + +## Attributes with unsupported datatypes are silently dropped + +**Status:** fixed 2026-09-19. `attrs()` omitted booleans, complex, compound +and reference attributes without notice, and cast `u64` arrays to `i64`. +Booleans now decode as 0/1, `AttrValue::U64Array` keeps unsigned arrays, and +`AttrValue::Raw` carries anything else. (Multi-dimensional numeric +attributes are still flattened; see +[HDF5 features still unsupported](#hdf5-features-still-unsupported).) + +## Compound datatype versions 1 and 2 are mis-parsed (default libver files) + +**Status:** fixed 2026-09-19. Every compound dataset written with default +libver bounds (plain `h5py.File(path, 'w')`) failed to read: the v1 legacy +array fields are 28 bytes, not 24, and v2 keeps the name padding. + +## `clawhdf5-gpu` `gpu_tests` can hang under the default parallel test runner + +**Status:** fixed 2026-09-19 (`706189c`), in v2.3.0. Tests now hold a +process-wide lock while they own a device, and `GpuAccelerator` readback +times out after 30 s with `GpuError::BufferMap`. + +## Revised reference datatype (class 7, version 4) is not parsed + +**Status:** fixed 2026-09-19 for object references. `H5T_STD_REF` object +references (`ReferenceType::Object2`) decode, tested against a file made +with libhdf5 through ctypes (`tests/fixtures/gen_std_ref.py`). Region and +attribute references are recognised but not decoded — see +[HDF5 features still unsupported](#hdf5-features-still-unsupported). + +## Native complex datatype (class 11) is mis-parsed (HDF5 2.0) + +**Status:** fixed 2026-09-18 (`b55b7db`), in v2.2.0. HDF5 2.0 native complex +types were parsed as a compound member list (garbage or `UnexpectedEof`); +class 11 now parses its base type and surfaces as an `{r, i}` compound. +h5py's default complex mapping (a compound) was never affected. + +## Compound datatype message version 5 is not parsed (HDF5 2.0) + +**Status:** fixed on `main` in `a13ff51` (2026-06-03), in v2.2.0; **not in +v2.1.0**, which was cut five commits earlier. Reported by M. Scot +Breitenfeld (The HDF Group), 2026-09-08, against v2.1.0. + +v2.1.0 rejects any compound dataset written by HDF5 2.0 with +`libver='latest'` (`InvalidDatatypeVersion { class: 6, version: 5 }`). +Versions 3-5 are now accepted for compound and array datatypes, and data +layout message version 5 too (every chunked dataset HDF5 2.0 writes). Tests: +`test_compound_v5_from_hdf5_2_0`, `test_array_v5_from_hdf5_2_0`, +`writer_h5py_tests.rs::read_h5py_generated_compound`. diff --git a/docs/openclaw.md b/docs/openclaw.md index 6e76a08..e6ac2f0 100644 --- a/docs/openclaw.md +++ b/docs/openclaw.md @@ -71,4 +71,6 @@ Building blocks, usable as a library today, but not an OpenClaw plugin: and export rewrites every heading as `##`. - `crates/clawhdf5-napi` and `packages/clawhdf5-node` — Node bindings and a TypeScript wrapper. **Not published, not built or tested in CI, and known to - be broken**; see `docs/known-issues.md`. + be broken**; see + [`docs/known-issues.md`](known-issues.md#the-nodejs-package-packagesclawhdf5-node-does-not-work) + (re-checked 2026-09-28: unchanged).