diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 7910b60..75024aa 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -63,15 +63,15 @@ average below 2. |---|---|---|---|---| | Agent memory search, `HDF5Memory::hybrid_search` p50 | 0.49 ms at 10K, 4.69 ms at 100K records | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` | [Current: search harness](#current-search-harness-2026-09-24) | | LongMemEval `longmemeval_s` (full haystack), default hybrid 0.4/0.6, turn-level retrieval Hit@5 (not QA accuracy) | 81.4% | 2026-09-27, tank, search code of `7a8fae0` | `longmemeval_bench … --embeddings weights/all-minilm-l6-v2` | [Re-run with real embeddings](#re-run-with-real-embeddings-2026-09-27-tank), [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500) | -| Loaded store memory, 100K × 384 | 399 MiB (2.72x raw) with the `f32` index; 256 MiB (1.74x) with the int8 index (int8 side not re-run since it was first measured) | 2026-09-24, tank, `5c8323c` | `search_harness -- --footprint --full [--int8]` | [Memory footprint](#memory-footprint), [Quantising the index copy](#quantising-the-index-copy-quantized_index) | -| int8 index vs `f32` index, QPS at equal recall | 1.63x (x86-64 AVX2), 1.18x (Raspberry Pi 5, `SDOT`) | undated; not re-checked against the 2026-09-24 `f32` figure | `search_harness -- --full` | [Quantising the index copy](#quantising-the-index-copy-quantized_index), [On ARM](#on-arm-raspberry-pi-5-cortex-a76) | +| Loaded store memory, 100K × 384 | 399 MiB (2.72x raw) with the `f32` index; 256 MiB (1.74x) with the int8 index (int8 side not re-run since it was first measured) | `f32`: 2026-09-24, tank, `5c8323c`; int8: 2026-09-19 (`c0a9206`), machine not recorded | `search_harness -- --footprint --full [--int8]` | [Memory footprint](#memory-footprint), [Quantising the index copy](#quantising-the-index-copy-quantized_index) | +| int8 index vs `f32` index, QPS at equal recall | 1.63x (x86-64 AVX2), 1.18x (Raspberry Pi 5, `SDOT`) | x86: 2026-09-20 (`dea02f5`), machine not recorded; Pi 5: 2026-09-21 (`114a2df`); not re-checked against the 2026-09-24 `f32` figure | `search_harness -- --full` | [Quantising the index copy](#quantising-the-index-copy-quantized_index), [On ARM](#on-arm-raspberry-pi-5-cortex-a76) | | `float16` store file size, 100K × 384 | 80.8 MiB vs 154.0 MiB `f32` (48% smaller) | 2026-09-23, tank | `search_harness -- --float16-study --full` | [float16 embedding storage](#float16-embedding-storage-memoryconfigfloat16) | | Full reads of chunked deflate data, 16 threads on one `File` | 4944 MB/s, 1.58x 16 h5py processes (noisy run: compare ratios, not MB/s) | 2026-09-26, tank, `c5334b1` | `concurrent_read` + `concurrent_read_h5py.py` | [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) | | Same, clawhdf5 only, against the build before range-read M2/M3 | 8525 MB/s vs 6258 (+36%); contiguous and metadata reads at parity | 2026-09-27, tank (idle), `7a8fae0` vs `8f59b2e` | `concurrent_read --decode-threads 1 --reps 3` | [Local metadata and data reads after range-read M2/M3](#local-metadata-and-data-reads-after-range-read-m2m3-2026-09-27-tank) | | `ObjectHeader::parse` (401 headers) | 23.5–23.6 µs, 1.0–2.6% below `8f59b2e` | 2026-09-27, tank (idle), `96086ad` | `cargo bench -p clawhdf5 --bench local_metadata_bench` | [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) | | Selection reads, 64 MB chunked + deflate `f64` | full 63.2 ms; one 64 × 64 window 0.18 ms | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin read_harness` | [Current: read harness](#current-read-harness-2026-09-24) | | Deflate backend, zlib-rs (default) vs zlib-ng | within 6% on every HDF5 read/write path | 2026-09-23, tank | `cargo bench -p clawhdf5-filters --bench deflate_bench` (and the two commands with it) | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng) | -| vs libhdf5 1.14.6: chunked deflate-6 write 512×512 / 128 attributes / 64 groups | 45.3x / 10.3x / 10.6x | 2026-08-03, tank | `cargo bench -p clawhdf5-bench --features libhdf5-compare` | [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) | +| vs libhdf5 1.14.6: chunked deflate-6 write 512×512 / 128 attributes / 64 groups | 35x (1.46 vs 51.4 ms, pure-Rust deflate) / 10.3x / 10.6x | write 2026-09-23, tank; attributes and groups 2026-08-03, tank | `cargo bench -p clawhdf5-bench --bench h5bench_write --features libhdf5-compare -- '^write_2d_chunked/'`; `cargo bench -p clawhdf5-bench --features libhdf5-compare` | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng), [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) | | Signed checkpoints | about 20% of a checkpoint (598 vs 495 ms at 100K) | 2026-09-25, tank | `search_harness -- --signing-study --full` | [Signed checkpoints](#signed-checkpoints) | --- @@ -2505,7 +2505,8 @@ global file mutex and flushes to disk on every attribute write or group creation > 2026-06-30). The newest run of this table is > [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) > (2026-08-03), which reproduces every row within ~15% except chunked -> write (45.3x on tank). +> write (45.3x on tank; 35x on 2026-09-23 with the pure-Rust deflate, see +> [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng)). | Workload | clawhdf5 | libhdf5 | Speedup | |----------|----------|---------|---------| diff --git a/CLAUDE.md b/CLAUDE.md index 428c6e2..ea1bf31 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -72,7 +72,7 @@ and `docs/design/swmr.md`, `CHANGELOG.md` (full detail of every fix). ## HDF5 library: invariants and gotchas - **Remote/range reads** (`docs/design/range-reads.md`, M0-M5 merged in PRs - #17-#21): every format-crate read path goes through `Storage` + #17-#19, M4 listing costs cut in #21): every format-crate read path goes through `Storage` (`read_at`/`read_ranges`/`hint`). `File::open_storage` takes any `Storage`; `clawhdf5_remote::open_url` wraps HTTP (`HttpStorage`, ureq) or `ObjectStoreStorage` in `BlockCache` (1 MiB blocks, LRU budget, in-flight @@ -197,8 +197,8 @@ bash scripts/ci-test.sh # everything CI runs (see below) `cargo bench --no-run`; `check-nostd.sh`; an optional fuzz smoke run (`CLAWHDF5_FUZZ_SECONDS`). - **`ci.yml` `test-arm64`** (`linux_arm64`; `vision-01` host mode, - `vision-02` Docker — steps must work in both): clippy and tests of - `clawhdf5-accel`, `-ann`, `-format`; the only place the NEON kernels build. + `vision-02` Docker — steps must work in both): clippy of + `clawhdf5-accel`, tests of `-accel`, `-ann`, `-format`; the only place the NEON kernels build. - **`conformance.yml`** (nightly 03:17 UTC and manual): probe unit tests, `conformance/test_ref.py`, then `conformance/run.sh` (gate: `conformance/check.py` against `baseline.json`). @@ -214,7 +214,7 @@ frozen at 0.6.1). CLAWHDF5_PYTHON=.venv/bin/python bash conformance/run.sh --no-fetch # writes CONFORMANCE.md ``` Reads 697 files of eight pinned corpora with clawhdf5 and h5py and compares -them object by object (602 ok as of 2026-09-27). `CONFORMANCE.md` is +them object by object (602 ok in the run of 2026-09-28). `CONFORMANCE.md` is generated — never hand-edit it (its wording lives in `conformance/report.py`). Use `--update-baseline` only after an intended change in results. `CONFORMANCE_CACHE` points at an existing corpus cache (`conformance/.cache`, @@ -261,7 +261,7 @@ python -m pytest crates/clawhdf5-py/tests # compares with h5py; editing tes ### CLI ```bash cargo run -p clawhdf5-cli -- --help -# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot +# create, save, search, recall, stats, flush-wal, agents-md, export, snapshot, keygen, verify ``` ## Integration diff --git a/README.md b/README.md index 889e631..1b85342 100644 --- a/README.md +++ b/README.md @@ -57,7 +57,9 @@ the end of a chunk, short unfiltered chunks, an N-Bit parameter list one value short) that HDF5 2.0 returns only by reading past a buffer; clawhdf5 refuses them, as libhdf5's development branch and its own `test_filter_bad_params` do. Details and evidence in -[CONFORMANCE.md § Reference bugs](CONFORMANCE.md#reference-bugs). The +[CONFORMANCE.md § Reference bugs](CONFORMANCE.md#reference-bugs); the +N-Bit file is counted as ref-bug or our-error depending on the run +([known-issues](docs/known-issues.md#conformance-bad_nbit_parms_walkh5-flips-between-ref-bug-and-our-error)). The run is a nightly CI job (`.gitea/workflows/conformance.yml`) that fails on any panic, hang or crash, or on an ok file that stops being ok. @@ -114,10 +116,10 @@ Limits and open issues, with dates, are in | **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | | **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex (HDF5 2.0 class 11) | Variable-length strings and sequences, object references | Writing variable-length data; decoding region and attribute references; x87 long double and binary128 | | **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does); fill values; resizable datasets; virtual datasets (read limits in known-issues) | Chunk indexes v1 B-tree and implicit (the editor also changes them) | External raw data files (explicit error) | -| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4, Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | +| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | | **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; filters this build cannot encode (refused before any write) | | **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | -| **Bindings** | Python (read, `'w'` for numeric arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound or VL-sequence datasets | Node.js (the package does not work; see known-issues) | +| **Bindings** | Python (read, `'w'` for numeric arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) | Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`, `blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure @@ -127,13 +129,15 @@ that only clawhdf5 reads. **C dependencies, precisely.** The core crates build no C by default: no libhdf5, and deflate is [zlib-rs](https://github.com/trifectatechfoundation/zlib-rs), -which produces output byte-identical to zlib-ng and matches its HDF5 read and -write speed within 6% -([BENCHMARKS.md](BENCHMARKS.md#deflate-backend-zlib-rs-vs-zlib-ng)). CI +whose output was byte-identical to zlib-ng's at levels 1, 6 and 9 on the +benchmark inputs and which matched its HDF5 read and write speed within 6% +(tank, 2026-09-23, [BENCHMARKS.md](BENCHMARKS.md#deflate-backend-zlib-rs-vs-zlib-ng)). CI fails if a C-building crate enters their default dependency tree. C comes in only when you ask: `fast-deflate` (zlib-ng, needs cmake), `zstd`, `szip`, `https` and the cloud stores (ring / aws-lc-rs), the BLAS backends, -`clawhdf5-migrate` (bundled SQLite) and the Node.js bindings. +`clawhdf5-migrate` (bundled SQLite) and the Node.js bindings. One exception +links rather than builds C: on macOS the default `system-zlib-decompress` +feature inflates with the system libz first. ## Install @@ -334,10 +338,12 @@ a knowledge graph — in one HDF5 file (readable by h5py), with: MiniLM embeddings) turn-level Hit@5 is **81.4%** — retrieval recall, not the official QA-accuracy metric (tank, re-run 2026-09-27, [BENCHMARKS.md](BENCHMARKS.md#longmemeval-results)). -- **Compact by default**: float16 embeddings on disk (48% smaller at 100K) - and an int8 index copy with exact re-scoring — 1.74x the raw vectors in - memory at 100K instead of 2.72x, and 1.63x the QPS at equal recall on - AVX2 (paired runs; see BENCHMARKS.md for which rows were re-run). +- **Compact by default**: float16 embeddings on disk (48% smaller at 100K; + tank, 2026-09-23) and an int8 index copy with exact re-scoring — 1.74x + the raw vectors in memory at 100K instead of 2.72x, and 1.63x the QPS at + equal recall on AVX2 (int8 side measured 2026-09-19/20, machine not + recorded, not re-run since; see + [BENCHMARKS.md](BENCHMARKS.md#quantising-the-index-copy-quantized_index)). - **Durability**: a write-ahead log with a chained CRC per entry, crash-safe checkpoints, a single-writer lock and a read-only open. WAL appends are not fsynced: saves since the last checkpoint can be lost on diff --git a/ROADMAP.md b/ROADMAP.md index 6ebf48e..7461ad8 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -42,7 +42,7 @@ Details per release: [`CHANGELOG.md`](CHANGELOG.md). | #18 | 2026-09-27 | range reads M2/M3 (`File::open_storage`; `clawhdf5-remote`: HTTP(S), S3, GCS, Azure), in-place editing of every chunk index, shrinking, dense attributes | | #19 | 2026-09-27 | remote files in the browser (`openUrl`, M4), SWMR reader (`File::open_swmr`, M5), Python remote reads and `'r+'` editing | | #20 | 2026-09-28 | benchmarks re-measured: LongMemEval with real MiniLM embeddings, local reads on an idle machine | -| #21 | 2026-09-28 | remote files open in a few requests (group lookups down the B-tree, `Storage::hint`), `ObjectHeader::parse` back to its earlier speed, the last conformance mismatches resolved: 602/697 ok, 0 our-error, 0 mismatch | +| #21 | 2026-09-28 | remote files open in a few requests (group lookups down the B-tree, `Storage::hint`), `ObjectHeader::parse` back to its earlier speed, the last conformance mismatches resolved: 602/697 ok, 0 mismatch (the run of 2026-09-28 in [`CONFORMANCE.md`](CONFORMANCE.md) still counts 1 our-error, a corrupt N-Bit file libhdf5's own tests refuse) | ### Range reads (design: [`docs/design/range-reads.md`](docs/design/range-reads.md)) @@ -76,9 +76,7 @@ Not scheduled; listed roughly by how much they unblock. None has a date. - [ ] **Publish the crates to crates.io.** Nothing is published; the READMEs say to depend on git. Before publishing: no `publish` settings exist - (only `clawhdf5-wasm` has `publish = false`), and several `Cargo.toml` - descriptions still name the old `rustyhdf5`/`edgehdf5` (`clawhdf5-accel`, - `-derive`, `-gpu`, `-io`, `-netcdf4`, `-android`). + (only `clawhdf5-wasm` has `publish = false`). - [ ] **Publish Python wheels to PyPI.** `crates/clawhdf5-py` builds with maturin and is tested in CI, but no wheel is published. The default wheel reads plain `http://` only; `https`/`s3`/`gcs`/`azure` wheels compile C diff --git a/conformance/README.md b/conformance/README.md index b4b682c..99c957d 100644 --- a/conformance/README.md +++ b/conformance/README.md @@ -10,9 +10,11 @@ conformance/run.sh --no-fetch # use the cache conformance/run.sh --update-baseline # after an intended change in results ``` -Latest result (tank, 2026-09-27, `conformance/run.sh --no-fetch`): 602 of -697 files ok, 0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read, and -no panic, hang, crash or out-of-memory. The report with every file is +Latest result (tank, 2026-09-28 04:29 UTC, `conformance/run.sh --no-fetch +--update-baseline`): 602 of 697 files ok, 1 our-error, 0 mismatch, 2 +ref-bug, 92 h5py-cannot-read, and no panic, hang, crash or out-of-memory. +The our-error file is `bad_nbit_parms_walk.h5`, which flips between ref-bug +and our-error from run to run (see `docs/known-issues.md`). The report with every file is [`CONFORMANCE.md`](../CONFORMANCE.md). ## Classes @@ -32,8 +34,8 @@ Where h5py itself returns wrong values through a known h5py bug (the big-endian variable-length bug: elements returned with the file's bytes under a little-endian dtype), `ref.py` checks that the installed h5py has the bug, corrects the values before hashing and marks them `ref_fix`, so -those objects are still compared. The evidence for the three current -ref-bug files is under "Conformance: the last non-ok files" in +those objects are still compared. The evidence for the three remaining +non-ok files (ref-bug or, for one, our-error) is under "Conformance: the last non-ok files" in [`docs/known-issues.md`](../docs/known-issues.md). Needs Rust, `git`, `h5dump` (Debian/Ubuntu `hdf5-tools`), `libaec` (for the diff --git a/crates/clawhdf5-accel/README.md b/crates/clawhdf5-accel/README.md index 26bacfc..bba4452 100644 --- a/crates/clawhdf5-accel/README.md +++ b/crates/clawhdf5-accel/README.md @@ -42,7 +42,8 @@ AVX2 and on NEON — with the `SDOT` instruction (through inline assembly, since the intrinsic is unstable) on cores that have dotprod, such as the Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8 index answers 1.63x the queries per second of the f32 one on x86-64 -(AVX2) and 1.18x on a Raspberry Pi 5 ([`BENCHMARKS.md`](../../BENCHMARKS.md)). +(AVX2; 2026-09-20, machine not recorded, not re-run) and 1.18x on a +Raspberry Pi 5 (2026-09-21) ([`BENCHMARKS.md` § Quantising the index copy](../../BENCHMARKS.md#quantising-the-index-copy-quantized_index)). The aarch64 code is compiled out on x86, so only the `test-arm64` CI job builds and tests it. diff --git a/crates/clawhdf5-agent/README.md b/crates/clawhdf5-agent/README.md index 8eeafe7..c807550 100644 --- a/crates/clawhdf5-agent/README.md +++ b/crates/clawhdf5-agent/README.md @@ -43,7 +43,7 @@ let hits = mem.search(&query, "deploy key", &SearchOptions::new(5).with_sources( for h in &hits { println!("{:.3} {}", h.score, h.chunk); } -mem.flush_wal()?; // checkpoint now; otherwise one is made after 500 WAL entries (wal_max_entries) +mem.flush_wal()?; // checkpoint now; otherwise one is made once the WAL holds more than 500 entries (wal_max_entries) # Ok::<(), clawhdf5_agent::MemoryError>(()) ``` diff --git a/crates/clawhdf5-format/README.md b/crates/clawhdf5-format/README.md index 3910da6..2223b55 100644 --- a/crates/clawhdf5-format/README.md +++ b/crates/clawhdf5-format/README.md @@ -77,7 +77,7 @@ assert!(!hdr.messages.is_empty()); | `checksum` | yes | verify Jenkins lookup3 checksums | no | | `deflate` | yes | deflate through flate2 | no | | `zlib-rs` | yes | flate2's pure-Rust zlib-rs backend, with `runtime_detection` (without it zlib-rs loses SIMD and inflates 3.5x slower) | no | -| `system-zlib-decompress` | yes | no effect (nothing reads it; kept so existing feature lists still build) | no | +| `system-zlib-decompress` | yes | macOS only: inflate with the system libz first, falling back to flate2; no effect elsewhere | no (links the system libz on macOS) | | `provenance` | yes | SHA-256 provenance hashes | no | | `lzf` | yes | LZF (32000) | no | | `parallel` | no | rayon-parallel chunk decoding | no | diff --git a/crates/clawhdf5-io/README.md b/crates/clawhdf5-io/README.md index a372453..ce95b72 100644 --- a/crates/clawhdf5-io/README.md +++ b/crates/clawhdf5-io/README.md @@ -24,7 +24,7 @@ clawhdf5-io = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5", features = |---|---| | `HDF5Read`, `HDF5ReadWrite` | byte-level read/write traits; `MemoryReader`, `BorrowedReader`, `FileReader`, `FileWriter` implement them | | `MmapReader`, `MmapReadWrite` (`mmap`) | memory-mapped files through `memmap2`; `HDF5Read::private_copy` gives a copy-on-write view | -| `prefetch::PrefetchReader`, `sweep::SweepDetector` | read-ahead (`madvise(MADV_WILLNEED)` on mappings) and chunk-sweep prediction | +| `prefetch::PrefetchReader`, `prefetch::SweepDetector`, `sweep` | read-ahead (`madvise(MADV_WILLNEED)` on mappings) and chunk-sweep prediction | | `ParallelConfig` | lane partitioning for parallel chunk decoding | | `vol::VirtualObjectLayer`, `vol::NativeVol` | a backend-agnostic object-layer trait (modelled on libhdf5's VOL) | | `async_read` (`async`) | tokio-based `AsyncHDF5Read` and `AsyncHDF5File` | diff --git a/crates/clawhdf5-netcdf4/README.md b/crates/clawhdf5-netcdf4/README.md index 9e82b97..e52e131 100644 --- a/crates/clawhdf5-netcdf4/README.md +++ b/crates/clawhdf5-netcdf4/README.md @@ -39,7 +39,7 @@ let values: Vec = temp.read_f64()?; | `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | | `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) | | `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` | -| `Dimension` | `name`, `size`, `is_unlimited` | +| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is wrongly 0 when it holds records; use the variables' shapes — [known issue](../../docs/known-issues.md#netcdf-4-an-unlimited-dimension-reports-size-0)) | | `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` | | `NcType` | the NetCDF type of a variable | diff --git a/docs/README.md b/docs/README.md index 8643533..0677abc 100644 --- a/docs/README.md +++ b/docs/README.md @@ -29,7 +29,7 @@ Every document in the repository, one line each. Start with the | Document | What it covers | |---|---| -| [design/range-reads.md](design/range-reads.md) | Reading through a `Storage` trait: milestones M1–M5 (storage, raw data, remote files, the browser, SWMR) | +| [design/range-reads.md](design/range-reads.md) | Reading through a `Storage` trait: milestones M0–M5 (indexed lookups, storage, raw data, remote files, the browser, SWMR) | | [design/swmr.md](design/swmr.md) | Reading files a libhdf5 SWMR writer is appending to (M5) | | [design/tools/](design/tools/) | Scripts behind the range-read design's measurements (`inventory.py`, `libhdf5_reads.py`, `range-trace`) | @@ -46,19 +46,23 @@ Every document in the repository, one line each. Start with the | [crates/clawhdf5-derive](../crates/clawhdf5-derive/README.md) | Derive macros | | [crates/clawhdf5-tools](../crates/clawhdf5-tools/README.md) | `h5rs` | | [crates/clawhdf5-py](../crates/clawhdf5-py/README.md) | Python bindings | +| [crates/clawhdf5-wasm](../crates/clawhdf5-wasm/README.md) | The browser reader crate | | [examples/wasm-viewer](../examples/wasm-viewer/README.md) | Browser viewer and the `clawhdf5-wasm` JavaScript API | -| [packages/clawhdf5-node](../packages/clawhdf5-node/README.md) | Node.js package (unpublished, does not work) | +| [crates/clawhdf5-napi](../crates/clawhdf5-napi/README.md), [packages/clawhdf5-node](../packages/clawhdf5-node/README.md) | Node.js bindings and package (unpublished, does not work) | +| [crates/clawhdf5-android](../crates/clawhdf5-android/README.md) | Android JNI bindings for the agent store | | [crates/clawhdf5-agent](../crates/clawhdf5-agent/README.md) | Agent memory (full guide: [agent-memory.md](agent-memory.md)) | | [crates/clawhdf5-ann](../crates/clawhdf5-ann/README.md) | HNSW index | | [crates/clawhdf5-accel](../crates/clawhdf5-accel/README.md) | SIMD kernels | | [crates/clawhdf5-gpu](../crates/clawhdf5-gpu/README.md) | GPU vector distances | | [crates/clawhdf5-migrate](../crates/clawhdf5-migrate/README.md) | SQLite migration | +| [crates/clawhdf5-cli](../crates/clawhdf5-cli/README.md) | The agent-memory CLI | +| [crates/clawhdf5-bench](../crates/clawhdf5-bench/README.md) | Benchmarks and harnesses | ## Project history and working notes | Document | What it covers | |---|---| -| [ROADMAP.md](../ROADMAP.md) | Agent-memory roadmap and implementation tracker | +| [ROADMAP.md](../ROADMAP.md) | What has shipped (releases and PRs since v2.7.0) and what is next | | [CLAUDE.md](../CLAUDE.md) | Architecture and workflow notes for contributors and coding agents | | [archive/IMPROVEMENT_LOG.md](archive/IMPROVEMENT_LOG.md), [archive/IMPROVEMENT_SCAN.md](archive/IMPROVEMENT_SCAN.md) | Logs of earlier automated improvement passes (archived, historical) | | [archive/plans/](archive/plans/) | Implementation plans from June 2026 (filter codecs, format write extensions, MPI-IO); archived, historical | diff --git a/docs/USE_CASES.md b/docs/USE_CASES.md index 8c23fc3..66c6fdb 100644 --- a/docs/USE_CASES.md +++ b/docs/USE_CASES.md @@ -48,8 +48,10 @@ not the whole download. error rather than mixed data. - In the browser, `clawhdf5-wasm`'s `openUrl` does the same from the page's main thread; the [viewer](../examples/wasm-viewer/README.md) is a working - example. Opening one dataset of a 3000-dataset, 198 MB h5py file took - 5 requests and 5.2 MB at 1 MiB blocks (tank, 2026-09-27, CHANGELOG). + example. Opening and reading one dataset of a 3000-dataset, 198 MB h5py + file took 5 requests and 5.2 MB at 1 MiB blocks (h5py's default + `libver="earliest"`; 7 requests and 6.7 MB with `"latest"`) (tank, + 2026-09-27, CHANGELOG "Unreleased"). - Design and measured request counts: [design/range-reads.md](design/range-reads.md). ### Files you did not write and do not trust @@ -95,8 +97,8 @@ An assistant accumulates preferences, decisions and context over months. - `clawhdf5-agent` keeps records, sessions and a knowledge graph in one `.h5` file with a write-ahead log: back it up or move it with the agent. - Hybrid search (HNSW + BM25) reaches 81.4% turn-level Hit@5 on the full - LongMemEval haystack — retrieval recall, not QA accuracy - ([BENCHMARKS.md](../BENCHMARKS.md#longmemeval-results)). + LongMemEval haystack with real MiniLM embeddings — retrieval recall, not + QA accuracy (tank, 2026-09-27; [BENCHMARKS.md](../BENCHMARKS.md#longmemeval-results)). - The consolidation engine (Working → Episodic → Semantic) and the knowledge graph are library components you drive; see [agent-memory.md](agent-memory.md#library-components). @@ -123,8 +125,8 @@ A Raspberry Pi or another ARM board, no server, no network. - Pure Rust, no database server, one file. - The int8 index uses NEON `SDOT` on cores with the dot-product extension (plain NEON elsewhere); on a Raspberry Pi 5 it - was 1.18x the `f32` index's QPS at equal recall - ([BENCHMARKS.md](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)). + was 1.18x the `f32` index's QPS at equal recall (2026-09-21, `114a2df`, + not re-run since; [BENCHMARKS.md](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)). CI builds and tests the aarch64 code on an ARM runner. - WAL appends are not fsynced: on power loss, saves since the last checkpoint can be lost, while checkpoints themselves are made durable as @@ -175,6 +177,6 @@ the one verified consumer of clawhdf5. | zlib-ng instead of zlib-rs | `fast-deflate` (needs cmake) | | Remote files | `clawhdf5-remote` (`http` default; `https`, `s3`, `gcs`, `azure`) | | Agent memory | `clawhdf5-agent` (defaults: `float16`, `hnsw`, `parallel`) | -| BLAS for the agent's brute-force paths | `fast-math`, `openblas`, or `accelerate` (macOS) | -| GPU distance computation | `gpu` (wgpu) | -| Async wrapper | `async` (Tokio) | +| Faster brute-force paths in the agent | `clawhdf5-agent`'s `fast-math` (matrixmultiply, pure Rust), or BLAS: `openblas`, `accelerate` (macOS) | +| GPU distance computation | `clawhdf5-agent`'s `gpu` (wgpu) | +| Async wrapper | `clawhdf5-agent`'s `async` (Tokio) | diff --git a/docs/agent-memory.md b/docs/agent-memory.md index 63ef06b..af9a3b5 100644 --- a/docs/agent-memory.md +++ b/docs/agent-memory.md @@ -278,11 +278,14 @@ N = 100K, M = 16, ef_construction = 64, ef = 64, recall against an exact scan | `f32` | 0.9945 | 13 399 | 3.2 s | | `i8` + exact re-score (**default for new stores**) | 0.9940 | **21 848** | **1.8 s** | -A paired comparison (medians of alternating runs, same binary), not re-run -on 2026-09-24: a single `f32` run that day measured recall 0.9945, 19 001 -QPS and a 2.7 s build, so the 1.63x ratio has not been re-checked. On a -Raspberry Pi 5 (NEON `SDOT`) the int8 index is 1.18x the `f32` QPS at equal -recall. Before the v2.4.0 neighbour-selection fix, recall@10 at 100K was +A paired comparison (medians of alternating runs, same binary: the int8 +index answers 1.63x the queries per second at equal recall), recorded +2026-09-20 with the machine not recorded, and not re-run since: a single +`f32` run on 2026-09-24 (tank) measured recall 0.9945, 19 001 QPS and a +2.7 s build, so the 1.63x ratio has not been re-checked. On a Raspberry +Pi 5 (NEON `SDOT`) the int8 index is 1.18x the `f32` QPS at equal recall +(2026-09-21; [§ On ARM](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)). +Before the v2.4.0 neighbour-selection fix, recall@10 at 100K was 0.31. **Operations:** @@ -344,9 +347,11 @@ The benchmark's vector stage needs `clawhdf5-bench`'s `embeddings` feature. text, `footprint_bench`: 810.4 KB at 1K records, 7.8 MB at 10K, 76.7 MB at 100K (803–829 bytes per record). The synthetic text is far more repetitive than real text (40 distinct strings, deflated), so real records will be -larger; the embeddings alone are 768 B per record. On the same data, 100K × -384 takes 80.8 MiB as `float16` and 154.0 MiB as `f32` -([§ Memory Footprint](../BENCHMARKS.md#memory-footprint-1)). +larger; the embeddings alone are 768 B per record +([§ Memory Footprint](../BENCHMARKS.md#memory-footprint-1), 2026-09-24). In the +float16 study (clustered data, 2026-09-23), 100K × 384 takes 80.8 MiB as +`float16` and 154.0 MiB as `f32` +([§ float16 embedding storage](../BENCHMARKS.md#float16-embedding-storage-memoryconfigfloat16)). **In memory** — a store reopened from disk, counting allocator ([§ Memory footprint](../BENCHMARKS.md#memory-footprint)): @@ -357,7 +362,9 @@ larger; the embeddings alone are 768 B per record. On the same data, 100K × | 10K | 15 MiB | 44 MiB (3.03x) | 27 MiB (1.81x) | | 100K | 146 MiB | 399 MiB (2.72x) | 256 MiB (1.74x) | -The `f32` column was re-measured on 2026-09-24; the `i8` column was not. +The `f32` column was re-measured on 2026-09-24 (tank); the `i8` column was +first measured 2026-09-19 (commit c0a9206, machine not recorded) and not +re-run ([§ Quantising the index copy](../BENCHMARKS.md#quantising-the-index-copy-quantized_index)). ## Feature flags and settings @@ -385,8 +392,8 @@ Settings stored in the file (`MemoryConfig`): above. Opt out with `quantized_index = false` or `create --f32-index`. - `hnsw_m`, `hnsw_ef_construction`, `hnsw_ef_search`: 16 / 64 / scaled with `k` by default. -- `compression` (off): deflate (or Zstd) for embeddings; text of 4 KiB or - more is always deflated. +- `compression` (off): deflate (or Zstd) for embeddings; string datasets + (text, channels, tags, ...) of 4 KiB or more are always deflated. - `wal_enabled` (on), `wal_max_entries`, `hebbian_boost`, `decay_factor`. ## File schema @@ -442,9 +449,9 @@ clawhdf5 --path agent.h5 snapshot backup.h5 clawhdf5 keygen --out signing.key # then --signing-key signing.key; verify --public-key ``` -Output is JSON. The CLI's `search` defaults to weights 0.7 / 0.3, not the -library's 0.4 / 0.6, so pass them. `recall`, `stats`, `agents-md` and -`export` open the store read-only. +Output is JSON (Markdown for `agents-md`). The CLI's `search` defaults to +weights 0.7 / 0.3, not the library's 0.4 / 0.6, so pass them. `recall`, +`stats`, `agents-md` and `export` open the store read-only. ## Migrating from SQLite diff --git a/docs/known-issues.md b/docs/known-issues.md index 9c5bf01..b50deff 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -138,7 +138,8 @@ before anything is written. On top of them: **Status:** open (documented 2026-09-26; re-checked 2026-09-28 in `Dataset::read_selection`, `crates/clawhdf5/src/reader.rs`). A selection -read (and so the Python `ds[...]`) materialises only the selection's +read (and so the Python `ds[...]`) of contiguous data copies just the +selected runs; of chunked data it materialises only the selection's bounding box when that box covers at most half the dataset (`partial_read`). It decodes the whole dataset and extracts the selection instead when: @@ -577,7 +578,8 @@ buffers per chunk and a second copy of the output, 4 KiB page faults on the output buffer, and element-by-element hyperslab copies. Re-measured at `c5334b1` (`BENCHMARKS.md`, "Results after in-place chunk decoding"): 16-thread chunked deflate reads 4944 MB/s against 3135 for 16 h5py -processes (1.58x), contiguous reads 1.29x h5py on one thread. Tests: +processes (1.58x), contiguous reads 6718 MB/s against h5py's 5545 on one +thread (1.21x; 1.29x h5py processes). Tests: `single_thread_decode_pool.rs`, `busy_decode_pool.rs`. ## Scale-offset data read back wrong values