docs: fact-check the refreshed documentation against its sources

Numbers, API names, feature defaults and PR references checked against
CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history.

- int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives
  them (2026-09-19/20, machine not recorded, not re-run) instead of none;
  the Pi 5 1.18x carries 2026-09-21.
- BENCHMARKS headline: the libhdf5 chunked-write figure is the newest
  measurement (35x, 2026-09-23), not 45.3x (2026-08-03).
- Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug)
  in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to
  the bad_nbit_parms_walk.h5 flip.
- README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/
  time datasets too; zlib-rs byte-identity scoped to what was measured;
  macOS default links the system libz for inflate.
- Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector
  lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4
  unlimited-dimension size warning.
- agent-memory.md: string-dataset compression threshold, agents-md prints
  Markdown, float16 file sizes linked to their study.
- known-issues.md: contiguous selection reads, 1.21x vs h5py threads.
- docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read
  milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify,
  dated figures, fast-math is not BLAS.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 11:23:51 -05:00
co-authored by Claude Opus 5.5
parent a48cb9f1a4
commit c27a478e44
14 changed files with 85 additions and 62 deletions
+7 -3
View File
@@ -29,7 +29,7 @@ Every document in the repository, one line each. Start with the
| Document | What it covers |
|---|---|
| [design/range-reads.md](design/range-reads.md) | Reading through a `Storage` trait: milestones M1–M5 (storage, raw data, remote files, the browser, SWMR) |
| [design/range-reads.md](design/range-reads.md) | Reading through a `Storage` trait: milestones M0–M5 (indexed lookups, storage, raw data, remote files, the browser, SWMR) |
| [design/swmr.md](design/swmr.md) | Reading files a libhdf5 SWMR writer is appending to (M5) |
| [design/tools/](design/tools/) | Scripts behind the range-read design's measurements (`inventory.py`, `libhdf5_reads.py`, `range-trace`) |
@@ -46,19 +46,23 @@ Every document in the repository, one line each. Start with the
| [crates/clawhdf5-derive](../crates/clawhdf5-derive/README.md) | Derive macros |
| [crates/clawhdf5-tools](../crates/clawhdf5-tools/README.md) | `h5rs` |
| [crates/clawhdf5-py](../crates/clawhdf5-py/README.md) | Python bindings |
| [crates/clawhdf5-wasm](../crates/clawhdf5-wasm/README.md) | The browser reader crate |
| [examples/wasm-viewer](../examples/wasm-viewer/README.md) | Browser viewer and the `clawhdf5-wasm` JavaScript API |
| [packages/clawhdf5-node](../packages/clawhdf5-node/README.md) | Node.js package (unpublished, does not work) |
| [crates/clawhdf5-napi](../crates/clawhdf5-napi/README.md), [packages/clawhdf5-node](../packages/clawhdf5-node/README.md) | Node.js bindings and package (unpublished, does not work) |
| [crates/clawhdf5-android](../crates/clawhdf5-android/README.md) | Android JNI bindings for the agent store |
| [crates/clawhdf5-agent](../crates/clawhdf5-agent/README.md) | Agent memory (full guide: [agent-memory.md](agent-memory.md)) |
| [crates/clawhdf5-ann](../crates/clawhdf5-ann/README.md) | HNSW index |
| [crates/clawhdf5-accel](../crates/clawhdf5-accel/README.md) | SIMD kernels |
| [crates/clawhdf5-gpu](../crates/clawhdf5-gpu/README.md) | GPU vector distances |
| [crates/clawhdf5-migrate](../crates/clawhdf5-migrate/README.md) | SQLite migration |
| [crates/clawhdf5-cli](../crates/clawhdf5-cli/README.md) | The agent-memory CLI |
| [crates/clawhdf5-bench](../crates/clawhdf5-bench/README.md) | Benchmarks and harnesses |
## Project history and working notes
| Document | What it covers |
|---|---|
| [ROADMAP.md](../ROADMAP.md) | Agent-memory roadmap and implementation tracker |
| [ROADMAP.md](../ROADMAP.md) | What has shipped (releases and PRs since v2.7.0) and what is next |
| [CLAUDE.md](../CLAUDE.md) | Architecture and workflow notes for contributors and coding agents |
| [archive/IMPROVEMENT_LOG.md](archive/IMPROVEMENT_LOG.md), [archive/IMPROVEMENT_SCAN.md](archive/IMPROVEMENT_SCAN.md) | Logs of earlier automated improvement passes (archived, historical) |
| [archive/plans/](archive/plans/) | Implementation plans from June 2026 (filter codecs, format write extensions, MPI-IO); archived, historical |
+11 -9
View File
@@ -48,8 +48,10 @@ not the whole download.
error rather than mixed data.
- In the browser, `clawhdf5-wasm`'s `openUrl` does the same from the page's
main thread; the [viewer](../examples/wasm-viewer/README.md) is a working
example. Opening one dataset of a 3000-dataset, 198 MB h5py file took
5 requests and 5.2 MB at 1 MiB blocks (tank, 2026-09-27, CHANGELOG).
example. Opening and reading one dataset of a 3000-dataset, 198 MB h5py
file took 5 requests and 5.2 MB at 1 MiB blocks (h5py's default
`libver="earliest"`; 7 requests and 6.7 MB with `"latest"`) (tank,
2026-09-27, CHANGELOG "Unreleased").
- Design and measured request counts: [design/range-reads.md](design/range-reads.md).
### Files you did not write and do not trust
@@ -95,8 +97,8 @@ An assistant accumulates preferences, decisions and context over months.
- `clawhdf5-agent` keeps records, sessions and a knowledge graph in one
`.h5` file with a write-ahead log: back it up or move it with the agent.
- Hybrid search (HNSW + BM25) reaches 81.4% turn-level Hit@5 on the full
LongMemEval haystack — retrieval recall, not QA accuracy
([BENCHMARKS.md](../BENCHMARKS.md#longmemeval-results)).
LongMemEval haystack with real MiniLM embeddings — retrieval recall, not
QA accuracy (tank, 2026-09-27; [BENCHMARKS.md](../BENCHMARKS.md#longmemeval-results)).
- The consolidation engine (Working → Episodic → Semantic) and the
knowledge graph are library components you drive; see
[agent-memory.md](agent-memory.md#library-components).
@@ -123,8 +125,8 @@ A Raspberry Pi or another ARM board, no server, no network.
- Pure Rust, no database server, one file.
- The int8 index uses NEON `SDOT` on cores with the dot-product extension
(plain NEON elsewhere); on a Raspberry Pi 5 it
was 1.18x the `f32` index's QPS at equal recall
([BENCHMARKS.md](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)).
was 1.18x the `f32` index's QPS at equal recall (2026-09-21, `114a2df`,
not re-run since; [BENCHMARKS.md](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)).
CI builds and tests the aarch64 code on an ARM runner.
- WAL appends are not fsynced: on power loss, saves since the last
checkpoint can be lost, while checkpoints themselves are made durable as
@@ -175,6 +177,6 @@ the one verified consumer of clawhdf5.
| zlib-ng instead of zlib-rs | `fast-deflate` (needs cmake) |
| Remote files | `clawhdf5-remote` (`http` default; `https`, `s3`, `gcs`, `azure`) |
| Agent memory | `clawhdf5-agent` (defaults: `float16`, `hnsw`, `parallel`) |
| BLAS for the agent's brute-force paths | `fast-math`, `openblas`, or `accelerate` (macOS) |
| GPU distance computation | `gpu` (wgpu) |
| Async wrapper | `async` (Tokio) |
| Faster brute-force paths in the agent | `clawhdf5-agent`'s `fast-math` (matrixmultiply, pure Rust), or BLAS: `openblas`, `accelerate` (macOS) |
| GPU distance computation | `clawhdf5-agent`'s `gpu` (wgpu) |
| Async wrapper | `clawhdf5-agent`'s `async` (Tokio) |
+21 -14
View File
@@ -278,11 +278,14 @@ N = 100K, M = 16, ef_construction = 64, ef = 64, recall against an exact scan
| `f32` | 0.9945 | 13 399 | 3.2 s |
| `i8` + exact re-score (**default for new stores**) | 0.9940 | **21 848** | **1.8 s** |
A paired comparison (medians of alternating runs, same binary), not re-run
on 2026-09-24: a single `f32` run that day measured recall 0.9945, 19 001
QPS and a 2.7 s build, so the 1.63x ratio has not been re-checked. On a
Raspberry Pi 5 (NEON `SDOT`) the int8 index is 1.18x the `f32` QPS at equal
recall. Before the v2.4.0 neighbour-selection fix, recall@10 at 100K was
A paired comparison (medians of alternating runs, same binary: the int8
index answers 1.63x the queries per second at equal recall), recorded
2026-09-20 with the machine not recorded, and not re-run since: a single
`f32` run on 2026-09-24 (tank) measured recall 0.9945, 19 001 QPS and a
2.7 s build, so the 1.63x ratio has not been re-checked. On a Raspberry
Pi 5 (NEON `SDOT`) the int8 index is 1.18x the `f32` QPS at equal recall
(2026-09-21; [§ On ARM](../BENCHMARKS.md#on-arm-raspberry-pi-5-cortex-a76)).
Before the v2.4.0 neighbour-selection fix, recall@10 at 100K was
0.31.
**Operations:**
@@ -344,9 +347,11 @@ The benchmark's vector stage needs `clawhdf5-bench`'s `embeddings` feature.
text, `footprint_bench`: 810.4 KB at 1K records, 7.8 MB at 10K, 76.7 MB at
100K (803–829 bytes per record). The synthetic text is far more repetitive
than real text (40 distinct strings, deflated), so real records will be
larger; the embeddings alone are 768 B per record. On the same data, 100K ×
384 takes 80.8 MiB as `float16` and 154.0 MiB as `f32`
([§ Memory Footprint](../BENCHMARKS.md#memory-footprint-1)).
larger; the embeddings alone are 768 B per record
([§ Memory Footprint](../BENCHMARKS.md#memory-footprint-1), 2026-09-24). In the
float16 study (clustered data, 2026-09-23), 100K × 384 takes 80.8 MiB as
`float16` and 154.0 MiB as `f32`
([§ float16 embedding storage](../BENCHMARKS.md#float16-embedding-storage-memoryconfigfloat16)).
**In memory** — a store reopened from disk, counting allocator
([§ Memory footprint](../BENCHMARKS.md#memory-footprint)):
@@ -357,7 +362,9 @@ larger; the embeddings alone are 768 B per record. On the same data, 100K ×
| 10K | 15 MiB | 44 MiB (3.03x) | 27 MiB (1.81x) |
| 100K | 146 MiB | 399 MiB (2.72x) | 256 MiB (1.74x) |
The `f32` column was re-measured on 2026-09-24; the `i8` column was not.
The `f32` column was re-measured on 2026-09-24 (tank); the `i8` column was
first measured 2026-09-19 (commit c0a9206, machine not recorded) and not
re-run ([§ Quantising the index copy](../BENCHMARKS.md#quantising-the-index-copy-quantized_index)).
## Feature flags and settings
@@ -385,8 +392,8 @@ Settings stored in the file (`MemoryConfig`):
above. Opt out with `quantized_index = false` or `create --f32-index`.
- `hnsw_m`, `hnsw_ef_construction`, `hnsw_ef_search`: 16 / 64 / scaled with
`k` by default.
- `compression` (off): deflate (or Zstd) for embeddings; text of 4 KiB or
more is always deflated.
- `compression` (off): deflate (or Zstd) for embeddings; string datasets
(text, channels, tags, ...) of 4 KiB or more are always deflated.
- `wal_enabled` (on), `wal_max_entries`, `hebbian_boost`, `decay_factor`.
## File schema
@@ -442,9 +449,9 @@ clawhdf5 --path agent.h5 snapshot backup.h5
clawhdf5 keygen --out signing.key # then --signing-key signing.key; verify --public-key <hex>
```
Output is JSON. The CLI's `search` defaults to weights 0.7 / 0.3, not the
library's 0.4 / 0.6, so pass them. `recall`, `stats`, `agents-md` and
`export` open the store read-only.
Output is JSON (Markdown for `agents-md`). The CLI's `search` defaults to
weights 0.7 / 0.3, not the library's 0.4 / 0.6, so pass them. `recall`,
`stats`, `agents-md` and `export` open the store read-only.
## Migrating from SQLite
+4 -2
View File
@@ -138,7 +138,8 @@ before anything is written. On top of them:
**Status:** open (documented 2026-09-26; re-checked 2026-09-28 in
`Dataset::read_selection`, `crates/clawhdf5/src/reader.rs`). A selection
read (and so the Python `ds[...]`) materialises only the selection's
read (and so the Python `ds[...]`) of contiguous data copies just the
selected runs; of chunked data it materialises only the selection's
bounding box when that box covers at most half the dataset
(`partial_read`). It decodes the whole dataset and extracts the selection
instead when:
@@ -577,7 +578,8 @@ buffers per chunk and a second copy of the output, 4 KiB page faults on the
output buffer, and element-by-element hyperslab copies. Re-measured at
`c5334b1` (`BENCHMARKS.md`, "Results after in-place chunk decoding"):
16-thread chunked deflate reads 4944 MB/s against 3135 for 16 h5py
processes (1.58x), contiguous reads 1.29x h5py on one thread. Tests:
processes (1.58x), contiguous reads 6718 MB/s against h5py's 5545 on one
thread (1.21x; 1.29x h5py processes). Tests:
`single_thread_decode_pool.rs`, `busy_decode_pool.rs`.
## Scale-offset data read back wrong values