From 06d8e2ee4563a76341e0530e1a28e86ec0a170e3 Mon Sep 17 00:00:00 2001 From: osobh Date: Mon, 28 Sep 2026 11:04:38 -0500 Subject: [PATCH] docs: benchmark headline numbers, superseded sections labelled A "Current headline numbers" table gives each figure's newest dated measurement with its machine, command and section. Sections a later run replaced are marked superseded with a link to the newer one; stale "still open" notes and cross-references now point at the fixes. No measured value changed. Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 112 ++++++++++++++++++++++++++++++++++++++++++++------ 1 file changed, 100 insertions(+), 12 deletions(-) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 621f912..7910b60 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -51,6 +51,31 @@ target: Criterion stretched it where 5 s could not hold the samples it needed --- +## Current headline numbers + +The newest dated measurement of each headline figure, as of 2026-09-28. +Everything below this section is the dated record behind them; sections whose +figures a later run replaced are marked *Superseded*. Machine "tank" is an AMD +Ryzen 7 7800X3D (8C/16T); rows marked idle were run with the 1-minute load +average below 2. + +| Figure | Value | Measured | Command | Details | +|---|---|---|---|---| +| Agent memory search, `HDF5Memory::hybrid_search` p50 | 0.49 ms at 10K, 4.69 ms at 100K records | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` | [Current: search harness](#current-search-harness-2026-09-24) | +| LongMemEval `longmemeval_s` (full haystack), default hybrid 0.4/0.6, turn-level retrieval Hit@5 (not QA accuracy) | 81.4% | 2026-09-27, tank, search code of `7a8fae0` | `longmemeval_bench … --embeddings weights/all-minilm-l6-v2` | [Re-run with real embeddings](#re-run-with-real-embeddings-2026-09-27-tank), [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500) | +| Loaded store memory, 100K × 384 | 399 MiB (2.72x raw) with the `f32` index; 256 MiB (1.74x) with the int8 index (int8 side not re-run since it was first measured) | 2026-09-24, tank, `5c8323c` | `search_harness -- --footprint --full [--int8]` | [Memory footprint](#memory-footprint), [Quantising the index copy](#quantising-the-index-copy-quantized_index) | +| int8 index vs `f32` index, QPS at equal recall | 1.63x (x86-64 AVX2), 1.18x (Raspberry Pi 5, `SDOT`) | undated; not re-checked against the 2026-09-24 `f32` figure | `search_harness -- --full` | [Quantising the index copy](#quantising-the-index-copy-quantized_index), [On ARM](#on-arm-raspberry-pi-5-cortex-a76) | +| `float16` store file size, 100K × 384 | 80.8 MiB vs 154.0 MiB `f32` (48% smaller) | 2026-09-23, tank | `search_harness -- --float16-study --full` | [float16 embedding storage](#float16-embedding-storage-memoryconfigfloat16) | +| Full reads of chunked deflate data, 16 threads on one `File` | 4944 MB/s, 1.58x 16 h5py processes (noisy run: compare ratios, not MB/s) | 2026-09-26, tank, `c5334b1` | `concurrent_read` + `concurrent_read_h5py.py` | [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) | +| Same, clawhdf5 only, against the build before range-read M2/M3 | 8525 MB/s vs 6258 (+36%); contiguous and metadata reads at parity | 2026-09-27, tank (idle), `7a8fae0` vs `8f59b2e` | `concurrent_read --decode-threads 1 --reps 3` | [Local metadata and data reads after range-read M2/M3](#local-metadata-and-data-reads-after-range-read-m2m3-2026-09-27-tank) | +| `ObjectHeader::parse` (401 headers) | 23.5–23.6 µs, 1.0–2.6% below `8f59b2e` | 2026-09-27, tank (idle), `96086ad` | `cargo bench -p clawhdf5 --bench local_metadata_bench` | [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) | +| Selection reads, 64 MB chunked + deflate `f64` | full 63.2 ms; one 64 × 64 window 0.18 ms | 2026-09-24, tank, `5c8323c` | `cargo run --release -p clawhdf5-bench --bin read_harness` | [Current: read harness](#current-read-harness-2026-09-24) | +| Deflate backend, zlib-rs (default) vs zlib-ng | within 6% on every HDF5 read/write path | 2026-09-23, tank | `cargo bench -p clawhdf5-filters --bench deflate_bench` (and the two commands with it) | [Deflate backend](#deflate-backend-zlib-rs-vs-zlib-ng) | +| vs libhdf5 1.14.6: chunked deflate-6 write 512×512 / 128 attributes / 64 groups | 45.3x / 10.3x / 10.6x | 2026-08-03, tank | `cargo bench -p clawhdf5-bench --features libhdf5-compare` | [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) | +| Signed checkpoints | about 20% of a checkpoint (598 vs 495 ms at 100K) | 2026-09-25, tank | `search_harness -- --signing-study --full` | [Signed checkpoints](#signed-checkpoints) | + +--- + ## Memory footprint `cargo run --release -p clawhdf5-bench --bin search_harness -- --footprint --full`, @@ -64,6 +89,10 @@ change at all. Measured that way a store holding the corpus twice and one holding it once came out *identical* (1.00x both), which is how the first attempt at this measurement went. +> *Superseded* by the current figures below (2026-09-24): this table is the +> record of the double-copy fix (commit 2e7e045, undated); the store measured +> 2.72x, not 2.43x, by the time the int8 index landed. + | N | vectors (raw) | reopened, before | reopened, after | |---:|---:|---:|---:| | 1 000 | 1 MiB | 5 MiB (3.41x) | 4 MiB (2.39x) | @@ -379,6 +408,9 @@ point: does a selection cost what the *selection* costs? ### Baseline (v2.4.0): every selection decodes the whole dataset +> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24) +> (2026-09-24). Kept as the before picture. + 4096 x 2048 f64 (64 MB per dataset), chunks 256 x 256, file 129 MB | layout | read | selected | time ms | MB/s of selection | vs full read | @@ -404,6 +436,9 @@ point: does a selection cost what the *selection* costs? ### After: partial reads +> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24) +> (2026-09-24). + Only the rows of a contiguous dataset, or the chunks, that overlap the selection's bounding box are read/decoded. A 64 x 64 window of the compressed dataset: **105 -> 0.39 ms**; one row: **106 -> 2.7 ms**; one column: @@ -435,6 +470,9 @@ because the machine's speed drifted; compare the *vs full read* column.) ### After: parallel cached decode, fewer copies (full reads) +> *Superseded* by [Current: read harness](#current-read-harness-2026-09-24) +> (2026-09-24). + Full-read times, old and new binaries run alternately at the same moment (this machine's absolute speed drifts over a long session, so only same-moment comparisons mean anything): @@ -542,6 +580,10 @@ rounds; run 2 also alternated `main` `425585e`. ### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) +> The `object_header_parse_x401` row (+4.2%) is *superseded* by +> [`ObjectHeader::parse` back at 8f59b2e's speed](#objectheaderparse-back-at-8f59b2es-speed-2026-09-27-tank) +> (2026-09-27, `96086ad`); the other rows are current. + `main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main` `7a8fae0` (PRs #18 and #19), each built in its own worktree and run as separate binaries, alternating base and candidate. Machine: tank (AMD Ryzen @@ -585,8 +627,8 @@ What this shows: - **`ObjectHeader::parse` alone is 4.2% slower** (about 2.5 ns per header; the base and candidate ranges do not overlap). It is the cost of reading continuation chunks from a bounded queue (the fix for unbounded reads on - crafted headers) and does not show in the listing. Kept open in - `docs/known-issues.md`. + crafted headers) and does not show in the listing. (Fixed later the + same day; see the section above and `docs/known-issues.md`.) - **Full reads of deflate data got faster** after #18 (in-place chunk decoding into the typed output and per-thread scratch buffers): +1.7% on one thread, +36% at 16. @@ -639,6 +681,10 @@ saturate memory bandwidth (1.05x). ### Results after the read fixes (2026-09-26, tank, `408f69e`) +> *Superseded* by [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) +> (2026-09-26, `c5334b1`), which closed the 16-thread gap listed at the end +> of this section. + Same machine, files and commands as the first run below, re-run on an idle tank (load average 1.60 at the start; the 1-minute figure rose to about 5 during the clawhdf5 runs, mostly their own threads) after two fixes: @@ -679,11 +725,17 @@ Read with care: - At 16 threads every tool dropped in this run (h5py threads on contiguous data from 8002 to 2285 MB/s, processes from 12846 to 6942), so the 16-thread rows are noisier than the others. -- Still behind: full reads of chunked data at 16 threads (0.69x-0.76x h5py - processes). See `docs/known-issues.md`. +- Still behind at this commit: full reads of chunked data at 16 threads + (0.69x-0.76x h5py processes); fixed by `c5334b1` (above), recorded as + fixed in `docs/known-issues.md`. ### First run, before the read fixes (2026-09-26, tank, `91644d8`) +> *Superseded* results: the tables and "What this shows" are the before +> picture for [Results after in-place chunk decoding](#results-after-in-place-chunk-decoding-2026-09-26-tank-c5334b1) +> (2026-09-26). The workload description and the **Run** box below are +> still how every `concurrent_read` figure in this file is produced. + Measured on tank (AMD Ryzen 7 7800X3D, 8 cores / 16 threads, 61 GiB, Linux 7.0) at commit `91644d8`, load average 1.84 when the run started (the 1-minute figure rose to 3.7 during the runs; that is mostly the benchmark's @@ -720,8 +772,8 @@ What this shows: - **clawhdf5 threads on one `File` do, for hyperslab reads of compressed data:** 1244 MB/s at 16 threads, 9.7x h5py threads and 0.89x h5py processes, without a process pool. -- **Where clawhdf5 is behind** (open performance bugs, see - `docs/known-issues.md`): +- **Where clawhdf5 was behind** at `91644d8` (both since fixed; see + `docs/known-issues.md`, "Concurrent and contiguous read performance"): - *Full reads of chunked datasets stop scaling at about 4 threads* (about 880 MB/s) while h5py processes reach 4424 MB/s. Hyperslab reads, which bypass the `File`'s chunk cache, keep scaling, so the @@ -811,6 +863,11 @@ Other flags (both harnesses): `--threads`, `--reps`, `--slab`, `--slabs`, ## Search harness baseline (v2.3.0) +> *Historical.* This baseline and the "After: …" subsections that follow +> record each step of the search work; they are *superseded* by +> [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24), +> the last subsection of this part. + Produced by `cargo run --release -p clawhdf5-bench --bin search_harness -- --full` on deterministic **clustered** synthetic data (384-dim, unit-normalised; points = cluster centre + noise — uniform random vectors are nearly equidistant in high @@ -873,10 +930,11 @@ build: 9752.6 ms (10254 vectors/s) · exact scan: 40 QPS, p50 24648 µs | 1000 | 11 | 3.9 | 0.9 | 68.1 | 5.48 | 5.57 | 182.5 | | 10000 | 114 | 32.2 | 10.9 | 845.0 | 48.56 | 78.65 | 19.8 | | 100000 | 1486 | 713.0 | 354.5 | 10486.5 | 883.51 | 975.23 | 1.1 | -wrote /tmp/claude-1000/-home-osobh-projects-clawhdf5/422f755e-dd25-4c35-8613-5439087e3aaa/scratchpad/baseline_full.json ### After: HNSW neighbour-selection heuristic +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + Same harness, same data, after replacing closest-M neighbour selection with the HNSW paper's diversity heuristic (Algorithm 4, keeping pruned connections) for both new links and back-link pruning. Recall@10 at `ef = 64`: **0.87 → 1.00** @@ -922,6 +980,8 @@ build: 36472.8 ms (2742 vectors/s) · exact scan: 40 QPS, p50 24644 µs ### After: persistent keyword index, no store rewrite per query +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + `hybrid_search` used to rebuild the BM25 index from scratch (re-tokenising every record) and rewrite the whole `.h5` file on **every query**. The index is now kept for the life of the store and updated incrementally, and activation boosts @@ -942,6 +1002,8 @@ index removes that. ### After: vector index persisted with the checkpoint +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + The HNSW graph (not the vectors, which the store already holds) is saved to `.h5.ann` at each checkpoint and reloaded by `open()`, tied to that checkpoint by a generation id. The index is now built once per store (the *cold @@ -959,6 +1021,8 @@ index incrementally. ### After: unit-vector dot product, reusable visited set +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + Cosine distance recomputed both vector norms on every evaluation; the index now stores unit vectors and uses a plain dot product. The per-call `HashSet` of visited nodes became a reusable epoch-stamped array. Recall is unchanged. @@ -1004,6 +1068,8 @@ build: 21084.6 ms (4743 vectors/s) · exact scan: 39 QPS, p50 24739 µs ### After: unranked keyword scores, top-k merge (rankings unchanged) +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + A fusion study (`search_harness --fusion-study`) showed that capping the keyword candidate pool is **not** a safe optimisation: against the current full-corpus normalisation the final top-10 overlap is only 0.83-0.92 and the @@ -1026,6 +1092,8 @@ results. ### After: batched bulk build (optionally parallel); deletions handled in search +> *Superseded* by [Current: search harness](#current-search-harness-2026-09-24) (2026-09-24). + Profiling showed **90% of a build's distance evaluations are in back-link pruning**. The bulk build now inserts in batches: plan each node's neighbours against the graph as it stood at the start of the batch, link, then prune every @@ -1588,6 +1656,11 @@ MRR, or a one-question change in recency, is within this variation. ### Full haystack — `longmemeval_s`, n=500 (the number to cite) +This table is **BM25-only** (zero embeddings). With real embeddings and the +default hybrid 0.4/0.6 the same corpus gives turn Hit@5 **81.4%** (2026-09-27; +see [Fusion method](#fusion-method--weighted-vs-rrf-full-haystack-n500)), +which is the headline figure. + 47.7 sessions and 493.5 turns per question; 4.0% of haystack sessions are evidence sessions, so retrieval has to actually discriminate. @@ -1695,7 +1768,9 @@ over rank-1 precision. activation. Until now its combined score contained **no relevance term at all** — `RerankInput` did not carry the retrieval score — so a caller that re-ranked its candidates threw the retriever's ordering away and returned them -ordered by age. The OpenClaw backend did exactly that on every search. +ordered by age. `ClawhdfBackend` (the `openclaw` module) did exactly that on +every search. (OpenClaw itself never integrated clawhdf5; see +`docs/openclaw.md`.) Measuring that is unambiguous. "Recency" below is the share of `knowledge-update` questions where the newest gold session outranked the stale @@ -1793,7 +1868,7 @@ worth stating plainly rather than hiding: LongMemEval questions share substantia vocabulary with their evidence turns, which is close to the best case for lexical matching, and MiniLM at 384 dimensions is a small embedding model. -> **Run:** `cargo run --release --bin longmemeval_bench --features embeddings -- \ +> **Run:** `cargo run --release -p clawhdf5-bench --bin longmemeval_bench --features embeddings -- \ > benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings weights/all-minilm-l6-v2` > For the GPU path use `--features embeddings-cuda`. That requires `nvcc` on > `PATH` at *build* time — cudarc's build script shells out to it. The toolkit @@ -2212,7 +2287,8 @@ The tank row was measured 2026-09-24 on tank (AMD Ryzen 7 7800X3D), commit ### Reproducibility ```bash -rustup override set nightly +# Any stable toolchain at or above the MSRV (1.92) works; the original +# 2026-07-01 run used a nightly, later runs stable. # Latency benchmarks (Criterion) cargo bench -p clawhdf5-agent @@ -2315,7 +2391,9 @@ libhdf5 reads from a temp file including `open` + `read` + `close` overhead. | clawhdf5 hyperslab (f64, 10% slice) | — | 4.09 µs / **1.8 GiB/s** | 50.1 µs / **1.5 GiB/s** | libhdf5 f64 comparison excluded — clawhdf5's datatype encoding differs from libhdf5's (known -gap), making cross-format reads unreliable for comparison. +gap), making cross-format reads unreliable for comparison. (That gap was the float sign-bit +bug, fixed 2026-09-23: `docs/known-issues.md`, "Every `f32` dataset we wrote was unreadable by +h5py / libhdf5". The comparison has not been re-run since.) ### Chunked Read Throughput @@ -2423,6 +2501,12 @@ global file mutex and flushes to disk on every attribute write or group creation ## vs libhdf5 Summary +> Measured on the original i7-12650H (clawhdf5 2026-07-01, libhdf5 +> 2026-06-30). The newest run of this table is +> [Independent Validation: tank](#independent-validation-tank-ryzen-7-7800x3d-2026-08-03) +> (2026-08-03), which reproduces every row within ~15% except chunked +> write (45.3x on tank). + | Workload | clawhdf5 | libhdf5 | Speedup | |----------|----------|---------|---------| | Sequential read, 1K f32 | 634 ns | 45.2 µs | **71×** | @@ -2454,7 +2538,7 @@ to the page cache. There is no algorithmic headroom above ~1.7 GiB/s on this har ### Caveats -- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap). f64 results are clawhdf5-only. +- libhdf5 f64 read comparison excluded — clawhdf5's f32 datatype encoding differs from libhdf5's (known compatibility gap at the time; fixed 2026-09-23, see [Sequential Read Throughput](#sequential-read-throughput)). f64 results are clawhdf5-only. - Serial benchmarks. clawhdf5 uses Rayon for chunk compression when > 2 chunks; that parallelism is already reflected in the chunked write numbers. - clawhdf5 reads from `Vec` (zero-copy from mmap in production); libhdf5 reads from a temp file. This gives clawhdf5 a structural read advantage that reflects realistic API usage. @@ -2651,6 +2735,10 @@ Same not-like-for-like caveat as the "Comparison to MemX" section at the top of file applies — MemX's figure is end-to-end, these are a single component. Ratios are an order-of-magnitude indication, not a benchmark result. +> The Ratio column below was retracted afterwards: see +> [Comparison to MemX](#comparison-to-memx-arxiv260316171). Kept as recorded +> on 2026-08-05; do not cite it. + | Metric | MemX (claimed, end-to-end) | ClawhDF5 (tank, component only) | Ratio | |--------|----------------------------|----------------------------------|-------| | 100K flat search | <90 ms | 6.60 ms | ~14x |