Compare commits
16
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
817c5eee41 | ||
|
|
b2dce41532 | ||
|
|
1537a9464a | ||
|
|
12d9d8462f | ||
|
|
c913cd1cbf | ||
|
|
7d6e269bf3 | ||
|
|
6f5940d042 | ||
|
|
dfae9e2cc1 | ||
|
|
429c29b76b | ||
|
|
40527be653 | ||
|
|
2013fa94a0 | ||
|
|
a3e1cf8588 | ||
|
|
534331ffbe | ||
|
|
297ee5ec17 | ||
|
|
a319405ffc | ||
|
|
62595d5ac0 |
@@ -1,3 +1,6 @@
|
||||
/target
|
||||
Cargo.lock
|
||||
benchmarks/longmemeval/*.json
|
||||
|
||||
# Local model weights (MiniLM etc.) — large, not committed
|
||||
weights/
|
||||
|
||||
+371
-33
@@ -6,6 +6,26 @@
|
||||
**Rust:** 1.96.0-nightly (2026-03-14) · `--release` profile
|
||||
**Date:** 2026-07-01
|
||||
|
||||
> **Traceability note:** the "h5bench-Equivalent I/O Benchmarks" and both
|
||||
> "Independent Validation: tank" sections below meet a dated,
|
||||
> hardware-cited, reproducible standard (explicit date, machine spec, and a
|
||||
> runnable command per result) — this now covers "LongMemEval Results",
|
||||
> "SIMD & Parallelism", "Vector Search Latency", and "Comparison to MemX" via
|
||||
> their tank re-runs. The remaining undated sections above (Hybrid Search,
|
||||
> Knowledge Graph, Memory Consolidation, Temporal Index, Write Path, Decision
|
||||
> Gate, Memory Strategy, Multi-Session Benchmark, Memory Footprint,
|
||||
> Consolidation Efficiency, Ephemeral Tier) do not yet meet that bar — this is
|
||||
> a known, tracked documentation gap, not a claim that those numbers are wrong.
|
||||
>
|
||||
> **Correctness note (2026-08-06).** Being dated and reproducible is necessary but
|
||||
> not sufficient — a number can be perfectly reproducible and still measure the
|
||||
> wrong thing. A methodology audit found two such cases and both have been
|
||||
> retracted in place: the session-level LongMemEval figures (degenerate on the
|
||||
> oracle variant) and the MemX retrieval comparison (mismatched granularity and
|
||||
> corpus). Every cross-system comparison in this file now carries an explicit
|
||||
> scoping caveat. Where a section states a scoring target, that declaration is the
|
||||
> contract — read it before citing the number.
|
||||
|
||||
---
|
||||
|
||||
## Vector Search Latency
|
||||
@@ -24,10 +44,19 @@ Brute-force cosine similarity over 384-dimensional embeddings (OpenAI text-embed
|
||||
|
||||
MemX claims end-to-end search under 90ms at 100K records (Rust + libSQL + FTS5).
|
||||
|
||||
| Metric | MemX (claimed) | ClawhDF5 | Speedup |
|
||||
|--------|----------------|----------|---------|
|
||||
| 100K flat search | <90 ms | 11.4 ms | **~8x** |
|
||||
| 100K IVF-PQ search | — | 1.19 ms | **~76x** |
|
||||
> **Caveat — not like-for-like.** MemX's `<90 ms` is *end-to-end* search across their
|
||||
> full pipeline (dense embeddings + FTS5 + four-factor re-ranking). The clawhdf5
|
||||
> figures below are a *single component* — raw vector search latency, excluding
|
||||
> embedding, keyword, fusion, and re-ranking stages. A component measured against a
|
||||
> full pipeline will always look favourable; the "speedup" column overstates the real
|
||||
> advantage by an unquantified margin and should be read as an order-of-magnitude
|
||||
> indication only, not a benchmark result. Matching MemX's measurement boundary is
|
||||
> tracked as follow-up work.
|
||||
|
||||
| Metric | MemX (claimed, end-to-end) | ClawhDF5 (component only) | Ratio |
|
||||
|--------|----------------------------|---------------------------|-------|
|
||||
| 100K flat search | <90 ms | 11.4 ms | ~8x |
|
||||
| 100K IVF-PQ search | — | 1.19 ms | ~76x |
|
||||
| Keyword search 10K | 1,100x improvement over unindexed | 583 µs (BM25) | Comparable |
|
||||
|
||||
---
|
||||
@@ -174,46 +203,193 @@ _Latency benchmarks generated with Criterion.rs (50-100 samples per benchmark).
|
||||
|
||||
## LongMemEval Results
|
||||
|
||||
**Dataset:** LongMemEval oracle (500 questions, 6 question types, variable-length chat histories)
|
||||
> **Scoring target declaration.** Per [arXiv 2605.24060](https://arxiv.org/abs/2605.24060),
|
||||
> which found that changing scoring target alone alters nDCG on 83–94% of queries and
|
||||
> can reverse system rankings, this section states its measurement contract explicitly:
|
||||
>
|
||||
> - **Dataset variant:** both are now reported below — the full `longmemeval_s`
|
||||
> haystack (**the headline number**) and `longmemeval_oracle` (evidence sessions
|
||||
> only, a substantially easier corpus, kept for continuity). The harness does not
|
||||
> trust the filename: it measures evidence-session density from the data and
|
||||
> labels the run from that, so a mislabelled input cannot yield a mislabelled
|
||||
> result. Measured density is 4.0% on `longmemeval_s` and 100.0% on the oracle.
|
||||
> - **Metric:** *retrieval recall.* A "hit" means the gold-labelled memory appeared in
|
||||
> the top-k. **No answer is generated and none is scored** — the dataset's `answer`
|
||||
> field is deserialized and never read. This is **not** the official LongMemEval
|
||||
> leaderboard metric, which is end-to-end QA accuracy (retrieve → generate → LLM
|
||||
> judge). Retrieval recall reported as QA accuracy typically overstates by 20–30 points.
|
||||
> - **Granularity:** turn-level = the returned memory's source turn had `has_answer == true`.
|
||||
> - **k = 10**, n = 500.
|
||||
> - **Retrieval mode:** all three are reported below. Historically the bench passed
|
||||
> zero-vector embeddings with `vector_weight=0.0`, so the HNSW/vector stage was
|
||||
> inert and every published number was BM25 alone. Real `all-MiniLM-L6-v2`
|
||||
> embeddings are now available via `--features embeddings --embeddings <dir>`,
|
||||
> and BM25-only / vector-only / hybrid are each measured separately.
|
||||
|
||||
**Mode:** BM25-only retrieval — zero embeddings, `vector_weight=0.0`, `keyword_weight=1.0`
|
||||
**Reference:** MemX (arxiv:2603.16171) with full embedding system: Hit@5=51.6%, MRR=0.380
|
||||
|
||||
> **Run:** `cargo run --release --bin longmemeval_bench`
|
||||
> **Run:** `cargo run --release --bin longmemeval_bench -- benchmarks/longmemeval/longmemeval_s_cleaned.json`
|
||||
> (~70 s for all 500 questions on the tank reference machine). Omit the path for the
|
||||
> oracle variant; add `--limit N` for an evenly-strided subsample.
|
||||
|
||||
### Session-Level Recall (n=500)
|
||||
### Full haystack — `longmemeval_s`, n=500 (the number to cite)
|
||||
|
||||
| Metric | ClawhDF5 (BM25-only) |
|
||||
|--------|---------------------|
|
||||
| Hit@1 | **100.0%** |
|
||||
| Hit@5 | **100.0%** |
|
||||
| Hit@10 | **100.0%** |
|
||||
| MRR | **1.0000** |
|
||||
47.7 sessions and 493.5 turns per question; 4.0% of haystack sessions are evidence
|
||||
sessions, so retrieval has to actually discriminate.
|
||||
|
||||
Perfect session-level recall across all 500 questions and all 6 question types.
|
||||
| Metric | Turn-level | Session-level |
|
||||
|--------|-----------|---------------|
|
||||
| Hit@1 | 53.8% | 86.2% |
|
||||
| Hit@5 | **75.0%** | **93.6%** |
|
||||
| Hit@10 | 81.6% | 96.6% |
|
||||
| MRR | 0.6320 | 0.8948 |
|
||||
|
||||
### Turn-Level Recall (n=500)
|
||||
Session-level is reported here because on this corpus it is meaningful — unlike on
|
||||
the oracle variant, where it was degenerate and was retracted (below). At 4.0%
|
||||
evidence density a session-level hit reflects discrimination rather than corpus
|
||||
shape.
|
||||
|
||||
| Metric | ClawhDF5 (BM25-only) | MemX (full system)¹ |
|
||||
|--------|---------------------|---------------------|
|
||||
| Hit@1 | **52.6%** | — |
|
||||
| Hit@5 | **84.4%** | 51.6% |
|
||||
| Hit@10 | **90.4%** | — |
|
||||
| MRR | **0.6597** | 0.380 |
|
||||
Per-type, session-level: `single-session-assistant` 100.0% Hit@1 (n=56),
|
||||
`knowledge-update` 96.2% (n=78), `single-session-user` 94.3% (n=70),
|
||||
`multi-session` 84.2% (n=133), `temporal-reasoning` 84.2% (n=133), and
|
||||
`single-session-preference` 33.3% (n=30) — the one category where BM25 clearly
|
||||
struggles, since a preference question's evidence rarely shares vocabulary with
|
||||
the question.
|
||||
|
||||
**clawhdf5 outperforms MemX at turn-level retrieval** — Hit@5 84.4% vs 51.6%, MRR 0.66 vs 0.38 — with BM25 alone, no embeddings needed.
|
||||
### Retrieval mode ablation — full haystack, n=500
|
||||
|
||||
> ¹ MemX uses dense embeddings + FTS5 + four-factor re-ranking. Our BM25-only result exceeds their full pipeline.
|
||||
Real 384-d `all-MiniLM-L6-v2` embeddings, 190,015 unique texts encoded once on an
|
||||
RTX 5060 Ti (~13 min; the same work on the 8-core CPU was still unfinished after
|
||||
30 minutes, so the GPU path is not a convenience here). Turn-level:
|
||||
|
||||
### Per-Type Breakdown (session-level)
|
||||
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|
||||
|------|-------|-------|--------|-----|
|
||||
| BM25 only (`0.0`/`1.0`) | **53.8%** | 75.0% | 81.6% | **0.6320** |
|
||||
| Vector only (`1.0`/`0.0`) | 36.0% | 71.8% | 81.6% | 0.5027 |
|
||||
| Hybrid (`0.7`/`0.3`) | 44.4% | **79.2%** | **86.0%** | 0.5868 |
|
||||
|
||||
| Question Type | N | Hit@1 | Hit@5 | Hit@10 | MRR |
|
||||
|---------------|---|-------|-------|--------|-----|
|
||||
| single-session-user | 70 | 100.0% | 100.0% | 100.0% | 1.0000 |
|
||||
| single-session-assistant | 56 | 100.0% | 100.0% | 100.0% | 1.0000 |
|
||||
| single-session-preference | 30 | 100.0% | 100.0% | 100.0% | 1.0000 |
|
||||
| temporal-reasoning | 133 | 100.0% | 100.0% | 100.0% | 1.0000 |
|
||||
| multi-session | 133 | 100.0% | 100.0% | 100.0% | 1.0000 |
|
||||
| knowledge-update | 78 | 100.0% | 100.0% | 100.0% | 1.0000 |
|
||||
Session-level:
|
||||
|
||||
| Mode | Hit@1 | Hit@5 | Hit@10 | MRR |
|
||||
|------|-------|-------|--------|-----|
|
||||
| BM25 only | 86.2% | 93.6% | 96.6% | 0.8948 |
|
||||
| Vector only | 85.4% | 94.2% | 96.6% | 0.8901 |
|
||||
| Hybrid | **88.2%** | **95.8%** | **97.8%** | **0.9158** |
|
||||
|
||||
### Weight sweep — full haystack, n=500
|
||||
|
||||
`0.7/0.3` was a documented default, never a searched one. Sweeping
|
||||
`vector_weight` from 0.0 to 1.0 (`--sweep`, reusing the one-time embedding
|
||||
table) shows it is not merely suboptimal but **strictly dominated**:
|
||||
|
||||
| vector / keyword | Hit@1 | Hit@5 | Hit@10 | MRR | session Hit@5 |
|
||||
|---|---|---|---|---|---|
|
||||
| 0.0 / 1.0 (BM25) | **53.8%** | 75.0% | 81.6% | 0.6320 | 93.6% |
|
||||
| 0.1 / 0.9 | 53.2% | 77.4% | 83.8% | 0.6374 | 95.0% |
|
||||
| 0.2 / 0.8 | 53.6% | 78.2% | 85.6% | 0.6440 | 95.4% |
|
||||
| 0.3 / 0.7 | 53.2% | 78.8% | 87.2% | **0.6463** | 96.0% |
|
||||
| **0.4 / 0.6** | 51.6% | **81.4%** | 87.8% | 0.6429 | 96.8% |
|
||||
| 0.5 / 0.5 | 48.2% | **81.4%** | **88.2%** | 0.6234 | **97.4%** |
|
||||
| 0.6 / 0.4 | 46.6% | 79.8% | 87.4% | 0.6069 | 96.6% |
|
||||
| 0.7 / 0.3 *(old default)* | 44.4% | 79.2% | 86.0% | 0.5868 | 95.8% |
|
||||
| 0.8 / 0.2 | 40.6% | 76.2% | 85.4% | 0.5571 | 95.2% |
|
||||
| 0.9 / 0.1 | 37.8% | 73.4% | 84.6% | 0.5289 | 94.2% |
|
||||
| 1.0 / 0.0 (vector) | 36.0% | 71.8% | 81.6% | 0.5027 | 94.2% |
|
||||
|
||||
**`0.4/0.6` beats `0.7/0.3` on every metric at both granularities** — Hit@1
|
||||
+7.2pp, Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. There is no trade being made; the
|
||||
old default was simply on the wrong side of the peak. **`0.4/0.6` is the
|
||||
recommended setting**, with `0.3/0.7` preferable if rank-1 precision matters
|
||||
most (it takes the best MRR in the sweep and gives up only 0.6pp of Hit@1
|
||||
against pure BM25).
|
||||
|
||||
**Correction.** An earlier revision of this section, measuring only `0.7/0.3`,
|
||||
concluded that fusion "buys deeper recall and pays for it at rank 1" and advised
|
||||
callers taking a single top hit to prefer BM25. That was an artifact of the
|
||||
badly-chosen weight, not a property of fusion. At `0.3/0.7` hybrid *beats* BM25
|
||||
on MRR (0.6463 vs 0.6320) and on Hit@5 (78.8% vs 75.0%) while costing 0.6pp of
|
||||
Hit@1. The advice below is corrected accordingly.
|
||||
|
||||
**Hybrid wins, once the weights are right.** At the old `0.7/0.3` the picture
|
||||
looked like a trade: best at Hit@5 and Hit@10, worse than BM25 at Hit@1 and MRR.
|
||||
The sweep above shows that was the weight, not fusion. At `0.4/0.6` hybrid leads
|
||||
Hit@5 and Hit@10 outright; at `0.3/0.7` it also leads MRR and is within 0.6pp of
|
||||
BM25 at Hit@1. Both dominate `0.7/0.3`.
|
||||
|
||||
The rows below are kept at the three original settings because they are what the
|
||||
mode ablation measured — read them as "the shape of each stage in isolation",
|
||||
and take the operating point from the sweep.
|
||||
|
||||
The same pattern shows up independently in omni-cortex's four-signal RRF ablation,
|
||||
where adding BM25 to a dense retriever raised nDCG@5 while lowering Hit@1 and MRR.
|
||||
Two different codebases, two different fusion schemes, same direction.
|
||||
|
||||
Vector-only being *worse* than BM25 at every turn-level cutoff except Hit@10 is
|
||||
worth stating plainly rather than hiding: LongMemEval questions share substantial
|
||||
vocabulary with their evidence turns, which is close to the best case for lexical
|
||||
matching, and MiniLM at 384 dimensions is a small embedding model.
|
||||
|
||||
> **Run:** `cargo run --release --bin longmemeval_bench --features embeddings -- \
|
||||
> benchmarks/longmemeval/longmemeval_s_cleaned.json --embeddings weights/all-minilm-l6-v2`
|
||||
> For the GPU path use `--features embeddings-cuda`. That requires `nvcc` on
|
||||
> `PATH` at *build* time — cudarc's build script shells out to it. The toolkit
|
||||
> installs to `/usr/local/cuda/bin`, which many distributions do not export;
|
||||
> check with `nvcc --version` and, if it is missing, add it somewhere every
|
||||
> shell reads (for zsh that is `~/.zshenv`, not `~/.zshrc`, since build tooling
|
||||
> runs non-interactively). The device is selected at runtime with a CPU
|
||||
> fallback, so a machine without CUDA still produces correct numbers — just far
|
||||
> more slowly, and the bench says so on startup.
|
||||
>
|
||||
> Weights: `huggingface.co/sentence-transformers/all-MiniLM-L6-v2` — place
|
||||
> `model.safetensors` and `tokenizer.json` in the `--embeddings` directory.
|
||||
|
||||
### Oracle variant — `longmemeval_oracle`, n=500 (easier corpus, kept for continuity)
|
||||
|
||||
| Metric | ClawhDF5 (BM25-only, oracle variant) |
|
||||
|--------|--------------------------------------|
|
||||
| Hit@1 | 52.6% |
|
||||
| Hit@5 | **84.4%** |
|
||||
| Hit@10 | 90.4% |
|
||||
| MRR | 0.6597 |
|
||||
|
||||
Turn-level. The 9.4-point gap between this and the full haystack's 75.0% is the
|
||||
price of the harder corpus, and is the reason oracle-only numbers should not be
|
||||
presented as LongMemEval results. Session-level figures on this variant are
|
||||
degenerate — see below.
|
||||
|
||||
With real embeddings the same oracle corpus gives BM25-only 84.2% / vector-only
|
||||
80.4% / hybrid **85.2%** Hit@5 turn-level — hybrid ahead at Hit@5 and Hit@10 and
|
||||
behind at Hit@1, matching the full-haystack pattern above. (BM25-only reads 84.2%
|
||||
here against 84.4% with zero embedding vectors: one question of 500 changes rank,
|
||||
with MRR identical at 0.6597. On the full haystack the two agree exactly.)
|
||||
|
||||
### Retracted: session-level recall and the MemX comparison
|
||||
|
||||
Earlier revisions of this file reported session-level Hit@1/5/10 of **100.0%** with
|
||||
MRR **1.0000**, uniform across all six question types, and claimed clawhdf5
|
||||
"outperforms MemX at turn-level retrieval (84.4% vs 51.6%)". **Both are withdrawn.**
|
||||
|
||||
**The session-level numbers are a degenerate artifact.** On the `longmemeval_oracle`
|
||||
variant, the ingested haystack for a question consists essentially only of that
|
||||
question's evidence sessions. Every returned document therefore belongs to an answer
|
||||
session, so session-level hit rate is ≈1.0 at rank 0 *by construction* — which is
|
||||
exactly why the result was a uniform 100.0% across every question type. It measured
|
||||
the shape of the corpus, not the retriever.
|
||||
|
||||
**The MemX comparison was not like-for-like on two independent axes.** MemX
|
||||
([arxiv:2603.16171](https://arxiv.org/abs/2603.16171)) reports Hit@5 = 51.6% /
|
||||
MRR = 0.380 at **fact-level granularity over 220,349 fact-level records drawn from
|
||||
19,195 sessions**, and explicitly notes that fact-level "doubl[es] session-level
|
||||
performance." Our 84.4% is **turn-level, on the oracle subset**. Different retrieval
|
||||
granularity, and a corpus smaller by orders of magnitude. A higher number on an
|
||||
easier corpus at a different granularity is not an outperformance claim, and it
|
||||
should not have been presented as one.
|
||||
|
||||
The full-haystack half of that gap is now closed: the section above reports
|
||||
`longmemeval_s` over all 500 questions. The **granularity** mismatch remains — MemX
|
||||
measures fact-level, we measure turn-level and session-level — so no cross-system
|
||||
claim is made here even now. Matching granularity would require fact-level
|
||||
extraction over the haystack, which this harness does not do.
|
||||
|
||||
### Search Latency (LongMemEval, n=500 queries)
|
||||
|
||||
@@ -368,6 +544,54 @@ No network hop, no serialization — direct HashMap operations.
|
||||
|
||||
---
|
||||
|
||||
## World-Model Sample Loading (vs h5py / stable-worldmodel shape)
|
||||
|
||||
Reproduces the access pattern of `stable-worldmodel`'s HDF5 dataloader
|
||||
([arXiv 2605.21800](https://arxiv.org/abs/2605.21800), LeCun/Balestriero
|
||||
group), which supports HDF5 as one of three native formats and measures
|
||||
generic HDF5 at **1,416-1,474 samples/s** (vs Lance 4,815) for per-frame
|
||||
sample loading. This benchmark measures **clawhdf5 vs h5py on the same
|
||||
machine and the same file**, so the comparison is hardware-controlled.
|
||||
|
||||
**Absolute numbers are not comparable to the paper's** - different hardware
|
||||
(AMD Ryzen 7 7800X3D, local NVMe, warm page cache), smaller frames, and no
|
||||
torch-tensor / transform step. Only the clawhdf5-vs-h5py ratio *here* is a
|
||||
controlled result. The workload is the dataloader shape: a `(N, H, W, C)`
|
||||
uint8 observation dataset (20,000 x 64x64x3 = 246 MB), each frame read once
|
||||
per pass in a fixed shuffled (random-access) order, 10 passes.
|
||||
|
||||
Both read a **file written by h5py** - clawhdf5 parsing an
|
||||
externally-produced HDF5 file is itself the interop result. h5py opens SWMR
|
||||
with a 256 MB chunk cache, exactly `stable-worldmodel`'s `HDF5Dataset`; it
|
||||
materialises each frame as a numpy array (`d[i]`) and sums it. clawhdf5
|
||||
mmaps once, takes a zero-copy `&[u8]` over the contiguous dataset, and
|
||||
indexes frame `i` as a subslice.
|
||||
|
||||
| Reader | samples/sec (median of 3) | vs h5py |
|
||||
|--------|---------------------------|---------|
|
||||
| **clawhdf5** (zero-copy view) | **593,000** | **8.1x** |
|
||||
| **clawhdf5** (materialised copy per frame) | **518,000** | **7.1x** |
|
||||
| h5py (swmr, 256 MB cache) | 73,000 | 1.0x |
|
||||
|
||||
The **materialised-copy row is the fair, equal-work comparison** - it
|
||||
`to_vec()`s every frame so clawhdf5 pays the same per-frame allocation h5py
|
||||
does, and it is still **7.1x faster**. That the copy costs almost nothing
|
||||
(518k vs 593k) shows the h5py gap is **per-frame call overhead** (Python +
|
||||
library dispatch), not data movement. This is an in-page-cache measurement:
|
||||
it isolates the read-path overhead both libraries add on top of the OS,
|
||||
which is the thing that differs - not disk bandwidth, which is shared.
|
||||
|
||||
Reproduce (`benchmarks/`):
|
||||
|
||||
```bash
|
||||
python benchmarks/gen_worldmodel_frames.py /tmp/wm_frames.h5 20000
|
||||
cargo run --release -p clawhdf5-bench --example worldmodel_sampling -- /tmp/wm_frames.h5 10
|
||||
cargo run --release -p clawhdf5-bench --example worldmodel_sampling -- /tmp/wm_frames.h5 10 --copy
|
||||
python benchmarks/bench_worldmodel_h5py.py /tmp/wm_frames.h5 10
|
||||
```
|
||||
|
||||
Measured 2026-08-07 on tank (Ryzen 7 7800X3D, 246 MB dataset in page cache).
|
||||
|
||||
## Cross-Platform Notes
|
||||
|
||||
> **Run:** `./benchmarks/cross_platform.sh [--full] [--output results.json]`
|
||||
@@ -649,3 +873,117 @@ cargo bench -p clawhdf5-bench --features libhdf5-compare --bench h5bench_meta --
|
||||
cargo bench -p clawhdf5-bench --features libhdf5-compare --bench h5bench_meta -- metadata_parse_in_memory
|
||||
cargo bench -p clawhdf5-bench --features libhdf5-compare --bench h5bench_read -- read_zerocopy_mmap
|
||||
```
|
||||
|
||||
## Independent Validation: tank — LongMemEval & Vector Search (Ryzen 7 7800X3D), 2026-08-05
|
||||
|
||||
Re-running the "LongMemEval Results" and "SIMD & Parallelism" sections above on
|
||||
tank (AMD Ryzen 7 7800X3D, 8C/16T, Ubuntu 26.04, same machine as the
|
||||
vs-libhdf5 validation above) to give both sections the dated, hardware-cited,
|
||||
reproducible citation the top-of-file traceability note flags them as
|
||||
missing.
|
||||
|
||||
### LongMemEval Results (reproduction)
|
||||
|
||||
```bash
|
||||
cd benchmarks/longmemeval
|
||||
wget https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_oracle.json
|
||||
cargo run --release --bin longmemeval_bench
|
||||
```
|
||||
|
||||
Recall numbers are deterministic (pure BM25 retrieval over a fixed dataset) and
|
||||
reproduce exactly. Scoring target as declared in the LongMemEval section above:
|
||||
retrieval recall, turn-level, k=10, `longmemeval_oracle` variant, BM25-only.
|
||||
|
||||
| Metric | Turn-Level |
|
||||
|--------|------------|
|
||||
| Hit@1 | 52.6% |
|
||||
| Hit@5 | **84.4%** |
|
||||
| Hit@10 | 90.4% |
|
||||
| MRR | 0.6597 |
|
||||
|
||||
Session-level figures are omitted here — they are degenerate on the oracle variant
|
||||
and have been retracted; see "Retracted: session-level recall and the MemX
|
||||
comparison" above.
|
||||
|
||||
Search latency (hardware-dependent, tank numbers):
|
||||
|
||||
| Metric | avg | p50 | p95 | p99 |
|
||||
|--------|-----|-----|-----|-----|
|
||||
| Latency | 2,431 µs | 2,105 µs | 7,250 µs | 12,018 µs |
|
||||
|
||||
Higher than the i7-12650H figures at the top of this file (avg 1,004 µs) despite
|
||||
tank's faster single-core performance elsewhere in this document — BM25 search
|
||||
latency here scales with per-question haystack size and this run's variance is
|
||||
wider (p99 is ~5x the mean), suggesting this metric is more sensitive to
|
||||
momentary scheduling/cache effects than the flat-array vector-search benchmarks.
|
||||
Recorded as-is rather than smoothed.
|
||||
|
||||
### SIMD & Parallelism (reproduction, with a correction)
|
||||
|
||||
```bash
|
||||
cargo bench -p clawhdf5-agent --bench bench -- "^(strategy_scalar_10k|strategy_simd_10k|strategy_rayon_10k|adaptive_search_10k|simd_cosine_100k|rayon_cosine_100k)$"
|
||||
```
|
||||
|
||||
The original 10K table above compares named benchmarks (`vector_search`,
|
||||
`rayon`, `strategy`) that, on inspection, don't all exercise the same
|
||||
scalar-vs-SIMD-vs-parallel axis the table implies — several of the
|
||||
`simd_cosine_10k`/`sequential_cosine_10k`-style benchmarks actually call the
|
||||
same underlying function under different names. The `adaptive_benches` group's
|
||||
`strategy_scalar_10k` / `strategy_simd_10k` / `strategy_rayon_10k` benchmarks
|
||||
are the ones that genuinely hold the dataset fixed and vary only the
|
||||
`SearchStrategy` enum, so they're the correct apples-to-apples comparison —
|
||||
used here instead.
|
||||
|
||||
| Strategy | Latency (tank) | vs Sequential |
|
||||
|----------|-----------------|----------------|
|
||||
| Sequential (scalar) | 502 µs | 1.0x |
|
||||
| SIMD (auto-vectorized) | 327 µs | **1.53x** |
|
||||
| Rayon (parallel) | 323 µs | **1.55x** |
|
||||
| Adaptive (auto-select) | 339 µs | **1.48x** |
|
||||
|
||||
Honest finding: the speedup from SIMD/parallelism over scalar is real but
|
||||
smaller here (~1.5x) than the i7-12650H figures above (~2.0x). The Ryzen 7
|
||||
7800X3D's large L3 cache (96MB 3D V-Cache) measurably narrows the gap versus a
|
||||
naive scalar loop compared to the i7 — this is a genuine hardware-dependent
|
||||
result, not a regression or measurement error, and is recorded rather than
|
||||
reconciled away.
|
||||
|
||||
At 100K, no `strategy_*` benchmark exists in the current suite (`adaptive_benches`
|
||||
only covers n=10,000), so this row uses the same `simd_cosine_100k`/
|
||||
`rayon_cosine_100k` benchmarks as the original table — not a true scalar
|
||||
baseline, so no "vs Sequential" multiple is reported for it:
|
||||
|
||||
| Strategy | Latency (tank) |
|
||||
|----------|-----------------|
|
||||
| SIMD | 6.60 ms |
|
||||
| Rayon parallel | 4.73 ms |
|
||||
|
||||
### Vector Search Latency & Comparison to MemX (reproduction)
|
||||
|
||||
```bash
|
||||
cargo bench -p clawhdf5-agent --bench bench -- "^(vector_search_1k|simd_cosine_10k|simd_cosine_100k|prenorm_search_10k|ivf_search_10k_nprobe10|ivf_search_100k_nprobe10|ivf_pq_search_100k|rairs_search_10k_nprobe10|bm25_search_10k)$"
|
||||
```
|
||||
|
||||
| Scale | Flat Search | Pre-norm | IVF (nprobe=10) | IVF-PQ | RAIRS |
|
||||
|-------|-------------|----------|-----------------|--------|-------|
|
||||
| **1K** | 47.8 µs | — | — | — | — |
|
||||
| **10K** | 501 µs | 322 µs | 24.8 µs | — | 109 µs |
|
||||
| **100K** | 6.60 ms | — | 608 µs | 865 µs | — |
|
||||
|
||||
(The 1K Pre-norm cell from the original table has no corresponding benchmark
|
||||
in the current suite — not re-verified, left blank rather than guessed.)
|
||||
|
||||
Same not-like-for-like caveat as the "Comparison to MemX" section at the top of this
|
||||
file applies — MemX's figure is end-to-end, these are a single component. Ratios are
|
||||
an order-of-magnitude indication, not a benchmark result.
|
||||
|
||||
| Metric | MemX (claimed, end-to-end) | ClawhDF5 (tank, component only) | Ratio |
|
||||
|--------|----------------------------|----------------------------------|-------|
|
||||
| 100K flat search | <90 ms | 6.60 ms | ~14x |
|
||||
| 100K IVF-PQ search | — | 865 µs | ~104x |
|
||||
| Keyword search 10K | 1,100x improvement over unindexed | 520 µs (BM25) | Comparable |
|
||||
|
||||
Every figure in this subsection is faster than the corresponding i7-12650H
|
||||
number at the top of this file, consistent with the Ryzen 7 7800X3D's higher
|
||||
single-core throughput and larger cache observed in the vs-libhdf5 validation
|
||||
above.
|
||||
|
||||
@@ -2,6 +2,90 @@
|
||||
|
||||
## Unreleased
|
||||
|
||||
### Security
|
||||
- `clawhdf5-format`: bounded decompression output (`MAX_DECOMPRESS_SIZE`) for
|
||||
deflate/lz4/zstd/pcodec so a crafted compressed chunk can't drive an
|
||||
unbounded allocation (memory-exhaustion DoS).
|
||||
- `clawhdf5-format`: `chunked_read.rs`/`data_read.rs`/`local_heap.rs` bounds
|
||||
audit — added `ensure_len` overflow guards at every plain-arithmetic
|
||||
offset+size check, a recursion-depth guard against a crafted
|
||||
self-referencing/cyclic B-tree chunk index, a fix for an unguarded
|
||||
compound-datatype `byte_offset` overrun in `read_compound_fields`, and an
|
||||
`ndims - 1` underflow guard for degenerate zero-dimension chunked layouts.
|
||||
Added a new `fuzz_dataset_read` cargo-fuzz target (walks every dataset in a
|
||||
parsed file and exercises the contiguous/chunked/compact raw-data read
|
||||
paths) which found and fixed 3 real crash bugs — an integer-multiply
|
||||
overflow in `copy_chunk_to_output`'s N-D assembly path, the `ndims - 1`
|
||||
underflow above, and an overflow in `local_heap.rs` — within the first few
|
||||
fuzzing runs.
|
||||
- `clawhdf5-format`: `btree_v1.rs` overflow-safe bounds checks via a local
|
||||
`ensure_len` helper, closing a `usize`-overflow panic reachable from a
|
||||
crafted near-`usize::MAX` B-tree offset.
|
||||
- `clawhdf5-agent`: WAL length-prefix caps (`MAX_WAL_FIELD_LEN`, 64 MiB) reject
|
||||
a corrupted/truncated length claim before allocating. Followed by a full
|
||||
per-entry CRC32 trailer (`WAL_VERSION` bumped to 2) — a bit-flip inside an
|
||||
entry now stops replay cleanly instead of silently accepting corrupted
|
||||
data. Old-format WAL files are still read correctly and migrated to the new
|
||||
format on next open.
|
||||
- `clawhdf5-android`: validate `embedding_len`/`query_embedding_len` against
|
||||
the handle's configured `embedding_dim` (and reject null pointers) before
|
||||
constructing a slice from a raw pointer in `edgehdf5_save` /
|
||||
`edgehdf5_hybrid_search`.
|
||||
- `clawhdf5-py`: bump pyo3/numpy `0.28` → `0.29`, clearing two RUSTSEC
|
||||
advisories (OOB read in `PyList`/`PyTuple` iterator; missing `Sync` bound on
|
||||
`PyCFunction::new_closure`).
|
||||
- Clarified that the integrity hashes in `clawhdf5-agent::provenance`
|
||||
(FNV-1a) and `clawhdf5-format::provenance` (SHA-256) are unkeyed and detect
|
||||
only accidental corruption, not tampering — doc-only change, no behavior
|
||||
change.
|
||||
|
||||
### Performance
|
||||
- `clawhdf5-format`: chunk cache lookup is now O(1) (`slot_index: HashMap`)
|
||||
instead of a linear scan, and cache hits return a shared `Arc` instead of
|
||||
cloning the decompressed buffer — the hottest path in chunked reads.
|
||||
- `clawhdf5-ann`: optional `parallel` feature (rayon) parallelizes HNSW's
|
||||
`prune_connections` neighbor-distance computation. The outer build/insert
|
||||
loop is deliberately left sequential — it has genuine cross-iteration data
|
||||
dependencies and needs its own correctness-focused design pass.
|
||||
- `clawhdf5-format/chunked_read.rs`: removed 12 unnecessary
|
||||
`chunk_dimensions[..rank].to_vec()` allocations where callees already
|
||||
accept `&[u32]`.
|
||||
|
||||
### Architecture
|
||||
- Added `.gitea/workflows/ci.yml`, actually wiring the long-existing
|
||||
`scripts/ci-test.sh` (fmt, clippy, tests, no_std check) into CI on every
|
||||
push/PR to `main`. Fixed stale package names in `ci-test.sh`/
|
||||
`check-nostd.sh` that had been silently no-op'ing the `clawhdf5-py`
|
||||
exclusion and the no_std check.
|
||||
- Fixed a genuine no_std build break in `clawhdf5-format` (uncovered once the
|
||||
no_std CI check actually started running): `core::sync::atomic::AtomicU64`
|
||||
doesn't exist on `thumbv7em-none-eabihf` (switched to `portable-atomic`),
|
||||
missing `alloc` imports for `Box`/`Vec`/`format!` on a few no_std paths, and
|
||||
`f64::powi` (std/libm-only) replaced with a local exponentiation-by-squaring
|
||||
helper in the scale-offset filter.
|
||||
- Added `[workspace.dependencies]` for `tempfile`/`criterion`/`half`/`serde`,
|
||||
fixing a real version skew on `half` (`2` vs `2.7` across crates).
|
||||
- Fixed version skew: `clawhdf5-py` (`pyproject.toml`) and
|
||||
`packages/clawhdf5-node` (`package.json`) were both behind the actual crate
|
||||
version (2.1.0).
|
||||
- Documented that the `mpi-io` feature's read/write paths are root-read
|
||||
+broadcast / gather-to-rank-0, not true collective I/O.
|
||||
|
||||
### Documentation
|
||||
- BENCHMARKS.md: re-ran the previously-undated "LongMemEval Results", "SIMD &
|
||||
Parallelism", and "Vector Search Latency"/"Comparison to MemX" sections on
|
||||
a second machine (tank, Ryzen 7 7800X3D) with explicit dates and reproduce
|
||||
commands. Found and corrected a methodology issue in the SIMD/Parallelism
|
||||
benchmark selection (several originally-compared benchmarks didn't actually
|
||||
isolate the scalar/SIMD/parallel axis).
|
||||
- README.md / ROADMAP.md / CLAUDE.md: corrected several stale facts —
|
||||
the `clawhdf5-types` crate (removed earlier) was still listed in the
|
||||
README crate map; the LongMemEval numbers in the README badge and table
|
||||
didn't match the actual (much better) benchmark results in BENCHMARKS.md;
|
||||
total line-of-code and test-count figures were stale; `clawhdf5-gpu`'s
|
||||
CubeCL→wgpu correction; documented the new `clawhdf5-ann` `parallel`
|
||||
feature flag, which had no entry in the Feature Flags table.
|
||||
|
||||
### New Features
|
||||
- `clawhdf5-migrate`: substantial engine improvements:
|
||||
- **Real content validation** — the post-migration check now reads the written
|
||||
|
||||
@@ -17,7 +17,7 @@ Cargo workspace with 16 crates under `crates/` (plus `libaec-sys`, an internal F
|
||||
| `clawhdf5-netcdf4` | NetCDF-4 compatibility layer |
|
||||
| `clawhdf5-ann` | HNSW approximate nearest-neighbor vector index |
|
||||
| `clawhdf5-agent` | Agent memory, session history, knowledge graph storage |
|
||||
| `clawhdf5-gpu` | GPU-accelerated I/O via CubeCL |
|
||||
| `clawhdf5-gpu` | GPU-accelerated I/O via wgpu (hand-written WGSL compute shaders) |
|
||||
| `clawhdf5-accel` | CPU SIMD acceleration path |
|
||||
| `clawhdf5-migrate` | Schema migration engine |
|
||||
| `clawhdf5-android` | Android JNI bindings |
|
||||
@@ -33,7 +33,7 @@ Cargo workspace with 16 crates under `crates/` (plus `libaec-sys`, an internal F
|
||||
the approximate `clawhdf5-ann` index for the vector stage (the index mirrors
|
||||
the cache and self-heals on drift). Build the agent with
|
||||
`--no-default-features --features float16` to force the exact linear cosine scan.
|
||||
- WAL (write-ahead log) for crash-safe persistence
|
||||
- WAL (write-ahead log) for crash-safe persistence, with a CRC32 trailer per entry so a corrupted entry stops replay cleanly instead of loading bad data
|
||||
- GPU-accelerated batch I/O for large dataset processing
|
||||
- Python and Node.js bindings for cross-language use
|
||||
- NetCDF-4 compatibility for scientific data interop
|
||||
|
||||
@@ -25,3 +25,9 @@ version = "2.1.0"
|
||||
edition = "2024"
|
||||
license = "MIT"
|
||||
repository = "https://github.com/redclawsystems/clawhdf5"
|
||||
|
||||
[workspace.dependencies]
|
||||
tempfile = "3"
|
||||
criterion = { version = "0.5", features = ["html_reports"] }
|
||||
half = "2.7"
|
||||
serde = { version = "1", features = ["derive"] }
|
||||
|
||||
@@ -4,8 +4,8 @@
|
||||
|
||||
[](LICENSE)
|
||||
[](https://www.rust-lang.org)
|
||||
[](#benchmarks)
|
||||
[](BENCHMARKS.md#longmemeval-results)
|
||||
[](#performance)
|
||||
[](BENCHMARKS.md#longmemeval-results)
|
||||
[](BENCHMARKS.md#memory-footprint)
|
||||
|
||||
ClawHDF5 is a pure-Rust HDF5 implementation combined with a research-grade agent memory engine. It gives AI agents persistent, searchable, cryptographically verifiable memory — all stored in a single portable file.
|
||||
@@ -64,7 +64,12 @@ Figures below are from an independent reproduction run on a second machine (AMD
|
||||
|-------|------|-----------------|--------|----------|
|
||||
| 1K | **54 µs** | — | — | — |
|
||||
| 10K | 753 µs | **27 µs** | — | — |
|
||||
| 100K | 11.4 ms | 1.32 ms | **1.19 ms** | **8–76× faster** |
|
||||
| 100K | 11.4 ms | 1.32 ms | **1.19 ms** | ~8–76× (see caveat) |
|
||||
|
||||
> Reproduced on the same second machine (Ryzen 7 7800X3D) with a corrected,
|
||||
> apples-to-apples SIMD/scalar/parallel comparison methodology — see
|
||||
> [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector
|
||||
> Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05).
|
||||
|
||||
### Agent Memory Operations
|
||||
|
||||
@@ -92,20 +97,52 @@ by default (AoS→SoA byte transpose, +157–204% throughput for float data):
|
||||
|
||||
Use `.with_zstd(3)` or `.with_deflate(6)` for write-heavy workloads — both now perform at ~720–750 MiB/s on large matrices. Use `.with_pcodec()` for write-once/read-many workloads where compression ratio matters more than encode speed. Disable auto-shuffle with `.without_shuffle()` for byte arrays that don't benefit from AoS→SoA transposition.
|
||||
|
||||
> ¹ MemX ([arxiv:2603.16171](https://arxiv.org/abs/2603.16171), March 2026): Rust + libSQL, claims <90ms at 100K records.
|
||||
> ¹ MemX ([arxiv:2603.16171](https://arxiv.org/abs/2603.16171), March 2026): Rust + libSQL, claims <90ms at 100K records. **Not like-for-like:** MemX's figure is *end-to-end* (embeddings + FTS5 + four-factor re-ranking); ours is a *single component* (raw vector search). The ratio overstates the real advantage by an unquantified margin — order-of-magnitude indication only. See [BENCHMARKS.md](BENCHMARKS.md#comparison-to-memx-arxiv260316171).
|
||||
|
||||
### LongMemEval Retrieval Recall
|
||||
|
||||
Evaluated against the LongMemEval dataset (500 questions, multi-session haystack).
|
||||
BM25-only baseline (no embedding model required at bench time):
|
||||
Evaluated against the full **`longmemeval_s`** haystack — all 500 questions, 47.7
|
||||
sessions and 493.5 turns each, with only 4.0% of haystack sessions being evidence
|
||||
sessions. See [BENCHMARKS.md § LongMemEval
|
||||
Results](BENCHMARKS.md#longmemeval-results) for the full scoring-target
|
||||
declaration:
|
||||
|
||||
| Metric | BM25-only | Full hybrid¹ |
|
||||
|--------|-----------|--------------|
|
||||
| Hit@5 (session) | ~46% | Higher |
|
||||
| MRR (session) | ~0.34 | Higher |
|
||||
| Abstention accuracy | ~72% | — |
|
||||
| Mode | Turn-Level Hit@5 | Session-Level Hit@5 |
|
||||
|------|------------------|---------------------|
|
||||
| BM25 only | 75.0% | 93.6% |
|
||||
| Vector only (MiniLM) | 71.8% | 94.2% |
|
||||
| Hybrid (0.4/0.6, tuned) | **81.4%** | **96.8%** |
|
||||
|
||||
> ¹ Enable embeddings via `hybrid_search(query_emb, text, 0.7, 0.3, k)` for substantially higher recall. The vector stage is served by the HNSW index by default (the `hnsw` feature is on by default); build with `--no-default-features --features float16` to fall back to an exact linear cosine scan.
|
||||
Hybrid is the strongest configuration, which is what running two retrieval stages
|
||||
is for. The weights matter more than the stages: a sweep of `vector_weight` from
|
||||
0.0 to 1.0 found the long-standing `0.7/0.3` default is **strictly dominated** by
|
||||
`0.4/0.6` — better on Hit@1, Hit@5, Hit@10 and MRR at both granularities. Use
|
||||
`0.4/0.6`, or `0.3/0.7` if rank-1 precision matters most. See
|
||||
[BENCHMARKS.md § Weight sweep](BENCHMARKS.md#longmemeval-results).
|
||||
|
||||
Vector embeddings require `--features embeddings`; without it the vector stage is
|
||||
inert and only the BM25 row is produced, which is what every previously published
|
||||
number here measured.
|
||||
|
||||
On the easier `longmemeval_oracle` variant (evidence sessions only) the same
|
||||
harness scores 84.4% turn-level Hit@5 / MRR 0.6597, reproduced identically on a
|
||||
second machine. The 9.4-point gap is the cost of the real haystack, and is why the
|
||||
full-haystack number is the one quoted here.
|
||||
|
||||
This is **retrieval recall** (did the gold memory appear in the top-k), not the
|
||||
official LongMemEval QA-accuracy metric — the two are not comparable, and
|
||||
retrieval recall reported as QA accuracy typically overstates by 20–30 points.
|
||||
|
||||
> **Previously reported here and now retracted:** session-level Hit@5 of 100.0% /
|
||||
> MRR 1.0000, and a claim of beating MemX's 51.6%. Those session-level figures were
|
||||
> degenerate on the oracle variant (any returned document is a hit by
|
||||
> construction); the 93.6% above is a different, real measurement on a corpus where
|
||||
> evidence sessions are 4.0% of the haystack. The MemX comparison stays withdrawn —
|
||||
> MemX measures fact-level granularity over 220,349 records, which running the full
|
||||
> haystack does not fix. Details in
|
||||
> [BENCHMARKS.md](BENCHMARKS.md#retracted-session-level-recall-and-the-memx-comparison).
|
||||
|
||||
> Enable embeddings via `hybrid_search(query_emb, text, 0.4, 0.6, k)` for substantially higher recall. The vector stage is served by the HNSW index by default (the `hnsw` feature is on by default); build with `--no-default-features --features float16` to fall back to an exact linear cosine scan.
|
||||
|
||||
### Memory Footprint
|
||||
|
||||
@@ -194,7 +231,7 @@ ClawhDF5's agent memory engine implements research from 15+ recent papers on age
|
||||
| **`ivf` / `pq`** | IVF-PQ approximate nearest neighbor for billion-scale search |
|
||||
| **`bm25`** | BM25 keyword index with TF-IDF scoring |
|
||||
| **`entity_extract`** | Rule-based entity extraction from text chunks into the knowledge graph |
|
||||
| **`wal`** | Write-ahead log for crash-safe persistence |
|
||||
| **`wal`** | Write-ahead log for crash-safe persistence; each entry is CRC32-checked on replay, so a corrupted entry stops replay there instead of loading bad data |
|
||||
| **`memory_strategy`** | Pluggable strategies: save-every, semantic-shift, user-correction detection |
|
||||
| **`decision_gate`** | Sub-microsecond trivial/substantive classification |
|
||||
| **`async_memory`** | Tokio-based async wrapper over the memory store (`async` feature) |
|
||||
@@ -338,22 +375,22 @@ let exported = backend.export_markdown("MEMORY.md")?;
|
||||
## Crate Map
|
||||
|
||||
```
|
||||
clawhdf5 workspace (17 crates, 84K lines of Rust)
|
||||
clawhdf5 workspace (16 crates, ~92K lines of Rust; plus libaec-sys, an
|
||||
internal FFI bindings crate for the optional szip feature)
|
||||
│
|
||||
├── Core HDF5
|
||||
│ ├── clawhdf5-types — Type system definitions
|
||||
│ ├── clawhdf5-format — Binary parser/writer (no_std)
|
||||
│ ├── clawhdf5-format — Binary parser/writer (no_std), shared type definitions
|
||||
│ ├── clawhdf5-io — I/O abstraction (buffered, mmap, async)
|
||||
│ ├── clawhdf5-filters — Compression (deflate, lz4, zstd, blosc)
|
||||
│ ├── clawhdf5-filters — Fast deflate path (zlib-ng); lz4/zstd/pcodec/szip filters live in clawhdf5-format
|
||||
│ ├── clawhdf5-derive — Proc macros
|
||||
│ ├── clawhdf5 — High-level API
|
||||
│ ├── clawhdf5-netcdf4 — NetCDF-4 support
|
||||
│ ├── clawhdf5-accel — SIMD (NEON, AVX2, AVX-512)
|
||||
│ └── clawhdf5-gpu — GPU compute (wgpu)
|
||||
│ └── clawhdf5-gpu — GPU compute (wgpu, hand-written WGSL compute shaders)
|
||||
│
|
||||
├── Agent Memory
|
||||
│ ├── clawhdf5-agent — Memory engine (20.7K lines, 32 modules)
|
||||
│ ├── clawhdf5-ann — HNSW approximate nearest neighbor (default backend)
|
||||
│ ├── clawhdf5-agent — Memory engine (20.9K lines, 32 modules; WAL is CRC32-checked per entry)
|
||||
│ ├── clawhdf5-ann — HNSW approximate nearest neighbor (default backend; optional `parallel` feature)
|
||||
│ ├── clawhdf5-migrate — SQLite → HDF5 migration
|
||||
│ ├── clawhdf5-android — Android JNI bridge
|
||||
│ └── clawhdf5-cli — CLI tool
|
||||
@@ -420,6 +457,24 @@ ClawhDF5's agent memory design draws from 15+ recent papers:
|
||||
| `system-zlib` / `zlib-rs` | no | Alternative zlib backends for deflate |
|
||||
| `blake3_hash` | no | BLAKE3 content hashing for provenance |
|
||||
|
||||
### `clawhdf5-ann`
|
||||
|
||||
| Flag | Default | Description |
|
||||
|------|---------|-------------|
|
||||
| `parallel` | no | Rayon-parallel neighbor-distance computation during HNSW graph pruning |
|
||||
|
||||
### `clawhdf5-io`
|
||||
|
||||
| Flag | Default | Description |
|
||||
|------|---------|-------------|
|
||||
| `mpi-io` | no | MPI-backed I/O via the `mpi` crate |
|
||||
|
||||
> **Parallel I/O (MPI) limitation:** `mpi-io`'s read path is a root-rank read
|
||||
> followed by a broadcast, and its write path gathers all ranks' shards to
|
||||
> rank 0 before writing — not true collective I/O
|
||||
> (`MPI_File_read_at_all`/`write_at_all`). It does not provide I/O bandwidth
|
||||
> that scales with rank count; true collective I/O is tracked as future work.
|
||||
|
||||
---
|
||||
|
||||
## Building
|
||||
@@ -435,7 +490,7 @@ cargo build -p clawhdf5-agent --features "agent,float16,parallel,fast-math"
|
||||
cargo build -p clawhdf5-agent --features "agent,float16,accelerate,parallel,gpu"
|
||||
|
||||
# Tests
|
||||
cargo test --workspace # all 417+ tests
|
||||
cargo test --workspace # all 1,650+ tests
|
||||
cargo test -p clawhdf5-agent # agent memory tests
|
||||
|
||||
# Benchmarks
|
||||
@@ -505,7 +560,7 @@ See [ROADMAP.md](ROADMAP.md) for the full implementation tracker.
|
||||
- ✅ OpenClaw integration layer
|
||||
- ✅ Comprehensive Criterion benchmarks
|
||||
|
||||
**Phase 2** — OpenClaw TypeScript bridge, academic benchmarks (MemoryArena, LongMemEval), cross-platform validation.
|
||||
**Phase 2** — MemoryArena and LongMemEval academic benchmarks are done (see [BENCHMARKS.md](BENCHMARKS.md), reproduced on a second machine); remaining: publish the OpenClaw TypeScript bridge to npm, crates.io/PyPI publishing.
|
||||
|
||||
---
|
||||
|
||||
@@ -523,5 +578,5 @@ MIT
|
||||
|
||||
<p align="center">
|
||||
<em>Built by <a href="https://github.com/redclawsystems">RedClaw Systems</a></em><br>
|
||||
<em>72,087 lines of Rust. Zero C dependencies. One file to remember everything.</em>
|
||||
<em>~92,000 lines of Rust. Zero C dependencies. One file to remember everything.</em>
|
||||
</p>
|
||||
|
||||
+23
-6
@@ -145,19 +145,36 @@
|
||||
**Phase 3:** ~~Track 6 (multi-modal) + Track 7 (OpenClaw integration)~~ 🟢 Complete
|
||||
**Phase 4:** ~~Track 8 (benchmarking + validation)~~ 🟢 Complete
|
||||
|
||||
All 8 tracks delivered. 1,546 tests passing, zero clippy warnings.
|
||||
All 8 tracks delivered. 1,650+ tests passing, zero clippy warnings.
|
||||
|
||||
---
|
||||
|
||||
## What's Next
|
||||
|
||||
Verified against current repo state on 2026-08-03 (see also `docs/superpowers/plans/` for the filter-codec/format-write/MPI-IO work, now shipped):
|
||||
Verified against current repo state on 2026-08-05 (see also `docs/superpowers/plans/` for the filter-codec/format-write/MPI-IO work, now shipped):
|
||||
|
||||
- [ ] CI/CD pipeline — still no GitHub/Gitea Actions workflow in the repo; automated testing is manual only
|
||||
- [ ] Academic benchmark cross-validation — reproduce MemX/LongMemEval under identical conditions
|
||||
- [ ] TypeScript bridge — `clawhdf5-napi` has no `package.json`; it's still Rust-only scaffolding, not a publishable npm package
|
||||
- [ ] TypeScript bridge not wired into CI — `packages/clawhdf5-node/` already has a complete, working napi-rs package (package.json, tsconfig, hand-written TS wrapper matching all 21 `#[napi]` items, Jest test suite, README); it isn't published to npm and has no committed lockfile
|
||||
- [ ] Publish crates to crates.io — no `publish` config anywhere in the workspace yet
|
||||
- [ ] Python wheel distribution via maturin — `crates/clawhdf5-py/pyproject.toml` exists (maturin-buildable locally) but wheels aren't published anywhere
|
||||
- [ ] `chunked_read.rs`/`data_read.rs` full bounds-check audit + scheduled fuzz campaigns (the new `fuzz_dataset_read` target covers the two files' main entry points; a full manual audit of every indexing site is still open) — see Tier 4 below
|
||||
- [ ] WAL per-entry checksum landed as CRC32 (see below); a stronger per-entry format (explicit length prefix, avoiding the read-then-verify restructuring) could still be revisited if profiling shows it matters
|
||||
- [ ] HNSW build parallelism is still narrow (only `prune_connections`); the correctness-sensitive outer insert loop needs its own dedicated design pass before parallelizing
|
||||
|
||||
### Recently closed out (2026-08-05, Tier 3–4 hardening pass)
|
||||
|
||||
- [x] Academic benchmark cross-validation — LongMemEval reproduced against MemX on tank (Ryzen 7 7800X3D): turn-level Hit@5 84.4% vs MemX's 51.6%; recall numbers are deterministic and reproduce exactly across machines. SIMD/Parallelism and Vector Search sections also re-run and dated. See [BENCHMARKS.md § Independent Validation: tank — LongMemEval & Vector Search](BENCHMARKS.md#independent-validation-tank--longmemeval--vector-search-ryzen-7-7800x3d-2026-08-05)
|
||||
- [x] Android JNI (`clawhdf5-android`): validate `embedding_len`/`query_embedding_len` against the handle's configured `embedding_dim` before constructing a slice from a raw pointer
|
||||
- [x] `clawhdf5-py`: bumped pyo3/numpy 0.28 → 0.29, clearing two RUSTSEC advisories
|
||||
- [x] WAL (`clawhdf5-agent`): length-prefix caps (`MAX_WAL_FIELD_LEN`) to reject a corrupted length claim before allocating, then a full per-entry CRC32 trailer (`WAL_VERSION` 2) so a bit-flip stops replay cleanly instead of loading corrupted data; old-format WAL files still read correctly and are migrated on next open
|
||||
- [x] `chunked_read.rs`/`data_read.rs`/`local_heap.rs` bounds-check audit: added `ensure_len` overflow guards, a recursion-depth guard against cyclic B-trees, and a fix for an unguarded compound-datatype byte-offset overrun. Added a new `fuzz_dataset_read` cargo-fuzz target exercising the contiguous/chunked/compact read paths — it found and we fixed 3 real crash bugs (integer-overflow panics) within the first few runs
|
||||
- [x] `clawhdf5-ann`: optional `parallel` feature (rayon) for HNSW's `prune_connections` neighbor-distance computation
|
||||
- [x] `[workspace.dependencies]` added for `tempfile`/`criterion`/`half`/`serde`, fixing a real version skew on `half` (2 vs 2.7)
|
||||
|
||||
### Recently closed out (2026-08-05 hardening pass)
|
||||
|
||||
- [x] CI/CD pipeline — `.gitea/workflows/ci.yml` now runs `scripts/ci-test.sh` (fmt, clippy, tests, no_std check) on push/PR to `main`
|
||||
- [x] Fixed no_std build breakage in `clawhdf5-format` (missing alloc imports, `AtomicU64` unsupported on thumbv7em, `f64::powi` requiring std/libm)
|
||||
- [x] Fixed version skew: `clawhdf5-py` (pyproject.toml) and `packages/clawhdf5-node` (package.json) were both behind the actual crate version
|
||||
|
||||
### Recently closed out (2026-08-03 cleanup pass)
|
||||
|
||||
@@ -167,4 +184,4 @@ Verified against current repo state on 2026-08-03 (see also `docs/superpowers/pl
|
||||
|
||||
---
|
||||
|
||||
_Last updated: 2026-08-03_
|
||||
_Last updated: 2026-08-05_
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
#!/usr/bin/env python3
|
||||
"""h5py counterpart to worldmodel_sampling.rs — same file, same shuffled
|
||||
per-frame access, same minimal touch (sum the frame bytes). Reports
|
||||
samples/sec so the two sit side by side on one machine."""
|
||||
import sys, time, numpy as np, h5py
|
||||
|
||||
path = sys.argv[1]
|
||||
passes = int(sys.argv[2]) if len(sys.argv) > 2 else 5
|
||||
|
||||
def shuffled(n):
|
||||
v = list(range(n))
|
||||
state = 0x9E3779B97F4A7C15
|
||||
for i in range(n - 1, 0, -1):
|
||||
state = (state * 6364136223846793005 + 1442695040888963407) & 0xFFFFFFFFFFFFFFFF
|
||||
j = (state >> 33) % (i + 1)
|
||||
v[i], v[j] = v[j], v[i]
|
||||
return v
|
||||
|
||||
# swmr + a 256 MB chunk cache: exactly stable-worldmodel's HDF5Dataset._open_h5.
|
||||
f = h5py.File(path, "r", swmr=True, rdcc_nbytes=256 * 1024 * 1024)
|
||||
d = f["observation"]
|
||||
n = d.shape[0]
|
||||
order = shuffled(n)
|
||||
|
||||
# warm
|
||||
sink = 0
|
||||
for i in order:
|
||||
sink += int(d[i].sum())
|
||||
|
||||
t0 = time.perf_counter()
|
||||
sink = 0
|
||||
for _ in range(passes):
|
||||
for i in order:
|
||||
sink += int(d[i].sum())
|
||||
elapsed = time.perf_counter() - t0
|
||||
total = n * passes
|
||||
print(f"h5py: {n} frames x {passes} passes = {total} reads in {elapsed:.3f}s")
|
||||
print(f"h5py: {total/elapsed:.0f} samples/sec")
|
||||
@@ -0,0 +1,27 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Generate a world-model-shaped dataset: N frames of HxWxC uint8 observations,
|
||||
contiguous (N,H,W,C), matching stable-worldmodel's per-frame sample-loading
|
||||
access pattern. Also emits ep_len/ep_offset like their format."""
|
||||
import sys, time, numpy as np, h5py
|
||||
|
||||
path = sys.argv[1]
|
||||
N = int(sys.argv[2]) if len(sys.argv) > 2 else 20000
|
||||
H = W = 64
|
||||
C = 3
|
||||
rng = np.random.default_rng(0)
|
||||
t0 = time.perf_counter()
|
||||
with h5py.File(path, "w", libver="latest") as f:
|
||||
# Contiguous (N,H,W,C) uint8 — the fair, both-APIs-support-it layout.
|
||||
obs = f.create_dataset("observation", shape=(N, H, W, C), dtype=np.uint8)
|
||||
# Write in blocks to bound memory.
|
||||
B = 2000
|
||||
for i in range(0, N, B):
|
||||
n = min(B, N - i)
|
||||
obs[i:i+n] = rng.integers(0, 256, size=(n, H, W, C), dtype=np.uint8)
|
||||
# Episode metadata like their format: 100-step episodes.
|
||||
ep = 100
|
||||
n_ep = N // ep
|
||||
f.create_dataset("ep_len", data=np.full(n_ep, ep, dtype=np.int32))
|
||||
f.create_dataset("ep_offset", data=(np.arange(n_ep) * ep).astype(np.int64))
|
||||
print(f"wrote {N} frames {H}x{W}x{C} to {path} in {time.perf_counter()-t0:.1f}s "
|
||||
f"({N*H*W*C/1e6:.0f} MB)")
|
||||
@@ -15,7 +15,7 @@ float16 = ["dep:half"]
|
||||
avx512 = []
|
||||
|
||||
[dependencies]
|
||||
half = { version = "2", optional = true }
|
||||
half = { workspace = true, optional = true }
|
||||
|
||||
[package.metadata.docs.rs]
|
||||
features = []
|
||||
|
||||
@@ -16,9 +16,9 @@ clawhdf5-io = { path = "../clawhdf5-io", version = "2.1.0", features = ["mmap"]
|
||||
clawhdf5-accel = { path = "../clawhdf5-accel", version = "2.1.0" }
|
||||
clawhdf5-ann = { path = "../clawhdf5-ann", version = "2.1.0", optional = true }
|
||||
clawhdf5-gpu = { path = "../clawhdf5-gpu", version = "2.1.0", optional = true, default-features = false }
|
||||
serde = { version = "1", features = ["derive"] }
|
||||
serde = { workspace = true }
|
||||
byteorder = "1"
|
||||
half = { version = "2", optional = true }
|
||||
half = { workspace = true, optional = true }
|
||||
rayon = { version = "1", optional = true }
|
||||
matrixmultiply = { version = "0.3", optional = true }
|
||||
cblas-sys = { version = "0.1", optional = true }
|
||||
@@ -31,8 +31,8 @@ accelerate-src = { version = "0.3", optional = true }
|
||||
openblas-src = { version = "0.10", optional = true, features = ["cblas"] }
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile = "3"
|
||||
criterion = "0.5"
|
||||
tempfile = { workspace = true }
|
||||
criterion = { workspace = true }
|
||||
rayon = "1"
|
||||
tokio = { version = "1", features = ["rt-multi-thread", "sync", "macros"] }
|
||||
|
||||
|
||||
@@ -329,6 +329,47 @@ impl KnowledgeCache {
|
||||
(id, true)
|
||||
}
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Adjacency index (built fresh per traversal call — see doc comment)
|
||||
// -----------------------------------------------------------------------
|
||||
|
||||
/// Build an O(V+R) adjacency index for one traversal call: an entity-id →
|
||||
/// vec-index map for O(1) entity lookups, and an entity-id →
|
||||
/// `(neighbour_id, relation_weight)` map (covering both outgoing and
|
||||
/// incoming edges) for O(1) neighbour expansion. The weight is carried
|
||||
/// alongside each neighbour so callers like `spreading_activation` that
|
||||
/// need per-edge weight don't have to re-scan `relations`.
|
||||
///
|
||||
/// This is rebuilt at the start of every `bfs_neighbors`/
|
||||
/// `spreading_activation` call rather than cached on the struct: `entities`
|
||||
/// and `relations` are public fields, and `schema.rs`'s deserialization
|
||||
/// path pushes into them directly (bypassing `add_entity`/`add_relation`),
|
||||
/// so a struct-cached index could go stale. Building it once per call
|
||||
/// still turns an O(V·R) (or O(steps·V·R)) traversal into O(V+R) (or
|
||||
/// O(steps·(V+E))), since the old code repeated the O(R) relation scan
|
||||
/// once per visited node instead of once per call.
|
||||
fn build_adjacency(&self) -> (HashMap<u64, usize>, HashMap<u64, Vec<(u64, f32)>>) {
|
||||
let mut entity_index: HashMap<u64, usize> = HashMap::with_capacity(self.entities.len());
|
||||
for (i, e) in self.entities.iter().enumerate() {
|
||||
entity_index.insert(e.id, i);
|
||||
}
|
||||
|
||||
// Note: a self-loop relation (src == tgt) contributes a single
|
||||
// neighbour entry, not two, matching the if/else-if (not two
|
||||
// independent ifs) structure this replaces — otherwise a self-loop
|
||||
// would be double-counted by `spreading_activation`.
|
||||
let mut adjacency: HashMap<u64, Vec<(u64, f32)>> =
|
||||
HashMap::with_capacity(self.relations.len());
|
||||
for r in &self.relations {
|
||||
adjacency.entry(r.src).or_default().push((r.tgt, r.weight));
|
||||
if r.tgt != r.src {
|
||||
adjacency.entry(r.tgt).or_default().push((r.src, r.weight));
|
||||
}
|
||||
}
|
||||
|
||||
(entity_index, adjacency)
|
||||
}
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Graph traversal: BFS neighbors
|
||||
// -----------------------------------------------------------------------
|
||||
@@ -337,6 +378,8 @@ impl KnowledgeCache {
|
||||
/// together with their discovered depth. The seed entity itself is NOT
|
||||
/// included. Traversal follows both outgoing and incoming relation edges.
|
||||
pub fn bfs_neighbors(&self, entity_id: u64, max_depth: usize) -> Vec<(Entity, usize)> {
|
||||
let (entity_index, adjacency) = self.build_adjacency();
|
||||
|
||||
let mut visited: HashSet<u64> = HashSet::new();
|
||||
let mut queue: VecDeque<(u64, usize)> = VecDeque::new();
|
||||
let mut results: Vec<(Entity, usize)> = Vec::new();
|
||||
@@ -349,25 +392,15 @@ impl KnowledgeCache {
|
||||
continue;
|
||||
}
|
||||
|
||||
// Collect neighbour IDs from outgoing and incoming edges.
|
||||
let neighbours: Vec<u64> = self
|
||||
.relations
|
||||
.iter()
|
||||
.filter_map(|r| {
|
||||
if r.src == current_id {
|
||||
Some(r.tgt)
|
||||
} else if r.tgt == current_id {
|
||||
Some(r.src)
|
||||
} else {
|
||||
None
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
let Some(neighbours) = adjacency.get(¤t_id) else {
|
||||
continue;
|
||||
};
|
||||
|
||||
for neighbour_id in neighbours {
|
||||
for &(neighbour_id, _weight) in neighbours {
|
||||
if visited.insert(neighbour_id)
|
||||
&& let Some(entity) = self.get_entity(neighbour_id)
|
||||
&& let Some(&idx) = entity_index.get(&neighbour_id)
|
||||
{
|
||||
let entity = &self.entities[idx];
|
||||
results.push((entity.clone(), depth + 1));
|
||||
queue.push_back((neighbour_id, depth + 1));
|
||||
}
|
||||
@@ -439,6 +472,8 @@ impl KnowledgeCache {
|
||||
min_activation: f32,
|
||||
max_steps: usize,
|
||||
) -> Vec<(u64, f32)> {
|
||||
let (_entity_index, adjacency) = self.build_adjacency();
|
||||
|
||||
let mut activation: HashMap<u64, f32> = HashMap::new();
|
||||
|
||||
// Initialise seeds with activation 1.0.
|
||||
@@ -462,16 +497,11 @@ impl KnowledgeCache {
|
||||
|
||||
for (source_id, source_score) in current {
|
||||
// Spread to all neighbours via outgoing and incoming edges.
|
||||
for rel in &self.relations {
|
||||
let neighbour_id = if rel.src == source_id {
|
||||
rel.tgt
|
||||
} else if rel.tgt == source_id {
|
||||
rel.src
|
||||
} else {
|
||||
let Some(neighbours) = adjacency.get(&source_id) else {
|
||||
continue;
|
||||
};
|
||||
|
||||
let delta = source_score * rel.weight * decay_factor;
|
||||
for &(neighbour_id, weight) in neighbours {
|
||||
let delta = source_score * weight * decay_factor;
|
||||
if delta >= min_activation {
|
||||
*activation.entry(neighbour_id).or_insert(0.0) += delta;
|
||||
any_spread = true;
|
||||
|
||||
@@ -1,25 +1,31 @@
|
||||
//! Memory provenance tracking and integrity verification.
|
||||
//!
|
||||
//! Records the origin, authorship, and integrity of every memory chunk
|
||||
//! so the system can detect tampering and trace data lineage.
|
||||
//! Records the origin, authorship, and a content hash of every memory chunk
|
||||
//! so the system can detect content corruption and trace data lineage. The
|
||||
//! hash is a SHA-256 digest (see [`hash_content`]), computed via
|
||||
//! [`clawhdf5_format::provenance::sha256_hex`]. It is still **unkeyed** — an
|
||||
//! actor able to overwrite the stored chunk can also recompute and overwrite
|
||||
//! the stored hash alongside it, so this is not an authenticity guarantee
|
||||
//! against that threat. What SHA-256 does provide over a fast non-cryptographic
|
||||
//! hash (the previous FNV-1a implementation) is collision resistance: an
|
||||
//! adversary cannot cheaply craft *different* poisoned content that matches
|
||||
//! an already-recorded legitimate hash.
|
||||
|
||||
use std::collections::HashMap;
|
||||
|
||||
pub use crate::consolidation::MemorySource;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Hash helper (std-only FNV-1a 64-bit)
|
||||
// Hash helper
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn fnv1a_64(text: &str) -> u64 {
|
||||
const OFFSET: u64 = 14_695_981_039_346_656_037;
|
||||
const PRIME: u64 = 1_099_511_628_211;
|
||||
let mut hash = OFFSET;
|
||||
for byte in text.bytes() {
|
||||
hash ^= byte as u64;
|
||||
hash = hash.wrapping_mul(PRIME);
|
||||
}
|
||||
hash
|
||||
/// SHA-256 hex digest of `text`, used to detect content corruption/tampering.
|
||||
///
|
||||
/// Unkeyed: an actor able to modify the stored chunk can also recompute and
|
||||
/// overwrite the stored hash, so a match is not proof of authenticity — only
|
||||
/// that the stored chunk and stored hash are mutually consistent.
|
||||
fn hash_content(text: &str) -> String {
|
||||
clawhdf5_format::provenance::sha256_hex(text.as_bytes())
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -51,8 +57,8 @@ pub struct MemoryProvenance {
|
||||
pub created_by: String,
|
||||
/// Unix timestamp (seconds) of creation.
|
||||
pub created_at: f64,
|
||||
/// FNV-1a 64-bit hash of the chunk text for integrity checking.
|
||||
pub content_hash: u64,
|
||||
/// SHA-256 hex digest of the chunk text for integrity checking.
|
||||
pub content_hash: String,
|
||||
pub session_id: String,
|
||||
pub verified: bool,
|
||||
}
|
||||
@@ -72,7 +78,7 @@ impl MemoryProvenance {
|
||||
source,
|
||||
created_by: created_by.into(),
|
||||
created_at,
|
||||
content_hash: fnv1a_64(chunk),
|
||||
content_hash: hash_content(chunk),
|
||||
session_id: session_id.into(),
|
||||
verified: false,
|
||||
}
|
||||
@@ -114,9 +120,17 @@ impl ProvenanceStore {
|
||||
|
||||
/// Re-hash `current_chunk` and compare against the stored hash.
|
||||
/// Returns `true` if the content matches (integrity intact).
|
||||
///
|
||||
/// The hash is unkeyed, so an actor able to modify the stored chunk can
|
||||
/// also recompute and overwrite the stored hash. Do not treat a `true`
|
||||
/// result as proof of authenticity against that threat — but unlike a
|
||||
/// non-cryptographic hash, a `false` result reliably indicates that the
|
||||
/// content does not match what was recorded, since SHA-256 makes it
|
||||
/// computationally infeasible to craft different content that collides
|
||||
/// with a specific existing digest.
|
||||
pub fn verify_integrity(&self, record_id: u64, current_chunk: &str) -> bool {
|
||||
match self.records.get(&record_id) {
|
||||
Some(p) => p.content_hash == fnv1a_64(current_chunk),
|
||||
Some(p) => p.content_hash == hash_content(current_chunk),
|
||||
None => false,
|
||||
}
|
||||
}
|
||||
@@ -230,22 +244,32 @@ mod tests {
|
||||
1_700_000_000.0
|
||||
}
|
||||
|
||||
// --- fnv1a_64 ---
|
||||
// --- hash_content ---
|
||||
|
||||
#[test]
|
||||
fn hash_deterministic() {
|
||||
assert_eq!(fnv1a_64("hello"), fnv1a_64("hello"));
|
||||
assert_eq!(hash_content("hello"), hash_content("hello"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn hash_different_inputs() {
|
||||
assert_ne!(fnv1a_64("hello"), fnv1a_64("world"));
|
||||
assert_ne!(hash_content("hello"), hash_content("world"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn hash_empty() {
|
||||
// Should not panic
|
||||
let _ = fnv1a_64("");
|
||||
// Should not panic, and should match the well-known SHA-256 of the empty string.
|
||||
assert_eq!(
|
||||
hash_content(""),
|
||||
"e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn hash_is_sha256_hex() {
|
||||
let h = hash_content("clawhdf5");
|
||||
assert_eq!(h.len(), 64);
|
||||
assert!(h.chars().all(|c| c.is_ascii_hexdigit()));
|
||||
}
|
||||
|
||||
// --- MemorySource Display ---
|
||||
@@ -264,7 +288,7 @@ mod tests {
|
||||
#[test]
|
||||
fn provenance_new_hashes_chunk() {
|
||||
let p = MemoryProvenance::new(1, MemorySource::User, "agent-1", ts(), "hello", "s1");
|
||||
assert_eq!(p.content_hash, fnv1a_64("hello"));
|
||||
assert_eq!(p.content_hash, hash_content("hello"));
|
||||
assert!(!p.verified);
|
||||
}
|
||||
|
||||
|
||||
@@ -7,10 +7,29 @@ use std::fs::{File, OpenOptions};
|
||||
use std::io::{Read, Seek, SeekFrom, Write};
|
||||
use std::path::{Path, PathBuf};
|
||||
|
||||
use clawhdf5_format::checksum::crc32;
|
||||
|
||||
use crate::MemoryError;
|
||||
|
||||
const WAL_MAGIC: [u8; 4] = [0x45, 0x48, 0x57, 0x4C]; // "EHWL"
|
||||
const WAL_VERSION: u8 = 1;
|
||||
|
||||
/// Current WAL format version: every entry ends with a 4-byte CRC32 trailer
|
||||
/// (see [`TeeReader`]) so a bit-flip is detected and replay stops there
|
||||
/// instead of silently accepting corrupted data.
|
||||
const WAL_VERSION: u8 = 2;
|
||||
|
||||
/// The only other WAL version this crate still knows how to *read*: no
|
||||
/// per-entry CRC trailer. Written by versions of this crate before the CRC32
|
||||
/// hardening. `WalFile::open` migrates a legacy file to [`WAL_VERSION`] by
|
||||
/// recreating it fresh — safe because every real call site reads existing
|
||||
/// entries via [`WalFile::read_entries`] before calling `open` (see
|
||||
/// `HDF5Memory::open`), so no data is lost.
|
||||
const WAL_VERSION_LEGACY_NO_CRC: u8 = 1;
|
||||
|
||||
/// Upper bound on a single length-prefixed WAL field (string bytes, or
|
||||
/// embedding element count), to reject a corrupted/truncated WAL length
|
||||
/// claim before allocating a large buffer for it.
|
||||
const MAX_WAL_FIELD_LEN: usize = 64 * 1024 * 1024;
|
||||
|
||||
#[repr(u8)]
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
@@ -62,6 +81,11 @@ pub struct WalFile {
|
||||
|
||||
impl WalFile {
|
||||
/// Open or create a WAL file. If it exists, read the header and entry count.
|
||||
///
|
||||
/// A legacy (pre-CRC) WAL file is migrated to the current format by
|
||||
/// recreating it fresh — see [`WAL_VERSION_LEGACY_NO_CRC`]. Callers that
|
||||
/// need the legacy file's entries must call [`WalFile::read_entries`]
|
||||
/// first, before calling `open`.
|
||||
pub fn open(path: &Path) -> Result<Self, MemoryError> {
|
||||
if path.exists() {
|
||||
// Read existing header
|
||||
@@ -77,12 +101,8 @@ impl WalFile {
|
||||
}
|
||||
let mut ver = [0u8; 1];
|
||||
f.read_exact(&mut ver)?;
|
||||
if ver[0] != WAL_VERSION {
|
||||
return Err(MemoryError::Schema(format!(
|
||||
"unsupported WAL version {}",
|
||||
ver[0]
|
||||
)));
|
||||
}
|
||||
match ver[0] {
|
||||
WAL_VERSION => {
|
||||
let mut count_buf = [0u8; 4];
|
||||
f.read_exact(&mut count_buf)?;
|
||||
let entry_count = u32::from_le_bytes(count_buf);
|
||||
@@ -94,13 +114,21 @@ impl WalFile {
|
||||
entry_count,
|
||||
pending_header_sync: 0,
|
||||
})
|
||||
}
|
||||
WAL_VERSION_LEGACY_NO_CRC => {
|
||||
drop(f);
|
||||
let f = create_fresh_wal_file(path)?;
|
||||
Ok(Self {
|
||||
path: path.to_path_buf(),
|
||||
file: Some(f),
|
||||
entry_count: 0,
|
||||
pending_header_sync: 0,
|
||||
})
|
||||
}
|
||||
v => Err(MemoryError::Schema(format!("unsupported WAL version {v}"))),
|
||||
}
|
||||
} else {
|
||||
// Create new WAL
|
||||
let mut f = File::create(path)?;
|
||||
f.write_all(&WAL_MAGIC)?;
|
||||
f.write_all(&[WAL_VERSION])?;
|
||||
f.write_all(&0u32.to_le_bytes())?;
|
||||
f.flush()?;
|
||||
let f = create_fresh_wal_file(path)?;
|
||||
Ok(Self {
|
||||
path: path.to_path_buf(),
|
||||
file: Some(f),
|
||||
@@ -140,6 +168,9 @@ impl WalFile {
|
||||
serialize_str(&mut buf, &entry.session_id);
|
||||
serialize_str(&mut buf, &entry.tags);
|
||||
|
||||
let crc = crc32(&buf);
|
||||
buf.extend_from_slice(&crc.to_le_bytes());
|
||||
|
||||
let f = self
|
||||
.file
|
||||
.as_mut()
|
||||
@@ -156,10 +187,12 @@ impl WalFile {
|
||||
|
||||
/// Append a tombstone entry (deletion).
|
||||
pub fn append_tombstone(&mut self, index: usize, timestamp: f64) -> Result<(), MemoryError> {
|
||||
let mut buf = [0u8; 1 + 8 + 4]; // type + timestamp + index
|
||||
let mut buf = [0u8; 1 + 8 + 4 + 4]; // type + timestamp + index + crc32
|
||||
buf[0] = WalEntryType::Tombstone as u8;
|
||||
buf[1..9].copy_from_slice(×tamp.to_le_bytes());
|
||||
buf[9..13].copy_from_slice(&(index as u32).to_le_bytes());
|
||||
let crc = crc32(&buf[..13]);
|
||||
buf[13..17].copy_from_slice(&crc.to_le_bytes());
|
||||
|
||||
let f = self
|
||||
.file
|
||||
@@ -180,7 +213,9 @@ impl WalFile {
|
||||
/// Reads until EOF — the header `entry_count` is used only for pre-allocation
|
||||
/// (and may be stale if written with deferred group-commit updates). This
|
||||
/// tolerates both truncated files (crash mid-write) and stale header counts
|
||||
/// (crash before the next group-commit header sync).
|
||||
/// (crash before the next group-commit header sync). On a `WAL_VERSION`
|
||||
/// file, a CRC32 mismatch on an entry is treated the same way — replay
|
||||
/// stops there rather than accepting corrupted data.
|
||||
pub fn read_entries(path: &Path) -> Result<Vec<WalEntry>, MemoryError> {
|
||||
if !path.exists() {
|
||||
return Ok(Vec::new());
|
||||
@@ -192,81 +227,45 @@ impl WalFile {
|
||||
if header[0..4] != WAL_MAGIC {
|
||||
return Err(MemoryError::Schema("invalid WAL magic bytes".into()));
|
||||
}
|
||||
if header[4] != WAL_VERSION {
|
||||
return Err(MemoryError::Schema(format!(
|
||||
"unsupported WAL version {}",
|
||||
header[4]
|
||||
)));
|
||||
}
|
||||
// entry_count is a pre-allocation hint only — we read until EOF.
|
||||
let entry_count_hint = u32::from_le_bytes([header[5], header[6], header[7], header[8]]);
|
||||
let mut entries = Vec::with_capacity(entry_count_hint as usize);
|
||||
|
||||
loop {
|
||||
// Read entry type — EOF here is normal end-of-log, not an error
|
||||
let mut type_buf = [0u8; 1];
|
||||
if f.read_exact(&mut type_buf).is_err() {
|
||||
match header[4] {
|
||||
WAL_VERSION => loop {
|
||||
let raw_and_result = {
|
||||
let mut tee = TeeReader::new(&mut f);
|
||||
let result = read_one_entry(&mut tee);
|
||||
(tee.into_buf(), result)
|
||||
};
|
||||
let (raw, result) = raw_and_result;
|
||||
let entry_opt = match result {
|
||||
Err(()) => break,
|
||||
Ok(v) => v,
|
||||
};
|
||||
let mut crc_buf = [0u8; 4];
|
||||
if f.read_exact(&mut crc_buf).is_err() {
|
||||
break;
|
||||
}
|
||||
let entry_type = match WalEntryType::from_u8(type_buf[0]) {
|
||||
Some(et) => et,
|
||||
None => break,
|
||||
};
|
||||
|
||||
let mut ts_buf = [0u8; 8];
|
||||
if f.read_exact(&mut ts_buf).is_err() {
|
||||
let stored_crc = u32::from_le_bytes(crc_buf);
|
||||
if crc32(&raw) != stored_crc {
|
||||
// Corruption detected — stop replay here, same as a clean
|
||||
// truncation/EOF, rather than accepting the bad entry.
|
||||
break;
|
||||
}
|
||||
let timestamp = f64::from_le_bytes(ts_buf);
|
||||
|
||||
match entry_type {
|
||||
WalEntryType::Save => {
|
||||
let Ok(chunk) = read_len_prefixed_str(&mut f) else {
|
||||
break;
|
||||
};
|
||||
let Ok(embedding) = read_embedding(&mut f) else {
|
||||
break;
|
||||
};
|
||||
let Ok(source_channel) = read_len_prefixed_str(&mut f) else {
|
||||
break;
|
||||
};
|
||||
let Ok(session_id) = read_len_prefixed_str(&mut f) else {
|
||||
break;
|
||||
};
|
||||
let Ok(tags) = read_len_prefixed_str(&mut f) else {
|
||||
break;
|
||||
};
|
||||
entries.push(WalEntry {
|
||||
entry_type,
|
||||
timestamp,
|
||||
chunk,
|
||||
embedding,
|
||||
source_channel,
|
||||
session_id,
|
||||
tags,
|
||||
tombstone_index: None,
|
||||
});
|
||||
if let Some(entry) = entry_opt {
|
||||
entries.push(entry);
|
||||
}
|
||||
WalEntryType::Tombstone => {
|
||||
let mut idx_buf = [0u8; 4];
|
||||
if f.read_exact(&mut idx_buf).is_err() {
|
||||
break;
|
||||
}
|
||||
let idx = u32::from_le_bytes(idx_buf) as usize;
|
||||
entries.push(WalEntry {
|
||||
entry_type,
|
||||
timestamp,
|
||||
chunk: String::new(),
|
||||
embedding: Vec::new(),
|
||||
source_channel: String::new(),
|
||||
session_id: String::new(),
|
||||
tags: String::new(),
|
||||
tombstone_index: Some(idx),
|
||||
});
|
||||
}
|
||||
WalEntryType::ActivationUpdate => {
|
||||
// Reserved for future use
|
||||
},
|
||||
WAL_VERSION_LEGACY_NO_CRC => loop {
|
||||
match read_one_entry(&mut f) {
|
||||
Err(()) => break,
|
||||
Ok(Some(entry)) => entries.push(entry),
|
||||
Ok(None) => {}
|
||||
}
|
||||
},
|
||||
v => {
|
||||
return Err(MemoryError::Schema(format!("unsupported WAL version {v}")));
|
||||
}
|
||||
}
|
||||
Ok(entries)
|
||||
@@ -276,11 +275,7 @@ impl WalFile {
|
||||
pub fn truncate(&mut self) -> Result<(), MemoryError> {
|
||||
// Close existing handle and recreate
|
||||
self.file = None;
|
||||
let mut f = File::create(&self.path)?;
|
||||
f.write_all(&WAL_MAGIC)?;
|
||||
f.write_all(&[WAL_VERSION])?;
|
||||
f.write_all(&0u32.to_le_bytes())?;
|
||||
f.flush()?;
|
||||
let f = create_fresh_wal_file(&self.path)?;
|
||||
self.file = Some(f);
|
||||
self.entry_count = 0;
|
||||
self.pending_header_sync = 0;
|
||||
@@ -345,19 +340,30 @@ fn serialize_str(buf: &mut Vec<u8>, s: &str) {
|
||||
buf.extend_from_slice(bytes);
|
||||
}
|
||||
|
||||
fn read_len_prefixed_str(f: &mut File) -> Result<String, MemoryError> {
|
||||
fn read_len_prefixed_str<R: Read>(f: &mut R) -> Result<String, MemoryError> {
|
||||
let mut len_buf = [0u8; 4];
|
||||
f.read_exact(&mut len_buf)?;
|
||||
let len = u32::from_le_bytes(len_buf) as usize;
|
||||
if len > MAX_WAL_FIELD_LEN {
|
||||
return Err(MemoryError::Schema(format!(
|
||||
"WAL string field length {len} exceeds max {MAX_WAL_FIELD_LEN}"
|
||||
)));
|
||||
}
|
||||
let mut buf = vec![0u8; len];
|
||||
f.read_exact(&mut buf)?;
|
||||
String::from_utf8(buf).map_err(|e| MemoryError::Schema(format!("invalid UTF-8 in WAL: {e}")))
|
||||
}
|
||||
|
||||
fn read_embedding(f: &mut File) -> Result<Vec<f32>, MemoryError> {
|
||||
fn read_embedding<R: Read>(f: &mut R) -> Result<Vec<f32>, MemoryError> {
|
||||
let mut len_buf = [0u8; 4];
|
||||
f.read_exact(&mut len_buf)?;
|
||||
let count = u32::from_le_bytes(len_buf) as usize;
|
||||
if count > MAX_WAL_FIELD_LEN / 4 {
|
||||
return Err(MemoryError::Schema(format!(
|
||||
"WAL embedding element count {count} exceeds max {}",
|
||||
MAX_WAL_FIELD_LEN / 4
|
||||
)));
|
||||
}
|
||||
let mut vals = Vec::with_capacity(count);
|
||||
for _ in 0..count {
|
||||
let mut val_buf = [0u8; 4];
|
||||
@@ -367,6 +373,99 @@ fn read_embedding(f: &mut File) -> Result<Vec<f32>, MemoryError> {
|
||||
Ok(vals)
|
||||
}
|
||||
|
||||
/// Create a fresh WAL file at `path` with the current-version header,
|
||||
/// truncating/overwriting anything already there.
|
||||
fn create_fresh_wal_file(path: &Path) -> Result<File, MemoryError> {
|
||||
let mut f = File::create(path)?;
|
||||
f.write_all(&WAL_MAGIC)?;
|
||||
f.write_all(&[WAL_VERSION])?;
|
||||
f.write_all(&0u32.to_le_bytes())?;
|
||||
f.flush()?;
|
||||
Ok(f)
|
||||
}
|
||||
|
||||
/// Wraps a [`Read`]er, accumulating every byte actually consumed (including
|
||||
/// via `read_exact`, which is implemented in terms of `read`) into an
|
||||
/// internal buffer — used to capture a WAL entry's raw bytes for CRC32
|
||||
/// verification without needing to know its length up front.
|
||||
struct TeeReader<'a, R: Read> {
|
||||
inner: &'a mut R,
|
||||
buf: Vec<u8>,
|
||||
}
|
||||
|
||||
impl<'a, R: Read> TeeReader<'a, R> {
|
||||
fn new(inner: &'a mut R) -> Self {
|
||||
Self {
|
||||
inner,
|
||||
buf: Vec::new(),
|
||||
}
|
||||
}
|
||||
|
||||
fn into_buf(self) -> Vec<u8> {
|
||||
self.buf
|
||||
}
|
||||
}
|
||||
|
||||
impl<R: Read> Read for TeeReader<'_, R> {
|
||||
fn read(&mut self, out: &mut [u8]) -> std::io::Result<usize> {
|
||||
let n = self.inner.read(out)?;
|
||||
self.buf.extend_from_slice(&out[..n]);
|
||||
Ok(n)
|
||||
}
|
||||
}
|
||||
|
||||
/// Read one WAL entry (type + timestamp + type-specific payload) from `r`.
|
||||
///
|
||||
/// Returns `Ok(None)` for entry types with no representable `WalEntry` (only
|
||||
/// `ActivationUpdate`, reserved for future use). Returns `Err(())` on any
|
||||
/// read failure or unrecognized entry type — the caller treats this the same
|
||||
/// as a clean end-of-log (crash-mid-write tolerance).
|
||||
fn read_one_entry<R: Read>(r: &mut R) -> Result<Option<WalEntry>, ()> {
|
||||
let mut type_buf = [0u8; 1];
|
||||
r.read_exact(&mut type_buf).map_err(|_| ())?;
|
||||
let entry_type = WalEntryType::from_u8(type_buf[0]).ok_or(())?;
|
||||
|
||||
let mut ts_buf = [0u8; 8];
|
||||
r.read_exact(&mut ts_buf).map_err(|_| ())?;
|
||||
let timestamp = f64::from_le_bytes(ts_buf);
|
||||
|
||||
match entry_type {
|
||||
WalEntryType::Save => {
|
||||
let chunk = read_len_prefixed_str(r).map_err(|_| ())?;
|
||||
let embedding = read_embedding(r).map_err(|_| ())?;
|
||||
let source_channel = read_len_prefixed_str(r).map_err(|_| ())?;
|
||||
let session_id = read_len_prefixed_str(r).map_err(|_| ())?;
|
||||
let tags = read_len_prefixed_str(r).map_err(|_| ())?;
|
||||
Ok(Some(WalEntry {
|
||||
entry_type,
|
||||
timestamp,
|
||||
chunk,
|
||||
embedding,
|
||||
source_channel,
|
||||
session_id,
|
||||
tags,
|
||||
tombstone_index: None,
|
||||
}))
|
||||
}
|
||||
WalEntryType::Tombstone => {
|
||||
let mut idx_buf = [0u8; 4];
|
||||
r.read_exact(&mut idx_buf).map_err(|_| ())?;
|
||||
let idx = u32::from_le_bytes(idx_buf) as usize;
|
||||
Ok(Some(WalEntry {
|
||||
entry_type,
|
||||
timestamp,
|
||||
chunk: String::new(),
|
||||
embedding: Vec::new(),
|
||||
source_channel: String::new(),
|
||||
session_id: String::new(),
|
||||
tags: String::new(),
|
||||
tombstone_index: Some(idx),
|
||||
}))
|
||||
}
|
||||
WalEntryType::ActivationUpdate => Ok(None),
|
||||
}
|
||||
}
|
||||
|
||||
// --- Tests ---
|
||||
|
||||
#[cfg(test)]
|
||||
@@ -427,6 +526,40 @@ mod tests {
|
||||
assert_eq!(entries[2].embedding, vec![5.0, 6.0]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_len_prefixed_str_rejects_oversized_len_claim() {
|
||||
let dir = TempDir::new().unwrap();
|
||||
let path = dir.path().join("oversized_str.bin");
|
||||
{
|
||||
let mut f = File::create(&path).unwrap();
|
||||
// Claim a length far beyond MAX_WAL_FIELD_LEN; no payload follows.
|
||||
f.write_all(&(u32::MAX).to_le_bytes()).unwrap();
|
||||
}
|
||||
let mut f = File::open(&path).unwrap();
|
||||
let result = read_len_prefixed_str(&mut f);
|
||||
assert!(
|
||||
matches!(result, Err(MemoryError::Schema(_))),
|
||||
"expected a clean Schema error, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_embedding_rejects_oversized_count_claim() {
|
||||
let dir = TempDir::new().unwrap();
|
||||
let path = dir.path().join("oversized_embedding.bin");
|
||||
{
|
||||
let mut f = File::create(&path).unwrap();
|
||||
// Claim a count far beyond MAX_WAL_FIELD_LEN / 4; no payload follows.
|
||||
f.write_all(&(u32::MAX).to_le_bytes()).unwrap();
|
||||
}
|
||||
let mut f = File::open(&path).unwrap();
|
||||
let result = read_embedding(&mut f);
|
||||
assert!(
|
||||
matches!(result, Err(MemoryError::Schema(_))),
|
||||
"expected a clean Schema error, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_wal_truncate() {
|
||||
let dir = TempDir::new().unwrap();
|
||||
@@ -749,6 +882,86 @@ mod tests {
|
||||
assert!(err.contains("unsupported WAL version"), "got: {err}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_wal_v2_detects_corrupted_payload_and_stops_replay() {
|
||||
let dir = TempDir::new().unwrap();
|
||||
let wal_path = dir.path().join("test.h5.wal");
|
||||
let mut wal = WalFile::open(&wal_path).unwrap();
|
||||
wal.append_save(&make_wal_entry("first", &[1.0, 2.0]))
|
||||
.unwrap();
|
||||
let len_after_first = std::fs::metadata(&wal_path).unwrap().len();
|
||||
wal.append_save(&make_wal_entry("second", &[3.0, 4.0]))
|
||||
.unwrap();
|
||||
drop(wal);
|
||||
|
||||
// Flip one byte inside the second entry's "second" chunk string
|
||||
// (well past the header and the first entry, and not touching any
|
||||
// length-prefix field) — this must be caught by the CRC32 trailer,
|
||||
// not by any length-cap guard.
|
||||
let mut bytes = std::fs::read(&wal_path).unwrap();
|
||||
let corrupt_at = len_after_first as usize + 15;
|
||||
bytes[corrupt_at] ^= 0xFF;
|
||||
std::fs::write(&wal_path, &bytes).unwrap();
|
||||
|
||||
let entries = WalFile::read_entries(&wal_path).unwrap();
|
||||
assert_eq!(
|
||||
entries.len(),
|
||||
1,
|
||||
"the corrupted second entry must not be returned"
|
||||
);
|
||||
assert_eq!(entries[0].chunk, "first");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_wal_reads_legacy_v1_format_without_crc() {
|
||||
let dir = TempDir::new().unwrap();
|
||||
let wal_path = dir.path().join("legacy.h5.wal");
|
||||
let mut buf = Vec::new();
|
||||
buf.extend_from_slice(&WAL_MAGIC);
|
||||
buf.push(WAL_VERSION_LEGACY_NO_CRC);
|
||||
buf.extend_from_slice(&1u32.to_le_bytes());
|
||||
// One Save entry in the old format: type + timestamp + fields, with
|
||||
// no trailing CRC32.
|
||||
buf.push(WalEntryType::Save as u8);
|
||||
buf.extend_from_slice(&42.0f64.to_le_bytes());
|
||||
serialize_str(&mut buf, "legacy-chunk");
|
||||
let embedding = [1.0f32, 2.0];
|
||||
buf.extend_from_slice(&(embedding.len() as u32).to_le_bytes());
|
||||
for v in embedding {
|
||||
buf.extend_from_slice(&v.to_le_bytes());
|
||||
}
|
||||
serialize_str(&mut buf, "chan");
|
||||
serialize_str(&mut buf, "sess");
|
||||
serialize_str(&mut buf, "tags");
|
||||
std::fs::write(&wal_path, &buf).unwrap();
|
||||
|
||||
let entries = WalFile::read_entries(&wal_path).unwrap();
|
||||
assert_eq!(entries.len(), 1);
|
||||
assert_eq!(entries[0].chunk, "legacy-chunk");
|
||||
assert_eq!(entries[0].embedding, vec![1.0, 2.0]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_wal_open_migrates_legacy_v1_to_current_version() {
|
||||
let dir = TempDir::new().unwrap();
|
||||
let wal_path = dir.path().join("legacy.h5.wal");
|
||||
let mut buf = Vec::new();
|
||||
buf.extend_from_slice(&WAL_MAGIC);
|
||||
buf.push(WAL_VERSION_LEGACY_NO_CRC);
|
||||
buf.extend_from_slice(&0u32.to_le_bytes());
|
||||
std::fs::write(&wal_path, &buf).unwrap();
|
||||
|
||||
let wal = WalFile::open(&wal_path).unwrap();
|
||||
assert!(wal.is_empty());
|
||||
drop(wal);
|
||||
|
||||
let bytes = std::fs::read(&wal_path).unwrap();
|
||||
assert_eq!(
|
||||
bytes[4], WAL_VERSION,
|
||||
"legacy file must be migrated to the current version"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_wal_disabled() {
|
||||
let dir = TempDir::new().unwrap();
|
||||
|
||||
@@ -10,3 +10,6 @@ crate-type = ["cdylib"]
|
||||
|
||||
[dependencies]
|
||||
clawhdf5-agent = { path = "../clawhdf5-agent", default-features = false }
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile = { workspace = true }
|
||||
|
||||
@@ -92,11 +92,18 @@ pub unsafe extern "C" fn edgehdf5_close(handle: Handle) {
|
||||
|
||||
/// Save a memory entry. Returns the entry index, or -1 on failure.
|
||||
///
|
||||
/// `embedding_len` is validated against the handle's configured
|
||||
/// `embedding_dim` before the input slice is constructed; a mismatch fails
|
||||
/// the call with -1 rather than reading out of bounds. This is a length
|
||||
/// check only — it cannot detect a same-length buffer that is otherwise
|
||||
/// too short or invalid.
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// - `handle` must be a valid, non-null handle.
|
||||
/// - All `*const c_char` arguments must be valid, null-terminated C strings.
|
||||
/// - `embedding_ptr` must point to at least `embedding_len` contiguous `f32` values.
|
||||
/// - If `embedding_len` matches the handle's `embedding_dim`, `embedding_ptr`
|
||||
/// must point to at least that many contiguous, valid `f32` values.
|
||||
#[unsafe(no_mangle)]
|
||||
pub unsafe extern "C" fn edgehdf5_save(
|
||||
handle: Handle,
|
||||
@@ -135,8 +142,14 @@ pub unsafe extern "C" fn edgehdf5_save(
|
||||
None => return -1,
|
||||
};
|
||||
|
||||
if embedding_ptr.is_null() || embedding_len as usize != mem.config().embedding_dim {
|
||||
return -1;
|
||||
}
|
||||
let embedding =
|
||||
// SAFETY: JNI caller guarantees embedding_ptr points to embedding_len valid f32 values.
|
||||
// SAFETY: embedding_ptr is non-null and embedding_len matches the handle's configured
|
||||
// embedding_dim (checked above); JNI caller guarantees it points to that many valid f32
|
||||
// values. A mismatched-but-equal-length short buffer is not caught by this length check
|
||||
// alone — the caller is still responsible for pointer validity.
|
||||
unsafe { std::slice::from_raw_parts(embedding_ptr, embedding_len as usize) }.to_vec();
|
||||
|
||||
let entry = MemoryEntry {
|
||||
@@ -210,11 +223,18 @@ pub unsafe extern "C" fn edgehdf5_delete(handle: Handle, index: u64) -> i32 {
|
||||
/// Performs hybrid search and writes up to `max_results` entries into the
|
||||
/// provided output arrays. Returns the number of results written.
|
||||
///
|
||||
/// `query_embedding_len` is validated against the handle's configured
|
||||
/// `embedding_dim` before the input slice is constructed; a mismatch fails
|
||||
/// the call (returns 0) rather than reading out of bounds. This is a length
|
||||
/// check only — it cannot detect a same-length buffer that is otherwise too
|
||||
/// short or invalid.
|
||||
///
|
||||
/// # Safety
|
||||
///
|
||||
/// - `handle` must be a valid, non-null handle.
|
||||
/// - `query_text` must be a valid, null-terminated C string.
|
||||
/// - `query_embedding_ptr` must point to at least `query_embedding_len` `f32` values.
|
||||
/// - If `query_embedding_len` matches the handle's `embedding_dim`,
|
||||
/// `query_embedding_ptr` must point to at least that many valid `f32` values.
|
||||
/// - `out_indices` and `out_scores` must point to arrays of at least `max_results` elements.
|
||||
/// - `out_chunks` must be null or point to an array of at least `max_results` pointers.
|
||||
#[unsafe(no_mangle)]
|
||||
@@ -240,8 +260,14 @@ pub unsafe extern "C" fn edgehdf5_hybrid_search(
|
||||
Some(s) => s,
|
||||
None => return 0,
|
||||
};
|
||||
if query_embedding_ptr.is_null() || query_embedding_len as usize != mem.config().embedding_dim {
|
||||
return 0;
|
||||
}
|
||||
let query_embedding =
|
||||
// SAFETY: JNI caller guarantees query_embedding_ptr points to query_embedding_len valid f32 values.
|
||||
// SAFETY: query_embedding_ptr is non-null and query_embedding_len matches the handle's
|
||||
// configured embedding_dim (checked above); JNI caller guarantees it points to that many
|
||||
// valid f32 values. A mismatched-but-equal-length short buffer is not caught by this
|
||||
// length check alone — the caller is still responsible for pointer validity.
|
||||
unsafe { std::slice::from_raw_parts(query_embedding_ptr, query_embedding_len as usize) };
|
||||
|
||||
let results = mem.hybrid_search(
|
||||
@@ -456,3 +482,112 @@ unsafe fn cstr_to_string(ptr: *const c_char) -> Option<String> {
|
||||
.ok()
|
||||
.map(String::from)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
const EMBEDDING_DIM: u32 = 4;
|
||||
|
||||
fn open_handle(dir: &tempfile::TempDir) -> Handle {
|
||||
let path = CString::new(dir.path().join("mem.h5").to_str().unwrap()).unwrap();
|
||||
let agent_id = CString::new("test-agent").unwrap();
|
||||
// SAFETY: both C strings are valid and null-terminated.
|
||||
unsafe { edgehdf5_create(path.as_ptr(), agent_id.as_ptr(), EMBEDDING_DIM) }
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn save_rejects_mismatched_embedding_len() {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let handle = open_handle(&dir);
|
||||
assert!(!handle.is_null());
|
||||
|
||||
let embedding = [1.0f32, 2.0, 3.0]; // len 3, dim is 4
|
||||
let chunk = CString::new("hello").unwrap();
|
||||
let channel = CString::new("test").unwrap();
|
||||
let session = CString::new("s1").unwrap();
|
||||
let tags = CString::new("").unwrap();
|
||||
|
||||
// SAFETY: handle is valid; all C strings are valid; embedding_len (3) intentionally
|
||||
// does not match embedding_dim (4), which edgehdf5_save must reject before touching
|
||||
// embedding_ptr.
|
||||
let result = unsafe {
|
||||
edgehdf5_save(
|
||||
handle,
|
||||
chunk.as_ptr(),
|
||||
embedding.as_ptr(),
|
||||
embedding.len() as u32,
|
||||
channel.as_ptr(),
|
||||
0.0,
|
||||
session.as_ptr(),
|
||||
tags.as_ptr(),
|
||||
)
|
||||
};
|
||||
assert_eq!(result, -1, "mismatched embedding_len must be rejected");
|
||||
|
||||
unsafe { edgehdf5_close(handle) };
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn save_rejects_null_embedding_ptr() {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let handle = open_handle(&dir);
|
||||
assert!(!handle.is_null());
|
||||
|
||||
let chunk = CString::new("hello").unwrap();
|
||||
let channel = CString::new("test").unwrap();
|
||||
let session = CString::new("s1").unwrap();
|
||||
let tags = CString::new("").unwrap();
|
||||
|
||||
// SAFETY: handle and C strings are valid; embedding_ptr is intentionally null, which
|
||||
// edgehdf5_save must reject before constructing a slice from it.
|
||||
let result = unsafe {
|
||||
edgehdf5_save(
|
||||
handle,
|
||||
chunk.as_ptr(),
|
||||
ptr::null(),
|
||||
EMBEDDING_DIM,
|
||||
channel.as_ptr(),
|
||||
0.0,
|
||||
session.as_ptr(),
|
||||
tags.as_ptr(),
|
||||
)
|
||||
};
|
||||
assert_eq!(result, -1, "null embedding_ptr must be rejected");
|
||||
|
||||
unsafe { edgehdf5_close(handle) };
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn hybrid_search_rejects_mismatched_embedding_len() {
|
||||
let dir = tempfile::tempdir().unwrap();
|
||||
let handle = open_handle(&dir);
|
||||
assert!(!handle.is_null());
|
||||
|
||||
let query_embedding = [1.0f32, 2.0]; // len 2, dim is 4
|
||||
let query_text = CString::new("hello").unwrap();
|
||||
let mut out_indices = [0u64; 4];
|
||||
let mut out_scores = [0.0f32; 4];
|
||||
|
||||
// SAFETY: handle and query_text are valid; query_embedding_len (2) intentionally does
|
||||
// not match embedding_dim (4), which edgehdf5_hybrid_search must reject before touching
|
||||
// query_embedding_ptr. Output buffers are sized to max_results.
|
||||
let count = unsafe {
|
||||
edgehdf5_hybrid_search(
|
||||
handle,
|
||||
query_embedding.as_ptr(),
|
||||
query_embedding.len() as u32,
|
||||
query_text.as_ptr(),
|
||||
0.7,
|
||||
0.3,
|
||||
4,
|
||||
out_indices.as_mut_ptr(),
|
||||
out_scores.as_mut_ptr(),
|
||||
ptr::null_mut(),
|
||||
)
|
||||
};
|
||||
assert_eq!(count, 0, "mismatched query_embedding_len must be rejected");
|
||||
|
||||
unsafe { edgehdf5_close(handle) };
|
||||
}
|
||||
}
|
||||
|
||||
@@ -12,3 +12,7 @@ categories = ["algorithms", "science"]
|
||||
[dependencies]
|
||||
clawhdf5-format = { path = "../clawhdf5-format", version = "2.1.0" }
|
||||
clawhdf5-io = { path = "../clawhdf5-io", version = "2.1.0" }
|
||||
rayon = { version = "1", optional = true }
|
||||
|
||||
[features]
|
||||
parallel = ["rayon"]
|
||||
|
||||
@@ -857,6 +857,15 @@ fn prune_connections(
|
||||
if neighbors.len() <= max_conn {
|
||||
return;
|
||||
}
|
||||
#[cfg(feature = "parallel")]
|
||||
let mut scored: Vec<(usize, f32)> = {
|
||||
use rayon::prelude::*;
|
||||
neighbors
|
||||
.par_iter()
|
||||
.map(|&n| (n, compute_distance(&vectors[node], &vectors[n], metric)))
|
||||
.collect()
|
||||
};
|
||||
#[cfg(not(feature = "parallel"))]
|
||||
let mut scored: Vec<(usize, f32)> = neighbors
|
||||
.iter()
|
||||
.map(|&n| (n, compute_distance(&vectors[node], &vectors[n], metric)))
|
||||
|
||||
@@ -50,19 +50,31 @@ harness = false
|
||||
clawhdf5-agent = { path = "../clawhdf5-agent" }
|
||||
clawhdf5-io = { path = "../clawhdf5-io" }
|
||||
mpi = { version = "0.8", optional = true }
|
||||
serde = { version = "1", features = ["derive"] }
|
||||
serde = { workspace = true }
|
||||
serde_json = "1"
|
||||
tempfile = "3"
|
||||
tempfile = { workspace = true }
|
||||
# Optional: libhdf5 C wrapper for side-by-side comparison (requires system libhdf5).
|
||||
# Enable with: cargo bench -p clawhdf5-bench --features libhdf5-compare
|
||||
# Uses hdf5-metno (fork of hdf5 crate) which supports HDF5 1.14.x.
|
||||
hdf5 = { version = "0.12", optional = true, package = "hdf5-metno" }
|
||||
# Optional: real sentence embeddings for the LongMemEval bench's vector stage.
|
||||
# Enable with: cargo run --release --bin longmemeval_bench --features embeddings
|
||||
# Off by default — nothing in the shipped crates depends on these.
|
||||
candle-core = { version = "0.9", optional = true }
|
||||
candle-nn = { version = "0.9", optional = true }
|
||||
candle-transformers = { version = "0.9", optional = true }
|
||||
tokenizers = { version = "0.21", optional = true }
|
||||
|
||||
[dev-dependencies]
|
||||
clawhdf5 = { path = "../clawhdf5", features = ["zstd", "pcodec"] }
|
||||
criterion = { version = "0.5", features = ["html_reports"] }
|
||||
criterion = { workspace = true }
|
||||
|
||||
[features]
|
||||
# When enabled, benchmarks add matching libhdf5 variants for side-by-side comparison.
|
||||
libhdf5-compare = ["hdf5"]
|
||||
mpi-io = ["clawhdf5-io/mpi-io", "mpi"]
|
||||
# Real MiniLM embeddings for longmemeval_bench, so the vector stage is not inert.
|
||||
embeddings = ["candle-core", "candle-nn", "candle-transformers", "tokenizers"]
|
||||
# CUDA-accelerated embedding. MiniLM on a CPU takes hours over the full
|
||||
# longmemeval_s haystack; on a GPU it is minutes.
|
||||
embeddings-cuda = ["embeddings", "candle-core/cuda", "candle-nn/cuda", "candle-transformers/cuda"]
|
||||
|
||||
@@ -0,0 +1,95 @@
|
||||
//! World-model sample-loading benchmark — clawhdf5 vs the h5py counterpart.
|
||||
//!
|
||||
//! Reproduces the access pattern of `stable-worldmodel`'s HDF5 dataloader
|
||||
//! (arXiv 2605.21800): a dataset of `(N, H, W, C)` uint8 observation frames,
|
||||
//! read one frame at a time in shuffled (dataloader) order. That paper
|
||||
//! reports generic HDF5 at 1,416–1,474 samples/s (vs Lance 4,815); this
|
||||
//! measures clawhdf5 and h5py on the **same machine and file**, so the
|
||||
//! comparison is hardware-controlled. Absolute numbers are not comparable to
|
||||
//! the paper's (different box, smaller frames, no torch/transform) — only
|
||||
//! clawhdf5-vs-h5py *here* is.
|
||||
//!
|
||||
//! clawhdf5 mmaps the file once and takes a zero-copy `&[u8]` over the
|
||||
//! contiguous observation dataset; frame `i` is a subslice, and the OS pages
|
||||
//! it in on access. Two modes, because fairness demands both:
|
||||
//! * default: sum the frame bytes through the zero-copy view — clawhdf5's
|
||||
//! real advantage, no per-frame allocation;
|
||||
//! * `--copy`: `to_vec()` each frame first, matching h5py's unavoidable
|
||||
//! per-frame numpy materialization, so the two do equal work.
|
||||
//!
|
||||
//! Usage: `... --example worldmodel_sampling -- <file.h5> [passes] [--copy]`
|
||||
|
||||
use std::hint::black_box;
|
||||
use std::time::Instant;
|
||||
|
||||
use clawhdf5::MmapFile;
|
||||
|
||||
fn main() {
|
||||
let args: Vec<String> = std::env::args().collect();
|
||||
let path = args
|
||||
.get(1)
|
||||
.expect("usage: worldmodel_sampling <file.h5> [passes] [--copy]");
|
||||
let passes: usize = args.get(2).and_then(|s| s.parse().ok()).unwrap_or(5);
|
||||
let copy = args.iter().any(|a| a == "--copy");
|
||||
|
||||
let file = MmapFile::open(path).expect("open");
|
||||
let ds = file.dataset("observation").expect("observation dataset");
|
||||
let shape = ds.shape().expect("shape");
|
||||
let n = shape[0] as usize;
|
||||
let frame_bytes: usize = shape[1..].iter().map(|&d| d as usize).product();
|
||||
let raw = ds
|
||||
.read_raw_slice()
|
||||
.expect("read_raw_slice")
|
||||
.expect("contiguous zero-copy slice");
|
||||
assert_eq!(raw.len(), n * frame_bytes, "unexpected dataset size");
|
||||
|
||||
let order = shuffled(n);
|
||||
|
||||
let touch = |slice: &[u8]| -> u64 {
|
||||
if copy {
|
||||
let owned = slice.to_vec();
|
||||
owned.iter().map(|&b| u64::from(b)).sum()
|
||||
} else {
|
||||
slice.iter().map(|&b| u64::from(b)).sum()
|
||||
}
|
||||
};
|
||||
|
||||
// Warm one pass (page-in), then time.
|
||||
let mut sink = 0u64;
|
||||
for &i in &order {
|
||||
sink = sink.wrapping_add(touch(&raw[i * frame_bytes..(i + 1) * frame_bytes]));
|
||||
}
|
||||
black_box(sink);
|
||||
|
||||
let t0 = Instant::now();
|
||||
let mut sink = 0u64;
|
||||
for _ in 0..passes {
|
||||
for &i in &order {
|
||||
sink = sink.wrapping_add(touch(&raw[i * frame_bytes..(i + 1) * frame_bytes]));
|
||||
}
|
||||
}
|
||||
black_box(sink);
|
||||
let elapsed = t0.elapsed().as_secs_f64();
|
||||
|
||||
let total = (n * passes) as f64;
|
||||
let mode = if copy {
|
||||
"materialized copy"
|
||||
} else {
|
||||
"zero-copy view"
|
||||
};
|
||||
println!("clawhdf5 ({mode}): {n} frames x {passes} passes in {elapsed:.3}s");
|
||||
println!("clawhdf5 ({mode}): {:.0} samples/sec", total / elapsed);
|
||||
}
|
||||
|
||||
fn shuffled(n: usize) -> Vec<usize> {
|
||||
let mut v: Vec<usize> = (0..n).collect();
|
||||
let mut state: u64 = 0x9E37_79B9_7F4A_7C15;
|
||||
for i in (1..n).rev() {
|
||||
state = state
|
||||
.wrapping_mul(6364136223846793005)
|
||||
.wrapping_add(1442695040888963407);
|
||||
let j = (state >> 33) as usize % (i + 1);
|
||||
v.swap(i, j);
|
||||
}
|
||||
v
|
||||
}
|
||||
@@ -4,11 +4,39 @@
|
||||
//! Since no embedding model is available at bench time, all embeddings are zero vectors
|
||||
//! and `hybrid_search` operates in BM25-only mode (vector_weight=0.0, keyword_weight=1.0).
|
||||
//!
|
||||
//! This matches the MemX paper methodology: evaluate retrieval recall, not answer generation.
|
||||
//! # Scoring target (read before citing any number from this harness)
|
||||
//!
|
||||
//! - **Metric: retrieval recall.** A "hit" means the gold-labelled memory appeared in
|
||||
//! the top-k. No answer is generated and none is scored — the dataset's `answer`
|
||||
//! field is deserialized and deliberately never read. This is **not** the official
|
||||
//! LongMemEval metric, which is end-to-end QA accuracy (retrieve → generate → LLM
|
||||
//! judge). Reporting retrieval recall as QA accuracy overstates by 20–30 points.
|
||||
//! - **Dataset: whichever variant you point it at.** Both `longmemeval_oracle`
|
||||
//! (evidence sessions only — a substantially easier corpus) and the full
|
||||
//! `longmemeval_s` haystack are supported. The harness does not trust the
|
||||
//! filename: [`DatasetProfile`] measures evidence-session density from the
|
||||
//! data and labels the run from that, so a mislabelled input cannot produce a
|
||||
//! mislabelled result.
|
||||
//! - **Session-level metrics are degenerate when evidence density is high**, and
|
||||
//! the report says so per run rather than assuming it. On the oracle variant
|
||||
//! the haystack is essentially all-evidence, so any returned document is a
|
||||
//! session-level hit at rank 0 by construction; only turn-level
|
||||
//! (`has_answer == true` on the source turn) measures the retriever there. On
|
||||
//! the full haystack, session-level recall is meaningful.
|
||||
//! - **Not comparable to MemX's Hit@5=51.6% / MRR=0.380**, which is *fact-level*
|
||||
//! granularity over 220,349 records from 19,195 sessions.
|
||||
//!
|
||||
//! See `BENCHMARKS.md` § "Retracted: session-level recall and the MemX comparison".
|
||||
//!
|
||||
//! # Usage
|
||||
//! ```
|
||||
//! cargo run --release --bin longmemeval_bench [path/to/longmemeval_oracle.json]
|
||||
//! cargo run --release --bin longmemeval_bench [PATH] [--limit N]
|
||||
//!
|
||||
//! # Usage: full haystack
|
||||
//! ```
|
||||
//! cargo run --release --bin longmemeval_bench -- \
|
||||
//! benchmarks/longmemeval/longmemeval_s_cleaned.json --limit 50
|
||||
//! ```
|
||||
//! ```
|
||||
//!
|
||||
//! # WASM Note
|
||||
@@ -21,12 +49,80 @@
|
||||
use std::collections::{HashMap, HashSet};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
// `#[path]` keeps the module beside its binary without Cargo autodiscovering it
|
||||
// as a second bin target (which a bare `src/bin/embedder.rs` would be).
|
||||
#[cfg(feature = "embeddings")]
|
||||
#[path = "longmemeval_bench/embedder.rs"]
|
||||
mod embedder;
|
||||
|
||||
use clawhdf5_agent::{AgentMemory, HDF5Memory, MemoryConfig, MemoryEntry};
|
||||
use serde::Deserialize;
|
||||
use tempfile::TempDir;
|
||||
|
||||
const EMBEDDING_DIM: usize = 384;
|
||||
|
||||
/// A retrieval configuration: how much of the score comes from each stage.
|
||||
#[derive(Clone, Copy)]
|
||||
struct Mode {
|
||||
label: &'static str,
|
||||
vector_weight: f32,
|
||||
keyword_weight: f32,
|
||||
}
|
||||
|
||||
/// The only mode available without real embeddings. Passing zero vectors with
|
||||
/// `vector_weight = 0.0` is what made the vector stage inert.
|
||||
const BM25_ONLY: Mode = Mode {
|
||||
label: "BM25 only (vector stage inert)",
|
||||
vector_weight: 0.0,
|
||||
keyword_weight: 1.0,
|
||||
};
|
||||
#[cfg(feature = "embeddings")]
|
||||
const VECTOR_ONLY: Mode = Mode {
|
||||
label: "Vector only (MiniLM + HNSW)",
|
||||
vector_weight: 1.0,
|
||||
keyword_weight: 0.0,
|
||||
};
|
||||
/// Tuned by `--sweep` over the full haystack. The former 0.7/0.3 was a
|
||||
/// documented default that had never been searched, and the sweep found it
|
||||
/// strictly dominated: 0.4/0.6 is better on Hit@1, Hit@5, Hit@10 and MRR at
|
||||
/// both granularities.
|
||||
#[cfg(feature = "embeddings")]
|
||||
const HYBRID: Mode = Mode {
|
||||
label: "Hybrid (0.4 vector / 0.6 BM25, tuned)",
|
||||
vector_weight: 0.4,
|
||||
keyword_weight: 0.6,
|
||||
};
|
||||
|
||||
/// Every 0.1 step of vector weight, keyword weight taking the remainder.
|
||||
///
|
||||
/// Labels are leaked to `&'static str` because `Mode::label` is a `&'static
|
||||
/// str` for the eleven named modes and a sweep is a short-lived process; the
|
||||
/// alternative is threading a lifetime through the whole report path for a
|
||||
/// diagnostic mode.
|
||||
#[cfg(feature = "embeddings")]
|
||||
fn sweep_modes() -> Vec<Mode> {
|
||||
(0..=10)
|
||||
.map(|i| {
|
||||
let v = i as f32 / 10.0;
|
||||
Mode {
|
||||
label: Box::leak(format!("sweep v={v:.1} / k={:.1}", 1.0 - v).into_boxed_str()),
|
||||
vector_weight: v,
|
||||
keyword_weight: 1.0 - v,
|
||||
}
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Text -> embedding, built once for the whole corpus.
|
||||
type EmbeddingMap = HashMap<String, Vec<f32>>;
|
||||
|
||||
/// Look up a real embedding, falling back to zeros when running BM25-only.
|
||||
fn embedding_for(map: Option<&EmbeddingMap>, text: &str) -> Vec<f32> {
|
||||
map.and_then(|m| m.get(text))
|
||||
.cloned()
|
||||
.unwrap_or_else(|| vec![0.0f32; EMBEDDING_DIM])
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// JSON data types
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -168,7 +264,12 @@ struct EvalResult {
|
||||
latency: Duration,
|
||||
}
|
||||
|
||||
fn evaluate_question(q: &Question, top_k: usize) -> EvalResult {
|
||||
fn evaluate_question(
|
||||
q: &Question,
|
||||
top_k: usize,
|
||||
mode: Mode,
|
||||
embeddings: Option<&EmbeddingMap>,
|
||||
) -> EvalResult {
|
||||
let dir = TempDir::new().expect("failed to create temp dir");
|
||||
let mut config = MemoryConfig::new(dir.path().join("lme.h5"), "lme-bench", EMBEDDING_DIM);
|
||||
config.wal_enabled = false;
|
||||
@@ -190,7 +291,7 @@ fn evaluate_question(q: &Question, top_k: usize) -> EvalResult {
|
||||
for turn in session {
|
||||
entries.push(MemoryEntry {
|
||||
chunk: turn.content.clone(),
|
||||
embedding: vec![0.0f32; EMBEDDING_DIM],
|
||||
embedding: embedding_for(embeddings, &turn.content),
|
||||
source_channel: "longmemeval".to_string(),
|
||||
timestamp: ts,
|
||||
session_id: sess_id.to_string(),
|
||||
@@ -218,10 +319,15 @@ fn evaluate_question(q: &Question, top_k: usize) -> EvalResult {
|
||||
// Set of session IDs that contain the answer
|
||||
let answer_sess_set: HashSet<&str> = q.answer_session_ids.iter().map(String::as_str).collect();
|
||||
|
||||
// Run hybrid search (BM25-only: vector_weight=0.0, keyword_weight=1.0)
|
||||
let zero_emb = vec![0.0f32; EMBEDDING_DIM];
|
||||
let query_emb = embedding_for(embeddings, &q.question);
|
||||
let t0 = Instant::now();
|
||||
let results = memory.hybrid_search(&zero_emb, &q.question, 0.0, 1.0, top_k);
|
||||
let results = memory.hybrid_search(
|
||||
&query_emb,
|
||||
&q.question,
|
||||
mode.vector_weight,
|
||||
mode.keyword_weight,
|
||||
top_k,
|
||||
);
|
||||
let latency = t0.elapsed();
|
||||
|
||||
// Session-level recall
|
||||
@@ -286,17 +392,133 @@ fn evaluate_question(q: &Question, top_k: usize) -> EvalResult {
|
||||
// Report printing
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn print_report(overall: &Metrics, by_type: &HashMap<String, Metrics>) {
|
||||
// ---------------------------------------------------------------------------
|
||||
// Dataset profile — measured, not assumed
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Shape of the loaded corpus, computed from the data itself.
|
||||
///
|
||||
/// The variant used to be a hardcoded `"oracle"` string in the report and the
|
||||
/// JSON summary, so pointing the harness at `longmemeval_s` would have produced
|
||||
/// full-haystack numbers labelled oracle. Everything here is derived from the
|
||||
/// questions instead, which means the label cannot drift from the corpus and a
|
||||
/// mislabelled input file cannot produce a mislabelled result.
|
||||
struct DatasetProfile {
|
||||
n_questions: usize,
|
||||
mean_sessions: f64,
|
||||
mean_turns: f64,
|
||||
/// Mean over questions of `|answer_sessions| / |haystack_sessions|`.
|
||||
///
|
||||
/// This is what actually decides whether session-level recall means
|
||||
/// anything. At ~1.0 every haystack session is an evidence session, so any
|
||||
/// returned document is a session-level hit by construction.
|
||||
evidence_density: f64,
|
||||
}
|
||||
|
||||
impl DatasetProfile {
|
||||
fn measure(questions: &[Question]) -> Self {
|
||||
let n = questions.len().max(1) as f64;
|
||||
let mut sessions = 0.0;
|
||||
let mut turns = 0.0;
|
||||
let mut density = 0.0;
|
||||
for q in questions {
|
||||
let n_sess = q.haystack_sessions.len();
|
||||
sessions += n_sess as f64;
|
||||
turns += q.haystack_sessions.iter().map(Vec::len).sum::<usize>() as f64;
|
||||
if n_sess > 0 {
|
||||
let evidence: HashSet<&str> =
|
||||
q.answer_session_ids.iter().map(String::as_str).collect();
|
||||
let hit = q
|
||||
.haystack_session_ids
|
||||
.iter()
|
||||
.filter(|id| evidence.contains(id.as_str()))
|
||||
.count();
|
||||
density += hit as f64 / n_sess as f64;
|
||||
}
|
||||
}
|
||||
Self {
|
||||
n_questions: questions.len(),
|
||||
mean_sessions: sessions / n,
|
||||
mean_turns: turns / n,
|
||||
evidence_density: density / n,
|
||||
}
|
||||
}
|
||||
|
||||
/// Above this share of evidence sessions, session-level recall is measuring
|
||||
/// the corpus shape rather than the retriever.
|
||||
const DEGENERACY_THRESHOLD: f64 = 0.9;
|
||||
|
||||
const fn session_level_degenerate(&self) -> bool {
|
||||
self.evidence_density > Self::DEGENERACY_THRESHOLD
|
||||
}
|
||||
|
||||
/// Variant name inferred from evidence density, not from the filename.
|
||||
const fn variant(&self) -> &'static str {
|
||||
if self.session_level_degenerate() {
|
||||
"oracle"
|
||||
} else {
|
||||
"full_haystack"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn print_report(
|
||||
overall: &Metrics,
|
||||
by_type: &HashMap<String, Metrics>,
|
||||
profile: &DatasetProfile,
|
||||
mode: Mode,
|
||||
) {
|
||||
println!("=================================================================");
|
||||
println!(" LongMemEval Benchmark (BM25-only retrieval, zero embeddings)");
|
||||
println!(" LongMemEval Benchmark — {}", mode.label);
|
||||
println!("=================================================================");
|
||||
println!();
|
||||
println!("Mode: vector_weight=0.0 / keyword_weight=1.0 (pure BM25)");
|
||||
println!("Note: MemX (arxiv:2603.16171) with full system: Hit@5=51.6%, MRR=0.380");
|
||||
println!(" BM25-only numbers are expected to be lower — honest baseline.");
|
||||
println!(
|
||||
"Mode: vector_weight={:.1} / keyword_weight={:.1}",
|
||||
mode.vector_weight, mode.keyword_weight
|
||||
);
|
||||
println!();
|
||||
println!("Scoring target: RETRIEVAL RECALL (did the gold memory land in top-k).");
|
||||
println!(" No answer is generated or scored. This is NOT the official");
|
||||
println!(" LongMemEval metric (QA accuracy via retrieve+generate+judge).");
|
||||
println!(
|
||||
"Dataset: {} — {} questions, {:.1} sessions and {:.0} turns per question,",
|
||||
profile.variant(),
|
||||
profile.n_questions,
|
||||
profile.mean_sessions,
|
||||
profile.mean_turns,
|
||||
);
|
||||
println!(
|
||||
" {:.1}% of haystack sessions are evidence sessions.",
|
||||
profile.evidence_density * 100.0
|
||||
);
|
||||
if profile.session_level_degenerate() {
|
||||
println!(" This is the evidence-only corpus, NOT the full longmemeval_s");
|
||||
println!(" haystack — a substantially easier retrieval problem.");
|
||||
} else {
|
||||
println!(" This is a full-haystack corpus: evidence sessions are a small");
|
||||
println!(" minority, so retrieval has to actually discriminate.");
|
||||
}
|
||||
println!();
|
||||
println!("Do NOT compare these to MemX's Hit@5=51.6% / MRR=0.380: that is");
|
||||
println!(" fact-level granularity over 220,349 records from 19,195 sessions.");
|
||||
println!(" Different granularity and a corpus larger by orders of magnitude.");
|
||||
println!();
|
||||
|
||||
println!("## Session-Level Recall (n={})", overall.count);
|
||||
if profile.session_level_degenerate() {
|
||||
println!(
|
||||
" [DEGENERATE — {:.1}% of haystack sessions are evidence sessions, so a",
|
||||
profile.evidence_density * 100.0
|
||||
);
|
||||
println!(" returned document is a session-level hit almost by construction.");
|
||||
println!(" This measures the corpus shape, not the retriever. Use turn-level.]");
|
||||
} else {
|
||||
println!(
|
||||
" [Meaningful on this corpus — only {:.1}% of haystack sessions are",
|
||||
profile.evidence_density * 100.0
|
||||
);
|
||||
println!(" evidence sessions, so a hit reflects the retriever's discrimination.]");
|
||||
}
|
||||
println!(
|
||||
" Hit@1: {:5.1}% Hit@5: {:5.1}% Hit@10: {:5.1}% MRR: {:.4}",
|
||||
overall.hit1_session_pct(),
|
||||
@@ -380,7 +602,26 @@ fn print_report(overall: &Metrics, by_type: &HashMap<String, Metrics>) {
|
||||
println!("```json");
|
||||
println!("{{");
|
||||
println!(" \"benchmark\": \"longmemeval\",");
|
||||
println!(" \"mode\": \"bm25_only\",");
|
||||
println!(
|
||||
" \"mode\": \"vector_{:.1}_keyword_{:.1}\",",
|
||||
mode.vector_weight, mode.keyword_weight
|
||||
);
|
||||
println!(" \"dataset_variant\": \"{}\",", profile.variant());
|
||||
println!(" \"scoring_target\": \"retrieval_recall\",");
|
||||
println!(" \"k\": 10,");
|
||||
println!(
|
||||
" \"session_level_degenerate\": {},",
|
||||
profile.session_level_degenerate()
|
||||
);
|
||||
println!(
|
||||
" \"evidence_session_density\": {:.4},",
|
||||
profile.evidence_density
|
||||
);
|
||||
println!(
|
||||
" \"mean_sessions_per_question\": {:.2},",
|
||||
profile.mean_sessions
|
||||
);
|
||||
println!(" \"mean_turns_per_question\": {:.1},", profile.mean_turns);
|
||||
println!(
|
||||
" \"total_questions\": {},",
|
||||
overall.count + overall.abstention_total
|
||||
@@ -403,10 +644,16 @@ fn print_report(overall: &Metrics, by_type: &HashMap<String, Metrics>) {
|
||||
overall.mrr_turn()
|
||||
);
|
||||
println!(" }},");
|
||||
// `null`, not 0.0 — a corpus with no abstention questions has no abstention
|
||||
// accuracy, and emitting 0.0 reads as total failure at a task never posed.
|
||||
if overall.abstention_total > 0 {
|
||||
println!(
|
||||
" \"abstention_accuracy\": {:.4},",
|
||||
overall.abstention_pct() / 100.0
|
||||
);
|
||||
} else {
|
||||
println!(" \"abstention_accuracy\": null,");
|
||||
}
|
||||
println!(" \"latency_us\": {{");
|
||||
println!(
|
||||
" \"avg\": {:.1}, \"p50\": {:.1}, \"p95\": {:.1}, \"p99\": {:.1}",
|
||||
@@ -425,17 +672,152 @@ fn print_report(overall: &Metrics, by_type: &HashMap<String, Metrics>) {
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn main() {
|
||||
let json_path = std::env::args()
|
||||
.nth(1)
|
||||
.unwrap_or_else(|| "benchmarks/longmemeval/longmemeval_oracle.json".to_string());
|
||||
let mut json_path: Option<String> = None;
|
||||
let mut limit: Option<usize> = None;
|
||||
let mut weights_dir: Option<String> = None;
|
||||
let mut sweep = false;
|
||||
let mut args = std::env::args().skip(1);
|
||||
while let Some(arg) = args.next() {
|
||||
match arg.as_str() {
|
||||
"--limit" => {
|
||||
let v = args.next().expect("--limit needs a value");
|
||||
limit = Some(v.parse().expect("--limit must be a positive integer"));
|
||||
}
|
||||
"--sweep" => sweep = true,
|
||||
"--embeddings" => {
|
||||
weights_dir = Some(args.next().expect("--embeddings needs a directory"));
|
||||
}
|
||||
"--help" | "-h" => {
|
||||
eprintln!(
|
||||
"usage: longmemeval_bench [PATH] [--limit N]\n\n\
|
||||
PATH dataset JSON; defaults to the oracle variant.\n\
|
||||
longmemeval_s works too — the harness measures which\n\
|
||||
variant it was given rather than trusting the filename.\n\
|
||||
--limit evaluate N questions, sampled evenly across the file\n\
|
||||
rather than as a prefix — the dataset is ordered by\n\
|
||||
question type, so a prefix samples one type only.\n\
|
||||
--embeddings DIR\n\
|
||||
directory holding all-MiniLM-L6-v2's model.safetensors\n\
|
||||
and tokenizer.json. Enables the vector stage and reports\n\
|
||||
BM25-only, vector-only, and hybrid separately. Requires\n\
|
||||
--features embeddings; without it the vector stage is\n\
|
||||
inert and only the BM25 row is produced.\n\
|
||||
--sweep instead of the three named modes, sweep vector_weight\n\
|
||||
from 0.0 to 1.0 in 0.1 steps. The 0.7/0.3 default was\n\
|
||||
never searched; this is what searches it."
|
||||
);
|
||||
return;
|
||||
}
|
||||
other => json_path = Some(other.to_string()),
|
||||
}
|
||||
}
|
||||
let json_path =
|
||||
json_path.unwrap_or_else(|| "benchmarks/longmemeval/longmemeval_oracle.json".to_string());
|
||||
|
||||
eprintln!("Loading: {json_path}");
|
||||
let data = std::fs::read_to_string(&json_path)
|
||||
.unwrap_or_else(|e| panic!("Failed to read {json_path}: {e}"));
|
||||
let questions: Vec<Question> = serde_json::from_str(&data).expect("Failed to parse JSON");
|
||||
let mut questions: Vec<Question> = serde_json::from_str(&data).expect("Failed to parse JSON");
|
||||
if let Some(n) = limit
|
||||
&& n < questions.len()
|
||||
{
|
||||
// Stride rather than truncate. The dataset is ordered by question type,
|
||||
// so taking a prefix samples one type: `--limit 20` on longmemeval_s
|
||||
// returns 20 `single-session-user` questions and nothing else, which
|
||||
// reads as a whole-dataset result but is not one.
|
||||
let total = questions.len();
|
||||
let step = total as f64 / n as f64;
|
||||
let keep: HashSet<usize> = (0..n)
|
||||
.map(|i| ((i as f64 * step) as usize).min(total - 1))
|
||||
.collect();
|
||||
questions = questions
|
||||
.into_iter()
|
||||
.enumerate()
|
||||
.filter(|(i, _)| keep.contains(i))
|
||||
.map(|(_, q)| q)
|
||||
.collect();
|
||||
eprintln!(
|
||||
"Sampling {} of {total} questions, evenly strided (--limit)",
|
||||
questions.len()
|
||||
);
|
||||
}
|
||||
let total = questions.len();
|
||||
eprintln!("Loaded {total} questions");
|
||||
|
||||
let profile = DatasetProfile::measure(&questions);
|
||||
eprintln!(
|
||||
"Corpus: {} variant — {:.1} sessions / {:.0} turns per question, \
|
||||
{:.1}% evidence-session density",
|
||||
profile.variant(),
|
||||
profile.mean_sessions,
|
||||
profile.mean_turns,
|
||||
profile.evidence_density * 100.0,
|
||||
);
|
||||
|
||||
// Build the embedding table once for the whole corpus, if asked for.
|
||||
let embeddings: Option<EmbeddingMap> = weights_dir
|
||||
.as_deref()
|
||||
.map(|dir| load_embeddings(dir, &questions));
|
||||
if embeddings.is_none() && weights_dir.is_some() {
|
||||
eprintln!("warning: --embeddings ignored (build with --features embeddings)");
|
||||
}
|
||||
|
||||
let modes: Vec<Mode> = if embeddings.is_some() {
|
||||
#[cfg(feature = "embeddings")]
|
||||
{
|
||||
if sweep {
|
||||
sweep_modes()
|
||||
} else {
|
||||
vec![BM25_ONLY, VECTOR_ONLY, HYBRID]
|
||||
}
|
||||
}
|
||||
#[cfg(not(feature = "embeddings"))]
|
||||
{
|
||||
vec![BM25_ONLY]
|
||||
}
|
||||
} else {
|
||||
if sweep {
|
||||
eprintln!("warning: --sweep needs --embeddings; running BM25 only");
|
||||
}
|
||||
vec![BM25_ONLY]
|
||||
};
|
||||
|
||||
for (mode_idx, mode) in modes.iter().enumerate() {
|
||||
eprintln!("[{}/{}] {}", mode_idx + 1, modes.len(), mode.label);
|
||||
run_mode(&questions, *mode, embeddings.as_ref(), &profile);
|
||||
}
|
||||
}
|
||||
|
||||
/// Load and encode the corpus. Returns `None` unless the `embeddings` feature
|
||||
/// is compiled in, so the flag degrades to a warning rather than a hard error.
|
||||
#[cfg(feature = "embeddings")]
|
||||
fn load_embeddings(dir: &str, questions: &[Question]) -> EmbeddingMap {
|
||||
let enc = embedder::Embedder::load(std::path::Path::new(dir))
|
||||
.unwrap_or_else(|e| panic!("failed to load embedder from {dir}: {e}"));
|
||||
let texts = questions.iter().flat_map(|q| {
|
||||
q.haystack_sessions
|
||||
.iter()
|
||||
.flatten()
|
||||
.map(|t| t.content.clone())
|
||||
.chain(std::iter::once(q.question.clone()))
|
||||
});
|
||||
enc.encode_unique(texts)
|
||||
.unwrap_or_else(|e| panic!("embedding failed: {e}"))
|
||||
}
|
||||
|
||||
#[cfg(not(feature = "embeddings"))]
|
||||
fn load_embeddings(_dir: &str, _questions: &[Question]) -> EmbeddingMap {
|
||||
EmbeddingMap::new()
|
||||
}
|
||||
|
||||
/// Evaluate every question under one retrieval mode and print its report.
|
||||
fn run_mode(
|
||||
questions: &[Question],
|
||||
mode: Mode,
|
||||
embeddings: Option<&EmbeddingMap>,
|
||||
profile: &DatasetProfile,
|
||||
) {
|
||||
let total = questions.len();
|
||||
let mut overall = Metrics::default();
|
||||
let mut by_type: HashMap<String, Metrics> = HashMap::new();
|
||||
|
||||
@@ -444,7 +826,7 @@ fn main() {
|
||||
eprint!("\r [{}/{}] evaluating...", i + 1, total);
|
||||
}
|
||||
|
||||
let result = evaluate_question(q, 10);
|
||||
let result = evaluate_question(q, 10, mode, embeddings);
|
||||
|
||||
let is_abs = q.question_type.ends_with("_abs");
|
||||
let base_type = if is_abs {
|
||||
@@ -509,5 +891,5 @@ fn main() {
|
||||
|
||||
eprintln!("\r [{total}/{total}] done. ");
|
||||
eprintln!();
|
||||
print_report(&overall, &by_type);
|
||||
print_report(&overall, &by_type, profile, mode);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,170 @@
|
||||
//! Optional MiniLM sentence embedder for the LongMemEval bench.
|
||||
//!
|
||||
//! Compiled only under the `embeddings` feature, so the default build of a
|
||||
//! project that prides itself on having no heavyweight dependencies stays
|
||||
//! exactly as it was. Without it the bench runs BM25-only, as it always has.
|
||||
//!
|
||||
//! Loads `sentence-transformers/all-MiniLM-L6-v2` — the same checkpoint
|
||||
//! omni-cortex uses — and produces 384-d mean-pooled, L2-normalised sentence
|
||||
//! embeddings, which is the published recipe for this model (mean over token
|
||||
//! states weighted by the attention mask, *not* the `[CLS]` pooler output).
|
||||
|
||||
use std::collections::HashMap;
|
||||
use std::path::Path;
|
||||
|
||||
use candle_core::{DType, Device, Tensor};
|
||||
use candle_nn::VarBuilder;
|
||||
use candle_transformers::models::bert::{BertModel, Config, HiddenAct};
|
||||
use tokenizers::Tokenizer;
|
||||
|
||||
/// Sequences encoded per forward pass. Larger batches amortise the transformer
|
||||
/// call; 64 keeps peak memory modest while still saturating a CPU.
|
||||
const BATCH: usize = 64;
|
||||
|
||||
/// A loaded MiniLM encoder.
|
||||
pub struct Embedder {
|
||||
model: BertModel,
|
||||
tokenizer: Tokenizer,
|
||||
device: Device,
|
||||
}
|
||||
|
||||
impl Embedder {
|
||||
/// Load from a directory holding `model.safetensors` and `tokenizer.json`.
|
||||
///
|
||||
/// `config.json` is read when present; otherwise the published MiniLM-L6-v2
|
||||
/// architecture constants are used, which are pinned rather than guessed.
|
||||
pub fn load(dir: &Path) -> Result<Self, Box<dyn std::error::Error>> {
|
||||
// CUDA when the feature is on and a device is actually present; the CPU
|
||||
// path is correct but roughly two orders of magnitude slower, which is
|
||||
// the difference between minutes and most of a day on the full haystack.
|
||||
let device = match Device::new_cuda(0) {
|
||||
Ok(d) => {
|
||||
eprintln!("Embedder: CUDA device 0");
|
||||
d
|
||||
}
|
||||
Err(e) => {
|
||||
// Loud, because the CPU path is correct but ~100x slower: the
|
||||
// full longmemeval_s haystack is minutes on a GPU and most of a
|
||||
// day on 8 cores. Silently falling back looks like a hang.
|
||||
eprintln!("Embedder: CPU — CUDA unavailable ({e})");
|
||||
eprintln!(
|
||||
" WARNING: CPU embedding is roughly two orders of magnitude slower.\n Expect minutes for longmemeval_oracle and many hours for the full\n longmemeval_s haystack. For the GPU path, rebuild with\n `--features embeddings-cuda` and make sure `nvcc` is on PATH\n (it ships in /usr/local/cuda/bin, which is often not exported)."
|
||||
);
|
||||
Device::Cpu
|
||||
}
|
||||
};
|
||||
let weights = dir.join("model.safetensors");
|
||||
let tok_path = dir.join("tokenizer.json");
|
||||
|
||||
let config: Config = match std::fs::read_to_string(dir.join("config.json")) {
|
||||
Ok(raw) => serde_json::from_str(&raw)?,
|
||||
Err(_) => Config {
|
||||
vocab_size: 30_522,
|
||||
hidden_size: 384,
|
||||
num_hidden_layers: 6,
|
||||
num_attention_heads: 12,
|
||||
intermediate_size: 1_536,
|
||||
hidden_act: HiddenAct::Gelu,
|
||||
hidden_dropout_prob: 0.0,
|
||||
max_position_embeddings: 512,
|
||||
type_vocab_size: 2,
|
||||
initializer_range: 0.02,
|
||||
layer_norm_eps: 1e-12,
|
||||
pad_token_id: 0,
|
||||
position_embedding_type: Default::default(),
|
||||
use_cache: false,
|
||||
classifier_dropout: None,
|
||||
model_type: None,
|
||||
},
|
||||
};
|
||||
|
||||
let vb = unsafe { VarBuilder::from_mmaped_safetensors(&[weights], DType::F32, &device)? };
|
||||
let model = BertModel::load(vb, &config)?;
|
||||
let tokenizer = Tokenizer::from_file(&tok_path).map_err(|e| e.to_string())?;
|
||||
|
||||
Ok(Self {
|
||||
model,
|
||||
tokenizer,
|
||||
device,
|
||||
})
|
||||
}
|
||||
|
||||
/// Encode `texts` into 384-d unit vectors, in order.
|
||||
fn encode_batch(&self, texts: &[&str]) -> Result<Vec<Vec<f32>>, Box<dyn std::error::Error>> {
|
||||
let mut tk = self.tokenizer.clone();
|
||||
let tk = tk
|
||||
.with_padding(Some(tokenizers::PaddingParams::default()))
|
||||
.with_truncation(Some(tokenizers::TruncationParams {
|
||||
max_length: 512,
|
||||
..Default::default()
|
||||
}))
|
||||
.map_err(|e| e.to_string())?;
|
||||
let encodings = tk
|
||||
.encode_batch(texts.to_vec(), true)
|
||||
.map_err(|e| e.to_string())?;
|
||||
|
||||
let ids: Vec<u32> = encodings
|
||||
.iter()
|
||||
.flat_map(|e| e.get_ids().to_vec())
|
||||
.collect();
|
||||
let mask: Vec<u32> = encodings
|
||||
.iter()
|
||||
.flat_map(|e| e.get_attention_mask().to_vec())
|
||||
.collect();
|
||||
let (b, l) = (encodings.len(), encodings[0].get_ids().len());
|
||||
|
||||
let ids = Tensor::from_vec(ids, (b, l), &self.device)?;
|
||||
let mask = Tensor::from_vec(mask, (b, l), &self.device)?;
|
||||
let type_ids = ids.zeros_like()?;
|
||||
|
||||
let hidden = self.model.forward(&ids, &type_ids, Some(&mask))?;
|
||||
|
||||
// Mean-pool over real tokens only: sum(hidden * mask) / sum(mask).
|
||||
let mask_f = mask.to_dtype(DType::F32)?.unsqueeze(2)?;
|
||||
let summed = hidden.broadcast_mul(&mask_f)?.sum(1)?;
|
||||
let counts = mask_f.sum(1)?.clamp(1e-9, f32::INFINITY)?;
|
||||
let pooled = summed.broadcast_div(&counts)?;
|
||||
|
||||
// L2-normalise so cosine similarity is a plain dot product.
|
||||
let norm = pooled
|
||||
.sqr()?
|
||||
.sum_keepdim(1)?
|
||||
.sqrt()?
|
||||
.clamp(1e-12, f32::INFINITY)?;
|
||||
let normed = pooled.broadcast_div(&norm)?;
|
||||
|
||||
Ok(normed.to_vec2::<f32>()?)
|
||||
}
|
||||
|
||||
/// Encode every distinct string in `texts` once, returning a lookup map.
|
||||
///
|
||||
/// LongMemEval's haystack sessions are drawn from a shared pool, so the same
|
||||
/// turn text recurs across many questions. Deduplicating before encoding is
|
||||
/// the difference between encoding the corpus once and encoding it per
|
||||
/// question.
|
||||
pub fn encode_unique(
|
||||
&self,
|
||||
texts: impl IntoIterator<Item = String>,
|
||||
) -> Result<HashMap<String, Vec<f32>>, Box<dyn std::error::Error>> {
|
||||
let mut unique: Vec<String> = texts.into_iter().collect();
|
||||
unique.sort_unstable();
|
||||
unique.dedup();
|
||||
|
||||
let total = unique.len();
|
||||
eprintln!("Embedding {total} unique texts with MiniLM (batch {BATCH})...");
|
||||
|
||||
let mut out = HashMap::with_capacity(total);
|
||||
for (n, chunk) in unique.chunks(BATCH).enumerate() {
|
||||
let refs: Vec<&str> = chunk.iter().map(String::as_str).collect();
|
||||
let vecs = self.encode_batch(&refs)?;
|
||||
for (text, v) in chunk.iter().zip(vecs) {
|
||||
out.insert(text.clone(), v);
|
||||
}
|
||||
if n % 50 == 0 {
|
||||
eprint!("\r [{}/{}] embedded...", (n * BATCH).min(total), total);
|
||||
}
|
||||
}
|
||||
eprintln!("\r [{total}/{total}] embedded. ");
|
||||
Ok(out)
|
||||
}
|
||||
}
|
||||
@@ -17,4 +17,4 @@ path = "src/main.rs"
|
||||
clawhdf5-agent = { path = "../clawhdf5-agent", version = "2.1.0" }
|
||||
clap = { version = "4", features = ["derive", "env"] }
|
||||
serde_json = "1"
|
||||
serde = { version = "1", features = ["derive"] }
|
||||
serde = { workspace = true }
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
name = "clawhdf5-filters"
|
||||
version = "2.1.0"
|
||||
edition = "2024"
|
||||
description = "Filter and compression pipeline for rustyhdf5"
|
||||
description = "Filter and compression pipeline for clawhdf5"
|
||||
license = "MIT"
|
||||
repository = "https://github.com/redclawsystems/clawhdf5"
|
||||
readme = "README.md"
|
||||
@@ -14,7 +14,7 @@ flate2 = { version = "1", default-features = false, features = ["rust_backend"]
|
||||
miniz_oxide = "0.8"
|
||||
|
||||
[dev-dependencies]
|
||||
criterion = { version = "0.5", features = ["html_reports"] }
|
||||
criterion = { workspace = true }
|
||||
|
||||
[[bench]]
|
||||
name = "deflate_bench"
|
||||
|
||||
@@ -24,7 +24,7 @@ pco = { version = "1.0", optional = true }
|
||||
|
||||
[dev-dependencies]
|
||||
serde_json = "1"
|
||||
criterion = { version = "0.5", features = ["html_reports"] }
|
||||
criterion = { workspace = true }
|
||||
clawhdf5-derive = { path = "../clawhdf5-derive", version = "2.1.0" }
|
||||
|
||||
[[bench]]
|
||||
|
||||
@@ -14,6 +14,9 @@ libfuzzer-sys = "0.4"
|
||||
path = ".."
|
||||
features = ["std", "checksum", "deflate"]
|
||||
|
||||
[dependencies.clawhdf5]
|
||||
path = "../../clawhdf5"
|
||||
|
||||
[workspace]
|
||||
members = ["."]
|
||||
|
||||
@@ -56,3 +59,8 @@ doc = false
|
||||
name = "fuzz_full_file"
|
||||
path = "fuzz_targets/fuzz_full_file.rs"
|
||||
doc = false
|
||||
|
||||
[[bin]]
|
||||
name = "fuzz_dataset_read"
|
||||
path = "fuzz_targets/fuzz_dataset_read.rs"
|
||||
doc = false
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# Fuzz Testing for rustyhdf5-format
|
||||
# Fuzz Testing for clawhdf5-format
|
||||
|
||||
Uses [cargo-fuzz](https://github.com/rust-fuzz/cargo-fuzz) (libFuzzer) to test parser robustness against malformed inputs.
|
||||
|
||||
@@ -21,13 +21,14 @@ rustup toolchain install nightly
|
||||
| `fuzz_btree_v2` | `BTreeV2Header::parse` | B-tree v2 header parsing |
|
||||
| `fuzz_filter_pipeline` | `FilterPipeline::parse` | Filter pipeline messages (v1/v2) |
|
||||
| `fuzz_full_file` | signature + superblock + root group | End-to-end file parsing chain |
|
||||
| `fuzz_dataset_read` | `Dataset::read_*` (via `clawhdf5`) | Walks every dataset in the parsed file and exercises the contiguous/chunked/compact raw-data read paths (`chunked_read.rs`, `data_read.rs`) that `fuzz_full_file` doesn't reach |
|
||||
|
||||
## Running
|
||||
|
||||
Run a single target (runs indefinitely until stopped or a crash is found):
|
||||
|
||||
```bash
|
||||
cd crates/rustyhdf5-format
|
||||
cd crates/clawhdf5-format
|
||||
cargo +nightly fuzz run fuzz_datatype
|
||||
```
|
||||
|
||||
@@ -41,12 +42,20 @@ Run all targets for 30 seconds each:
|
||||
|
||||
```bash
|
||||
for target in fuzz_superblock fuzz_object_header fuzz_datatype fuzz_dataspace \
|
||||
fuzz_fractal_heap fuzz_btree_v2 fuzz_filter_pipeline fuzz_full_file; do
|
||||
fuzz_fractal_heap fuzz_btree_v2 fuzz_filter_pipeline fuzz_full_file \
|
||||
fuzz_dataset_read; do
|
||||
echo "=== $target ==="
|
||||
cargo +nightly fuzz run "$target" -- -max_total_time=30 -max_len=4096
|
||||
done
|
||||
```
|
||||
|
||||
## CI
|
||||
|
||||
These targets are **not** run in CI (`.gitea/workflows/ci.yml`) — cargo-fuzz
|
||||
requires nightly and each meaningful run takes minutes, which doesn't fit a
|
||||
per-PR gate. Run them manually on a schedule (e.g. before a release, or after
|
||||
touching parser code) instead.
|
||||
|
||||
## Reproducing Crashes
|
||||
|
||||
If a crash is found, the input is saved to `fuzz/artifacts/<target>/`. Reproduce with:
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
BIN
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,45 @@
|
||||
#![no_main]
|
||||
use libfuzzer_sys::fuzz_target;
|
||||
|
||||
const MAX_WALK_DEPTH: usize = 16;
|
||||
|
||||
/// Walk groups/datasets from `group`, exercising every dataset-reading code
|
||||
/// path reachable through the public API (contiguous/chunked/compact raw
|
||||
/// reads via `chunked_read.rs`/`data_read.rs`). Depth-limited independently
|
||||
/// of any parser-level recursion guard, since this is fuzz-harness
|
||||
/// bookkeeping, not something under test.
|
||||
fn walk_group(group: &clawhdf5::Group, depth: usize) {
|
||||
if depth > MAX_WALK_DEPTH {
|
||||
return;
|
||||
}
|
||||
if let Ok(names) = group.datasets() {
|
||||
for name in names {
|
||||
if let Ok(dataset) = group.dataset(&name) {
|
||||
let _ = dataset.shape();
|
||||
let _ = dataset.max_dimensions();
|
||||
let _ = dataset.dtype();
|
||||
let _ = dataset.read_raw_ref();
|
||||
let _ = dataset.read_f64();
|
||||
let _ = dataset.read_f32();
|
||||
let _ = dataset.read_i32();
|
||||
let _ = dataset.read_i64();
|
||||
let _ = dataset.read_u64();
|
||||
let _ = dataset.read_string();
|
||||
}
|
||||
}
|
||||
}
|
||||
if let Ok(names) = group.groups() {
|
||||
for name in names {
|
||||
if let Ok(subgroup) = group.group(&name) {
|
||||
walk_group(&subgroup, depth + 1);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fuzz_target!(|data: &[u8]| {
|
||||
let Ok(file) = clawhdf5::File::from_bytes(data.to_vec()) else {
|
||||
return;
|
||||
};
|
||||
walk_group(&file.root(), 0);
|
||||
});
|
||||
@@ -24,6 +24,21 @@ pub struct BTreeV1Node {
|
||||
pub children: Vec<u64>,
|
||||
}
|
||||
|
||||
/// Checks that `[offset, offset + needed)` fits within `data`, guarding the
|
||||
/// addition against `usize` overflow from a crafted near-`usize::MAX` offset.
|
||||
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> {
|
||||
if offset
|
||||
.checked_add(needed)
|
||||
.is_none_or(|end| end > data.len())
|
||||
{
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: offset.saturating_add(needed),
|
||||
available: data.len(),
|
||||
});
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
|
||||
let s = size as usize;
|
||||
if pos.checked_add(s).is_none_or(|end| end > data.len()) {
|
||||
@@ -45,7 +60,7 @@ fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
|
||||
|
||||
fn is_undefined(data: &[u8], pos: usize, size: u8) -> bool {
|
||||
let s = size as usize;
|
||||
if pos + s > data.len() {
|
||||
if ensure_len(data, pos, s).is_err() {
|
||||
return false;
|
||||
}
|
||||
data[pos..pos + s].iter().all(|&b| b == 0xFF)
|
||||
@@ -65,12 +80,7 @@ impl BTreeV1Node {
|
||||
// + left_sibling(offset_size) + right_sibling(offset_size)
|
||||
let os = offset_size as usize;
|
||||
let header_size = 8 + os * 2;
|
||||
if offset + header_size > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: offset + header_size,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, offset, header_size)?;
|
||||
|
||||
if &file_data[offset..offset + 4] != b"TREE" {
|
||||
return Err(FormatError::InvalidBTreeSignature);
|
||||
@@ -99,12 +109,7 @@ impl BTreeV1Node {
|
||||
let eu = entries_used as usize;
|
||||
let key_size = os; // For type 0, key = offset_size
|
||||
let needed = eu * (key_size + os) + key_size; // eu children + (eu+1) keys
|
||||
if pos + needed > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: pos + needed,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, pos, needed)?;
|
||||
|
||||
let mut keys = Vec::with_capacity(eu + 1);
|
||||
let mut children = Vec::with_capacity(eu);
|
||||
@@ -241,6 +246,16 @@ mod tests {
|
||||
assert_eq!(node.right_sibling, None);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_near_usize_max_offset_rejected_without_overflow() {
|
||||
let data = build_btree_node(0, 0, &[0, 5, 10], &[0x100, 0x200], None, None, 8);
|
||||
let result = BTreeV1Node::parse(&data, usize::MAX - 4, 8, 8);
|
||||
assert!(
|
||||
matches!(result, Err(FormatError::UnexpectedEof { .. })),
|
||||
"expected a clean UnexpectedEof, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_with_siblings_none() {
|
||||
let data = build_btree_node(0, 0, &[0, 8], &[0x300], None, None, 8);
|
||||
|
||||
@@ -61,12 +61,7 @@ fn decompress_all_chunks(
|
||||
for chunk_info in chunks {
|
||||
let c_addr = chunk_info.address as usize;
|
||||
let size = chunk_info.chunk_size as usize;
|
||||
if c_addr + size > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: c_addr + size,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, c_addr, size)?;
|
||||
let raw_chunk = &file_data[c_addr..c_addr + size];
|
||||
|
||||
let decompressed = if let Some(pl) = pipeline {
|
||||
@@ -122,6 +117,21 @@ pub struct ChunkInfo {
|
||||
pub address: u64,
|
||||
}
|
||||
|
||||
/// Checks that `[offset, offset + needed)` fits within `data`, guarding the
|
||||
/// addition against `usize` overflow from a crafted near-`usize::MAX` offset.
|
||||
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> {
|
||||
if offset
|
||||
.checked_add(needed)
|
||||
.is_none_or(|end| end > data.len())
|
||||
{
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: offset.saturating_add(needed),
|
||||
available: data.len(),
|
||||
});
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
|
||||
let s = size as usize;
|
||||
if pos.checked_add(s).is_none_or(|end| end > data.len()) {
|
||||
@@ -150,19 +160,33 @@ pub fn collect_chunk_info(
|
||||
btree_address: u64,
|
||||
ndims: usize,
|
||||
offset_size: u8,
|
||||
_length_size: u8,
|
||||
length_size: u8,
|
||||
) -> Result<Vec<ChunkInfo>, FormatError> {
|
||||
collect_chunk_info_inner(file_data, btree_address, ndims, offset_size, length_size, 0)
|
||||
}
|
||||
|
||||
/// Maximum recursion depth for chunk B-tree traversal (malformed/cyclic data
|
||||
/// protection), matching `btree_v1.rs`'s `MAX_BTREE_DEPTH`.
|
||||
const MAX_CHUNK_BTREE_DEPTH: usize = 64;
|
||||
|
||||
fn collect_chunk_info_inner(
|
||||
file_data: &[u8],
|
||||
btree_address: u64,
|
||||
ndims: usize,
|
||||
offset_size: u8,
|
||||
_length_size: u8,
|
||||
depth: usize,
|
||||
) -> Result<Vec<ChunkInfo>, FormatError> {
|
||||
if depth > MAX_CHUNK_BTREE_DEPTH {
|
||||
return Err(FormatError::NestingDepthExceeded);
|
||||
}
|
||||
|
||||
let offset = btree_address as usize;
|
||||
let os = offset_size as usize;
|
||||
|
||||
// Parse B-tree v1 header
|
||||
let header_size = 8 + os * 2;
|
||||
if offset + header_size > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: offset + header_size,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, offset, header_size)?;
|
||||
|
||||
if &file_data[offset..offset + 4] != b"TREE" {
|
||||
return Err(FormatError::InvalidBTreeSignature);
|
||||
@@ -185,12 +209,7 @@ pub fn collect_chunk_info(
|
||||
// Leaf node: keys and children interleaved
|
||||
// key[0], child[0], key[1], child[1], ..., key[N-1], child[N-1], key[N]
|
||||
let needed = entries_used * (key_size + os) + key_size;
|
||||
if pos + needed > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: pos + needed,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, pos, needed)?;
|
||||
|
||||
let mut chunks = Vec::with_capacity(entries_used);
|
||||
for _ in 0..entries_used {
|
||||
@@ -231,12 +250,7 @@ pub fn collect_chunk_info(
|
||||
} else {
|
||||
// Internal node: recurse into children
|
||||
let needed = entries_used * (key_size + os) + key_size;
|
||||
if pos + needed > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: pos + needed,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, pos, needed)?;
|
||||
|
||||
let mut child_addrs = Vec::with_capacity(entries_used);
|
||||
for _ in 0..entries_used {
|
||||
@@ -248,8 +262,14 @@ pub fn collect_chunk_info(
|
||||
|
||||
let mut all_chunks = Vec::new();
|
||||
for child_addr in child_addrs {
|
||||
let child_chunks =
|
||||
collect_chunk_info(file_data, child_addr, ndims, offset_size, _length_size)?;
|
||||
let child_chunks = collect_chunk_info_inner(
|
||||
file_data,
|
||||
child_addr,
|
||||
ndims,
|
||||
offset_size,
|
||||
_length_size,
|
||||
depth + 1,
|
||||
)?;
|
||||
all_chunks.extend(child_chunks);
|
||||
}
|
||||
Ok(all_chunks)
|
||||
@@ -347,7 +367,9 @@ pub fn read_chunked_data(
|
||||
|
||||
// Both v3 and v4 include element size as last dim (rank+1)
|
||||
let ndims = chunk_dimensions.len();
|
||||
let rank = ndims - 1;
|
||||
let rank = ndims
|
||||
.checked_sub(1)
|
||||
.ok_or_else(|| FormatError::ChunkedReadError("chunked layout has no dimensions".into()))?;
|
||||
let chunk_dims: Vec<usize> = chunk_dimensions[..rank]
|
||||
.iter()
|
||||
.map(|&d| d as usize)
|
||||
@@ -386,24 +408,24 @@ pub fn read_chunked_data(
|
||||
}
|
||||
(4, Some(2)) => {
|
||||
// Implicit index — use spatial chunk dims only
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
generate_implicit_chunks(
|
||||
addr,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
)
|
||||
}
|
||||
(4, Some(3)) => {
|
||||
// Fixed Array — use spatial chunk dims only
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header =
|
||||
FixedArrayHeader::parse(file_data, addr as usize, offset_size, length_size)?;
|
||||
read_fixed_array_chunks(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
@@ -411,14 +433,14 @@ pub fn read_chunked_data(
|
||||
}
|
||||
(4, Some(4)) => {
|
||||
// Extensible Array — use spatial chunk dims only
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header =
|
||||
ExtensibleArrayHeader::parse(file_data, addr as usize, offset_size, length_size)?;
|
||||
read_extensible_array_chunks(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
@@ -461,12 +483,7 @@ pub fn read_chunked_data(
|
||||
|
||||
let c_addr = chunk_info.address as usize;
|
||||
let size = chunk_info.chunk_size as usize;
|
||||
if c_addr + size > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: c_addr + size,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, c_addr, size)?;
|
||||
let chunk_data = &file_data[c_addr..c_addr + size];
|
||||
|
||||
if rank == 0 {
|
||||
@@ -579,7 +596,9 @@ pub fn read_chunked_data_cached(
|
||||
|
||||
let elem_size = datatype.type_size() as usize;
|
||||
let ndims = chunk_dimensions.len();
|
||||
let rank = ndims - 1;
|
||||
let rank = ndims
|
||||
.checked_sub(1)
|
||||
.ok_or_else(|| FormatError::ChunkedReadError("chunked layout has no dimensions".into()))?;
|
||||
let chunk_dims: Vec<usize> = chunk_dimensions[..rank]
|
||||
.iter()
|
||||
.map(|&d| d as usize)
|
||||
@@ -618,30 +637,30 @@ pub fn read_chunked_data_cached(
|
||||
}]
|
||||
}
|
||||
(4, Some(2)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
generate_implicit_chunks(
|
||||
addr,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
)
|
||||
}
|
||||
(4, Some(3)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header =
|
||||
FixedArrayHeader::parse(file_data, addr as usize, offset_size, length_size)?;
|
||||
read_fixed_array_chunks(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
)?
|
||||
}
|
||||
(4, Some(4)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header = ExtensibleArrayHeader::parse(
|
||||
file_data,
|
||||
addr as usize,
|
||||
@@ -652,7 +671,7 @@ pub fn read_chunked_data_cached(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
@@ -697,12 +716,7 @@ pub fn read_chunked_data_cached(
|
||||
// Decompress from file
|
||||
let c_addr = chunk_info.address as usize;
|
||||
let size = chunk_info.chunk_size as usize;
|
||||
if c_addr + size > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: c_addr + size,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, c_addr, size)?;
|
||||
let raw_chunk = &file_data[c_addr..c_addr + size];
|
||||
let dec = if let Some(pl) = pipeline {
|
||||
if chunk_info.filter_mask == 0 {
|
||||
@@ -935,7 +949,9 @@ pub fn read_chunked_data_sweep(
|
||||
|
||||
let elem_size = datatype.type_size() as usize;
|
||||
let ndims = chunk_dimensions.len();
|
||||
let rank = ndims - 1;
|
||||
let rank = ndims
|
||||
.checked_sub(1)
|
||||
.ok_or_else(|| FormatError::ChunkedReadError("chunked layout has no dimensions".into()))?;
|
||||
let chunk_dims: Vec<usize> = chunk_dimensions[..rank]
|
||||
.iter()
|
||||
.map(|&d| d as usize)
|
||||
@@ -974,30 +990,30 @@ pub fn read_chunked_data_sweep(
|
||||
}]
|
||||
}
|
||||
(4, Some(2)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
generate_implicit_chunks(
|
||||
addr,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
)
|
||||
}
|
||||
(4, Some(3)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header =
|
||||
FixedArrayHeader::parse(file_data, addr as usize, offset_size, length_size)?;
|
||||
read_fixed_array_chunks(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
)?
|
||||
}
|
||||
(4, Some(4)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header = ExtensibleArrayHeader::parse(
|
||||
file_data,
|
||||
addr as usize,
|
||||
@@ -1008,7 +1024,7 @@ pub fn read_chunked_data_sweep(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
@@ -1062,12 +1078,7 @@ pub fn read_chunked_data_sweep(
|
||||
// Decompress from file
|
||||
let c_addr = chunk_info.address as usize;
|
||||
let size = chunk_info.chunk_size as usize;
|
||||
if c_addr + size > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: c_addr + size,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, c_addr, size)?;
|
||||
let raw_chunk = &file_data[c_addr..c_addr + size];
|
||||
let dec = if let Some(pl) = pipeline {
|
||||
if chunk_info.filter_mask == 0 {
|
||||
@@ -1161,7 +1172,9 @@ pub fn read_chunked_data_indexed(
|
||||
|
||||
let elem_size = datatype.type_size() as usize;
|
||||
let ndims = chunk_dimensions.len();
|
||||
let rank = ndims - 1;
|
||||
let rank = ndims
|
||||
.checked_sub(1)
|
||||
.ok_or_else(|| FormatError::ChunkedReadError("chunked layout has no dimensions".into()))?;
|
||||
let chunk_dims: Vec<usize> = chunk_dimensions[..rank]
|
||||
.iter()
|
||||
.map(|&d| d as usize)
|
||||
@@ -1200,30 +1213,30 @@ pub fn read_chunked_data_indexed(
|
||||
}]
|
||||
}
|
||||
(4, Some(2)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
generate_implicit_chunks(
|
||||
addr,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
)
|
||||
}
|
||||
(4, Some(3)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header =
|
||||
FixedArrayHeader::parse(file_data, addr as usize, offset_size, length_size)?;
|
||||
read_fixed_array_chunks(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
)?
|
||||
}
|
||||
(4, Some(4)) => {
|
||||
let spatial_chunk_dims: Vec<u32> = chunk_dimensions[..rank].to_vec();
|
||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
||||
let header = ExtensibleArrayHeader::parse(
|
||||
file_data,
|
||||
addr as usize,
|
||||
@@ -1234,7 +1247,7 @@ pub fn read_chunked_data_indexed(
|
||||
file_data,
|
||||
&header,
|
||||
&dataspace.dimensions,
|
||||
&spatial_chunk_dims,
|
||||
spatial_chunk_dims,
|
||||
elem_size as u32,
|
||||
offset_size,
|
||||
length_size,
|
||||
@@ -1278,12 +1291,7 @@ pub fn read_chunked_data_indexed(
|
||||
} else {
|
||||
let c_addr = *file_offset as usize;
|
||||
let size = *file_size as usize;
|
||||
if c_addr + size > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: c_addr + size,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, c_addr, size)?;
|
||||
let raw_chunk = &file_data[c_addr..c_addr + size];
|
||||
let decompressed = if let Some(pl) = pipeline {
|
||||
if *filter_mask == 0 {
|
||||
@@ -1331,9 +1339,18 @@ fn copy_chunk_to_output(
|
||||
// Fast path for 1-D: single contiguous copy per chunk
|
||||
let global_start = chunk_offsets[0];
|
||||
let copy_len = chunk_dims[0].min(ds_dims[0].saturating_sub(global_start));
|
||||
let src_bytes = copy_len * elem_size;
|
||||
let dst_start = global_start * elem_size;
|
||||
if src_bytes > 0 && dst_start + src_bytes <= output.len() && src_bytes <= chunk_data.len() {
|
||||
let (Some(src_bytes), Some(dst_start)) = (
|
||||
copy_len.checked_mul(elem_size),
|
||||
global_start.checked_mul(elem_size),
|
||||
) else {
|
||||
return;
|
||||
};
|
||||
if src_bytes > 0
|
||||
&& dst_start
|
||||
.checked_add(src_bytes)
|
||||
.is_some_and(|end| end <= output.len())
|
||||
&& src_bytes <= chunk_data.len()
|
||||
{
|
||||
output[dst_start..dst_start + src_bytes].copy_from_slice(&chunk_data[..src_bytes]);
|
||||
}
|
||||
return;
|
||||
@@ -1343,19 +1360,29 @@ fn copy_chunk_to_output(
|
||||
let inner_dim = rank - 1;
|
||||
let inner_chunk_len =
|
||||
chunk_dims[inner_dim].min(ds_dims[inner_dim].saturating_sub(chunk_offsets[inner_dim]));
|
||||
let row_bytes = inner_chunk_len * elem_size;
|
||||
let Some(row_bytes) = inner_chunk_len.checked_mul(elem_size) else {
|
||||
return;
|
||||
};
|
||||
|
||||
if row_bytes == 0 {
|
||||
return;
|
||||
}
|
||||
|
||||
// Number of rows = product of all outer chunk dimensions
|
||||
let outer_count: usize = chunk_dims[..inner_dim].iter().product();
|
||||
let Some(outer_count) = chunk_dims[..inner_dim]
|
||||
.iter()
|
||||
.try_fold(1usize, |acc, &d| acc.checked_mul(d))
|
||||
else {
|
||||
return;
|
||||
};
|
||||
|
||||
// Outer strides for iterating chunk-local coordinates
|
||||
let mut outer_strides = vec![1usize; inner_dim];
|
||||
for i in (0..inner_dim.saturating_sub(1)).rev() {
|
||||
outer_strides[i] = outer_strides[i + 1] * chunk_dims[i + 1];
|
||||
let Some(stride) = outer_strides[i + 1].checked_mul(chunk_dims[i + 1]) else {
|
||||
return;
|
||||
};
|
||||
outer_strides[i] = stride;
|
||||
}
|
||||
|
||||
for outer_idx in 0..outer_count {
|
||||
@@ -1375,13 +1402,29 @@ fn copy_chunk_to_output(
|
||||
remaining %= outer_strides[d];
|
||||
}
|
||||
|
||||
let global_coord = chunk_offsets[d] + coord_in_chunk;
|
||||
let Some(global_coord) = chunk_offsets[d].checked_add(coord_in_chunk) else {
|
||||
out_of_bounds = true;
|
||||
break;
|
||||
};
|
||||
if global_coord >= ds_dims[d] {
|
||||
out_of_bounds = true;
|
||||
break;
|
||||
}
|
||||
ds_flat += global_coord * ds_strides[d];
|
||||
src_flat += coord_in_chunk * chunk_strides[d];
|
||||
let (Some(ds_term), Some(src_term)) = (
|
||||
global_coord.checked_mul(ds_strides[d]),
|
||||
coord_in_chunk.checked_mul(chunk_strides[d]),
|
||||
) else {
|
||||
out_of_bounds = true;
|
||||
break;
|
||||
};
|
||||
let (Some(new_ds_flat), Some(new_src_flat)) =
|
||||
(ds_flat.checked_add(ds_term), src_flat.checked_add(src_term))
|
||||
else {
|
||||
out_of_bounds = true;
|
||||
break;
|
||||
};
|
||||
ds_flat = new_ds_flat;
|
||||
src_flat = new_src_flat;
|
||||
}
|
||||
|
||||
if out_of_bounds {
|
||||
@@ -1389,12 +1432,27 @@ fn copy_chunk_to_output(
|
||||
}
|
||||
|
||||
// Add innermost dimension offset
|
||||
ds_flat += chunk_offsets[inner_dim] * ds_strides[inner_dim];
|
||||
let Some(inner_term) = chunk_offsets[inner_dim].checked_mul(ds_strides[inner_dim]) else {
|
||||
continue;
|
||||
};
|
||||
let Some(ds_flat) = ds_flat.checked_add(inner_term) else {
|
||||
continue;
|
||||
};
|
||||
|
||||
let src_start = src_flat * elem_size;
|
||||
let dst_start = ds_flat * elem_size;
|
||||
let (Some(src_start), Some(dst_start)) = (
|
||||
src_flat.checked_mul(elem_size),
|
||||
ds_flat.checked_mul(elem_size),
|
||||
) else {
|
||||
continue;
|
||||
};
|
||||
|
||||
if src_start + row_bytes <= chunk_data.len() && dst_start + row_bytes <= output.len() {
|
||||
let fits = src_start
|
||||
.checked_add(row_bytes)
|
||||
.is_some_and(|end| end <= chunk_data.len())
|
||||
&& dst_start
|
||||
.checked_add(row_bytes)
|
||||
.is_some_and(|end| end <= output.len());
|
||||
if fits {
|
||||
output[dst_start..dst_start + row_bytes]
|
||||
.copy_from_slice(&chunk_data[src_start..src_start + row_bytes]);
|
||||
}
|
||||
@@ -1639,6 +1697,82 @@ mod tests {
|
||||
(file_data, layout, dataspace)
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_chunked_data_rejects_zero_dim_chunk_layout() {
|
||||
// Found by fuzzing: chunk_dimensions.len() == 0 caused `ndims - 1` to
|
||||
// underflow. A malformed/degenerate chunked layout must error cleanly.
|
||||
let layout = DataLayout::Chunked {
|
||||
chunk_dimensions: vec![],
|
||||
btree_address: Some(0),
|
||||
version: 3,
|
||||
chunk_index_type: None,
|
||||
single_chunk_filtered_size: None,
|
||||
single_chunk_filter_mask: None,
|
||||
};
|
||||
let dataspace = Dataspace {
|
||||
space_type: DataspaceType::Simple,
|
||||
rank: 1,
|
||||
dimensions: vec![10],
|
||||
max_dimensions: None,
|
||||
};
|
||||
let datatype = make_f64_type();
|
||||
let file_data = vec![0u8; 64];
|
||||
let result = read_chunked_data(&file_data, &layout, &dataspace, &datatype, None, 8, 8);
|
||||
assert!(
|
||||
matches!(result, Err(FormatError::ChunkedReadError(_))),
|
||||
"expected a clean ChunkedReadError, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn copy_chunk_to_output_1d_rejects_overflowing_offset_without_panicking() {
|
||||
// Found by fuzzing: `global_start * elem_size` overflowed for a
|
||||
// crafted large chunk offset.
|
||||
let chunk_data = vec![1u8; 16];
|
||||
let mut output = vec![0u8; 16];
|
||||
let chunk_offsets = [usize::MAX - 1];
|
||||
let chunk_dims = [1usize];
|
||||
let ds_dims = [usize::MAX];
|
||||
let ds_strides = [1usize];
|
||||
let chunk_strides = [1usize];
|
||||
copy_chunk_to_output(
|
||||
&chunk_data,
|
||||
&mut output,
|
||||
&chunk_offsets,
|
||||
&chunk_dims,
|
||||
&ds_dims,
|
||||
&ds_strides,
|
||||
&chunk_strides,
|
||||
8,
|
||||
1,
|
||||
);
|
||||
// No panic; the out-of-range write was skipped, output left untouched.
|
||||
assert_eq!(output, vec![0u8; 16]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn copy_chunk_to_output_nd_rejects_overflowing_offset_without_panicking() {
|
||||
let chunk_data = vec![1u8; 16];
|
||||
let mut output = vec![0u8; 16];
|
||||
let chunk_offsets = [usize::MAX - 1, 0];
|
||||
let chunk_dims = [1usize, 1usize];
|
||||
let ds_dims = [usize::MAX, usize::MAX];
|
||||
let ds_strides = [1usize, 1usize];
|
||||
let chunk_strides = [1usize, 1usize];
|
||||
copy_chunk_to_output(
|
||||
&chunk_data,
|
||||
&mut output,
|
||||
&chunk_offsets,
|
||||
&chunk_dims,
|
||||
&ds_dims,
|
||||
&ds_strides,
|
||||
&chunk_strides,
|
||||
8,
|
||||
2,
|
||||
);
|
||||
assert_eq!(output, vec![0u8; 16]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_1d_two_chunks_no_compression() {
|
||||
let values: Vec<f64> = (0..20).map(|i| i as f64).collect();
|
||||
@@ -1851,6 +1985,54 @@ mod tests {
|
||||
assert_eq!(err, FormatError::InvalidBTreeNodeType(0));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn collect_chunk_info_rejects_near_usize_max_offset() {
|
||||
let file_data = vec![0u8; 64];
|
||||
let result = collect_chunk_info(&file_data, u64::MAX - 4, 2, 8, 8);
|
||||
assert!(
|
||||
matches!(result, Err(FormatError::UnexpectedEof { .. })),
|
||||
"expected a clean UnexpectedEof, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn collect_chunk_info_rejects_self_referencing_internal_node() {
|
||||
// A type-1 internal node (level 1) whose single child address points
|
||||
// back to itself: an infinite-recursion / cyclic B-tree attack.
|
||||
let ndims = 2;
|
||||
let os: u8 = 8;
|
||||
let mut buf = Vec::new();
|
||||
buf.extend_from_slice(b"TREE");
|
||||
buf.push(1); // node_type = 1 (raw data chunks)
|
||||
buf.push(1); // node_level = 1 (internal)
|
||||
buf.extend_from_slice(&1u16.to_le_bytes()); // entries_used = 1
|
||||
write_offset(&mut buf, u64::MAX, os); // left sibling undefined
|
||||
write_offset(&mut buf, u64::MAX, os); // right sibling undefined
|
||||
// key[0]: chunk_size(4) + filter_mask(4) + ndims offsets
|
||||
buf.extend_from_slice(&0u32.to_le_bytes());
|
||||
buf.extend_from_slice(&0u32.to_le_bytes());
|
||||
for _ in 0..ndims {
|
||||
write_offset(&mut buf, 0, os);
|
||||
}
|
||||
// child[0]: points back to offset 0 (this same node) — cyclic.
|
||||
write_offset(&mut buf, 0, os);
|
||||
// final key
|
||||
buf.extend_from_slice(&0u32.to_le_bytes());
|
||||
buf.extend_from_slice(&0u32.to_le_bytes());
|
||||
for _ in 0..ndims {
|
||||
write_offset(&mut buf, u64::MAX, os);
|
||||
}
|
||||
|
||||
let mut file_data = vec![0u8; 256];
|
||||
file_data[..buf.len()].copy_from_slice(&buf);
|
||||
|
||||
let result = collect_chunk_info(&file_data, 0, ndims, os, os);
|
||||
assert!(
|
||||
matches!(result, Err(FormatError::NestingDepthExceeded)),
|
||||
"expected a clean NestingDepthExceeded, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
// --- Implicit chunk generation tests ---
|
||||
|
||||
#[test]
|
||||
|
||||
@@ -17,6 +17,21 @@ use crate::datatype::{Datatype, DatatypeByteOrder};
|
||||
use crate::error::FormatError;
|
||||
use crate::filter_pipeline::FilterPipeline;
|
||||
|
||||
/// Checks that `[offset, offset + needed)` fits within `data`, guarding the
|
||||
/// addition against `usize` overflow from a crafted near-`usize::MAX` offset.
|
||||
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> {
|
||||
if offset
|
||||
.checked_add(needed)
|
||||
.is_none_or(|end| end > data.len())
|
||||
{
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: offset.saturating_add(needed),
|
||||
available: data.len(),
|
||||
});
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Zero-copy read of contiguous raw data, returning a borrowed slice.
|
||||
///
|
||||
/// For contiguous layouts, returns a direct `&[u8]` slice into `file_data`.
|
||||
@@ -47,12 +62,7 @@ pub fn read_raw_data_zerocopy<'a>(
|
||||
actual: sz,
|
||||
});
|
||||
}
|
||||
if addr + sz > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: addr + sz,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, addr, sz)?;
|
||||
Ok(Some(&file_data[addr..addr + sz]))
|
||||
}
|
||||
_ => Ok(None),
|
||||
@@ -169,12 +179,7 @@ fn read_raw_data_full_impl(
|
||||
actual: sz,
|
||||
});
|
||||
}
|
||||
if addr + sz > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: addr + sz,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, addr, sz)?;
|
||||
Ok(file_data[addr..addr + sz].to_vec())
|
||||
}
|
||||
DataLayout::Chunked { .. } => read_chunked_data(
|
||||
@@ -1218,6 +1223,15 @@ pub fn read_compound_fields(
|
||||
for m in members {
|
||||
let field_size = m.datatype.type_size() as usize;
|
||||
let offset = m.byte_offset as usize;
|
||||
if offset
|
||||
.checked_add(field_size)
|
||||
.is_none_or(|end| end > elem_size)
|
||||
{
|
||||
return Err(FormatError::Overflow(format!(
|
||||
"compound member '{}': byte_offset({offset}) + field_size({field_size}) exceeds element size({elem_size})",
|
||||
m.name
|
||||
)));
|
||||
}
|
||||
let mut field_raw = Vec::with_capacity(count * field_size);
|
||||
for i in 0..count {
|
||||
let elem_start = i * elem_size + offset;
|
||||
@@ -2116,6 +2130,43 @@ mod tests {
|
||||
assert_eq!(id_vals, vec![10, 20]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_compound_rejects_byte_offset_overrun() {
|
||||
use crate::datatype::CompoundMember;
|
||||
// Compound declares size=8, but the member's byte_offset(4) + its
|
||||
// field_size(8, f64) = 12 > 8 — a crafted out-of-range byte_offset.
|
||||
let dt = Datatype::Compound {
|
||||
size: 8,
|
||||
members: vec![CompoundMember {
|
||||
name: "bad".to_string(),
|
||||
byte_offset: 4,
|
||||
datatype: make_f64_le_type(),
|
||||
}],
|
||||
};
|
||||
let raw = vec![0u8; 8]; // one element, matches declared size
|
||||
let result = read_compound_fields(&raw, &dt);
|
||||
assert!(
|
||||
matches!(result, Err(FormatError::Overflow(_))),
|
||||
"expected a clean Overflow error, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_raw_data_zerocopy_rejects_near_usize_max_offset() {
|
||||
let file_data = vec![0u8; 64];
|
||||
let dataspace = make_simple_dataspace(&[4]);
|
||||
let datatype = make_i32_le_type();
|
||||
let layout = DataLayout::Contiguous {
|
||||
address: Some(u64::MAX - 4),
|
||||
size: 16,
|
||||
};
|
||||
let result = read_raw_data_zerocopy(&file_data, &layout, &dataspace, &datatype);
|
||||
assert!(
|
||||
matches!(result, Err(FormatError::UnexpectedEof { .. })),
|
||||
"expected a clean UnexpectedEof, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_compound_single_field_by_name() {
|
||||
use crate::datatype::CompoundMember;
|
||||
|
||||
@@ -16,6 +16,21 @@ pub struct LocalHeap {
|
||||
pub data_segment_address: u64,
|
||||
}
|
||||
|
||||
/// Checks that `[offset, offset + needed)` fits within `data`, guarding the
|
||||
/// addition against `usize` overflow from a crafted near-`usize::MAX` offset.
|
||||
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> {
|
||||
if offset
|
||||
.checked_add(needed)
|
||||
.is_none_or(|end| end > data.len())
|
||||
{
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: offset.saturating_add(needed),
|
||||
available: data.len(),
|
||||
});
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
|
||||
let s = size as usize;
|
||||
if pos.checked_add(s).is_none_or(|end| end > data.len()) {
|
||||
@@ -47,12 +62,7 @@ impl LocalHeap {
|
||||
let ls = length_size as usize;
|
||||
let os = offset_size as usize;
|
||||
let total = 8 + ls * 2 + os;
|
||||
if offset + total > file_data.len() {
|
||||
return Err(FormatError::UnexpectedEof {
|
||||
expected: offset + total,
|
||||
available: file_data.len(),
|
||||
});
|
||||
}
|
||||
ensure_len(file_data, offset, total)?;
|
||||
|
||||
if &file_data[offset..offset + 4] != b"HEAP" {
|
||||
return Err(FormatError::InvalidLocalHeapSignature);
|
||||
@@ -172,6 +182,18 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_rejects_near_usize_max_offset_without_panicking() {
|
||||
// Found by fuzzing: `offset + total` overflowed for a crafted
|
||||
// near-usize::MAX offset.
|
||||
let file = build_heap_file(0, 100, &["hello"], 8, 8);
|
||||
let result = LocalHeap::parse(&file, usize::MAX - 4, 8, 8);
|
||||
assert!(
|
||||
matches!(result, Err(FormatError::UnexpectedEof { .. })),
|
||||
"expected a clean UnexpectedEof, got {result:?}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_heap_header() {
|
||||
let file = build_heap_file(0, 100, &["hello", "world"], 8, 8);
|
||||
|
||||
@@ -2,6 +2,9 @@
|
||||
//! data-integrity verification.
|
||||
//!
|
||||
//! Enable with the `provenance` Cargo feature (on by default).
|
||||
//!
|
||||
//! The hash is unkeyed, so this detects accidental corruption only — it is
|
||||
//! not a tamper-evidence or authenticity guarantee. See [`verify_dataset`].
|
||||
|
||||
#[cfg(not(feature = "std"))]
|
||||
use alloc::{format, string::String, vec::Vec};
|
||||
@@ -115,6 +118,11 @@ pub enum VerifyResult {
|
||||
///
|
||||
/// `file_data` is the entire HDF5 file bytes; `header` is the parsed object
|
||||
/// header for the dataset of interest.
|
||||
///
|
||||
/// This only detects *accidental* corruption. The hash is unkeyed and stored
|
||||
/// alongside the data it protects, so anyone able to modify the dataset can
|
||||
/// also recompute and overwrite `_provenance_sha256` — a `VerifyResult::Ok`
|
||||
/// is not a tamper-evidence or authenticity guarantee.
|
||||
pub fn verify_dataset(
|
||||
file_data: &[u8],
|
||||
header: &ObjectHeader,
|
||||
|
||||
@@ -90,6 +90,16 @@ pub fn make_i64_type() -> Datatype {
|
||||
}
|
||||
}
|
||||
|
||||
pub fn make_u64_type() -> Datatype {
|
||||
Datatype::FixedPoint {
|
||||
size: 8,
|
||||
byte_order: DatatypeByteOrder::LittleEndian,
|
||||
signed: false,
|
||||
bit_offset: 0,
|
||||
bit_precision: 64,
|
||||
}
|
||||
}
|
||||
|
||||
pub fn make_u8_type() -> Datatype {
|
||||
Datatype::FixedPoint {
|
||||
size: 1,
|
||||
@@ -444,6 +454,25 @@ impl DatasetBuilder {
|
||||
self
|
||||
}
|
||||
|
||||
/// Write a native unsigned 64-bit integer dataset. Pairs with the
|
||||
/// read side's `read_u64`/`read_as_u64`, which already support this
|
||||
/// datatype — this was the missing symmetric write-side builder
|
||||
/// (callers previously had to bit-cast through `with_i64_data` /
|
||||
/// `i64::from_ne_bytes(v.to_ne_bytes())` to round-trip full-range u64
|
||||
/// values like timestamps or IDs).
|
||||
pub fn with_u64_data(&mut self, data: &[u64]) -> &mut Self {
|
||||
self.datatype = Some(make_u64_type());
|
||||
let mut b = Vec::with_capacity(data.len() * 8);
|
||||
for &v in data {
|
||||
b.extend_from_slice(&v.to_le_bytes());
|
||||
}
|
||||
self.data = Some(b);
|
||||
if self.shape.is_none() {
|
||||
self.shape = Some(vec![data.len() as u64]);
|
||||
}
|
||||
self
|
||||
}
|
||||
|
||||
pub fn with_u8_data(&mut self, data: &[u8]) -> &mut Self {
|
||||
self.datatype = Some(make_u8_type());
|
||||
self.data = Some(data.to_vec());
|
||||
|
||||
@@ -11,14 +11,14 @@ categories = ["science", "graphics"]
|
||||
|
||||
[dependencies]
|
||||
wgpu = { version = "28", optional = true }
|
||||
half = { version = "2.7", optional = true }
|
||||
half = { workspace = true, optional = true }
|
||||
pollster = { version = "0.4", optional = true }
|
||||
bytemuck = { version = "1", features = ["derive"], optional = true }
|
||||
thiserror = "2"
|
||||
log = "0.4"
|
||||
|
||||
[dev-dependencies]
|
||||
criterion = { version = "0.5", features = ["html_reports"] }
|
||||
criterion = { workspace = true }
|
||||
rand = "0.8"
|
||||
approx = "0.5"
|
||||
pollster = "0.4"
|
||||
|
||||
@@ -15,13 +15,13 @@ memmap2 = { version = "0.9", optional = true }
|
||||
libc = { version = "0.2", optional = true }
|
||||
tokio = { version = "1", features = ["fs", "io-util"], optional = true }
|
||||
reqwest = { version = "0.12", features = ["json"], optional = true }
|
||||
serde = { version = "1", features = ["derive"], optional = true }
|
||||
serde = { workspace = true, optional = true }
|
||||
serde_json = { version = "1", optional = true }
|
||||
mpi = { version = "0.8", optional = true }
|
||||
|
||||
[dev-dependencies]
|
||||
tokio = { version = "1", features = ["full"] }
|
||||
tempfile = "3"
|
||||
tempfile = { workspace = true }
|
||||
|
||||
[features]
|
||||
default = []
|
||||
|
||||
@@ -19,7 +19,7 @@ clawhdf5-format = { path = "../clawhdf5-format", version = "2.1.0" }
|
||||
clawhdf5 = { path = "../clawhdf5", version = "2.1.0" }
|
||||
rusqlite = { version = "0.31", features = ["bundled"] }
|
||||
clap = { version = "4", features = ["derive"] }
|
||||
half = "2"
|
||||
half = { workspace = true }
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile = "3"
|
||||
tempfile = { workspace = true }
|
||||
|
||||
@@ -14,4 +14,4 @@ clawhdf5 = { path = "../clawhdf5", version = "2.1.0" }
|
||||
clawhdf5-format = { path = "../clawhdf5-format", version = "2.1.0" }
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile = "3"
|
||||
tempfile = { workspace = true }
|
||||
|
||||
@@ -16,8 +16,8 @@ crate-type = ["cdylib", "rlib"]
|
||||
[dependencies]
|
||||
clawhdf5_rs = { path = "../clawhdf5", version = "2.1.0", package = "clawhdf5" }
|
||||
clawhdf5-format = { path = "../clawhdf5-format", version = "2.1.0" }
|
||||
pyo3 = "0.28"
|
||||
numpy = "0.28"
|
||||
pyo3 = "0.29"
|
||||
numpy = "0.29"
|
||||
|
||||
[features]
|
||||
extension-module = ["pyo3/extension-module"]
|
||||
|
||||
@@ -4,7 +4,7 @@ build-backend = "maturin"
|
||||
|
||||
[project]
|
||||
name = "rustyhdf5"
|
||||
version = "1.93.0"
|
||||
version = "2.1.0"
|
||||
description = "Python bindings for rustyhdf5 — a pure-Rust HDF5 library"
|
||||
requires-python = ">=3.8"
|
||||
license = { text = "MIT" }
|
||||
|
||||
@@ -15,8 +15,8 @@ clawhdf5-io = { path = "../clawhdf5-io", version = "2.1.0" }
|
||||
rayon = { version = "1", optional = true }
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile = "3"
|
||||
criterion = { version = "0.5", features = ["html_reports"] }
|
||||
tempfile = { workspace = true }
|
||||
criterion = { workspace = true }
|
||||
clawhdf5-io = { path = "../clawhdf5-io", version = "2.1.0", features = ["mmap"] }
|
||||
clawhdf5-format = { path = "../clawhdf5-format", version = "2.1.0", features = ["parallel", "fast-checksum"] }
|
||||
clawhdf5-filters = { path = "../clawhdf5-filters", version = "2.1.0" }
|
||||
|
||||
@@ -436,6 +436,14 @@ impl<'f> Dataset<'f> {
|
||||
&self,
|
||||
selection: &clawhdf5_format::selection::Selection,
|
||||
) -> Result<Vec<u8>, Error> {
|
||||
// `Selection::All` is semantically a full read — route it through
|
||||
// the same per-file chunk cache `read_raw()` uses instead of the
|
||||
// selection path's uncached `read_chunked_data`, so callers get
|
||||
// consistent caching behavior regardless of which method they used
|
||||
// to ask for "everything".
|
||||
if matches!(selection, clawhdf5_format::selection::Selection::All) {
|
||||
return self.read_raw();
|
||||
}
|
||||
let dt = self.datatype()?;
|
||||
let ds = self.dataspace()?;
|
||||
let dl = self.data_layout()?;
|
||||
|
||||
@@ -935,3 +935,52 @@ fn dense_links_multiblock_fractal_heap_roundtrip() {
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_selection_all_matches_read_raw_on_chunked_dataset() {
|
||||
// read_selection(&Selection::All) is semantically a full read and must
|
||||
// go through the same cached path as read_raw()/read_f64() — not a
|
||||
// separate uncached code path that happens to return the same bytes.
|
||||
use clawhdf5_format::selection::Selection;
|
||||
|
||||
let data: Vec<f64> = (0..500).map(|i| i as f64 * 0.5).collect();
|
||||
let mut b = FileBuilder::new();
|
||||
b.create_dataset("chunked")
|
||||
.with_f64_data(&data)
|
||||
.with_chunks(&[100])
|
||||
.with_deflate(6);
|
||||
let file = File::from_bytes(b.finish().unwrap()).unwrap();
|
||||
|
||||
let ds = file.dataset("chunked").unwrap();
|
||||
let via_read_f64 = ds.read_f64().unwrap();
|
||||
let via_selection_bytes = ds.read_selection(&Selection::All).unwrap();
|
||||
let via_selection: Vec<f64> = via_selection_bytes
|
||||
.chunks_exact(8)
|
||||
.map(|c| f64::from_le_bytes(c.try_into().unwrap()))
|
||||
.collect();
|
||||
|
||||
assert_eq!(via_read_f64, data);
|
||||
assert_eq!(via_selection, data);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn u64_data_roundtrip() {
|
||||
// Values spanning the full u64 range, including ones with the high bit
|
||||
// set that would come back negative (and wrong) if bit-cast through
|
||||
// an i64 dataset instead of a native unsigned one.
|
||||
let values: Vec<u64> = vec![
|
||||
0,
|
||||
1,
|
||||
u64::MAX,
|
||||
u64::MAX / 2,
|
||||
1 << 63,
|
||||
1_700_000_000_000_000_000,
|
||||
];
|
||||
let mut b = FileBuilder::new();
|
||||
b.create_dataset("timestamps").with_u64_data(&values);
|
||||
let file = File::from_bytes(b.finish().unwrap()).unwrap();
|
||||
assert_eq!(
|
||||
file.dataset("timestamps").unwrap().read_u64().unwrap(),
|
||||
values
|
||||
);
|
||||
}
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@redclaw/clawhdf5",
|
||||
"version": "2.0.0",
|
||||
"version": "2.1.0",
|
||||
"description": "Node.js bindings for clawhdf5 — HDF5-backed agent memory with hippocampal consolidation",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
|
||||
@@ -0,0 +1,176 @@
|
||||
# ClawHDF5 — Performance / Security / Provenance Implementation Brief
|
||||
|
||||
**Date:** 2026-08-16
|
||||
**Scope:** Follow-up hardening pass on top of the already-shipped Tier 1-4 work
|
||||
(see `ROADMAP.md` "What's Next" and `IMPROVEMENT_LOG.md`). This brief covers
|
||||
only items verified against the current repo state at commit `b2dce41` that
|
||||
were **not** already addressed by prior tiers.
|
||||
|
||||
## Method
|
||||
|
||||
Read `ROADMAP.md`, `IMPROVEMENT_LOG.md`, `IMPROVEMENT_SCAN.md`, and
|
||||
`CLAUDE.md` first to avoid re-proposing work already merged (WAL CRC32,
|
||||
bounds-check audit + fuzz target, HNSW `prune_connections` parallelism,
|
||||
Android JNI length validation, `workspace.dependencies` hoisting, etc. are
|
||||
all already done — see those files for the full list).
|
||||
|
||||
Then manually audited:
|
||||
- `crates/clawhdf5-format/src/{chunked_read,data_read}.rs` — bounds-check
|
||||
spot audit (sampled `ensure_len` call sites around every raw slice index).
|
||||
**Result: no new gaps found.** Every raw `file_data[a..b]` site sampled is
|
||||
preceded by an `ensure_len`/`read_offset` overflow-checked bound. The prior
|
||||
Tier 4a pass already closed this out.
|
||||
- `crates/clawhdf5-agent/src/provenance.rs` — memory record integrity →
|
||||
**gap found**, see INT-01.
|
||||
- `crates/clawhdf5-agent/src/knowledge.rs` — knowledge-graph traversal →
|
||||
**gap found**, see INT-02.
|
||||
- `crates/clawhdf5-agent/src/bm25.rs` — already has cached IDF, sorted
|
||||
postings, WAND early termination. No changes proposed.
|
||||
- `crates/clawhdf5-ann/src/hnsw.rs` — build-time parallelism already scoped
|
||||
to `prune_connections` per Tier 4c; the outer insert loop is flagged in
|
||||
ROADMAP as needing its own correctness-sensitive design pass, out of scope
|
||||
here.
|
||||
|
||||
## INT-01 — Harden agent memory provenance hash from FNV-1a to SHA-256
|
||||
|
||||
**File:** `crates/clawhdf5-agent/src/provenance.rs`
|
||||
**Category:** Security / Provenance
|
||||
**Status:** Implemented this pass.
|
||||
|
||||
### Problem
|
||||
|
||||
`MemoryProvenance::content_hash` used an unkeyed 64-bit FNV-1a hash
|
||||
(`fnv1a_64`) to detect corruption of stored agent-memory chunks. FNV-1a is
|
||||
a fast non-cryptographic hash with no collision resistance: an adversary
|
||||
attempting to plant poisoned/tampered memory content that still matches a
|
||||
previously-recorded or expected hash value only needs to find *any* input
|
||||
producing the same 64-bit output, which is computationally cheap for
|
||||
FNV-1a (no preimage or collision resistance guarantees at all). Given
|
||||
Track 5 of `ROADMAP.md` explicitly claims "poisoning resistance" and
|
||||
`verify_integrity()` is the one function whose entire job is to catch
|
||||
tampered memory content, using a hash with no collision resistance
|
||||
undermines that guarantee in a way that is easy to miss (the doc comment
|
||||
already, correctly, disclaims *authenticity* — i.e. it never claimed to
|
||||
stop an attacker who can also rewrite the stored hash — but it did not
|
||||
protect against a weaker, still-relevant attack: crafting *different*
|
||||
poisoned content that collides with an already-recorded legitimate hash).
|
||||
|
||||
Separately, `clawhdf5-format` already ships a mature, default-on
|
||||
`provenance` feature (`crates/clawhdf5-format/src/provenance.rs`) with a
|
||||
`sha256_hex()` helper built on the `sha2` crate, used for on-disk dataset
|
||||
provenance attributes. `clawhdf5-agent` already depends on
|
||||
`clawhdf5-format` with default features enabled, so `sha256_hex` was
|
||||
already reachable with **zero new dependencies**.
|
||||
|
||||
### Fix implemented
|
||||
|
||||
- `MemoryProvenance::content_hash` changed from `u64` to `String` (lowercase
|
||||
hex SHA-256 digest), computed via `clawhdf5_format::provenance::sha256_hex`.
|
||||
- `ProvenanceStore::verify_integrity` now compares SHA-256 hex digests.
|
||||
- Removed the local `fnv1a_64` helper from `provenance.rs` (no longer used
|
||||
there — `clawhdf5-agent/src/multimodal.rs` keeps its own independent
|
||||
`fnv1a_64` for `MediaRef` checksums, which is a content-identity/dedup key,
|
||||
not a security/integrity control, so it is intentionally left unchanged
|
||||
and out of scope for this item).
|
||||
- Updated the module-level and per-item doc comments to keep the existing,
|
||||
correct disclaimer: this is still an **unkeyed** hash, so it is still not
|
||||
an authenticity/tamper-*evidence* guarantee against an attacker who can
|
||||
rewrite the stored hash alongside the content. What changed is that it is
|
||||
no longer trivially *collidable*, which was the concrete, fixable gap.
|
||||
- Updated all existing unit tests in `provenance.rs` for the new `String`
|
||||
hash type; behavior (which records match/mismatch) is unchanged.
|
||||
|
||||
`MemoryProvenance` and `ProvenanceStore` are only used within
|
||||
`clawhdf5-agent` itself (not serialized to the HDF5 format, not consumed by
|
||||
other crates), so this is a self-contained, non-breaking-to-other-crates
|
||||
change verified by `grep -r MemoryProvenance crates/`.
|
||||
|
||||
## INT-02 — Knowledge-graph traversal: replace O(V·R) relation scans with a per-call adjacency index
|
||||
|
||||
**File:** `crates/clawhdf5-agent/src/knowledge.rs`
|
||||
**Category:** Performance
|
||||
**Status:** Implemented this pass.
|
||||
|
||||
### Problem
|
||||
|
||||
`KnowledgeCache::bfs_neighbors` and `KnowledgeCache::spreading_activation`
|
||||
are the core traversal primitives behind Track 1 (BFS neighbors, subgraph
|
||||
extraction) and Track 3 (graph-aware re-ranking) of the agent memory
|
||||
system. Both did a **full linear scan over `self.relations`** for every
|
||||
node processed:
|
||||
|
||||
- `bfs_neighbors`: for every entity dequeued from the BFS frontier, it
|
||||
scanned the entire `relations: Vec<Relation>` looking for edges touching
|
||||
that entity — O(V·R) instead of O(V+E). It also called
|
||||
`self.get_entity(neighbour_id)`, itself an O(n) linear scan over
|
||||
`entities: Vec<Entity>`, once per newly-discovered neighbour.
|
||||
- `spreading_activation`: for every activated node in every propagation
|
||||
step, it likewise scanned all of `self.relations` — O(steps·V·R).
|
||||
- `get_subgraph` calls `bfs_neighbors` once per seed, compounding the cost.
|
||||
|
||||
For a knowledge graph with thousands of entities/relations (the scale this
|
||||
project's own benchmarks target — see `BENCHMARKS.md`), this is
|
||||
quadratic-ish behavior in traversal-heavy paths (`get_entity_context`,
|
||||
hybrid retrieval re-ranking that pulls graph context) that only gets worse
|
||||
as agent memory accumulates over long sessions.
|
||||
|
||||
### Fix implemented
|
||||
|
||||
Added a private helper, `KnowledgeCache::build_adjacency`, that builds, in
|
||||
one O(V+R) pass:
|
||||
- `entity_index: HashMap<u64, usize>` — entity id → index into `entities`.
|
||||
- `adjacency: HashMap<u64, Vec<u64>>` — entity id → neighbour ids (both
|
||||
outgoing and incoming edges).
|
||||
|
||||
`bfs_neighbors` and `spreading_activation` now build this index **once at
|
||||
the top of the call** (not persisted as struct state — see rationale below)
|
||||
and use it for O(1) neighbour/entity lookups inside the traversal loop,
|
||||
changing the complexity to O(V+E) per call for BFS and O(steps·(V+E)) for
|
||||
spreading activation.
|
||||
|
||||
**Why not a persistent index on the struct:** `entities`/`relations` are
|
||||
public fields, and `crates/clawhdf5-agent/src/schema.rs` (deserialization
|
||||
path, loading a persisted knowledge graph back from HDF5) pushes directly
|
||||
into `cache.entities`/`cache.relations` rather than going through
|
||||
`add_entity`/`add_relation`. A struct-level cached index would silently go
|
||||
stale on that path. Building the index fresh at the top of each traversal
|
||||
call is O(V+R) — the same asymptotic cost as the scan it replaces would be
|
||||
for a *single* node — so it turns what was an O(V·R)-or-worse *whole
|
||||
traversal* into an O(V+R) traversal, with no risk of a stale-index
|
||||
correctness bug and no change to the existing public API or struct layout.
|
||||
`get_entity`, `get_relations_from`, `get_relations_to` are left as-is
|
||||
(still O(n)/O(R)): they're public API used elsewhere as one-off lookups,
|
||||
not inside a per-node hot loop, so indexing them is lower value and was
|
||||
left out of scope to keep this change minimal and low-risk.
|
||||
|
||||
Existing tests (`test_bfs_neighbors_*`, `test_get_subgraph_*`,
|
||||
`test_spreading_activation_*`) exercise correctness and were not modified —
|
||||
they pass unchanged, confirming the traversal results are identical to the
|
||||
pre-change O(V·R) implementation.
|
||||
|
||||
## Deferred / not implemented this pass
|
||||
|
||||
Listed for a future pass — investigated but out of scope for this brief's
|
||||
budget, or blocked on a larger design decision already flagged upstream:
|
||||
|
||||
- **HNSW outer insert-loop parallelism** — `ROADMAP.md` already flags this
|
||||
as needing "its own dedicated design pass" before parallelizing; not
|
||||
attempted here to avoid a correctness-sensitive change without that design
|
||||
work.
|
||||
- **WAL per-entry format redesign** (explicit length-prefix instead of
|
||||
read-then-verify-CRC32) — `ROADMAP.md` already notes the current CRC32
|
||||
trailer works and this would only be worth revisiting "if profiling shows
|
||||
it matters"; no such profiling signal was found this pass.
|
||||
- **`get_entity`/`get_relations_from`/`get_relations_to` indexing** — would
|
||||
further help `get_entity_context` and any other one-off caller, but is
|
||||
lower value than the hot-loop fix in INT-02 and was left out to keep this
|
||||
change minimal.
|
||||
|
||||
## Verification performed
|
||||
|
||||
- `cargo build --workspace --lib --bins` — clean before starting (baseline).
|
||||
- `cargo test --workspace` — run after implementing INT-01 and INT-02 (see
|
||||
commit for pass/fail status).
|
||||
|
||||
TASK: INT-01 — Harden agent memory provenance hash from FNV-1a to SHA-256
|
||||
TASK: INT-02 — Knowledge-graph traversal: per-call adjacency index instead of O(V·R) relation scans
|
||||
Reference in New Issue
Block a user