clawmates/mission-01a00d0f-407d84f0
20
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
122849b5a9 |
research: add implementation brief with 17 numbered INT items
Covers performance, security, and provenance findings across clawhdf5-format, clawhdf5-migrate, and memory/query crates. Each item lists target file, problem, and proposed change for the coding phase. |
||
|
|
1537a9464a |
bench: sweep the hybrid weights, and correct the recommendation
CI / test (push) Failing after 3s
Tier 4b reported hybrid retrieval at 0.7/0.3 and noted the weights were "the
documented default, not a searched optimum". `--sweep` searches them: 0.0 to 1.0
in 0.1 steps, reusing the one-time embedding table so eleven configurations cost
barely more than three.
The result is not a refinement. 0.7/0.3 is **strictly dominated**:
vector/keyword Hit@1 Hit@5 Hit@10 MRR sHit@5
0.0 / 1.0 53.8% 75.0% 81.6% 0.6320 93.6%
0.3 / 0.7 53.2% 78.8% 87.2% 0.6463 96.0%
0.4 / 0.6 51.6% 81.4% 87.8% 0.6429 96.8%
0.5 / 0.5 48.2% 81.4% 88.2% 0.6234 97.4%
0.7 / 0.3 44.4% 79.2% 86.0% 0.5868 95.8%
1.0 / 0.0 36.0% 71.8% 81.6% 0.5027 94.2%
0.4/0.6 beats 0.7/0.3 on every metric at both granularities — Hit@1 +7.2pp,
Hit@5 +2.2, Hit@10 +1.8, MRR +0.056. No trade is being made; the default simply
sat on the wrong side of the peak. It is now 0.4/0.6, and README's usage snippet
recommends the same.
This corrects a conclusion I published one commit ago. Measuring only 0.7/0.3, I
wrote that fusion "buys deeper recall and pays for it at rank 1" and advised
callers taking a single top hit to prefer BM25. That was an artifact of the bad
weight, not a property of fusion: at 0.3/0.7 hybrid *beats* BM25 on MRR (0.6463
vs 0.6320) and Hit@5 (78.8% vs 75.0%) while giving up 0.6pp of Hit@1. Both
BENCHMARKS.md and README carry the correction rather than a quiet edit, since
the old text told readers to configure their systems a particular way.
The three-mode ablation rows are kept at their original settings — they measure
the shape of each stage in isolation, and the operating point now comes from the
sweep instead.
|
||
|
|
12d9d8462f |
bench: make the CUDA embedding path discoverable when it is unavailable
CI / test (push) Failing after 2s
The GPU path worked but was effectively hidden. cudarc's build script shells out
to `nvcc`, which ships in /usr/local/cuda/bin — a directory the reference host
had installed but never exported to the login shell, so `--features
embeddings-cuda` failed with a bare "`nvcc --version` failed" panic from a
dependency's build script, and the runtime fallback then reported only
"Embedder: CPU (...)" before spending hours on work a GPU does in minutes.
Two changes, both about making the failure legible rather than changing what the
code does:
- The CPU fallback now says why it fell back and what that costs, with the
concrete fix. A run that silently takes two orders of magnitude longer reads
as a hang, not as a configuration choice.
- BENCHMARKS.md states the build-time nvcc requirement, where the toolkit
actually installs, and that a shell file read non-interactively is the place
to export it — `~/.zshenv` rather than `~/.zshrc`, because build scripts do
not run in an interactive shell.
Host-side, the reference machine's CUDA exports lived in ~/.bashrc below its
non-interactive guard while the login shell is zsh, so they never applied to
anything. Moved to ~/.zshenv with duplicate-prepend guards; `nvcc --version`
and `cargo build --features embeddings-cuda` now both work over a plain
non-interactive ssh with no manual export.
|
||
|
|
c913cd1cbf |
bench: make the vector stage real, and measure BM25 vs vector vs hybrid (Tier 4b)
CI / test (push) Failing after 2s
Every LongMemEval number this project has published measured BM25 alone. The
bench passed zero-vector embeddings with vector_weight=0.0, so the HNSW/vector
stage — the thing the README credits for retrieval quality — contributed
nothing and was never tested.
An optional `embeddings` feature loads all-MiniLM-L6-v2 via candle and encodes
the corpus for real. It is off by default and nothing in the shipped crates
depends on it, so a project that advertises no heavyweight dependencies keeps
that property; without the feature the bench behaves exactly as before.
Full haystack, n=500, turn-level:
Hit@1 Hit@5 Hit@10 MRR
BM25 only 53.8% 75.0% 81.6% 0.6320
Vector only 36.0% 71.8% 81.6% 0.5027
Hybrid 0.7/0.3 44.4% 79.2% 86.0% 0.5868
Session-level, hybrid leads outright: 88.2 / 95.8 / 97.8 / 0.9158.
The hybrid claim holds for depth and not for precision@1. Hybrid is the best
configuration at Hit@5 and Hit@10 at both granularities — turn-level Hit@5 gains
4.2 points over BM25 and 7.4 over vector-only, which is the result that justifies
running two stages at all. But BM25 alone still leads turn-level Hit@1 and MRR,
so fusing buys deeper recall and pays at rank 1. Callers assembling five memories
of context want hybrid; callers taking a single top hit are better served by BM25
today. The 0.7/0.3 weights are the documented default, not a searched optimum.
omni-cortex's four-signal ablation found the same direction independently — there,
adding BM25 to a dense retriever raised nDCG@5 while lowering Hit@1 and MRR. Two
codebases, two fusion schemes, same trade.
Vector-only trailing BM25 at every turn-level cutoff except Hit@10 is stated
plainly rather than buried: LongMemEval questions share heavy vocabulary with
their evidence turns, which is close to the best case for lexical matching, and
MiniLM at 384-d is a small model.
Implementation notes:
- Texts are deduplicated before encoding. The haystack sessions are drawn from
a shared pool, so 500 questions x 493.5 turns collapses to 190,015 unique
strings — the difference between encoding the corpus once and per question.
- `embeddings-cuda` adds the GPU path, and it is not a convenience: 190k texts
take ~13 min on an RTX 5060 Ti, while the same work on 8 CPU cores was still
unfinished after 30 minutes. The device is selected at runtime with a CPU
fallback, so a machine without CUDA still works.
- Mean-pooling is masked and the output L2-normalised, which is the published
recipe for this checkpoint (not the [CLS] pooler).
One measurement wrinkle, recorded rather than smoothed over: on the oracle
variant BM25-only reads 84.2% Hit@5 with real embedding vectors present against
84.4% with zero vectors — one question of 500 changes rank, MRR identical at
0.6597. On the full haystack the two agree exactly. Weight 0.0 evidently does not
make the vector stage bit-for-bit absent from candidate selection on a small
corpus.
Verified on the Linux dev host: 49 groups / 1659 passed / 0 failed, clippy clean
under -D warnings, fmt clean, with and without the feature.
|
||
|
|
7d6e269bf3 |
bench: run the full longmemeval_s haystack, and measure the variant (Tier 4a)
CI / test (push) Failing after 2s
The harness only ever ran longmemeval_oracle — evidence sessions only, which is
a substantially easier corpus than the dataset LongMemEval results are normally
quoted on. Worse, the variant was a hardcoded "oracle" string in both the report
header and the JSON summary, so pointing it at longmemeval_s would have produced
full-haystack numbers labelled oracle.
DatasetProfile now measures the corpus instead of asserting it: sessions and
turns per question, and evidence-session density (the mean share of a question's
haystack sessions that are answer sessions). The variant label and the
session-level degeneracy warning are both derived from that density, so a
mislabelled input file cannot produce a mislabelled result. Measured: 100.0%
density on the oracle variant, 4.0% on longmemeval_s.
The full haystack, all 500 questions, 47.7 sessions and 493.5 turns each:
turn-level session-level
Hit@1 53.8% 86.2%
Hit@5 75.0% 93.6%
Hit@10 81.6% 96.6%
MRR 0.6320 0.8948
Turn-level drops 84.4% -> 75.0% against the oracle variant. That 9.4-point gap
is the price of the real haystack and is exactly why oracle-only numbers should
not be presented as LongMemEval results.
Session-level is now reportable. It was retracted before because at 100% evidence
density every returned document is a hit by construction; at 4.0% density a hit
reflects discrimination, so 93.6% is a real measurement rather than a restatement
of the corpus shape. Per-type it also finally separates: single-session-assistant
100.0% Hit@1 against single-session-preference 33.3% — BM25 has nothing to grip
on a preference question whose evidence shares no vocabulary with the query.
The MemX comparison stays withdrawn. Running the full haystack closes the corpus
half of that mismatch but not the granularity half: MemX measures fact-level over
220,349 records, and this harness measures turn- and session-level.
Two smaller fixes found while running it:
- --limit samples evenly across the file rather than taking a prefix. The
dataset is ordered by question type, so `--limit 20` returned 20
single-session-user questions and nothing else while reading like a
whole-dataset result.
- abstention_accuracy emits null rather than 0.0 when a corpus poses no
abstention questions. longmemeval_s has none, and 0.0000 reads as total
failure at a task that was never asked.
README.md and BENCHMARKS.md now lead with the full-haystack numbers and keep the
oracle figures alongside, labelled as the easier corpus.
Verified on the Linux dev host: 49 groups / 1659 passed / 0 failed, clippy clean
under -D warnings, fmt clean. The full 500-question run takes ~70 s.
|
||
|
|
6f5940d042 |
docs: retract degenerate LongMemEval session-level numbers and the MemX comparison
CI / test (push) Failing after 4s
A methodology audit found that two benchmark claims published in this repo two days ago measure the wrong thing. Both are retracted in place rather than quietly edited, with the reasoning recorded. 1. Session-level LongMemEval recall (100.0% Hit@1/5/10, MRR 1.0000, uniform across all six question types) is a degenerate artifact. On the longmemeval_oracle variant the ingested haystack for a question is essentially only that question's evidence sessions, so every returned document belongs to an answer session and session-level hit rate is ~1.0 at rank 0 by construction. The uniform 100% across every question type was the tell. It measured the shape of the corpus, not the retriever. Only the turn-level figure (84.4% Hit@5) carries signal, and it is now the only retrieval number cited. 2. The "clawhdf5 outperforms MemX at turn-level retrieval (84.4% vs 51.6%)" claim was not like-for-like on two independent axes. Confirmed against arxiv:2603.16171: MemX's Hit@5=51.6% / MRR=0.380 is *fact-level* granularity over 220,349 fact-level records drawn from 19,195 sessions, and the paper explicitly notes fact-level "doubl[es] session-level performance". Ours is turn-level on the oracle subset — different granularity, and a corpus smaller by orders of magnitude. A higher number on an easier corpus at a different granularity is not an outperformance claim. Also caveats the vector-search "vs MemX" latency ratios, which compare a single clawhdf5 component (raw vector search) against MemX's end-to-end pipeline figure (embeddings + FTS5 + four-factor re-ranking). The numbers are real; the "speedup" framing overstated by an unquantified margin and is now labelled an order-of-magnitude indication. Adds an explicit scoring-target declaration to BENCHMARKS.md per arXiv 2605.24060, which found that changing scoring target alone alters nDCG on 83-94% of queries and can reverse system rankings. States dataset variant, metric (retrieval recall, NOT the official QA-accuracy metric), granularity, k, and that the vector stage is inert (zero embeddings, vector_weight=0.0). The harness itself now prints its scoring target, flags the session-level block as degenerate, warns against the MemX comparison, and emits dataset_variant/scoring_target/k/session_level_degenerate in its JSON summary, so the caveats travel with the numbers instead of living only in docs. |
||
|
|
dfae9e2cc1 |
feat: add with_u64_data builder; fix read_selection cache bypass
CI / test (push) Failing after 2s
Found via a real-world integration audit against omni-cortex (a JEPA-based cognitive architecture built on clawhdf5 as its tiered Working/Episodic/ Semantic memory store). - Add DatasetBuilder::with_u64_data (crates/clawhdf5-format/type_builders.rs). The read side already has read_u64/read_as_u64, but there was no symmetric write-side builder — only signed with_i32_data/with_i64_data existed. Every consumer needing full-range u64 (timestamps, IDs) had to bit-cast through i64 via `i64::from_ne_bytes(v.to_ne_bytes())` on write and reverse it on read. omni-cortex does this in at least 6 places across its writer/reader/mmap-reader/consolidate crates. Confirmed the new builder round-trips full-range u64 (including values with the high bit set) end-to-end in a standalone sanity check mirroring their usage. - Fix Dataset::read_selection(&Selection::All) to route through the same per-file chunk cache read_raw()/read_f64() etc. already use, instead of the uncached read_chunked_data path. Selection::All is semantically a full read; there's no reason two ways of asking for "everything" should have different caching behavior. Also gains read_raw()'s virtual-dataset resolver support for free. omni-cortex's Reader/mmap-reader/consolidate crates all call read_selection(&Selection::All) for their chunked/ compressed dataset reads, so this was a real, if currently low-traffic (single-pass read pattern), inconsistency in the public API's behavior. - README: fix a stale crate-map claim that clawhdf5-filters supports "blosc" compression — it never did (the crate only ever held fast_deflate.rs; lz4/zstd/pcodec/szip filters live in clawhdf5-format). New tests: u64_data_roundtrip, read_selection_all_matches_read_raw_on_chunked_dataset. |
||
|
|
429c29b76b |
docs: sync README/ROADMAP/CLAUDE/CHANGELOG with Tier 1-4 hardening work
CI / test (push) Failing after 3s
README.md: - Fix badly stale LongMemEval numbers (badge said Hit@5 46%, table showed fabricated ~46%/~0.34/~72% figures that never matched BENCHMARKS.md's actual results of Hit@5 100% session / 84.4% turn-level, MRR 1.0/0.6597) - Remove clawhdf5-types from the Crate Map — that crate was removed in an earlier cleanup pass but the README diagram was never updated; fix the crate count (16, not 17) and stale line-of-code figures (72,087/84K -> ~92K) - Fix a dead #benchmarks badge anchor (no such heading exists) -> #performance - Document the new clawhdf5-ann `parallel` feature (had no Feature Flags entry) - Note WAL's CRC32 per-entry check, link the new tank LongMemEval/SIMD/ vector-search reproduction section, update stale test-count comment (417+ -> 1,650+) and Phase 2 roadmap blurb (LongMemEval is now done) ROADMAP.md: - Check off "Academic benchmark cross-validation" (done via the tank LongMemEval re-run) and add a new "Recently closed out" section summarizing the Tier 3-4 hardening pass (Android JNI validation, pyo3 bump, WAL CRC32, bounds-check audit + fuzz harness that found 3 real bugs, HNSW optional parallel feature, workspace.dependencies) - Update stale test count (1,546 -> 1,650+) and last-updated date CLAUDE.md: mention WAL's per-entry CRC32 check CHANGELOG.md: add Security/Performance/Architecture/Documentation entries under Unreleased summarizing all of Tiers 1-4 (this had not been touched since 2026-06-04, predating the entire hardening pass) |
||
|
|
40527be653 |
docs: Tier 4e — dated tank re-run for LongMemEval, SIMD, and vector-search sections
CI / test (push) Failing after 2s
Re-ran the three previously-undated sections flagged by the top-of-file traceability note on tank (Ryzen 7 7800X3D, 2026-08-05), the same machine already used for the vs-libhdf5 validation: - LongMemEval Results: recall numbers reproduce exactly (deterministic BM25 retrieval), latency numbers are new/hardware-specific and higher than the i7 citation with much wider variance — recorded as-is. - SIMD & Parallelism: found that several of the originally-named benchmarks don't actually hold the dataset fixed while varying only the SIMD/scalar/parallel axis — several call the same underlying function under different names. Used adaptive_benches' strategy_* benchmarks instead, which genuinely do isolate that axis via the SearchStrategy enum. Real finding: the speedup on tank (~1.5x) is smaller than on the i7 (~2.0x), attributed to the Ryzen's large L3 cache narrowing the scalar-vs-SIMD gap — recorded rather than reconciled away. - Vector Search Latency / Comparison to MemX: re-run with tank numbers, all faster than the i7 citation as expected; the 1K Pre-norm cell has no corresponding benchmark in the current suite and is left blank rather than guessed. Updated the top-of-file traceability note to reflect that these three sections (plus Comparison to MemX) now meet the dated/hardware-cited/ reproducible bar, narrowing the list of sections that don't. |
||
|
|
2013fa94a0 |
security: Tier 4b — WAL per-entry CRC32 checksum (WAL_VERSION 2)
CI / test (push) Failing after 3s
Bump WAL_VERSION to 2: every entry (Save and Tombstone) now ends with a 4-byte CRC32 trailer computed over its type+timestamp+payload bytes, using the existing clawhdf5_format::checksum::crc32 (already available since clawhdf5-agent depends on clawhdf5-format with fast-checksum enabled). A bit-flip inside an entry is now detected and replay stops there, instead of silently accepting corrupted data as before. Write side needed no restructuring — append_save/append_tombstone already buffer an entry's bytes before a single write_all, so the CRC is just appended to that buffer first. Read side: read_len_prefixed_str/read_embedding are generalized from &mut File to R: Read, and a new TeeReader<R> wraps the file handle for one entry at a time, accumulating every byte actually consumed (via read_exact) into a buffer. This lets read_entries compute the CRC over exactly the bytes read for a Save entry without needing to know its length up front (its sub-fields are length-prefixed and interleaved with the length itself only becoming known as parsing proceeds). A new read_one_entry<R: Read> factors the per-entry-type field parsing shared by both the legacy and current read paths. Backward compatibility: WAL_VERSION_LEGACY_NO_CRC (1) files are still readable via WalFile::read_entries (old field-by-file-handle path, unchanged, no CRC expected). WalFile::open migrates a legacy file by recreating it fresh in the current format — safe because the only two real call sites (HDF5Memory::open/create) always call read_entries before open, so entries are already replayed by the time migration happens. New tests: a corrupted-payload-byte test confirming replay stops cleanly at the corrupted entry (no prior coverage existed for mid-entry bit-flip detection), a legacy-v1-format read test, and an open()-migration test. |
||
|
|
a3e1cf8588 |
perf: Tier 4c — optional rayon parallelism for HNSW prune_connections
CI / test (push) Failing after 3s
Add a default-off `parallel` feature to clawhdf5-ann (rayon optional dep), matching the convention already used in clawhdf5-format/clawhdf5-agent. Gate prune_connections' per-neighbor distance computation on it — a pure read-only map with no shared mutable state, sorted immediately after, so swapping to rayon's par_iter is low-risk. Deliberately not touching build_with_metric's outer insert loop per the original plan: it has genuine cross-iteration data dependencies (graph mutation, entry-point updates) and needs its own correctness-focused design pass. The win here is likely small since neighbor lists are bounded by m/m_max0 (typically small) — this is a low-risk completeness item, not a headline perf change. Verified identical results with default features and --features parallel across the full HNSW test suite (23/23 both ways), including the build+search end-to-end tests (build_small_index, search_accuracy_cosine, incremental_insert_matches_batch_recall). |
||
|
|
534331ffbe |
chore: Tier 4d — hoist tempfile/criterion/half/serde to workspace.dependencies
CI / test (push) Failing after 13s
Add [workspace.dependencies] to the root Cargo.toml for the four
duplicated-across-many-crates dependencies flagged by the earlier review:
tempfile (7 crates), criterion (6), half (4 — real version skew, clawhdf5-gpu
pinned 2.7 while others used bare 2), and serde (4). Update every consuming
crate to `dep = { workspace = true }`, preserving crate-local `optional =
true` where it already existed. half now resolves uniformly to 2.7.x
workspace-wide instead of two separate semver ranges.
Also fixed clawhdf5-filters/Cargo.toml's stale "rustyhdf5" description
while touching the file (same class of leftover rename as prior fixes).
Not touching rayon/byteorder/clap (no skew found, lower priority).
|
||
|
|
297ee5ec17 |
security: Tier 4a — bounds-check audit + new dataset-read fuzz target
CI / test (push) Failing after 4s
- Add ensure_len(data, offset, needed) helper to chunked_read.rs, data_read.rs, and local_heap.rs (matching the existing btree_v1.rs/ object_header.rs convention) and use it at every plain-arithmetic offset+size bounds check found in these files, closing usize-overflow panics reachable from crafted near-usize::MAX offsets/addresses. - collect_chunk_info: add a depth-limited internal wrapper (collect_chunk_info_inner, MAX_CHUNK_BTREE_DEPTH=64) to reject a crafted self-referencing/cyclic B-tree v1 chunk index instead of recursing unboundedly (stack-overflow DoS). - read_compound_fields: validate byte_offset+field_size against the compound's declared element size before slicing, instead of an unguarded out-of-bounds panic on a crafted member offset. - read_chunked_data/_cached/_sweep/_indexed: guard `ndims - 1` against underflow for a degenerate zero-dimension chunked layout. - copy_chunk_to_output: rewrite all offset/stride arithmetic (both the 1-D fast path and the general N-D path) to use checked_add/checked_mul, skipping an out-of-range row/chunk instead of panicking on overflow. Add a new cargo-fuzz target, fuzz_dataset_read, that walks every dataset in a parsed file via the clawhdf5 facade and exercises the contiguous/ chunked/compact raw-data read paths that the existing fuzz_full_file target doesn't reach. Seeded with the chunked/VDS/compound-relevant test fixtures plus two crash regressions found during this pass (the copy_chunk_to_output overflow and the ndims-1 underflow, both fixed above — this target found real bugs within the first couple of runs). Not wired into CI (nightly-only, multi-minute runs); documented in fuzz/README.md as a manual/scheduled check instead. Also fixed the README's stale rustyhdf5-format naming while touching this file. Added regression tests for every fix (near-usize::MAX offsets, the self-referencing B-tree case, the compound byte_offset overrun, the zero-dim layout, and both copy_chunk_to_output overflow paths) so these are caught by `cargo test`, not just the fuzz corpus. |
||
|
|
a319405ffc |
security: Tier 3 — Android JNI length validation, pyo3 bump, WAL caps
CI / test (push) Failing after 2s
- clawhdf5-android: validate embedding_len/query_embedding_len against the handle's configured embedding_dim (and reject null pointers) before constructing a slice via from_raw_parts in edgehdf5_save and edgehdf5_hybrid_search. Strengthen the # Safety docs to state the now-enforced invariant and its limits. Add unit tests covering mismatched length and null-pointer rejection. - clawhdf5-py: bump pyo3/numpy 0.28 -> 0.29, clearing RUSTSEC-2026-0176 (OOB read in PyList/PyTuple iterator) and RUSTSEC-2026-0177 (missing Sync bound on PyCFunction::new_closure). No source changes needed; confirmed via cargo audit that both advisories no longer appear. - clawhdf5-agent/wal.rs: cap read_len_prefixed_str/read_embedding's length claims at a new MAX_WAL_FIELD_LEN (64 MiB) before allocating, so a corrupted/truncated WAL length field fails cleanly instead of attempting a huge allocation. Add regression tests for both. - BENCHMARKS.md: add a top-of-file traceability note distinguishing the dated/hardware-cited/reproducible h5bench and tank-validation sections from the older sections that don't yet meet that bar. |
||
|
|
62595d5ac0 |
chore: Tier 2 quick wins — version skew, docs, cleanup, overflow-safe bounds
CI / test (push) Failing after 14s
- Fix version skew: clawhdf5-py (pyproject.toml 1.93.0 -> 2.1.0) and packages/clawhdf5-node (package.json 2.0.0 -> 2.1.0) were both behind the actual crate version. - Correct stale ROADMAP.md claims: the TypeScript bridge already has a complete napi-rs package (not "no package.json"); CI/CD is now wired up via .gitea/workflows/ci.yml. - Fix CLAUDE.md: clawhdf5-gpu uses wgpu with hand-written WGSL compute shaders, not CubeCL. - chunked_read.rs: drop 12 unnecessary chunk_dimensions[..rank].to_vec() allocations — all three callees already accept &[u32]. - btree_v1.rs: add an overflow-safe ensure_len(data, offset, needed) helper (checked_add) and use it at the two plain-arithmetic bounds guards, closing a usize-overflow edge case reachable from a crafted near-usize::MAX B-tree offset. Add a regression test. - Clarify that the integrity hashes in clawhdf5-agent/provenance.rs (FNV-1a) and clawhdf5-format/provenance.rs (SHA-256) are unkeyed and only detect accidental corruption, not tampering — doc-only change. - README.md: document that the mpi-io feature's read/write paths are root-read+broadcast / gather-to-rank-0, not true collective I/O. |
||
|
|
55959b4920 |
ci: wire up CI, fix no_std build, fix stale package names in scripts
CI / test (push) Failing after 15s
- Add .gitea/workflows/ci.yml running scripts/ci-test.sh (fmt, clippy,
test, no_std check) on push/PR to main.
- Fix stale rustyhdf5-py/rustyhdf5-format package names in
ci-test.sh/check-nostd.sh, which had been silently no-op'ing those
checks (cargo warns but doesn't fail on an unknown --exclude/-p
target).
- With those checks actually running, fix the real issues they surface:
- clippy: useless_conversion in chunked_write.rs, byte_char_slices in
global_heap.rs/object_header.rs.
- cargo fmt: apply formatting across the workspace (whitespace only).
- no_std (thumbv7em-none-eabihf) build errors in clawhdf5-format:
core::sync::atomic::AtomicU64 doesn't exist on that target (no
native 64-bit atomics) — switch profiling.rs's counters to
portable-atomic, which falls back to a CAS-based emulation there
and is a no-op wrapper elsewhere. Add missing alloc imports for
Box (filters.rs), Vec (filters_szip.rs), and format! (dict_encoding.rs)
on no_std paths. Replace f64::powi (std/libm-only) with a small
local exponentiation-by-squaring helper in the scale-offset filter.
|
||
|
|
b70d594c4f |
perf: O(1) chunk cache lookup with shared Arc buffers instead of O(n) scan+clone
The decompressed-chunk LRU cache was the hottest path in the read pipeline (every chunked-dataset read goes through it) but did a linear scan through up to 521 slots on every get/put, and a full buffer copy on every cache hit (to_vec()/clone() of the whole decompressed chunk). chunked_read.rs then cloned the buffer a second time just to insert it into the cache after already having it in hand. - Added a HashMap<ChunkCoord, usize> index alongside the LRU slots for O(1) lookup. Eviction uses swap_remove, so the swapped-in slot's index entry is fixed up on every eviction (covered by a dedicated test). - CachedChunk.data is now Arc<CacheAlignedBuffer> — a cache hit is a refcount bump, not a copy. CacheAlignedBuffer gained a Sync impl (same soundness argument as its existing Send impl: access is only ever through borrow-checked &/&mut, like Vec<u8>) so Arc<CacheAlignedBuffer> is itself Send/Sync. - put_decompressed/put_decompressed_aligned now return the Arc they just inserted (or the existing cached copy), so callers can reuse that allocation instead of holding a separate clone — eliminates the second copy in chunked_read.rs's three call sites, which now consume the Arc<CacheAlignedBuffer> (Deref's to &[u8], so downstream indexing/copy code is unchanged). - prefetch_hint's doc comment now leads with "bookkeeping only, does not prefetch" instead of describing behavior it doesn't have. Co-Authored-By: Claude Sonnet 5 <[email protected]> |
||
|
|
b9898c2a9c |
security: bound decompression output to prevent memory-exhaustion DoS
decompress_chunk() already threaded chunk_size (the pipeline's declared decompressed size) into the scale-offset/nbit/szip decoders to bound their output, but not into deflate/lz4/zstd/pcodec, all four of which allocated based on attacker-controlled input with no cap: - lz4: read a raw u32 "orig_size" straight from the compressed payload's first 4 bytes and passed it directly to lz4_flex::block::decompress with no upper bound — a 4-byte attacker-controlled field could request ~4 GiB. - deflate (non-macOS path): unbounded flate2 read_to_end into a fresh Vec. - zstd: zstd::decode_all with no output cap (classic decompression-bomb vector, ratios can exceed 1000:1). - pcodec: simple_decompress with no cap. All four now take the expected chunk size and reject output that exceeds it (or a 256 MiB absolute ceiling when the size is unavailable), matching the pattern the other three filters already used. Also fixes the same unbounded read_to_end in clawhdf5-filters' fast_deflate streaming fallback (used when no size hint is available). Added tests for each codec plus one exercising the actually-exploited path through the public decompress_chunk() entrypoint. Co-Authored-By: Claude Sonnet 5 <[email protected]> |
||
|
|
88195d1c33 |
docs: fix untraceable benchmark claims, add dual-audience framing, validate on second machine
- README's "HDF5 Core I/O" table claimed 19ns/2,080µs labeled 308× (real ratio ~109,000×) and a 313ns zero-copy mmap figure — neither traced to any dated benchmark in BENCHMARKS.md. Replaced the table wholesale with the existing "vs libhdf5 Summary" figures, relabeled from "h5py/C HDF5" to "libhdf5" (BENCHMARKS.md never benchmarks against h5py, only libhdf5 directly). - Added two new Criterion benchmarks to close the coverage gaps that produced the untraceable numbers: metadata_open_from_disk (I/O-inclusive, fair clawhdf5-vs-libhdf5 file-open comparison) and metadata_parse_in_memory (clawhdf5-only, explicitly labeled as excluding I/O) in h5bench_meta.rs; read_zerocopy_mmap in h5bench_read.rs (forces real page-ins by summing elements rather than just returning a slice length — the mmap path turns out to be slower than a plain copy at these sizes, an honest, unflattering but real result now documented instead of a fabricated 313ns). - Re-ran the full existing benchmark suite plus the two new ones on a second, independently administered machine (tank: Ryzen 7 7800X3D) to validate the numbers before publishing them. 5 of 6 rows landed within ~15% of the original i7-12650H figures; recorded both in BENCHMARKS.md's new "Independent Validation" section. README now cites the tank numbers. - Added a short top-of-file README callout naming both halves of the project (general-purpose HDF5 library vs. agent memory layer) with links to BENCHMARKS.md and the Crate Map, so a data-infra reader isn't 60% through a memory-store pitch before finding the part relevant to them. - Added one factual, no-names line noting benchmark numbers are being validated in collaboration with HDF5 Group engineers. - Fixed the same untraceable "2-300x faster than h5py/C HDF5" / "313 ns" claims in docs/QUICKSTART.md, one click from the README's own "New here?" link. Co-Authored-By: Claude Sonnet 5 <[email protected]> |
||
|
|
6b1ea450f5 |
chore: cleanup pass — remove empty types stub, implement superblock v4, reconcile plan docs
- Remove clawhdf5-types (empty 1-line stub crate; type defs already live in
clawhdf5-format). Update workspace Cargo.toml and CLAUDE.md accordingly.
- Implement HDF5 superblock v4 (page-buffer mode) read and write support in
clawhdf5-format: Superblock::parse_v4, page_size field, v4 serialize
branch, and FileWriter::with_page_size. This was the one task left
unimplemented from docs/superpowers/plans/2026-06-29-format-write-extensions.md.
- Reconcile the three docs/superpowers/plans/*.md docs (filter codecs,
format write extensions, MPI-IO VOL) against actual shipped code: they
were pre-work plans for
|