# Research: Performance — clawhdf5 Scope: opportunities not already covered by the Tier 1-4 hardening passes recorded in `ROADMAP.md`/`CHANGELOG.md`/`IMPROVEMENT_LOG.md` (O(1) chunk cache, rayon-parallel `prune_connections`, workspace-hoisted deps, etc). ## Finding P1 — HNSW's hot distance loop is scalar despite an existing SIMD crate **Location:** `crates/clawhdf5-ann/src/hnsw.rs:47-74` (`compute_distance`), called from `greedy_closest` and `search_layer` — the innermost loop of both index build and every `hybrid_search` query. **Problem:** `compute_distance` is a plain per-component `for i in 0..a.len()` scalar loop for both the `L2` and `Cosine` metrics. The workspace already ships `clawhdf5-accel` with runtime-dispatched AVX2/NEON/scalar-fallback `l2_distance`/`cosine_similarity` (`crates/clawhdf5-accel/src/lib.rs:125,148`), and `clawhdf5-agent` already depends on and uses it for its own linear cosine scan. `clawhdf5-ann/Cargo.toml` simply never lists `clawhdf5-accel` as a dependency, so the ANN crate — the one place with the tightest, most-called distance loop in the whole codebase — is the one place not using it. **Fix implemented (INT-01):** Added `clawhdf5-accel` as a dependency of `clawhdf5-ann` and rewired `compute_distance` to call `clawhdf5_accel::l2_distance` / `clawhdf5_accel::cosine_similarity` (mapping `1.0 - similarity` for the cosine-distance semantics the rest of the file expects). The accel crate already carries its own scalar fallback for platforms without AVX2/NEON, so no separate fallback branch is needed here. Existing `hnsw.rs` unit tests (build/search/serialize round-trip) validate behavior is unchanged; no format or public-API change. ## Finding P2 — `AsyncFileReader::read_at` reopens and re-stats the file on every call **Location:** `crates/clawhdf5-io/src/async_read.rs:90-104`. **Problem:** Each `read_at` call does `tokio::fs::File::open` + `.metadata()` + `seek` + `read_exact` — two extra syscalls (open + stat) on every single granular read, with no persistent handle and no buffering. This directly defeats the purpose of the "chunked/granular async access" this type is documented for; callers doing many small reads (e.g. chunked dataset iteration) pay file-open overhead per chunk. **Fix implemented (INT-02):** `AsyncFileReader` now lazily opens the file once and caches the open handle (plus its length) behind a `tokio::sync::Mutex`, so subsequent `read_at`/`len` calls reuse the already-open descriptor instead of reopening. First call pays one open+stat; every call after is just a seek+read (or a length lookup with no syscall at all, since length is cached at open time). Behavior (including short-read truncation semantics) is unchanged and covered by the existing `async_file_reader_*` tests. ## Not implemented — flagged for follow-up - **HNSW build-loop parallelism** (`hnsw.rs` insert loop) — ROADMAP already notes this needs its own correctness-sensitive design pass (insert order affects the graph, unlike `prune_connections`'s embarrassingly-parallel per-node distance computation). Left as-is; out of scope for this pass.