Numbers, API names, feature defaults and PR references checked against CONFORMANCE.md, BENCHMARKS.md, CHANGELOG.md, the code and git history. - int8 index figures (1.74x memory, 1.63x QPS) carry the dates git gives them (2026-09-19/20, machine not recorded, not re-run) instead of none; the Pi 5 1.18x carries 2026-09-21. - BENCHMARKS headline: the libhdf5 chunked-write figure is the newest measurement (35x, 2026-09-23), not 45.3x (2026-08-03). - Conformance counts follow the 2026-09-28 run (1 our-error, 2 ref-bug) in conformance/README.md, ROADMAP.md and CLAUDE.md, with a pointer to the bad_nbit_parms_walk.h5 flip. - README: LZ4 is opt-in; the browser refuses reference/opaque/bitfield/ time datasets too; zlib-rs byte-identity scoped to what was measured; macOS default links the system libz for inflate. - Crate READMEs: system-zlib-decompress does something (macOS), SweepDetector lives in prefetch, checkpoint after more than 500 WAL entries, NetCDF-4 unlimited-dimension size warning. - agent-memory.md: string-dataset compression threshold, agents-md prints Markdown, float16 file sizes linked to their study. - known-issues.md: contiguous selection reads, 1.21x vs h5py threads. - docs/README.md, USE_CASES.md, ROADMAP.md, CLAUDE.md: range-read milestones M0-M5 and PRs #17-#19, missing README rows, CLI keygen/verify, dated figures, fast-math is not BLAS. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
64 lines
2.3 KiB
Markdown
64 lines
2.3 KiB
Markdown
# clawhdf5-accel
|
|
|
|
CPU SIMD kernels for vector search: dot products, cosine similarity, L2
|
|
distance, norms and int8 dot products, dispatched at run time to the best
|
|
backend the CPU has, with a portable scalar fallback for every operation.
|
|
[`clawhdf5-ann`](../clawhdf5-ann/README.md) and
|
|
[`clawhdf5-agent`](../clawhdf5-agent/README.md) use it in their distance
|
|
loops; it has nothing to do with HDF5 file I/O.
|
|
|
|
Not on crates.io yet; depend on it from git:
|
|
|
|
```toml
|
|
[dependencies]
|
|
clawhdf5-accel = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
|
|
```
|
|
|
|
## API
|
|
|
|
```rust
|
|
use clawhdf5_accel::{cosine_similarity, detect_backend, dot_i8, dot_product, l2_distance};
|
|
|
|
let a = [1.0f32, 2.0, 3.0, 4.0];
|
|
let b = [4.0f32, 3.0, 2.0, 1.0];
|
|
assert_eq!(dot_product(&a, &b), 20.0);
|
|
let _cos = cosine_similarity(&a, &b);
|
|
let _l2 = l2_distance(&a, &b);
|
|
assert_eq!(dot_i8(&[1, -2, 3], &[4, 5, -6]), -24);
|
|
println!("{:?}", detect_backend()); // e.g. Avx2 on x86-64, Neon on aarch64
|
|
```
|
|
|
|
Also `vector_norm`, `batch_norms`, `batch_cosine`, `batch_cosine_prenorm`,
|
|
`f16_to_f32_batch`, `checksum_fletcher32` and `align_to_cache_line`.
|
|
|
|
## Backends
|
|
|
|
`detect_backend()` picks once per process: `Avx512` (with the `avx512`
|
|
feature), `Avx2` (AVX2 + FMA), `Neon` (every aarch64 CPU), or `Scalar`.
|
|
`Sse4` and `WasmSimd128` are reported when detected but run the scalar
|
|
kernels.
|
|
`dot_i8`, used by the agent's quantised (int8) HNSW index, runs on
|
|
AVX2 and on NEON — with the `SDOT` instruction (through inline assembly,
|
|
since the intrinsic is unstable) on cores that have dotprod, such as the
|
|
Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8
|
|
index answers 1.63x the queries per second of the f32 one on x86-64
|
|
(AVX2; 2026-09-20, machine not recorded, not re-run) and 1.18x on a
|
|
Raspberry Pi 5 (2026-09-21) ([`BENCHMARKS.md` § Quantising the index copy](../../BENCHMARKS.md#quantising-the-index-copy-quantized_index)).
|
|
|
|
The aarch64 code is compiled out on x86, so only the `test-arm64` CI job
|
|
builds and tests it.
|
|
|
|
## Features
|
|
|
|
| Feature | Default | What | Builds C |
|
|
|---|---|---|---|
|
|
| `avx512` | no | AVX-512F kernels | no |
|
|
| `float16` | no | `f16_to_f32_batch` through the `half` crate (a software conversion otherwise) | no |
|
|
|
|
The half-precision conversion used for stored embeddings is
|
|
`clawhdf5_format::float16`, not this crate's.
|
|
|
|
## License
|
|
|
|
MIT
|