Every crate under crates/ now has a README (android, bench, cli, napi and wasm had none), each saying what the crate is, its main types and functions (names checked against the code), its cargo features with defaults and which ones build C (checked with `cargo tree`), and links to the top-level docs. Corrections to the old stubs: - clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs clawhdf5-format as a dependency. - clawhdf5-filters: deflate backends only, and no library crate depends on it; the filter pipeline and every other codec are in -format. - clawhdf5-gpu: vector distance compute, not I/O; not used by HDF5Memory::search. - clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO. - clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and search(q, k, ef). - clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm backends are reported but run the scalar kernels. - clawhdf5-gpu: the old example called l2_distances, which does not exist (l2_search). - clawhdf5-agent: it described a "vector store" with "GPU acceleration"; it now covers HDF5Memory, search options, WAL, signing, the graph. - crates.io/docs.rs badges removed and `cargo install <crate>` replaced: nothing is published; depend on git. - fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh. - tools: the FileEditor interop tests that live in this crate. - remote, py: license, other front ends, limits, File.mode/flush/chunks. The Rust examples of the facade, format, filters, accel, ann, derive and agent READMEs were compiled and run as tests (netcdf4, gpu and remote compiled only) in a scratch crate; the CLI example was run. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2.2 KiB
clawhdf5-accel
CPU SIMD kernels for vector search: dot products, cosine similarity, L2
distance, norms and int8 dot products, dispatched at run time to the best
backend the CPU has, with a portable scalar fallback for every operation.
clawhdf5-ann and
clawhdf5-agent use it in their distance
loops; it has nothing to do with HDF5 file I/O.
Not on crates.io yet; depend on it from git:
[dependencies]
clawhdf5-accel = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
API
use clawhdf5_accel::{cosine_similarity, detect_backend, dot_i8, dot_product, l2_distance};
let a = [1.0f32, 2.0, 3.0, 4.0];
let b = [4.0f32, 3.0, 2.0, 1.0];
assert_eq!(dot_product(&a, &b), 20.0);
let _cos = cosine_similarity(&a, &b);
let _l2 = l2_distance(&a, &b);
assert_eq!(dot_i8(&[1, -2, 3], &[4, 5, -6]), -24);
println!("{:?}", detect_backend()); // e.g. Avx2 on x86-64, Neon on aarch64
Also vector_norm, batch_norms, batch_cosine, batch_cosine_prenorm,
f16_to_f32_batch, checksum_fletcher32 and align_to_cache_line.
Backends
detect_backend() picks once per process: Avx512 (with the avx512
feature), Avx2 (AVX2 + FMA), Neon (every aarch64 CPU), or Scalar.
Sse4 and WasmSimd128 are reported when detected but run the scalar
kernels.
dot_i8, used by the agent's quantised (int8) HNSW index, runs on
AVX2 and on NEON — with the SDOT instruction (through inline assembly,
since the intrinsic is unstable) on cores that have dotprod, such as the
Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8
index answers 1.63x the queries per second of the f32 one on x86-64
(AVX2) and 1.18x on a Raspberry Pi 5 (BENCHMARKS.md).
The aarch64 code is compiled out on x86, so only the test-arm64 CI job
builds and tests it.
Features
| Feature | Default | What | Builds C |
|---|---|---|---|
avx512 |
no | AVX-512F kernels | no |
float16 |
no | f16_to_f32_batch through the half crate (a software conversion otherwise) |
no |
The half-precision conversion used for stored embeddings is
clawhdf5_format::float16, not this crate's.
License
MIT