Files
clawhdf5/crates/clawhdf5-accel
osobhandClaude Opus 5.5 3100f0143b docs: fix cross-links after the refresh; remove the stale benchmark script
- docs/README.md links the improvement logs and June plans where the
  refresh archived them (docs/archive/).
- clawhdf5-py README: 'r+' creates and replaces attributes (compact or
  dense); only deleting them is unsupported.
- scripts/run-benchmarks.sh benchmarked the pre-rename rustyhdf5-format
  and overwrote BENCHMARKS.md; nothing referenced it. Removed.
- Cargo.toml descriptions no longer name rustyhdf5/edgehdf5; clawhdf5-gpu
  says it is not HDF5 I/O.
- benchmarks/cross_platform.sh pointed at a ROADMAP section that no longer
  exists.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:15:18 -05:00
..

clawhdf5-accel

CPU SIMD kernels for vector search: dot products, cosine similarity, L2 distance, norms and int8 dot products, dispatched at run time to the best backend the CPU has, with a portable scalar fallback for every operation. clawhdf5-ann and clawhdf5-agent use it in their distance loops; it has nothing to do with HDF5 file I/O.

Not on crates.io yet; depend on it from git:

[dependencies]
clawhdf5-accel = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }

API

use clawhdf5_accel::{cosine_similarity, detect_backend, dot_i8, dot_product, l2_distance};

let a = [1.0f32, 2.0, 3.0, 4.0];
let b = [4.0f32, 3.0, 2.0, 1.0];
assert_eq!(dot_product(&a, &b), 20.0);
let _cos = cosine_similarity(&a, &b);
let _l2 = l2_distance(&a, &b);
assert_eq!(dot_i8(&[1, -2, 3], &[4, 5, -6]), -24);
println!("{:?}", detect_backend()); // e.g. Avx2 on x86-64, Neon on aarch64

Also vector_norm, batch_norms, batch_cosine, batch_cosine_prenorm, f16_to_f32_batch, checksum_fletcher32 and align_to_cache_line.

Backends

detect_backend() picks once per process: Avx512 (with the avx512 feature), Avx2 (AVX2 + FMA), Neon (every aarch64 CPU), or Scalar. Sse4 and WasmSimd128 are reported when detected but run the scalar kernels. dot_i8, used by the agent's quantised (int8) HNSW index, runs on AVX2 and on NEON — with the SDOT instruction (through inline assembly, since the intrinsic is unstable) on cores that have dotprod, such as the Raspberry Pi 5, and plain NEON on older ones. At equal recall the int8 index answers 1.63x the queries per second of the f32 one on x86-64 (AVX2) and 1.18x on a Raspberry Pi 5 (BENCHMARKS.md).

The aarch64 code is compiled out on x86, so only the test-arm64 CI job builds and tests it.

Features

Feature Default What Builds C
avx512 no AVX-512F kernels no
float16 no f16_to_f32_batch through the half crate (a software conversion otherwise) no

The half-precision conversion used for stored embeddings is clawhdf5_format::float16, not this crate's.

License

MIT