Files
clawhdf5/crates/clawhdf5-gpu
osobhandClaude Opus 5.5 3100f0143b docs: fix cross-links after the refresh; remove the stale benchmark script
- docs/README.md links the improvement logs and June plans where the
  refresh archived them (docs/archive/).
- clawhdf5-py README: 'r+' creates and replaces attributes (compact or
  dense); only deleting them is unsupported.
- scripts/run-benchmarks.sh benchmarked the pre-rename rustyhdf5-format
  and overwrote BENCHMARKS.md; nothing referenced it. Removed.
- Cargo.toml descriptions no longer name rustyhdf5/edgehdf5; clawhdf5-gpu
  says it is not HDF5 I/O.
- benchmarks/cross_platform.sh pointed at a ROADMAP section that no longer
  exists.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-28 11:15:18 -05:00
..

clawhdf5-gpu

GPU vector distance computation through wgpu and hand-written WGSL compute shaders: upload a set of vectors once, then run cosine or L2 top-k searches, dot products, distance matrices and norms against them on Vulkan, Metal, DirectX 12 or OpenGL.

This crate does not read or write HDF5: dataset I/O in clawhdf5 is CPU-only. It is a vector-search accelerator used optionally by clawhdf5-agent (its gpu feature exposes gpu_search::GpuSearchBackend and a GPU arm of strategy::search_with_metrics; HDF5Memory::search itself uses the HNSW index on the CPU).

Not on crates.io yet; depend on it from git:

[dependencies]
clawhdf5-gpu = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }

Usage

use clawhdf5_gpu::GpuAccelerator;

// Fall back to a CPU path when there is no usable GPU.
let mut gpu = match GpuAccelerator::new() {
    Ok(g) => g,
    Err(_) => return,
};

let dim = 128;
let vectors = vec![0.5f32; 1000 * dim]; // 1000 vectors, row-major
gpu.upload_vectors(&vectors, dim).unwrap();
let norms = gpu.compute_norms_gpu(&vectors, dim).unwrap();
gpu.upload_norms(&norms).unwrap();

let query = vec![1.0f32; dim];
let top10 = gpu.cosine_search(&query, 10).unwrap(); // (index, similarity), best first
let near10 = gpu.l2_search(&query, 10).unwrap();    // (index, distance), nearest first

GpuAccelerator also has is_available, device_info, batch_cosine_search, batch_dot_product, distance_matrix, compute_norms, and f16_to_f32_batch/f32_to_f16_batch. Vector sets larger than the device's largest storage buffer binding are split into chunks and the results merged. A GPU→CPU readback waits at most 30 s and then fails with GpuError::BufferMap instead of hanging.

Features

Feature Default What
gpu-wgpu yes the wgpu implementation. Without it GpuAccelerator::new() returns GpuError::NotCompiled and is_available() is false.

No C is compiled, but wgpu talks to the system's graphics drivers at run time; the crate is exempt from CI's "no C in the default build" check for that reason.

License

MIT