- docs/README.md links the improvement logs and June plans where the refresh archived them (docs/archive/). - clawhdf5-py README: 'r+' creates and replaces attributes (compact or dense); only deleting them is unsupported. - scripts/run-benchmarks.sh benchmarked the pre-rename rustyhdf5-format and overwrote BENCHMARKS.md; nothing referenced it. Removed. - Cargo.toml descriptions no longer name rustyhdf5/edgehdf5; clawhdf5-gpu says it is not HDF5 I/O. - benchmarks/cross_platform.sh pointed at a ROADMAP section that no longer exists. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
clawhdf5-gpu
GPU vector distance computation through wgpu and hand-written WGSL compute shaders: upload a set of vectors once, then run cosine or L2 top-k searches, dot products, distance matrices and norms against them on Vulkan, Metal, DirectX 12 or OpenGL.
This crate does not read or write HDF5: dataset I/O in clawhdf5 is
CPU-only. It is a vector-search accelerator used optionally by
clawhdf5-agent (its gpu feature exposes
gpu_search::GpuSearchBackend and a GPU arm of strategy::search_with_metrics;
HDF5Memory::search itself uses the HNSW index on the CPU).
Not on crates.io yet; depend on it from git:
[dependencies]
clawhdf5-gpu = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
Usage
use clawhdf5_gpu::GpuAccelerator;
// Fall back to a CPU path when there is no usable GPU.
let mut gpu = match GpuAccelerator::new() {
Ok(g) => g,
Err(_) => return,
};
let dim = 128;
let vectors = vec![0.5f32; 1000 * dim]; // 1000 vectors, row-major
gpu.upload_vectors(&vectors, dim).unwrap();
let norms = gpu.compute_norms_gpu(&vectors, dim).unwrap();
gpu.upload_norms(&norms).unwrap();
let query = vec![1.0f32; dim];
let top10 = gpu.cosine_search(&query, 10).unwrap(); // (index, similarity), best first
let near10 = gpu.l2_search(&query, 10).unwrap(); // (index, distance), nearest first
GpuAccelerator also has is_available, device_info,
batch_cosine_search, batch_dot_product, distance_matrix,
compute_norms, and f16_to_f32_batch/f32_to_f16_batch. Vector sets
larger than the device's largest storage buffer binding are split into
chunks and the results merged. A GPU→CPU readback waits at most 30 s and
then fails with GpuError::BufferMap instead of hanging.
Features
| Feature | Default | What |
|---|---|---|
gpu-wgpu |
yes | the wgpu implementation. Without it GpuAccelerator::new() returns GpuError::NotCompiled and is_available() is false. |
No C is compiled, but wgpu talks to the system's graphics drivers at run time; the crate is exempt from CI's "no C in the default build" check for that reason.
License
MIT