Every crate under crates/ now has a README (android, bench, cli, napi and wasm had none), each saying what the crate is, its main types and functions (names checked against the code), its cargo features with defaults and which ones build C (checked with `cargo tree`), and links to the top-level docs. Corrections to the old stubs: - clawhdf5-derive: the derive is `H5Type`, not `HDF5Type`, and it needs clawhdf5-format as a dependency. - clawhdf5-filters: deflate backends only, and no library crate depends on it; the filter pipeline and every other codec are in -format. - clawhdf5-gpu: vector distance compute, not I/O; not used by HDF5Memory::search. - clawhdf5-io: MpiVol is root-read + broadcast, not collective MPI-IO. - clawhdf5-ann: from_hdf5/search(q, k) did not exist; load_from_hdf5 and search(q, k, ef). - clawhdf5-accel: checksum::crc32_simd did not exist; the SSE4 and wasm backends are reported but run the scalar kernels. - clawhdf5-gpu: the old example called l2_distances, which does not exist (l2_search). - clawhdf5-agent: it described a "vector store" with "GPU acceleration"; it now covers HDF5Memory, search options, WAL, signing, the graph. - crates.io/docs.rs badges removed and `cargo install <crate>` replaced: nothing is published; depend on git. - fuzz: the opt-in CLAWHDF5_FUZZ_SECONDS smoke run in ci-test.sh. - tools: the FileEditor interop tests that live in this crate. - remote, py: license, other front ends, limits, File.mode/flush/chunks. The Rust examples of the facade, format, filters, accel, ann, derive and agent READMEs were compiled and run as tests (netcdf4, gpu and remote compiled only) in a scratch crate; the CLI example was run. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
3.2 KiB
Fuzz Testing for clawhdf5-format
Uses cargo-fuzz (libFuzzer) to test parser robustness against malformed inputs.
Prerequisites
cargo install cargo-fuzz
rustup toolchain install nightly
Fuzz Targets
| Target | Parser | Description |
|---|---|---|
fuzz_superblock |
Superblock::parse |
Superblock parsing (v0-v3) with signature search |
fuzz_object_header |
ObjectHeader::parse |
Object header v1/v2 with various offset/length sizes |
fuzz_datatype |
Datatype::parse |
All 12 HDF5 datatype classes (recursive) |
fuzz_dataspace |
Dataspace::parse |
Dataspace messages with various length sizes |
fuzz_fractal_heap |
FractalHeapHeader::parse |
Fractal heap header parsing |
fuzz_btree_v2 |
BTreeV2Header::parse |
B-tree v2 header parsing |
fuzz_filter_pipeline |
FilterPipeline::parse |
Filter pipeline messages (v1/v2) |
fuzz_full_file |
signature + superblock + root group | End-to-end file parsing chain |
fuzz_dataset_read |
Dataset::read_* (via clawhdf5) |
Walks every dataset in the parsed file and exercises the contiguous/chunked/compact raw-data read paths (chunked_read.rs, data_read.rs) that fuzz_full_file doesn't reach |
Running
Run a single target (runs indefinitely until stopped or a crash is found):
cd crates/clawhdf5-format
cargo +nightly fuzz run fuzz_datatype
Run with a time limit (seconds):
cargo +nightly fuzz run fuzz_datatype -- -max_total_time=60
Run all targets for 30 seconds each:
for target in fuzz_superblock fuzz_object_header fuzz_datatype fuzz_dataspace \
fuzz_fractal_heap fuzz_btree_v2 fuzz_filter_pipeline fuzz_full_file \
fuzz_dataset_read; do
echo "=== $target ==="
cargo +nightly fuzz run "$target" -- -max_total_time=30 -max_len=4096
done
CI
These targets are not run by the CI workflows (.gitea/workflows/ci.yml)
— cargo-fuzz requires nightly and each meaningful run takes minutes, which
doesn't fit a per-PR gate. Run them by hand before a release or after
touching parser code. scripts/ci-test.sh has an opt-in smoke run: with
CLAWHDF5_FUZZ_SECONDS=N it runs every target of this crate and of
crates/clawhdf5-agent/fuzz (the WAL parser) for N seconds each.
Other robustness checks that do run: the nightly conformance sweep reads
the HDF Group's CVE reproducers and fails on any panic, hang, crash or
out-of-memory (conformance/README.md),
and scripts/h5rs-fuzz.sh runs every h5rs subcommand over them, optionally
on byte-flipped copies.
Reproducing Crashes
If a crash is found, the input is saved to fuzz/artifacts/<target>/. Reproduce with:
cargo +nightly fuzz run fuzz_datatype fuzz/artifacts/fuzz_datatype/crash-<hash>
Minimize the crashing input:
cargo +nightly fuzz tmin fuzz_datatype fuzz/artifacts/fuzz_datatype/crash-<hash>
Design
Each fuzz target feeds arbitrary bytes directly to a parser entry point. The parsers must never panic on any input -- they should return Err(FormatError) for malformed data. Any panic found by fuzzing is a bug that should be fixed with proper bounds checks and error returns.