Files
clawhdf5/conformance
osobhandClaude Opus 5.5 742ed4dfb8 fix(format): read a v1 chunk B-tree where libhdf5's lookup finds chunks
libhdf5 does not walk the chunk B-tree to read a dataset: it looks each
chunk up (H5B_find with H5D__btree_cmp3 and H5D__btree_found), asking for
the element-size coordinate as 0. collect_chunk_info_checked now parses
the tree with its keys and returns each stored chunk only when that
lookup, replayed over the scaled keys, finds it.

A key with a non-zero element-size coordinate is therefore found in a
1-D dataset (cmp3 compares only the first coordinate there, and found
compares with <=) and missed in a dataset of rank 2 or more, which reads
fill values. The previous commit refused every such key, which refused
1-D files libhdf5 reads correctly; before that, the rank-2 case read the
chunk's data where h5py reads fill values (cve-2025-44905
/Shuffle_float_data_le, now identical to h5py, so it leaves the
conformance report's list of libhdf5 bugs).

Test: chunk_keys_with_an_element_offset_read_as_libhdf5_reads_them
compares 1-D and 2-D files against h5py's values. It fails on the
previous commit (the 1-D file is refused) and with the refusal removed
(the 2-D file reads 0..23 where h5py reads fill values).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 11:35:13 -05:00
..

Conformance sweep

Reads every HDF5 file of eight public corpora with clawhdf5 and with h5py/libhdf5, compares the two readings object by object, and writes CONFORMANCE.md.

CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh      # ~30 s once the corpus is cached
conformance/run.sh --update-baseline                             # after an intended change in results

Needs Rust, git, h5dump (Debian/Ubuntu hdf5-tools), libaec (for the probe's szip feature; libaec-dev), and a Python with the packages in requirements.txt. The first run downloads about 450 MB of sparse checkouts.

file role
corpus.txt the corpora: git URL, pinned commit, swept root, sparse-checkout patterns
fetch-corpus.sh shallow, sparse, blob-filtered checkout of each pinned commit into .cache/src/ (gitignored); no-op when already there
list_files.py which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers)
probe/ the clawhdf5 side: a standalone crate (outside the workspace, so cargo test --workspace never builds it) that walks a file with clawhdf5-format and prints canonical JSON
ref.py the h5py side: the same JSON from h5py
run_one.sh runs both sides on one file (and h5dump on the CVE corpus) under a timeout and an address-space limit
compare.py classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes
report.py writes CONFORMANCE.md
check.py the gate: fails on any panic/hang/crash/oom, on an ok count below baseline.json, or on a baseline-ok file that is no longer ok
baseline.json the ok files the gate holds the line on
requirements.txt pinned h5py / numpy / hdf5plugin / netCDF4

Results for every file (both sides' JSON and stderr, results.csv, results.json, summary.md) are left in .cache/results/.

The nightly job is .gitea/workflows/conformance.yml; it prints the report into the job log.

The canonical value encoding both sides hash is documented at the top of probe/src/main.rs. Values are compared as libhdf5 presents them: a float with a non-IEEE bit layout (N-Bit) or an integer with a bit offset is compared as the converted number, not as raw file bytes.