ref.py and compare.py reported 13 files as mismatches or our-errors that were artefacts of the harness, not differences between the readers: - User-defined links (tall.h5, tudlink.h5, twithub*.h5, tmany.h5, ...): h5py's `get(name, getlink=True)` reports a user-defined link as a HardLink, so ref.py listed it as an object. Read the link type from H5Lget_info instead. - Objects h5py cannot open (cve-2019-8397/8398, cve-2021-46243, cve-2024-32618): the probe deduplicates by header address, ref.py by ObjectID, which an unopenable object does not have, so each extra hard link to it was listed again. Deduplicate those by link address. - Nested array types (tarray3.h5): h5py expands them into trailing dims; hash_values stripped one level and numpy broadcast every element into a whole subarray. Strip every level. compare.py no longer compares the attributes or links of an object h5py could not open at all (cve-2018-17438/17439, cve-2019-9151): h5py read none, so ours are neither extra nor errors against it. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Conformance sweep
Reads every HDF5 file of eight public corpora with clawhdf5 and with
h5py/libhdf5, compares the two readings object by object, and writes
CONFORMANCE.md.
CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh # ~30 s once the corpus is cached
conformance/run.sh --update-baseline # after an intended change in results
Needs Rust, git, h5dump (Debian/Ubuntu hdf5-tools), libaec (for the
probe's szip feature; libaec-dev), and a Python with the packages in
requirements.txt. The first run downloads about 450 MB of sparse checkouts.
| file | role |
|---|---|
corpus.txt |
the corpora: git URL, pinned commit, swept root, sparse-checkout patterns |
fetch-corpus.sh |
shallow, sparse, blob-filtered checkout of each pinned commit into .cache/src/ (gitignored); no-op when already there |
list_files.py |
which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers) |
probe/ |
the clawhdf5 side: a standalone crate (outside the workspace, so cargo test --workspace never builds it) that walks a file with clawhdf5-format and prints canonical JSON |
ref.py |
the h5py side: the same JSON from h5py |
run_one.sh |
runs both sides on one file (and h5dump on the CVE corpus) under a timeout and an address-space limit |
compare.py |
classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes |
report.py |
writes CONFORMANCE.md |
check.py |
the gate: fails on any panic/hang/crash/oom, on an ok count below baseline.json, or on a baseline-ok file that is no longer ok |
baseline.json |
the ok files the gate holds the line on |
requirements.txt |
pinned h5py / numpy / hdf5plugin / netCDF4 |
Results for every file (both sides' JSON and stderr, results.csv,
results.json, summary.md) are left in .cache/results/.
The nightly job is .gitea/workflows/conformance.yml; it prints the report
into the job log.
The canonical value encoding both sides hash is documented at the top of
probe/src/main.rs. Values are compared as libhdf5 presents them: a float
with a non-IEEE bit layout (N-Bit) or an integer with a bit offset is compared
as the converted number, not as raw file bytes.