The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us, on a heuristic (any big-endian VL mismatch) and
a fixed list.
- ref.py checks that the installed h5py returns big-endian VL elements
with the file's bytes under a little-endian dtype (writing and reading
a vlen('>f4') in memory) and, if so, relabels them with the file's byte
order before hashing, marking the object `ref_fix`. The values are now
compared: attr_datatypes.hdf5 /@vlen_uint64 and tcomplex_be.h5
/VariableLengthDatasetFloatComplex are identical to ours (h5dump 1.14.6
prints the same (1, 2), (3, 4, 5), (42)).
- ref_bugs.py re-reads each object h5py reads only through a libhdf5 bug
in six processes with different heaps (import order, MALLOC_PERTURB_).
Values the file determines are the same every time; these three change
(6, 6 and 3 distinct results), so they are over-read memory, not data
clawhdf5 could match. compare.py classifies a file `ref-bug` only when
every difference is such an object confirmed in the same run.
- report.py: the ref-bug class, the evidence table, the corrected
objects; test_ref.py covers both (run in the nightly job).
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2.7 KiB
Conformance sweep
Reads every HDF5 file of eight public corpora with clawhdf5 and with
h5py/libhdf5, compares the two readings object by object, and writes
CONFORMANCE.md.
CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh # ~30 s once the corpus is cached
conformance/run.sh --update-baseline # after an intended change in results
Needs Rust, git, h5dump (Debian/Ubuntu hdf5-tools), libaec (for the
probe's szip feature; libaec-dev), and a Python with the packages in
requirements.txt. The first run downloads about 450 MB of sparse checkouts.
| file | role |
|---|---|
corpus.txt |
the corpora: git URL, pinned commit, swept root, sparse-checkout patterns |
fetch-corpus.sh |
shallow, sparse, blob-filtered checkout of each pinned commit into .cache/src/ (gitignored); no-op when already there |
list_files.py |
which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers) |
probe/ |
the clawhdf5 side: a standalone crate (outside the workspace, so cargo test --workspace never builds it) that walks a file with clawhdf5-format and prints canonical JSON |
ref.py |
the h5py side: the same JSON from h5py (values corrected for a known h5py bug are marked ref_fix) |
ref_bugs.py |
re-reads the objects h5py reads only through a libhdf5 bug in six differently-set-up processes; an object whose values change is confirmed as a libhdf5 over-read |
test_ref.py |
tests of ref.py's correction and ref_bugs.py's confirmation (python conformance/test_ref.py) |
run_one.sh |
runs both sides on one file (and h5dump on the CVE corpus) under a timeout and an address-space limit |
compare.py |
classifies each file (ok / our-error / mismatch / h5py-cannot-read / ref-bug / panic / hang / crash / oom) and groups root causes |
report.py |
writes CONFORMANCE.md |
check.py |
the gate: fails on any panic/hang/crash/oom, on an ok count below baseline.json, or on a baseline-ok file that is no longer ok |
baseline.json |
the ok files the gate holds the line on |
requirements.txt |
pinned h5py / numpy / hdf5plugin / netCDF4 |
Results for every file (both sides' JSON and stderr, results.csv,
results.json, summary.md) are left in .cache/results/.
The nightly job is .gitea/workflows/conformance.yml; it prints the report
into the job log.
The canonical value encoding both sides hash is documented at the top of
probe/src/main.rs. Values are compared as libhdf5 presents them: a float
with a non-IEEE bit layout (N-Bit) or an integer with a bit offset is compared
as the converted number, not as raw file bytes.