Files
clawhdf5/conformance
osobhandClaude Opus 5.5 b22b15f00a fix: refuse at open the dataset storage libhdf5 refuses at open
libhdf5 checks a dataset's storage when it opens the dataset
(H5D__contig_check, H5D__compact_init): the element count times the
element size must not overflow, contiguous storage must end inside the
file, compact data must be the dataset's size. File::dataset opened
cve-2024-32624's /Dset_OBJREF (2^62 + 2 references of 8 bytes) and
reported its shape; only reading failed.

data_read::check_dataset_storage makes those checks (new
FormatError::InvalidDatasetStorage), and File, MmapFile and LazyFile run
it whenever they open a dataset (by path, by address, from a group), as
does the conformance probe. As before, a datatype, dataspace or layout
that does not decode is left for the read to report, so such a dataset
still opens and its attributes still read. An empty contiguous dataset at
a defined address, which libhdf5 refuses, is still accepted: clawhdf5 up
to v2.7.0 wrote them.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 10:22:04 -05:00
..

Conformance sweep

Reads every HDF5 file of eight public corpora with clawhdf5 and with h5py/libhdf5, compares the two readings object by object, and writes CONFORMANCE.md.

CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh      # ~30 s once the corpus is cached
conformance/run.sh --update-baseline                             # after an intended change in results

Needs Rust, git, h5dump (Debian/Ubuntu hdf5-tools), libaec (for the probe's szip feature; libaec-dev), and a Python with the packages in requirements.txt. The first run downloads about 450 MB of sparse checkouts.

file role
corpus.txt the corpora: git URL, pinned commit, swept root, sparse-checkout patterns
fetch-corpus.sh shallow, sparse, blob-filtered checkout of each pinned commit into .cache/src/ (gitignored); no-op when already there
list_files.py which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers)
probe/ the clawhdf5 side: a standalone crate (outside the workspace, so cargo test --workspace never builds it) that walks a file with clawhdf5-format and prints canonical JSON
ref.py the h5py side: the same JSON from h5py
run_one.sh runs both sides on one file (and h5dump on the CVE corpus) under a timeout and an address-space limit
compare.py classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes
report.py writes CONFORMANCE.md
check.py the gate: fails on any panic/hang/crash/oom, on an ok count below baseline.json, or on a baseline-ok file that is no longer ok
baseline.json the ok files the gate holds the line on
requirements.txt pinned h5py / numpy / hdf5plugin / netCDF4

Results for every file (both sides' JSON and stderr, results.csv, results.json, summary.md) are left in .cache/results/.

The nightly job is .gitea/workflows/conformance.yml; it prints the report into the job log.

The canonical value encoding both sides hash is documented at the top of probe/src/main.rs. Values are compared as libhdf5 presents them: a float with a non-IEEE bit layout (N-Bit) or an integer with a bit offset is compared as the converted number, not as raw file bytes.