Files
clawhdf5/conformance
osobhandClaude Opus 5.5 a6ed3a5c7d fix(format): resolve cache-image flush-dependency parents as libhdf5 does
The review suggested libhdf5 loads every image entry and resolves
flush-dependency parents afterwards. It does not:
H5C__reconstruct_cache_contents (HDF5 1.14.6 and 2.0.0, and develop)
inserts each entry and then searches the cache index for its parents in
the same loop, failing with "fd parent not in cache?!?" when one is
missing. So a parent must be an earlier image entry, as before, or
metadata cached before the image loads: the superblock (address 0) and
the superblock extension's object header, which libhdf5 reads to find
the image. Those two were refused as parents; they are now accepted.
A parent listed after its child is still refused, as libhdf5 refuses
it, and so is an entry that is its own parent ("Child entry flush
dependency parent can't be itself").

apply_cache_image takes the superblock to know the extension address.

Test: superblock_ext::tests::flush_dependency_parents_must_already_be_cached
(parent-first loads, child-first refused, extension header accepted,
self-parent refused); the extension-header case fails without the fix.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
2026-09-26 11:36:10 -05:00
..

Conformance sweep

Reads every HDF5 file of eight public corpora with clawhdf5 and with h5py/libhdf5, compares the two readings object by object, and writes CONFORMANCE.md.

CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh      # ~30 s once the corpus is cached
conformance/run.sh --update-baseline                             # after an intended change in results

Needs Rust, git, h5dump (Debian/Ubuntu hdf5-tools), libaec (for the probe's szip feature; libaec-dev), and a Python with the packages in requirements.txt. The first run downloads about 450 MB of sparse checkouts.

file role
corpus.txt the corpora: git URL, pinned commit, swept root, sparse-checkout patterns
fetch-corpus.sh shallow, sparse, blob-filtered checkout of each pinned commit into .cache/src/ (gitignored); no-op when already there
list_files.py which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers)
probe/ the clawhdf5 side: a standalone crate (outside the workspace, so cargo test --workspace never builds it) that walks a file with clawhdf5-format and prints canonical JSON
ref.py the h5py side: the same JSON from h5py
run_one.sh runs both sides on one file (and h5dump on the CVE corpus) under a timeout and an address-space limit
compare.py classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes
report.py writes CONFORMANCE.md
check.py the gate: fails on any panic/hang/crash/oom, on an ok count below baseline.json, or on a baseline-ok file that is no longer ok
baseline.json the ok files the gate holds the line on
requirements.txt pinned h5py / numpy / hdf5plugin / netCDF4

Results for every file (both sides' JSON and stderr, results.csv, results.json, summary.md) are left in .cache/results/.

The nightly job is .gitea/workflows/conformance.yml; it prints the report into the job log.

The canonical value encoding both sides hash is documented at the top of probe/src/main.rs. Values are compared as libhdf5 presents them: a float with a non-IEEE bit layout (N-Bit) or an integer with a bit offset is compared as the converted number, not as raw file bytes.