The review suggested libhdf5 loads every image entry and resolves
flush-dependency parents afterwards. It does not:
H5C__reconstruct_cache_contents (HDF5 1.14.6 and 2.0.0, and develop)
inserts each entry and then searches the cache index for its parents in
the same loop, failing with "fd parent not in cache?!?" when one is
missing. So a parent must be an earlier image entry, as before, or
metadata cached before the image loads: the superblock (address 0) and
the superblock extension's object header, which libhdf5 reads to find
the image. Those two were refused as parents; they are now accepted.
A parent listed after its child is still refused, as libhdf5 refuses
it, and so is an entry that is its own parent ("Child entry flush
dependency parent can't be itself").
apply_cache_image takes the superblock to know the extension address.
Test: superblock_ext::tests::flush_dependency_parents_must_already_be_cached
(parent-first loads, child-first refused, extension header accepted,
self-parent refused); the extension-header case fails without the fix.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Conformance sweep
Reads every HDF5 file of eight public corpora with clawhdf5 and with
h5py/libhdf5, compares the two readings object by object, and writes
CONFORMANCE.md.
CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh # ~30 s once the corpus is cached
conformance/run.sh --update-baseline # after an intended change in results
Needs Rust, git, h5dump (Debian/Ubuntu hdf5-tools), libaec (for the
probe's szip feature; libaec-dev), and a Python with the packages in
requirements.txt. The first run downloads about 450 MB of sparse checkouts.
| file | role |
|---|---|
corpus.txt |
the corpora: git URL, pinned commit, swept root, sparse-checkout patterns |
fetch-corpus.sh |
shallow, sparse, blob-filtered checkout of each pinned commit into .cache/src/ (gitignored); no-op when already there |
list_files.py |
which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers) |
probe/ |
the clawhdf5 side: a standalone crate (outside the workspace, so cargo test --workspace never builds it) that walks a file with clawhdf5-format and prints canonical JSON |
ref.py |
the h5py side: the same JSON from h5py |
run_one.sh |
runs both sides on one file (and h5dump on the CVE corpus) under a timeout and an address-space limit |
compare.py |
classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes |
report.py |
writes CONFORMANCE.md |
check.py |
the gate: fails on any panic/hang/crash/oom, on an ok count below baseline.json, or on a baseline-ok file that is no longer ok |
baseline.json |
the ok files the gate holds the line on |
requirements.txt |
pinned h5py / numpy / hdf5plugin / netCDF4 |
Results for every file (both sides' JSON and stderr, results.csv,
results.json, summary.md) are left in .cache/results/.
The nightly job is .gitea/workflows/conformance.yml; it prints the report
into the job log.
The canonical value encoding both sides hash is documented at the top of
probe/src/main.rs. Values are compared as libhdf5 presents them: a float
with a non-IEEE bit layout (N-Bit) or an integer with a bit offset is compared
as the converted number, not as raw file bytes.