apply_cache_image returned a copy of the whole file with the image's
entries written in, and File (mmap by default), MmapFile and LazyFile
used that copy for every read: opening a 1 GiB sparse file with an image
needed 2 GB of memory, and an 8 GiB one aborted the process, where
de2a53f (which ignored the image) opened them in a few MB.
The metadata parsers read one contiguous slice, so the image still has
to be laid over the file's bytes; it is now laid over a private copy
that costs only the pages it touches:
- clawhdf5_format::superblock_ext::CacheImage decodes the image into an
entry list (address, offset in the block, length) and applies it to
any destination; cache_image_state tells an opener whether the file
has no image, a loadable one, or one libhdf5 cannot load;
apply_cache_image_in_place is for readers that own their buffer.
apply_cache_image and metadata_view (which copied) are gone.
- clawhdf5_io::HDF5Read::private_copy returns a writable private copy
of a reader's bytes: MmapReader gives a MAP_PRIVATE copy-on-write
mapping (memmap2 map_copy), so only the pages the entries land on are
copied; the default copies the bytes (in-memory readers).
- File, MmapFile and LazyFile write the image into that mapping
(crate::cache_image). File::from_bytes / open_buffered patch their own
buffer in place, copying only the image block, as libhdf5 does. A
file without an image is read straight from the mapping, unchanged.
An image entry that runs past the end of file is now refused: libhdf5
checks only that it starts inside the file, and the images libhdf5
writes never do this, but those bytes have nowhere to go in a view of
the file.
Tests: tests/cache_image_memory.rs has libhdf5 (through ctypes) add an
image to a 1 GiB sparse file and bounds resident-memory growth for all
three openers at 256 MiB; it fails on the previous commit (File::open
grew 2,148,720,640 bytes). reader.rs zero_copy_tests check that a file
without an image is read from the mapping itself and that an image goes
into a copy-on-write mapping, not a heap copy; clawhdf5-io checks that
private_copy writes never reach the reader or the file.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Conformance sweep
Reads every HDF5 file of eight public corpora with clawhdf5 and with
h5py/libhdf5, compares the two readings object by object, and writes
CONFORMANCE.md.
CLAWHDF5_PYTHON=/path/to/venv/bin/python conformance/run.sh # ~30 s once the corpus is cached
conformance/run.sh --update-baseline # after an intended change in results
Needs Rust, git, h5dump (Debian/Ubuntu hdf5-tools), libaec (for the
probe's szip feature; libaec-dev), and a Python with the packages in
requirements.txt. The first run downloads about 450 MB of sparse checkouts.
| file | role |
|---|---|
corpus.txt |
the corpora: git URL, pinned commit, swept root, sparse-checkout patterns |
fetch-corpus.sh |
shallow, sparse, blob-filtered checkout of each pinned commit into .cache/src/ (gitignored); no-op when already there |
list_files.py |
which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers) |
probe/ |
the clawhdf5 side: a standalone crate (outside the workspace, so cargo test --workspace never builds it) that walks a file with clawhdf5-format and prints canonical JSON |
ref.py |
the h5py side: the same JSON from h5py |
run_one.sh |
runs both sides on one file (and h5dump on the CVE corpus) under a timeout and an address-space limit |
compare.py |
classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes |
report.py |
writes CONFORMANCE.md |
check.py |
the gate: fails on any panic/hang/crash/oom, on an ok count below baseline.json, or on a baseline-ok file that is no longer ok |
baseline.json |
the ok files the gate holds the line on |
requirements.txt |
pinned h5py / numpy / hdf5plugin / netCDF4 |
Results for every file (both sides' JSON and stderr, results.csv,
results.json, summary.md) are left in .cache/results/.
The nightly job is .gitea/workflows/conformance.yml; it prints the report
into the job log.
The canonical value encoding both sides hash is documented at the top of
probe/src/main.rs. Values are compared as libhdf5 presents them: a float
with a non-IEEE bit layout (N-Bit) or an integer with a bit offset is compared
as the converted number, not as raw file bytes.