Merge branch 'perf/header-parse-and-last-files' into feat/listing-header-last-files

This commit is contained in:
osobh
2026-09-27 23:16:23 -05:00
13 changed files with 532 additions and 117 deletions
+2
View File
@@ -40,6 +40,8 @@ jobs:
run: cargo test --release --manifest-path conformance/probe/Cargo.toml run: cargo test --release --manifest-path conformance/probe/Cargo.toml
env: env:
CARGO_TARGET_DIR: conformance/.cache/target CARGO_TARGET_DIR: conformance/.cache/target
- name: Reference-side tests
run: /opt/conformance/bin/python conformance/test_ref.py
- name: Sweep - name: Sweep
# The corpora come from GitHub (pinned commits, conformance/corpus.txt), # The corpora come from GitHub (pinned commits, conformance/corpus.txt),
# so this job needs a runner that reaches github.com. # so this job needs a runner that reaches github.com.
+40
View File
@@ -500,6 +500,46 @@ explain the slower windows.
## Local file speed after range reads ## Local file speed after range reads
### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank)
The remaining 4% (below) was the call to the version-1 message loop, which
`4313917` kept out of line with `#[inline(never)]`. Found with A/B builds
changing one piece at a time (perf is not available: `perf_event_paranoid`
4): `#[inline]` on `parse_v1_messages` alone brought
`object_header_parse_x401` from about 24.5–24.9 µs to 23.6–24.0 µs against
8f59b2e's 23.6–24.1 µs (short 4-second rounds); no attribute measured like
`#[inline(never)]`;
creating the chunk list only when a continuation is found measured no
faster on top and was not kept.
Same method as below: `8f59b2e` built in its own worktree and target
directory, separate binaries alternating, `taskset -c 5
local_metadata_bench --bench --warm-up-time 3 --measurement-time 10`, every
binary started with the 1-minute load average below 2 (0.19–1.86) and no
`rustc` running. Candidate: `96086ad` (this change). Median (range) of 3
rounds; run 2 also alternated `main` `425585e`.
| function | 8f59b2e | 425585e (main) | 96086ad | vs 8f59b2e |
|---|---:|---:|---:|---:|
| run 1: `object_header_parse_x401` | 23.81 µs (23.76–24.00) | | 23.57 µs (23.23–23.89) | **−1.0%** |
| run 1: `snod_parse_all` | 1.840 µs (1.837–1.854) | | 1.839 µs (1.837–1.873) | 0.0% |
| run 1: `btree_v1_walk` | 343 ns (338–349) | | 352 ns (344–363) | +2.7% |
| run 1: `facade_list_400_groups` | 8.03 ms (8.01–8.10) | | 8.05 ms (8.01–8.15) | +0.2% |
| run 2: `object_header_parse_x401` | 24.15 µs (23.74–24.41) | 24.93 µs (24.72–25.05) | 23.52 µs (23.34–23.68) | **−2.6%** |
| run 2: `snod_parse_all` | 1.843 µs (1.830–1.847) | 1.856 µs (1.850–1.865) | 1.861 µs (1.858–1.869) | +1.0% |
| run 2: `btree_v1_walk` | 346 ns (339–366) | 356 ns (347–356) | 357 ns (350–359) | +3.2% |
| run 2: `facade_list_400_groups` | 8.09 ms (8.02–8.10) | 8.12 ms (7.97–8.19) | 8.08 ms (7.98–8.12) | −0.1% |
- `ObjectHeader::parse` is at or below 8f59b2e (−1.0%, −2.6%) and 5.6%
faster than `main` in the same run.
- `btree_v1_walk` (one walk of a 350 ns B-tree) is 3% above 8f59b2e in
both runs, with overlapping ranges, and is the same on `main` (+0.2%
between `main` and this change): not from this change. The walk's code
changed in `e553153` (after a failed child the siblings are only read,
so the error returns after them; the fixture never takes that path, but
the loop carries the extra state); left as is.
- `snod_parse_all` and the facade listing are within noise.
### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank) ### Local metadata and data reads after range-read M2/M3 (2026-09-27, tank)
`main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main` `main` just before range-read M2/M3 (`8f59b2e`, PR #17) against `main`
+31
View File
@@ -2,6 +2,37 @@
## Unreleased ## Unreleased
### `ObjectHeader::parse` back at its pre-M2/M3 speed (2026-09-27)
- Parsing a version-1 object header was 4% slower than before range-read
M2/M3 (`docs/known-issues.md`). The cause was the call to the per-chunk
message loop, kept out of line since the allocation fix; it is inlined
again. `object_header_parse_x401` is now 1.0% and 2.6% below `8f59b2e`
and 5.6% below `main` (tank, idle, separate binaries alternating;
`BENCHMARKS.md`). The chunk queue's checks are unchanged (65,536 chunks,
cycle refusal, file-size budget, one chunk buffer at a time, libhdf5's
order, overlapping chunks allowed).
### Conformance: no our-errors or mismatches left (2026-09-27)
- The last 3 our-errors and 2 mismatches were documented as not ours but
still counted against us. Re-checked with evidence
(`docs/known-issues.md`, "Conformance: the last non-ok files"):
- The 2 mismatches were h5py's big-endian VL bug (elements returned with
the file's bytes under a little-endian dtype). `conformance/ref.py` now
checks the installed h5py has the bug and relabels such elements before
hashing, so their values are compared: `attr_datatypes.hdf5` and
`tcomplex_be.h5` are identical to clawhdf5's and now **ok**.
- The 3 our-errors are objects HDF5 2.0 reads by over-reading memory
(scale-offset codes past a chunk, unfiltered chunks shorter than a
chunk, an N-Bit parameter list one value short). h5py's values for them
change between runs, with `MALLOC_PERTURB_` and with import order, and
h5dump 1.14.6 prints others, so they are not the file's data and
clawhdf5 keeps refusing them. `conformance/ref_bugs.py` repeats that
check in every run, and such a file is the new class **ref-bug** only
while its values keep changing.
- Result (`conformance/run.sh --no-fetch`, tank, 2026-09-27): 602 of 697
ok (baseline 600), 0 our-error, 0 mismatch, 3 ref-bug, 92
h5py-cannot-read, no panic/hang/crash/oom.
### Deterministic errors on damaged chunked datasets (2026-09-27) ### Deterministic errors on damaged chunked datasets (2026-09-27)
- A read through the file's chunk cache listed a damaged dataset's chunks in - A read through the file's chunk cache listed a damaged dataset's chunks in
hash-map order, seeded per `File`, so two opens of the same file could hash-map order, seeded per `File`, so two opens of the same file could
+49 -39
View File
@@ -13,15 +13,15 @@ fatal. This file is generated by `conformance/run.sh`; do not edit it by hand.
| | | | | |
|---|---| |---|---|
| date | 2026-09-27 00:34 UTC | | date | 2026-09-28 03:33 UTC |
| clawhdf5 commit | `f37e7ae3263277319dba4bc39be5397194eb00c3` | | clawhdf5 commit | `ff0b2f8a4e4ef105d7ff4694349ca475abf3f611` |
| machine | `tank`: AMD Ryzen 7 7800X3D 8-Core Processor, 16 CPUs, 61 GiB, Linux 7.0.0-34-generic x86_64 | | machine | `tank`: AMD Ryzen 7 7800X3D 8-Core Processor, 16 CPUs, 61 GiB, Linux 7.0.0-34-generic x86_64 |
| command | `conformance/run.sh --no-fetch --update-baseline` | | command | `conformance/run.sh --no-fetch` |
| rustc | rustc 1.98.1 (48a229cea 2026-09-01) | | rustc | |
| reference | h5py 3.16.0, HDF5 2.0.0, numpy 2.5.3, hdf5plugin 7.1.0, Python 3.14.4 | | reference | h5py 3.16.0, HDF5 2.0.0, numpy 2.5.3, hdf5plugin 7.1.0, Python 3.14.4 |
| h5dump | Version 1.14.6 (CVE corpus only) | | h5dump | Version 1.14.6 (CVE corpus only) |
| limits | 20 s timeout (SIGKILL), 4096 MiB address space, per process; 16 files in parallel | | limits | 20 s timeout (SIGKILL), 4096 MiB address space, per process; 16 files in parallel |
| runtime | 21 s probing + comparing (0 s fetch/build before it) | | runtime | 21 s probing + comparing (18 s fetch/build before it) |
## Results ## Results
@@ -29,25 +29,24 @@ A file's class is the first that applies:
- **panic / hang / crash / oom** — clawhdf5 panicked (caught per object or not), hit the timeout, died on a signal, or failed an allocation. The CI gate fails on any of these. - **panic / hang / crash / oom** — clawhdf5 panicked (caught per object or not), hit the timeout, died on a signal, or failed an allocation. The CI gate fails on any of these.
- **h5py-cannot-read** — libhdf5 could not open the file (or itself crashed or hung). Nothing to compare against; most are the deliberately malformed CVE reproducers. - **h5py-cannot-read** — libhdf5 could not open the file (or itself crashed or hung). Nothing to compare against; most are the deliberately malformed CVE reproducers.
- **ref-bug** — every difference is an object clawhdf5 refuses that h5py reads only through a libhdf5 bug: the values h5py returns for it change with the reading process's heap, re-checked in every run (see *Reference bugs*).
- **our-error** — clawhdf5 returned an error for something h5py reads. - **our-error** — clawhdf5 returned an error for something h5py reads.
- **mismatch** — both read it, but the shapes, values, object set or attribute set differ. - **mismatch** — both read it, but the shapes, values, object set or attribute set differ.
- **ok** — every object h5py reads, clawhdf5 reads identically. - **ok** — every object h5py reads, clawhdf5 reads identically.
| corpus | files | ok | our-error | mismatch | h5py-cannot-read | panic | hang | crash | oom | | corpus | files | ok | our-error | mismatch | h5py-cannot-read | ref-bug | panic | hang | crash | oom |
|---|---|---|---|---|---|---|---|---|---| |---|---|---|---|---|---|---|---|---|---|---|
| NCAS-CMS_pyfive | 33 | 32 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | | NCAS-CMS_pyfive | 33 | 33 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| cve_hdf5 | 147 | 113 | 2 | 0 | 32 | 0 | 0 | 0 | 0 | | cve_hdf5 | 147 | 113 | 0 | 0 | 32 | 2 | 0 | 0 | 0 | 0 |
| h5py_data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | | h5py_data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| hdf5 | 466 | 404 | 1 | 1 | 60 | 0 | 0 | 0 | 0 | | hdf5 | 466 | 405 | 0 | 0 | 60 | 1 | 0 | 0 | 0 | 0 |
| netcdf-c | 20 | 20 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | | netcdf-c | 20 | 20 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| netcdf4-python | 18 | 18 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | | netcdf4-python | 18 | 18 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| usnistgov_h5wasm | 5 | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | | usnistgov_h5wasm | 5 | 5 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| xarray-data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | | xarray-data | 4 | 4 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| **all** | **697** | **600** | **3** | **2** | **92** | **0** | **0** | **0** | **0** | | **all** | **697** | **602** | **0** | **0** | **92** | **3** | **0** | **0** | **0** | **0** |
2 of the 2 mismatches are a known h5py bug, not ours (see *Known not-our-bug*). **Our errors and mismatches: 0.** Files not ok: 92 h5py-cannot-read, 3 ref-bug. 2 object(s) were compared against h5py's values corrected for a known h5py bug (2 identical to clawhdf5's; see *Reference bugs*).
3 of the 3 our-errors are corrupt data that HDF5 2.0 reads only through a bug and clawhdf5 refuses (see *Known not-our-bug*).
Corpora (fetched by `conformance/fetch-corpus.sh` into the gitignored `conformance/.cache/`): Corpora (fetched by `conformance/fetch-corpus.sh` into the gitignored `conformance/.cache/`):
@@ -68,18 +67,11 @@ None.
## Our-error root causes ## Our-error root causes
Grouped by normalised error message. *files* counts files whose class this cause affects. None.
| files | objects | error | examples |
|---:|---:|---|---|
| 3 | 3 | `ChunkedReadError("…")` | `cve_hdf5/cvefiles/cve-2025-2308.h5`, `cve_hdf5/cvefiles/cve-2025-44904.h5`, `hdf5/test/testfiles/bad_nbit_parms_walk.h5` |
## Mismatch root causes ## Mismatch root causes
| files | objects | cause | examples | None.
|---:|---:|---|---|
| 1 | 1 | `attr-values: ours=vlen(>u8) h5py=object layout=- filters=-` | `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` |
| 1 | 1 | `values: ours=vlen({r:>f4,i:>f4}8) h5py=object layout=contiguous filters=-` | `hdf5/tools/test/testfiles/tcomplex_be.h5` |
## CVE corpus: clawhdf5 vs h5dump vs h5py ## CVE corpus: clawhdf5 vs h5dump vs h5py
@@ -203,7 +195,7 @@ columns are.
| cvefiles/cve-2024-33876.h5 | ok | read 3 obj, 1 errors | read 3 obj, 1 errors | ok | | cvefiles/cve-2024-33876.h5 | ok | read 3 obj, 1 errors | read 3 obj, 1 errors | ok |
| cvefiles/cve-2024-33877.h5 | error exit | read 8 obj, 1 errors | read 8 obj, 1 errors | ok | | cvefiles/cve-2024-33877.h5 | error exit | read 8 obj, 1 errors | read 8 obj, 1 errors | ok |
| cvefiles/cve-2025-2153.h5 | error exit | open error | read 1 obj, 1 errors | h5py-cannot-read | | cvefiles/cve-2025-2153.h5 | error exit | open error | read 1 obj, 1 errors | h5py-cannot-read |
| cvefiles/cve-2025-2308.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | our-error | | cvefiles/cve-2025-2308.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | ref-bug |
| cvefiles/cve-2025-2309.h5 | ok | read 6 obj, 1 errors | read 6 obj | ok | | cvefiles/cve-2025-2309.h5 | ok | read 6 obj, 1 errors | read 6 obj | ok |
| cvefiles/cve-2025-2310.h5 | error exit | read 24 obj, 8 errors | read 24 obj, 8 errors | ok | | cvefiles/cve-2025-2310.h5 | error exit | read 24 obj, 8 errors | read 24 obj, 8 errors | ok |
| cvefiles/cve-2025-2912.h5 | error exit | open error | open error | h5py-cannot-read | | cvefiles/cve-2025-2912.h5 | error exit | open error | open error | h5py-cannot-read |
@@ -214,7 +206,7 @@ columns are.
| cvefiles/cve-2025-2924.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok | | cvefiles/cve-2025-2924.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
| cvefiles/cve-2025-2925.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok | | cvefiles/cve-2025-2925.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
| cvefiles/cve-2025-2926.h5 | error exit | open error | open error | h5py-cannot-read | | cvefiles/cve-2025-2926.h5 | error exit | open error | open error | h5py-cannot-read |
| cvefiles/cve-2025-44904.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | our-error | | cvefiles/cve-2025-44904.h5 | error exit | read 25 obj, 1 errors | read 25 obj, 2 errors | ref-bug |
| cvefiles/cve-2025-44905.h5 | error exit | read 25 obj, 3 errors | read 25 obj, 3 errors | ok | | cvefiles/cve-2025-44905.h5 | error exit | read 25 obj, 3 errors | read 25 obj, 3 errors | ok |
| cvefiles/cve-2025-6269-1.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok | | cvefiles/cve-2025-6269-1.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
| cvefiles/cve-2025-6269-2.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok | | cvefiles/cve-2025-6269-2.h5 | error exit | read 1 obj, 1 errors | read 1 obj, 1 errors | ok |
@@ -249,13 +241,36 @@ columns are.
</details> </details>
## Known not-our-bug ## Reference bugs
### Objects h5py reads only through a libhdf5 bug (*ref-bug*)
clawhdf5 refuses these objects; h5py 3.16 / HDF5 2.0 returns values for them. `conformance/ref_bugs.py`
re-reads each with h5py in six fresh processes whose heaps differ (h5py imported before numpy, three
times and twice more with `MALLOC_PERTURB_`, and numpy imported first). Values the file determines
come out the same every time; these do not, so they are memory libhdf5 over-reads, not the file's
data. A file is *ref-bug* only while every one of its differences is such an object confirmed in
the same run; an object that reads the same every time goes back to *our-error*. Reproducer:
`python conformance/ref_bugs.py conformance/.cache/corpus` (prints every read's outcome).
| file | object | distinct results in 6 reads | confirmed | what goes wrong |
|---|---|---:|---|---|
| `cve_hdf5/cvefiles/cve-2025-2308.h5` | `/Scale_offset_long_long_data_le` | 6 | yes | the first chunk records minbits 11: its 12 values need 17 bytes of codes, and the 26-byte chunk holds 5 after its 21-byte header; libhdf5's scale-offset decoder reads past its buffer, and develop refuses the chunk ("Buffer too short") |
| `cve_hdf5/cvefiles/cve-2025-44904.h5` | `/Scale_offset_float_data_le` | 6 | yes | unfiltered chunks stored as 38 and 37 bytes for 48-byte chunks: 1.14/2.0 read the stored bytes into a buffer of that size and use it as the whole chunk (H5D__chunk_lock), so the rest is heap memory; develop refuses them ("incorrect chunk size returned from index for unfiltered chunk") |
| `hdf5/test/testfiles/bad_nbit_parms_walk.h5` | `/Nbit_int_data_le` | 3 | yes | the N-Bit parameter list holds 7 values (cd_values[0] = 7) where an integer needs 8: the decoder takes the bit offset from cd_values[7], past the list; libhdf5's own test (`test_filter_bad_params`, test/dsets.c on develop) requires the read to fail |
### Values corrected for a known h5py bug
- **h5py big-endian variable-length sequences.** h5py returns the elements of a VL sequence - **h5py big-endian variable-length sequences.** h5py returns the elements of a VL sequence
whose base type is big-endian with the file's big-endian bytes but a native (little-endian) whose base type is big-endian with the file's big-endian bytes but a native (little-endian)
numpy dtype, so the values it reports are byte-swapped garbage; `h5dump` prints the values numpy dtype: a `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]` reads back as
clawhdf5 reads. Reproducer: `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]` `[4.6e-41, 9.0e-44]`; `h5dump` prints the file's values. `ref.py` checks that the installed
reads back in h5py as `[4.6e-41, 9.0e-44]`. Affected here: `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5`, `hdf5/tools/test/testfiles/tcomplex_be.h5`. h5py still does this (by writing and reading exactly that dataset in memory) and, if so,
relabels such elements with the file's byte order before hashing, so the values are still
compared. Corrected objects: `NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64` (same as clawhdf5), `hdf5/tools/test/testfiles/tcomplex_be.h5` `/VariableLengthDatasetFloatComplex` (same as clawhdf5).
## Other comparison rules
- **Non-IEEE floats and partial-precision integers (N-Bit).** libhdf5 converts a float whose - **Non-IEEE floats and partial-precision integers (N-Bit).** libhdf5 converts a float whose
bit layout is not IEEE (e.g. `H5Tset_precision` for the N-Bit filter) or an integer with a bit layout is not IEEE (e.g. `H5Tset_precision` for the N-Bit filter) or an integer with a
bit offset / reduced precision into the plain numpy type of the same size. The probe bit offset / reduced precision into the plain numpy type of the same size. The probe
@@ -264,11 +279,6 @@ columns are.
- **Types h5py widens.** Where h5py reads a type into a numpy type of a different size - **Types h5py widens.** Where h5py reads a type into a numpy type of a different size
(FP8 -> float16, bfloat16 -> float32, x87 long double -> float128) the values are not (FP8 -> float16, bfloat16 -> float32, x87 long double -> float128) the values are not
compared (shape and presence still are): dataset file type size 1 -> numpy float16 (2) (15x), attr file type size 1 -> numpy float16 (2) (15x), dataset file type size 2 -> numpy float32 (4) (2x), dataset file type size 8 -> numpy float128 (16) (1x), dataset file type size 12 -> numpy float128 (16) (1x), attr file type size 2 -> numpy float32 (4) (1x), dataset file type size 2 -> numpy >f4 (4) (1x), attr file type size 2 -> numpy >f4 (4) (1x). compared (shape and presence still are): dataset file type size 1 -> numpy float16 (2) (15x), attr file type size 1 -> numpy float16 (2) (15x), dataset file type size 2 -> numpy float32 (4) (2x), dataset file type size 8 -> numpy float128 (16) (1x), dataset file type size 12 -> numpy float128 (16) (1x), attr file type size 2 -> numpy float32 (4) (1x), dataset file type size 2 -> numpy >f4 (4) (1x), attr file type size 2 -> numpy >f4 (4) (1x).
- **Corrupt data HDF5 2.0 reads through a bug.** clawhdf5 refuses these objects; h5py 3.16 /
HDF5 2.0 returns values for them that the file does not hold:
- `cve_hdf5/cvefiles/cve-2025-2308.h5` `/Scale_offset_long_long_data_le`: scale-offset codes run past the end of the chunk: HDF5 2.0 reads past its buffer; libhdf5's develop branch refuses the chunk ("Buffer too short").
- `cve_hdf5/cvefiles/cve-2025-44904.h5` `/Scale_offset_float_data_le`: unfiltered chunks of 38 and 37 bytes for 48-byte chunks: HDF5 2.0 fills the rest with whatever its buffer held; libhdf5's develop branch refuses them ("incorrect chunk size returned from index for unfiltered chunk").
- `hdf5/test/testfiles/bad_nbit_parms_walk.h5` `/Nbit_int_data_le`: an N-Bit parameter list one value short: HDF5 2.0 reads past the list; libhdf5's own test (`test_filter_bad_params`, test/dsets.c) now requires the read to fail.
- **References** are compared by presence only (`R`), not by target. - **References** are compared by presence only (`R`), not by target.
## Objects h5py fails on but clawhdf5 reads ## Objects h5py fails on but clawhdf5 reads
+4 -2
View File
@@ -19,9 +19,11 @@ probe's `szip` feature; `libaec-dev`), and a Python with the packages in
| `fetch-corpus.sh` | shallow, sparse, blob-filtered checkout of each pinned commit into `.cache/src/` (gitignored); no-op when already there | | `fetch-corpus.sh` | shallow, sparse, blob-filtered checkout of each pinned commit into `.cache/src/` (gitignored); no-op when already there |
| `list_files.py` | which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers) | | `list_files.py` | which files are probed (HDF5/netCDF-4 extensions minus netCDF classic, plus the CVE reproducers) |
| `probe/` | the clawhdf5 side: a standalone crate (outside the workspace, so `cargo test --workspace` never builds it) that walks a file with `clawhdf5-format` and prints canonical JSON | | `probe/` | the clawhdf5 side: a standalone crate (outside the workspace, so `cargo test --workspace` never builds it) that walks a file with `clawhdf5-format` and prints canonical JSON |
| `ref.py` | the h5py side: the same JSON from h5py | | `ref.py` | the h5py side: the same JSON from h5py (values corrected for a known h5py bug are marked `ref_fix`) |
| `ref_bugs.py` | re-reads the objects h5py reads only through a libhdf5 bug in six differently-set-up processes; an object whose values change is confirmed as a libhdf5 over-read |
| `test_ref.py` | tests of `ref.py`'s correction and `ref_bugs.py`'s confirmation (`python conformance/test_ref.py`) |
| `run_one.sh` | runs both sides on one file (and `h5dump` on the CVE corpus) under a timeout and an address-space limit | | `run_one.sh` | runs both sides on one file (and `h5dump` on the CVE corpus) under a timeout and an address-space limit |
| `compare.py` | classifies each file (ok / our-error / mismatch / h5py-cannot-read / panic / hang / crash / oom) and groups root causes | | `compare.py` | classifies each file (ok / our-error / mismatch / h5py-cannot-read / ref-bug / panic / hang / crash / oom) and groups root causes |
| `report.py` | writes `CONFORMANCE.md` | | `report.py` | writes `CONFORMANCE.md` |
| `check.py` | the gate: fails on any panic/hang/crash/oom, on an ok count below `baseline.json`, or on a baseline-ok file that is no longer ok | | `check.py` | the gate: fails on any panic/hang/crash/oom, on an ok count below `baseline.json`, or on a baseline-ok file that is no longer ok |
| `baseline.json` | the ok files the gate holds the line on | | `baseline.json` | the ok files the gate holds the line on |
+30 -3
View File
@@ -5,6 +5,9 @@ Writes <results_dir>/results.csv, results.json and summary.md.
File classes (first match wins): File classes (first match wins):
hang, oom, crash, panic ours: timeout / allocation failure / signal / any panic (caught or not) hang, oom, crash, panic ours: timeout / allocation failure / signal / any panic (caught or not)
h5py-cannot-read libhdf5/h5py failed to open the file (or crashed/hung) h5py-cannot-read libhdf5/h5py failed to open the file (or crashed/hung)
ref-bug every issue is an object we refuse that h5py reads only through a
libhdf5 bug, confirmed in this run by ref_bugs.py (its values
change with the reading process's heap)
our-error we fail to open, list, or read something h5py reads our-error we fail to open, list, or read something h5py reads
mismatch we read something with different shape/values, or a different object set mismatch we read something with different shape/values, or a different object set
ok ok
@@ -19,6 +22,20 @@ import sys
R = sys.argv[1] R = sys.argv[1]
RUNS = os.path.join(R, "runs") RUNS = os.path.join(R, "runs")
# Objects ref_bugs.py confirmed in this run: h5py's values for them come from
# libhdf5 reading memory the file does not determine.
try:
REF_BUGS = {(b["file"], b["object"])
for b in json.load(open(os.path.join(R, "ref_bugs.json")))["read_bugs"] if b.get("confirmed")}
except (OSError, ValueError, KeyError):
REF_BUGS = set()
def is_ref_bug(rel, issue):
"""An our-error on reading an object that ref_bugs.py confirmed."""
kind, detail = issue[0], issue[1]
return kind == "our-error" and any(f == rel and detail.startswith(obj + ": error: ") for f, obj in REF_BUGS)
def load(d, name): def load(d, name):
rc_p = os.path.join(d, name + ".rc") rc_p = os.path.join(d, name + ".rc")
@@ -86,6 +103,8 @@ mismatch_causes = collections.defaultdict(lambda: {"files": set(), "count": 0, "
panics = [] panics = []
ref_only_errors = collections.Counter() ref_only_errors = collections.Counter()
incomparable = collections.Counter() incomparable = collections.Counter()
# Values ref.py corrected for a known h5py bug: (file, object, fixes, same as ours)
ref_fixes = []
def add(bucket, key, file, example): def add(bucket, key, file, example):
@@ -169,6 +188,8 @@ for rel in files:
elif a.get("hash") != b.get("hash"): elif a.get("hash") != b.get("hash"):
issues.append(("mismatch", f"{p}: values differ (h5py {a.get('dtype')} vs ours {b.get('dtype')})", "values", b | {"ref_head": a.get("head"), "ref_dtype": a.get("dtype")})) issues.append(("mismatch", f"{p}: values differ (h5py {a.get('dtype')} vs ours {b.get('dtype')})", "values", b | {"ref_head": a.get("head"), "ref_dtype": a.get("dtype")}))
ok = False ok = False
if a.get("ref_fix") and "hash" in b:
ref_fixes.append((rel, p, a["ref_fix"], a.get("hash") == b.get("hash")))
ra, oa = a.get("attrs") or {}, b.get("attrs") or {} ra, oa = a.get("attrs") or {}, b.get("attrs") or {}
if "attrs_error" not in b and "attrs_error" not in a and not ref_unopened: if "attrs_error" not in b and "attrs_error" not in a and not ref_unopened:
for an in sorted(set(ra) | set(oa)): for an in sorted(set(ra) | set(oa)):
@@ -187,6 +208,8 @@ for rel in files:
issues.append(("mismatch", f"{p}@{an}: attr shape {x.get('shape')} vs ours {y.get('shape')}", "attr-shape", y | {"ref_dtype": x.get("dtype")})) issues.append(("mismatch", f"{p}@{an}: attr shape {x.get('shape')} vs ours {y.get('shape')}", "attr-shape", y | {"ref_dtype": x.get("dtype")}))
elif x.get("hash") != y.get("hash"): elif x.get("hash") != y.get("hash"):
issues.append(("mismatch", f"{p}@{an}: attr values differ (h5py {x.get('dtype')} vs ours {y.get('dtype')})", "attr-values", y | {"ref_head": x.get("head"), "ref_dtype": x.get("dtype")})) issues.append(("mismatch", f"{p}@{an}: attr values differ (h5py {x.get('dtype')} vs ours {y.get('dtype')})", "attr-values", y | {"ref_head": x.get("head"), "ref_dtype": x.get("dtype")}))
if x.get("ref_fix") and "hash" in y:
ref_fixes.append((rel, f"{p}@{an}", x["ref_fix"], x.get("hash") == y.get("hash")))
if ok: if ok:
n_ok += 1 n_ok += 1
@@ -197,6 +220,8 @@ for rel in files:
cls = "panic" cls = "panic"
elif ref_open_fail: elif ref_open_fail:
cls = "h5py-cannot-read" cls = "h5py-cannot-read"
elif issues and all(is_ref_bug(rel, i) for i in issues):
cls = "ref-bug"
elif ours_open_err: elif ours_open_err:
cls = "our-error" cls = "our-error"
issues.append(("our-error", f"open: {ours_open_err}", ours_open_err, {})) issues.append(("our-error", f"open: {ours_open_err}", ours_open_err, {}))
@@ -214,7 +239,8 @@ for rel in files:
"caught": [(p, w, m[:2500]) for p, w, m in caught_panics[:3]], "caught": [(p, w, m[:2500]) for p, w, m in caught_panics[:3]],
"n_caught": len(caught_panics), "n_caught": len(caught_panics),
}) })
for kind, detail, key, rec in issues: # A ref-bug file's differences are listed with the evidence instead.
for kind, detail, key, rec in (issues if cls != "ref-bug" else []):
if kind == "our-error": if kind == "our-error":
add(root_causes, norm(key), rel, detail[:300]) add(root_causes, norm(key), rel, detail[:300])
else: else:
@@ -259,10 +285,11 @@ def ser(b):
json.dump({"rows": rows, "issues": issues_by_file, "root_causes": ser(root_causes), "mismatch_causes": ser(mismatch_causes), json.dump({"rows": rows, "issues": issues_by_file, "root_causes": ser(root_causes), "mismatch_causes": ser(mismatch_causes),
"panics": panics, "incomparable": incomparable.most_common(), "ref_only_errors": ref_only_errors.most_common()}, "panics": panics, "incomparable": incomparable.most_common(), "ref_only_errors": ref_only_errors.most_common(),
"ref_fixes": ref_fixes, "ref_bugs_confirmed": sorted(REF_BUGS)},
open(os.path.join(R, "results.json"), "w"), indent=1) open(os.path.join(R, "results.json"), "w"), indent=1)
classes = ["ok", "our-error", "mismatch", "h5py-cannot-read", "hang", "panic", "crash", "oom"] classes = ["ok", "our-error", "mismatch", "h5py-cannot-read", "ref-bug", "hang", "panic", "crash", "oom"]
by_corpus = collections.defaultdict(collections.Counter) by_corpus = collections.defaultdict(collections.Counter)
for r in rows: for r in rows:
by_corpus[r["corpus"]][r["class"]] += 1 by_corpus[r["corpus"]][r["class"]] += 1
+53 -1
View File
@@ -53,6 +53,49 @@ def packed(dt):
return dt return dt
# --- reference corrections ---------------------------------------------------
# Where h5py is known to return values the file does not hold, and the right
# values follow from what it returned, ref.py corrects them and records the
# correction on the object ("ref_fix"), so the comparison is still a real
# comparison and CONFORMANCE.md lists every corrected object. Each correction
# first checks that the installed h5py still has the bug.
# Corrections applied while encoding the current object.
FIXES = set()
_BE_VLEN_BUG = None
def be_vlen_bug():
"""h5py (3.16 / HDF5 2.0 at least) returns the elements of a
variable-length sequence whose base type is big-endian with the file's
big-endian bytes under a native (little-endian) dtype: a
`vlen_dtype('>f4')` dataset holding [1.0, 2.0] reads back as
[4.6e-41, 9.0e-44]. `h5dump` prints the file's values. Checked once per
process by writing and reading exactly that dataset in memory."""
global _BE_VLEN_BUG
if _BE_VLEN_BUG is None:
import io
try:
bio = io.BytesIO()
with h5py.File(bio, "w") as f:
d = f.create_dataset("v", (1,), dtype=h5py.vlen_dtype(np.dtype(">f4")))
d[0] = np.array([1.0, 2.0], dtype=">f4")
with h5py.File(bio, "r") as f:
got = np.asarray(f["v"][0])
_BE_VLEN_BUG = (got.dtype == np.dtype("<f4")
and got.view(">f4").tolist() == [1.0, 2.0]
and got.tolist() != [1.0, 2.0])
except Exception: # noqa: BLE001
_BE_VLEN_BUG = False
return _BE_VLEN_BUG
def unswapped(got, base):
"""`got` is `base` (big-endian somewhere) with every field in native
little-endian order instead: the shape of h5py's big-endian VL bug."""
return base.newbyteorder("<") == got and base != got
def canon_el(dt, val, out): def canon_el(dt, val, out):
if dt.fields: if dt.fields:
for n in dt.names: for n in dt.names:
@@ -79,7 +122,13 @@ def canon_el(dt, val, out):
base = h5py.check_vlen_dtype(dt) base = h5py.check_vlen_dtype(dt)
if base is None: if base is None:
raise TypeError(f"unhandled object dtype {dt!r}") raise TypeError(f"unhandled object dtype {dt!r}")
arr = np.asarray(val if val is not None else [], dtype=base).reshape(-1) arr = np.asarray(val if val is not None else [])
if arr.dtype != base and be_vlen_bug() and unswapped(arr.dtype, base):
# h5py's big-endian VL bug (see be_vlen_bug): the bytes are
# the file's, the dtype label is wrong. Relabel, don't convert.
arr = arr.view(base)
FIXES.add("h5py-be-vlen")
arr = np.asarray(arr, dtype=base).reshape(-1)
out += b"V" + struct.pack("<I", arr.shape[0]) out += b"V" + struct.pack("<I", arr.shape[0])
if simple(base): if simple(base):
out += arr.astype(packed(base)).tobytes() out += arr.astype(packed(base)).tobytes()
@@ -118,6 +167,7 @@ def hash_values(arr, dt, rec):
while dt.subdtype is not None: while dt.subdtype is not None:
dt = dt.subdtype[0] dt = dt.subdtype[0]
arr = np.asarray(arr, dtype=dt) arr = np.asarray(arr, dtype=dt)
FIXES.clear()
if simple(dt): if simple(dt):
c = np.ascontiguousarray(arr).astype(packed(dt)).tobytes() c = np.ascontiguousarray(arr).astype(packed(dt)).tobytes()
else: else:
@@ -125,6 +175,8 @@ def hash_values(arr, dt, rec):
for x in arr.reshape(-1): for x in arr.reshape(-1):
canon_el(dt, x, out) canon_el(dt, x, out)
c = bytes(out) c = bytes(out)
if FIXES:
rec["ref_fix"] = sorted(FIXES)
rec["hash"] = hashlib.sha256(c).hexdigest() rec["hash"] = hashlib.sha256(c).hexdigest()
rec["head"] = c[:48].hex() rec["head"] = c[:48].hex()
+105
View File
@@ -0,0 +1,105 @@
#!/usr/bin/env python3
"""ref_bugs.py <corpus_dir>: re-check the objects h5py reads only through a
libhdf5 bug.
For each object of READ_BUGS (below), h5py reads it in several fresh
processes whose heaps differ: h5py imported before numpy (three runs, plus
two with glibc's MALLOC_PERTURB_, which fills newly allocated and freed heap
blocks with a byte pattern) and numpy imported first. Values the file
determines come out the same every time. An object whose values differ
between those runs is read from memory the file does not determine — an
over-read or an uninitialised buffer in libhdf5 — so the values h5py reports
for it are not the file's, and clawhdf5 refusing the object is not a
clawhdf5 error. compare.py classifies a file as `ref-bug` only on objects
confirmed that way in the same run (`$OUT/ref_bugs.json`); an object whose
reading turns out stable stays an our-error.
Run by conformance/run.sh; on its own it is the reproducer (JSON on stdout).
"""
import concurrent.futures
import json
import os
import subprocess
import sys
# (file, object) -> what goes wrong. Checked 2026-09-27 against HDF5 2.0.0
# (h5py 3.16), h5dump 1.14.6 and the HDFGroup/hdf5 sources (tag hdf5_1_14_6
# and develop); see docs/known-issues.md, "Conformance: the last non-ok files".
READ_BUGS = {
("cve_hdf5/cvefiles/cve-2025-2308.h5", "/Scale_offset_long_long_data_le"):
"the first chunk records minbits 11: its 12 values need 17 bytes of codes, and the "
"26-byte chunk holds 5 after its 21-byte header; libhdf5's scale-offset decoder reads "
"past its buffer, and develop refuses the chunk (\"Buffer too short\")",
("cve_hdf5/cvefiles/cve-2025-44904.h5", "/Scale_offset_float_data_le"):
"unfiltered chunks stored as 38 and 37 bytes for 48-byte chunks: 1.14/2.0 read the "
"stored bytes into a buffer of that size and use it as the whole chunk "
"(H5D__chunk_lock), so the rest is heap memory; develop refuses them (\"incorrect chunk "
"size returned from index for unfiltered chunk\")",
("hdf5/test/testfiles/bad_nbit_parms_walk.h5", "/Nbit_int_data_le"):
"the N-Bit parameter list holds 7 values (cd_values[0] = 7) where an integer needs 8: "
"the decoder takes the bit offset from cd_values[7], past the list; libhdf5's own test "
"(`test_filter_bad_params`, test/dsets.c on develop) requires the read to fail",
}
# (which module is imported first, MALLOC_PERTURB_)
RUNS = [("h5py", None), ("h5py", None), ("h5py", None), ("h5py", "170"), ("h5py", "255"),
("numpy", None)]
READ = r"""
import hashlib, sys
if sys.argv[3] == "h5py":
import h5py, numpy as np
else:
import numpy as np, h5py
try:
import hdf5plugin # noqa: F401
except Exception:
pass
try:
with h5py.File(sys.argv[1], "r") as f:
a = np.ascontiguousarray(f[sys.argv[2]][()])
print("values " + hashlib.sha256(a.tobytes()).hexdigest()[:16])
except Exception as e:
print("error " + (str(e).splitlines() or [type(e).__name__])[0][:120])
"""
def read_once(path, obj, first, perturb):
env = dict(os.environ)
env.pop("MALLOC_PERTURB_", None)
if perturb:
env["MALLOC_PERTURB_"] = perturb
try:
p = subprocess.run([sys.executable, "-c", READ, path, obj, first], env=env,
capture_output=True, text=True, timeout=60)
out = p.stdout.strip().splitlines()
return out[-1] if out else f"exit {p.returncode}"
except subprocess.TimeoutExpired:
return "timeout"
def check(corpus, key):
f, obj = key
path = os.path.join(corpus, f)
rec = {"file": f, "object": obj, "why": READ_BUGS[key]}
if not os.path.exists(path):
return rec | {"missing": True, "confirmed": False}
runs = [{"first": a, "malloc_perturb": p, "outcome": read_once(path, obj, a, p)} for a, p in RUNS]
distinct = sorted({r["outcome"] for r in runs})
return rec | {
"runs": runs,
"distinct": len(distinct),
"confirmed": len(distinct) > 1 and any(o.startswith("values ") for o in distinct),
}
def main():
corpus = sys.argv[1]
keys = list(READ_BUGS)
with concurrent.futures.ThreadPoolExecutor(max_workers=len(keys)) as ex:
out = list(ex.map(lambda k: check(corpus, k), keys))
print(json.dumps({"read_bugs": out}, indent=1))
if __name__ == "__main__":
main()
+62 -69
View File
@@ -26,7 +26,7 @@ except Exception: # noqa: BLE001
R, OUT_MD, CORPUS = sys.argv[1], sys.argv[2], sys.argv[3] R, OUT_MD, CORPUS = sys.argv[1], sys.argv[2], sys.argv[3]
HERE = os.path.dirname(os.path.abspath(__file__)) HERE = os.path.dirname(os.path.abspath(__file__))
ROOT = os.path.dirname(HERE) ROOT = os.path.dirname(HERE)
CLASSES = ["ok", "our-error", "mismatch", "h5py-cannot-read", "panic", "hang", "crash", "oom"] CLASSES = ["ok", "our-error", "mismatch", "h5py-cannot-read", "ref-bug", "panic", "hang", "crash", "oom"]
def sh(*cmd, cwd=ROOT): def sh(*cmd, cwd=ROOT):
@@ -89,46 +89,15 @@ def ex_list(files, n=3):
return s + (f" (+{len(files) - n} more)" if len(files) > n else "") return s + (f" (+{len(files) - n} more)" if len(files) > n else "")
# --- known causes that are not clawhdf5 bugs -------------------------------- # --- reference bugs ---------------------------------------------------------
def is_h5py_be_vlen(i): # ref_bugs.py's re-check of the objects h5py reads only through a libhdf5 bug
"""h5py returns the elements of a VL sequence of a big-endian base type # (compare.py classifies on the confirmed ones), and the objects whose h5py
with their file (big-endian) bytes but a native-endian dtype.""" # values ref.py corrected (compare.py's ref_fixes).
return (i["kind"] == "mismatch" and i["key"] in ("values", "attr-values") try:
and (i.get("ref_dtype") == "object") and (i.get("ours_dtype") or "").startswith("vlen(") ref_bugs = json.load(open(os.path.join(R, "ref_bugs.json")))["read_bugs"]
and ">" in (i.get("ours_dtype") or "")) except (OSError, ValueError, KeyError):
ref_bugs = []
ref_fixes = res.get("ref_fixes", [])
# Objects the reference (h5py 3.16 / HDF5 2.0) reads only because of an
# HDF5 2.0 bug, and that clawhdf5 refuses: each one reads past a buffer or
# returns bytes the file does not hold, and libhdf5's develop branch refuses all
# three. (file, object) -> why. Checked 2026-09-26 against HDF5 2.0.0
# and HDFGroup/hdf5 develop sources; see docs/known-issues.md.
LIBHDF5_BUGS = {
("cve_hdf5/cvefiles/cve-2025-2308.h5", "/Scale_offset_long_long_data_le"):
"scale-offset codes run past the end of the chunk: HDF5 2.0 reads past its buffer; "
"libhdf5's develop branch refuses the chunk (\"Buffer too short\")",
("cve_hdf5/cvefiles/cve-2025-44904.h5", "/Scale_offset_float_data_le"):
"unfiltered chunks of 38 and 37 bytes for 48-byte chunks: HDF5 2.0 fills the rest with "
"whatever its buffer held; libhdf5's develop branch refuses them (\"incorrect chunk size returned "
"from index for unfiltered chunk\")",
("hdf5/test/testfiles/bad_nbit_parms_walk.h5", "/Nbit_int_data_le"):
"an N-Bit parameter list one value short: HDF5 2.0 reads past the list; libhdf5's own "
"test (`test_filter_bad_params`, test/dsets.c) now requires the read to fail",
}
def is_libhdf5_bug(rel, i):
return i["kind"] == "our-error" and any(
f == rel and i["detail"].startswith(obj + ":") for (f, obj) in LIBHDF5_BUGS)
known = collections.defaultdict(list)
for r in rows:
iss = issues.get(r["file"], [])
if r["class"] == "mismatch" and iss and all(is_h5py_be_vlen(i) for i in iss):
known["h5py-be-vlen"].append(r["file"])
if r["class"] == "our-error" and iss and all(is_libhdf5_bug(r["file"], i) for i in iss):
known["libhdf5-2.0"].append(r["file"])
# --- the CVE corpus: clawhdf5 vs h5dump vs h5py ------------------------------ # --- the CVE corpus: clawhdf5 vs h5dump vs h5py ------------------------------
@@ -243,6 +212,7 @@ w("A file's class is the first that applies:")
w("") w("")
w("- **panic / hang / crash / oom** — clawhdf5 panicked (caught per object or not), hit the timeout, died on a signal, or failed an allocation. The CI gate fails on any of these.") w("- **panic / hang / crash / oom** — clawhdf5 panicked (caught per object or not), hit the timeout, died on a signal, or failed an allocation. The CI gate fails on any of these.")
w("- **h5py-cannot-read** — libhdf5 could not open the file (or itself crashed or hung). Nothing to compare against; most are the deliberately malformed CVE reproducers.") w("- **h5py-cannot-read** — libhdf5 could not open the file (or itself crashed or hung). Nothing to compare against; most are the deliberately malformed CVE reproducers.")
w("- **ref-bug** — every difference is an object clawhdf5 refuses that h5py reads only through a libhdf5 bug: the values h5py returns for it change with the reading process's heap, re-checked in every run (see *Reference bugs*).")
w("- **our-error** — clawhdf5 returned an error for something h5py reads.") w("- **our-error** — clawhdf5 returned an error for something h5py reads.")
w("- **mismatch** — both read it, but the shapes, values, object set or attribute set differ.") w("- **mismatch** — both read it, but the shapes, values, object set or attribute set differ.")
w("- **ok** — every object h5py reads, clawhdf5 reads identically.") w("- **ok** — every object h5py reads, clawhdf5 reads identically.")
@@ -254,14 +224,12 @@ for c in sorted(by_corpus):
w(f"| {c} | {sum(cnt.values())} | " + " | ".join(str(cnt.get(k, 0)) for k in CLASSES) + " |") w(f"| {c} | {sum(cnt.values())} | " + " | ".join(str(cnt.get(k, 0)) for k in CLASSES) + " |")
w(f"| **all** | **{len(rows)}** | " + " | ".join(f"**{total.get(k, 0)}**" for k in CLASSES) + " |") w(f"| **all** | **{len(rows)}** | " + " | ".join(f"**{total.get(k, 0)}**" for k in CLASSES) + " |")
w("") w("")
if known["h5py-be-vlen"]: nonok = total.get("our-error", 0) + total.get("mismatch", 0)
w(f"{len(known['h5py-be-vlen'])} of the {total.get('mismatch', 0)} mismatches are a known h5py bug, " w(f"**Our errors and mismatches: {nonok}.** Files not ok: "
"not ours (see *Known not-our-bug*).") + (", ".join(f"{total[c]} {c}" for c in CLASSES if c != "ok" and total.get(c)) or "none") + "."
w("") + (f" {len(ref_fixes)} object(s) were compared against h5py's values corrected for a known h5py bug"
if known["libhdf5-2.0"]: f" ({sum(1 for x in ref_fixes if x[3])} identical to clawhdf5's; see *Reference bugs*)." if ref_fixes else ""))
w(f"{len(known['libhdf5-2.0'])} of the {total.get('our-error', 0)} our-errors are corrupt data that " w("")
"HDF5 2.0 reads only through a bug and clawhdf5 refuses (see *Known not-our-bug*).")
w("")
w("Corpora (fetched by `conformance/fetch-corpus.sh` into the gitignored `conformance/.cache/`):") w("Corpora (fetched by `conformance/fetch-corpus.sh` into the gitignored `conformance/.cache/`):")
w("") w("")
w("| corpus | source | commit |") w("| corpus | source | commit |")
@@ -281,19 +249,25 @@ w("")
w("## Our-error root causes") w("## Our-error root causes")
w("") w("")
w("Grouped by normalised error message. *files* counts files whose class this cause affects.") if res["root_causes"]:
w("") w("Grouped by normalised error message. *files* counts files whose class this cause affects.")
w("| files | objects | error | examples |") w("")
w("|---:|---:|---|---|") w("| files | objects | error | examples |")
for k, v in res["root_causes"].items(): w("|---:|---:|---|---|")
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |") for k, v in res["root_causes"].items():
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |")
else:
w("None.")
w("") w("")
w("## Mismatch root causes") w("## Mismatch root causes")
w("") w("")
w("| files | objects | cause | examples |") if res["mismatch_causes"]:
w("|---:|---:|---|---|") w("| files | objects | cause | examples |")
for k, v in res["mismatch_causes"].items(): w("|---:|---:|---|---|")
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |") for k, v in res["mismatch_causes"].items():
w(f"| {v['files']} | {v['count']} | `{k.replace('|', '/')}` | {ex_list(v['file_list'])} |")
else:
w("None.")
w("") w("")
w("## CVE corpus: clawhdf5 vs h5dump vs h5py") w("## CVE corpus: clawhdf5 vs h5dump vs h5py")
@@ -321,14 +295,38 @@ w("")
w("</details>") w("</details>")
w("") w("")
w("## Known not-our-bug") w("## Reference bugs")
w("")
w("### Objects h5py reads only through a libhdf5 bug (*ref-bug*)")
w("")
w("clawhdf5 refuses these objects; h5py 3.16 / HDF5 2.0 returns values for them. `conformance/ref_bugs.py`")
w("re-reads each with h5py in six fresh processes whose heaps differ (h5py imported before numpy, three")
w("times and twice more with `MALLOC_PERTURB_`, and numpy imported first). Values the file determines")
w("come out the same every time; these do not, so they are memory libhdf5 over-reads, not the file's")
w("data. A file is *ref-bug* only while every one of its differences is such an object confirmed in")
w("the same run; an object that reads the same every time goes back to *our-error*. Reproducer:")
w("`python conformance/ref_bugs.py conformance/.cache/corpus` (prints every read's outcome).")
w("")
w("| file | object | distinct results in 6 reads | confirmed | what goes wrong |")
w("|---|---|---:|---|---|")
for b in ref_bugs:
n = "missing" if b.get("missing") else b.get("distinct", "?")
w(f"| `{b['file']}` | `{b['object']}` | {n} | {'yes' if b.get('confirmed') else '**no**'} | {b['why']} |")
w("")
w("### Values corrected for a known h5py bug")
w("") w("")
w("- **h5py big-endian variable-length sequences.** h5py returns the elements of a VL sequence") w("- **h5py big-endian variable-length sequences.** h5py returns the elements of a VL sequence")
w(" whose base type is big-endian with the file's big-endian bytes but a native (little-endian)") w(" whose base type is big-endian with the file's big-endian bytes but a native (little-endian)")
w(" numpy dtype, so the values it reports are byte-swapped garbage; `h5dump` prints the values") w(" numpy dtype: a `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]` reads back as")
w(" clawhdf5 reads. Reproducer: `h5py.vlen_dtype(np.dtype('>f4'))` dataset holding `[1.0, 2.0]`") w(" `[4.6e-41, 9.0e-44]`; `h5dump` prints the file's values. `ref.py` checks that the installed")
w(" reads back in h5py as `[4.6e-41, 9.0e-44]`. Affected here: " w(" h5py still does this (by writing and reading exactly that dataset in memory) and, if so,")
+ (ex_list(sorted(known["h5py-be-vlen"]), 10) if known["h5py-be-vlen"] else "none") + ".") w(" relabels such elements with the file's byte order before hashing, so the values are still")
w(" compared. Corrected objects: "
+ (", ".join(f"`{f}` `{p}` ({'same as clawhdf5' if same else '**differs from clawhdf5**'})"
for f, p, _, same in ref_fixes) if ref_fixes else "none") + ".")
w("")
w("## Other comparison rules")
w("")
w("- **Non-IEEE floats and partial-precision integers (N-Bit).** libhdf5 converts a float whose") w("- **Non-IEEE floats and partial-precision integers (N-Bit).** libhdf5 converts a float whose")
w(" bit layout is not IEEE (e.g. `H5Tset_precision` for the N-Bit filter) or an integer with a") w(" bit layout is not IEEE (e.g. `H5Tset_precision` for the N-Bit filter) or an integer with a")
w(" bit offset / reduced precision into the plain numpy type of the same size. The probe") w(" bit offset / reduced precision into the plain numpy type of the same size. The probe")
@@ -339,11 +337,6 @@ if res["incomparable"]:
w(" (FP8 -> float16, bfloat16 -> float32, x87 long double -> float128) the values are not") w(" (FP8 -> float16, bfloat16 -> float32, x87 long double -> float128) the values are not")
w(" compared (shape and presence still are): " w(" compared (shape and presence still are): "
+ ", ".join(f"{k} ({n}x)" for k, n in res["incomparable"]) + ".") + ", ".join(f"{k} ({n}x)" for k, n in res["incomparable"]) + ".")
w("- **Corrupt data HDF5 2.0 reads through a bug.** clawhdf5 refuses these objects; h5py 3.16 /")
w(" HDF5 2.0 returns values for them that the file does not hold:")
for (f, obj), why in sorted(LIBHDF5_BUGS.items()):
here = "" if f in known["libhdf5-2.0"] else " (not an our-error in this run)"
w(f" - `{f}` `{obj}`: {why}{here}.")
w("- **References** are compared by presence only (`R`), not by target.") w("- **References** are compared by presence only (`R`), not by target.")
w("") w("")
if res.get("ref_only_errors"): if res.get("ref_only_errors"):
+2
View File
@@ -71,6 +71,8 @@ xargs -a "$OUT/files.txt" -d '\n' -P "$JOBS" -I{} bash -c '
f="$1"; d="$OUT/runs/${f//\//__}" f="$1"; d="$OUT/runs/${f//\//__}"
case "$f" in cve_hdf5/*) export WITH_H5DUMP=1 ;; esac case "$f" in cve_hdf5/*) export WITH_H5DUMP=1 ;; esac
"$HERE/run_one.sh" "$C/$f" "$d"' _ {} 2>"$OUT/probe.log" "$HERE/run_one.sh" "$C/$f" "$d"' _ {} 2>"$OUT/probe.log"
echo "== re-checking the objects h5py reads only through a libhdf5 bug"
"$PY" "$HERE/ref_bugs.py" "$C" > "$OUT/ref_bugs.json" 2> "$OUT/ref_bugs.err" || true
echo "== comparing" echo "== comparing"
"$PY" "$HERE/compare.py" "$OUT" >/dev/null "$PY" "$HERE/compare.py" "$OUT" >/dev/null
t2=$(date +%s) t2=$(date +%s)
+92
View File
@@ -0,0 +1,92 @@
#!/usr/bin/env python3
"""Tests of the reference side's corrections: `python conformance/test_ref.py`.
- ref.py compares a big-endian VL sequence by the file's values even though
h5py returns them byte-swapped (and records that it corrected them);
- ref_bugs.py confirms an object only when its reads disagree.
"""
import json
import os
import subprocess
import sys
import tempfile
import unittest
import h5py
import numpy as np
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, HERE)
import ref_bugs # noqa: E402
def ref_objects(path):
out = subprocess.run([sys.executable, os.path.join(HERE, "ref.py"), path],
capture_output=True, text=True, check=True).stdout
return {o["path"]: o for o in json.loads(out)["objects"]}
class BigEndianVlen(unittest.TestCase):
def test_be_vlen_compared_by_file_values(self):
with tempfile.TemporaryDirectory() as d:
path = os.path.join(d, "v.h5")
with h5py.File(path, "w") as f:
for name, order in (("be", ">"), ("le", "<")):
t = np.dtype(order + "f4")
ds = f.create_dataset(name, (2,), dtype=h5py.vlen_dtype(t))
ds[0] = np.array([1.0, 2.0], dtype=t)
ds[1] = np.array([3.0], dtype=t)
u = np.dtype(order + "u8")
f.attrs.create(name, [np.array([1, 2], dtype=u), np.array([42], dtype=u)],
dtype=h5py.vlen_dtype(u))
objs = ref_objects(path)
be, le = objs["/be"], objs["/le"]
# Same values, so the same canonical hash whatever the file's byte order.
self.assertEqual(be["hash"], le["hash"])
self.assertEqual(objs["/"]["attrs"]["be"]["hash"], objs["/"]["attrs"]["le"]["hash"])
self.assertNotIn("ref_fix", le)
# And the correction is recorded wherever h5py needed it.
import ref
if ref.be_vlen_bug():
self.assertEqual(be.get("ref_fix"), ["h5py-be-vlen"])
self.assertEqual(objs["/"]["attrs"]["be"].get("ref_fix"), ["h5py-be-vlen"])
class RefBugsConfirmation(unittest.TestCase):
def run_check(self, outcomes):
seq = iter(outcomes)
saved = ref_bugs.read_once
ref_bugs.read_once = lambda *a: next(seq)
try:
key = next(iter(ref_bugs.READ_BUGS))
with tempfile.TemporaryDirectory() as d:
p = os.path.join(d, key[0])
os.makedirs(os.path.dirname(p))
open(p, "wb").close()
return ref_bugs.check(d, key)
finally:
ref_bugs.read_once = saved
def test_stable_values_are_not_confirmed(self):
r = self.run_check(["values a"] * len(ref_bugs.RUNS))
self.assertFalse(r["confirmed"])
def test_changing_values_are_confirmed(self):
r = self.run_check(["values a"] * (len(ref_bugs.RUNS) - 1) + ["values b"])
self.assertTrue(r["confirmed"])
r = self.run_check(["values a"] * (len(ref_bugs.RUNS) - 1) + ["error filter failed"])
self.assertTrue(r["confirmed"])
def test_errors_only_are_not_confirmed(self):
# h5py cannot read it at all: nothing it reads, nothing to excuse.
r = self.run_check(["error x"] * (len(ref_bugs.RUNS) - 1) + ["error y"])
self.assertFalse(r["confirmed"])
def test_missing_file_is_not_confirmed(self):
key = next(iter(ref_bugs.READ_BUGS))
with tempfile.TemporaryDirectory() as d:
self.assertFalse(ref_bugs.check(d, key)["confirmed"])
if __name__ == "__main__":
unittest.main()
+7 -1
View File
@@ -295,7 +295,13 @@ impl ObjectHeader {
/// The messages of one version-1 chunk: each checked and appended to /// The messages of one version-1 chunk: each checked and appended to
/// `messages` (NIL ones dropped), each continuation added to `spans`. /// `messages` (NIL ones dropped), each continuation added to `spans`.
/// Returns how many messages (NIL ones included) the chunk holds. /// Returns how many messages (NIL ones included) the chunk holds.
#[inline(never)] ///
/// Inlined into the chunk loop: kept out of line (`#[inline(never)]`,
/// 4313917), the call cost `ObjectHeader::parse` about 2.5 ns per
/// header, 4% on small version-1 headers (A/B builds, 2026-09-27; see
/// `BENCHMARKS.md`). Without an attribute the compiler keeps it out of
/// line too.
#[inline]
fn parse_v1_messages( fn parse_v1_messages(
data: &[u8], data: &[u8],
offset_size: u8, offset_size: u8,
+55 -2
View File
@@ -9,7 +9,14 @@ deleting it.
## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27) ## `ObjectHeader::parse` 4% slower after range-read M2/M3 (measured 2026-09-27)
**Status:** open (speed only; values are correct). Measured on an idle tank **Status:** fixed 2026-09-27 (`96086ad`, branch
`perf/header-parse-and-last-files`): the remaining cost was the call to the
version-1 message loop, kept out of line by `#[inline(never)]`; inlined,
`object_header_parse_x401` is 1.0% and 2.6% below `8f59b2e` in two idle
A/B runs (`BENCHMARKS.md`, "`ObjectHeader::parse` back at 8f59b2e's
speed"), with every chunk-queue check unchanged.
Was: open (speed only; values are correct). Measured on an idle tank
(load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`) (load below 2 at every round), `main` before range-read M2/M3 (`8f59b2e`)
against `main` `7a8fae0`, separate binaries alternating, 3 rounds against `main` `7a8fae0`, separate binaries alternating, 3 rounds
(`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"): (`BENCHMARKS.md`, "Local metadata and data reads after range-read M2/M3"):
@@ -25,6 +32,50 @@ An earlier run the same day under load also listed single-thread contiguous
hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with hyperslab reads as 5.6% slower; the idle rerun puts them at −1.7% with
overlapping ranges (noise), so that item is withdrawn. overlapping ranges (noise), so that item is withdrawn.
## Conformance: the last non-ok files (checked 2026-09-27)
**Status:** classified with evidence — none is a clawhdf5 bug. The
conformance report had 3 our-errors and 2 mismatches left, each documented
as "not ours" but still counted against us. Re-checked on tank
(h5py 3.16 / HDF5 2.0.0, h5dump 1.14.6):
- **2 mismatches, h5py's big-endian VL bug:**
`NCAS-CMS_pyfive/tests/data/attr_datatypes.hdf5` `/@vlen_uint64` and
`hdf5/tools/test/testfiles/tcomplex_be.h5`
`/VariableLengthDatasetFloatComplex`. h5py returns the elements with the
file's big-endian bytes under a little-endian dtype (a `vlen('>f4')`
holding `[1.0, 2.0]` reads back as `[4.6e-41, 9.0e-44]`). h5dump 1.14.6
prints `(1, 2), (3, 4, 5), (42)` for `vlen_uint64`, which is what we read;
it cannot print `tcomplex_be.h5` (complex types are HDF5 2.0). The
reference side (`conformance/ref.py`) now checks that the installed h5py
has the bug and relabels such elements with the file's byte order, so the
values are compared rather than excused: both objects are identical to
ours, and both files are now *ok*.
- **3 our-errors, objects HDF5 2.0 reads only by over-reading memory:**
`cve-2025-2308.h5` `/Scale_offset_long_long_data_le` (first chunk records
`minbits` 11; its 12 values need 17 bytes of codes and the 26-byte chunk
holds 5 after its 21-byte header), `cve-2025-44904.h5`
`/Scale_offset_float_data_le` (unfiltered chunks stored as 38 and 37 bytes
for 48-byte chunks; 1.14.6's `H5D__chunk_lock` reads the stored bytes
into a buffer of that size and uses it as the whole chunk) and
`bad_nbit_parms_walk.h5` `/Nbit_int_data_le` (N-Bit parameters
`(7, 0, 40, 1, 4, 0, 20)`: an integer needs 8, and the decoder takes the
bit offset from `cd_values[7]`, past the list). **Could we match
libhdf5?** No: its output is not determined by the file. h5py's values for
the first two change from one run to the next (three plain runs gave three
different results), and all three change with `MALLOC_PERTURB_` or with
whether numpy is imported before h5py (`bad_nbit_parms_walk` then fails
with "filter returned failure during read" or reads all zeros); h5dump
1.14.6 prints yet other values for the over-read parts (`1280, 0, 0` where
h5py gave e.g. `1283, 749, 1713`). libhdf5's develop branch refuses all
three. We keep
refusing them. `conformance/ref_bugs.py` repeats the check in every
conformance run (six reads per object in differently set-up processes) and
such a file is classified *ref-bug* only while its values keep changing.
Result (`conformance/run.sh --no-fetch`, tank, 2026-09-27): 602 of 697 ok
(was 600), 0 our-error, 0 mismatch, 3 ref-bug, 92 h5py-cannot-read.
## Files a SWMR writer had open could not be read past a stale end of file ## Files a SWMR writer had open could not be read past a stale end of file
**Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any **Status:** fixed 2026-09-27 (branch `feat/p3-m5-swmr-reader`), before any
@@ -505,7 +556,9 @@ fill-value item that did is fixed).
short"); `bad_nbit_parms_walk.h5` has an N-Bit parameter list one value short"); `bad_nbit_parms_walk.h5` has an N-Bit parameter list one value
short (libhdf5's own `test_filter_bad_params` in `test/dsets.c` now short (libhdf5's own `test_filter_bad_params` in `test/dsets.c` now
requires that read to fail). We refuse both; `CONFORMANCE.md` lists them requires that read to fail). We refuse both; `CONFORMANCE.md` lists them
under *Known not-our-bug*. Scale-offset did decode three cases under *Known not-our-bug* (since 2026-09-27: the *ref-bug* class, with
the evidence re-checked every run — see *Conformance: the last non-ok
files* above). Scale-offset did decode three cases
differently from libhdf5 (codes after a `minval` of recorded size other differently from libhdf5 (codes after a `minval` of recorded size other
than 8, `minbits` 0 with a fill value, full-width `minbits`): fixed than 8, `minbits` 0 with a fill value, full-width `minbits`): fixed
2026-09-26, and the full-width case was silent wrong data on ordinary 2026-09-26, and the full-width case was silent wrong data on ordinary