Chunked reads beat an h5py process pool; unlimited writer B-trees; Blosc2; 599/697 conformance #16
+26
-6
@@ -61,12 +61,10 @@ ZFP filters; the 2 mismatches are the known h5py big-endian VL bug.
|
|||||||
- a simple dataspace of rank 0 holds one element (it held 0;
|
- a simple dataspace of rank 0 holds one element (it held 0;
|
||||||
`cve-2020-18494`), and contiguous storage larger than the dataset reads
|
`cve-2020-18494`), and contiguous storage larger than the dataset reads
|
||||||
(`cve-2024-32623`, `cve-2025-2309`; libhdf5 ignores the excess);
|
(`cve-2024-32623`, `cve-2025-2309`; libhdf5 ignores the excess);
|
||||||
- scale-offset: the packed codes start at byte 21 whatever size the chunk
|
- scale-offset returned wrong values for ordinary h5py files — see
|
||||||
records for `minval`, a chunk with `minbits` 0 and a fill value is all
|
*Correctness* below; E-scale is refused, as in libhdf5; codes past the
|
||||||
fill values, full-width `minbits` stores the elements as they are
|
end of the chunk stay an error (`cve-2025-2308`, where HDF5 2.0 reads
|
||||||
(these decoded differently from libhdf5); E-scale is refused, as in
|
past its buffer);
|
||||||
libhdf5; codes past the end of the chunk stay an error
|
|
||||||
(`cve-2025-2308`, where HDF5 2.0 reads past its buffer);
|
|
||||||
- shuffle uses its own parameter as the element size, as libhdf5 does
|
- shuffle uses its own parameter as the element size, as libhdf5 does
|
||||||
(`cve-2025-44905`);
|
(`cve-2025-44905`);
|
||||||
- an unfiltered chunk the index records at other than the chunk's size is
|
- an unfiltered chunk the index records at other than the chunk's size is
|
||||||
@@ -878,6 +876,28 @@ ZFP filters; the 2 mismatches are the known h5py big-endian VL bug.
|
|||||||
- CI keeps zlib-ng building and tested; the arm64 job no longer needs cmake.
|
- CI keeps zlib-ng building and tested; the arm64 job no longer needs cmake.
|
||||||
|
|
||||||
### Correctness
|
### Correctness
|
||||||
|
- **Scale-offset data read wrong values in every release that decoded it
|
||||||
|
(v2.2.0 to v2.7.0), silently, on ordinary h5py files** (fixed
|
||||||
|
2026-09-26). Of 1480 scale-offset datasets h5py writes across every
|
||||||
|
integer type (`i1` .. `u8`), `f4` and `f8`, both byte orders, with and
|
||||||
|
without a fill value, and `scaleoffset` from 0 to the full width, 332 did
|
||||||
|
not read as h5py reads them: **151 returned wrong values with no error**
|
||||||
|
and 181 failed to read. The common cause was a chunk libhdf5 stores at
|
||||||
|
full width (`minbits` equal to the type's width), which it does for any
|
||||||
|
full-width `scaleoffset` and on its own whenever a chunk's values span
|
||||||
|
most of the type's range: `scaleoffset=0` integer data with a wide range
|
||||||
|
(82 datasets, all wrong values), full-width `u4`/`i4`/`u8`/`i8` (51 wrong
|
||||||
|
values; the narrower types and the rest failed with "truncated minval" or
|
||||||
|
"implausible minbits"), and `f4` D-scale data with a large range (18,
|
||||||
|
wrong values). Such a chunk holds the elements as they are; they were
|
||||||
|
decoded as offsets from `minval`. Also fixed, found on crafted files: the
|
||||||
|
packed codes start at byte 21 whatever size the chunk records for
|
||||||
|
`minval` (`cve-2025-44905` `/Scale_offset_short_data_be`), and a chunk
|
||||||
|
with `minbits` 0 and a fill value is all fill values (it read as
|
||||||
|
`minval`). The whole matrix is now an interop test
|
||||||
|
(`crates/clawhdf5/tests/scaleoffset_interop.rs`, generated by h5py at
|
||||||
|
test time, every dataset compared); on v2.7.0's decoder it reports the
|
||||||
|
332. See `docs/known-issues.md`.
|
||||||
- **Corrupt files libhdf5 refuses are now refused instead of read.** On the
|
- **Corrupt files libhdf5 refuses are now refused instead of read.** On the
|
||||||
HDF Group's CVE reproducers, 18 objects that libhdf5 (HDF5 2.0, through
|
HDF Group's CVE reproducers, 18 objects that libhdf5 (HDF5 2.0, through
|
||||||
h5py) refuses to open were read by clawhdf5, some as wrong data (a chunk
|
h5py) refuses to open were read by clawhdf5, some as wrong data (a chunk
|
||||||
|
|||||||
@@ -0,0 +1,151 @@
|
|||||||
|
//! Scale-offset datasets written by h5py (libhdf5) read exactly as h5py
|
||||||
|
//! reads them.
|
||||||
|
//!
|
||||||
|
//! h5py writes the whole matrix the filter has: every integer type (`i1` ..
|
||||||
|
//! `u8`) and `f4`/`f8`, both byte orders, with and without a fill value
|
||||||
|
//! (the type's minimum, maximum, or another value), with random, constant,
|
||||||
|
//! all-fill and extreme data, and every interesting `scaleoffset` setting
|
||||||
|
//! (0 = let libhdf5 choose the bits, a few bits, full width less one, full
|
||||||
|
//! width; decimal scale factors 0..7 for floats) — about 1480 datasets. For
|
||||||
|
//! each one h5py's decoded values are stored uncompressed next to it, and
|
||||||
|
//! the raw bytes clawhdf5 decodes must equal them.
|
||||||
|
//!
|
||||||
|
//! Until 2026-09-26 clawhdf5 silently returned wrong values for 332 of
|
||||||
|
//! these, in every release that decoded scale-offset: ordinary `u8`, `u4`,
|
||||||
|
//! `i8` and `f4` data with a wide range (where libhdf5 stores the elements
|
||||||
|
//! as they are, at full width), chunks whose `minval` field was recorded at
|
||||||
|
//! another size than 8 bytes, and chunks with `minbits` 0 and a fill value.
|
||||||
|
//!
|
||||||
|
//! Skipped when python3 with h5py is unavailable, unless
|
||||||
|
//! `CLAWHDF5_REQUIRE_INTEROP=1`.
|
||||||
|
|
||||||
|
use std::process::Command;
|
||||||
|
|
||||||
|
use clawhdf5::File;
|
||||||
|
|
||||||
|
fn python() -> String {
|
||||||
|
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn python_available() -> bool {
|
||||||
|
Command::new(python())
|
||||||
|
.args(["-c", "import h5py, numpy"])
|
||||||
|
.output()
|
||||||
|
.map(|o| o.status.success())
|
||||||
|
.unwrap_or(false)
|
||||||
|
}
|
||||||
|
|
||||||
|
const GENERATE: &str = r#"
|
||||||
|
import sys, h5py, numpy as np
|
||||||
|
rng = np.random.default_rng(7)
|
||||||
|
out, expect = sys.argv[1], sys.argv[2]
|
||||||
|
made = []
|
||||||
|
with h5py.File(out, "w") as f:
|
||||||
|
def make(name, a, so, fv, desc):
|
||||||
|
try:
|
||||||
|
d = f.create_dataset(name, data=a, chunks=(16,), scaleoffset=so, fillvalue=fv)
|
||||||
|
except Exception:
|
||||||
|
return # a combination libhdf5 refuses to write
|
||||||
|
d.attrs["desc"] = desc
|
||||||
|
made.append(name)
|
||||||
|
i = 0
|
||||||
|
for dt in ["i1", "u1", "i2", "u2", "i4", "u4", "i8", "u8"]:
|
||||||
|
for bo in "<>":
|
||||||
|
t = np.dtype(bo + dt)
|
||||||
|
info = np.iinfo(t)
|
||||||
|
for so in [0, 1, 3, 8 * t.itemsize - 1, 8 * t.itemsize]:
|
||||||
|
for fill in [None, "min", "max", "mid"]:
|
||||||
|
for pat in ["rand", "const", "allfill", "extreme"]:
|
||||||
|
n = 64
|
||||||
|
fv = {None: None, "min": info.min, "max": info.max, "mid": t.type(7)}[fill]
|
||||||
|
if pat == "rand":
|
||||||
|
lo = max(info.min, -50) if so else info.min
|
||||||
|
hi = min(info.max, 50) if so else info.max
|
||||||
|
a = rng.integers(lo, hi, size=n, endpoint=True,
|
||||||
|
dtype=np.int64 if t.kind == "i" else np.uint64).astype(t)
|
||||||
|
elif pat == "const":
|
||||||
|
a = np.full(n, 3, t)
|
||||||
|
elif pat == "extreme":
|
||||||
|
a = np.array([info.min, info.max] * (n // 2), t)
|
||||||
|
else:
|
||||||
|
if fv is None:
|
||||||
|
continue
|
||||||
|
a = np.full(n, fv, t)
|
||||||
|
make(f"d{i}", a, so, fv, f"{bo}{dt} so={so} fill={fill} pat={pat}")
|
||||||
|
i += 1
|
||||||
|
for dt in ["f4", "f8"]:
|
||||||
|
for bo in "<>":
|
||||||
|
t = np.dtype(bo + dt)
|
||||||
|
for so in [0, 1, 2, 4, 7]:
|
||||||
|
for fill in [None, -1.5, 0.0]:
|
||||||
|
for pat in ["rand", "const", "allfill", "neg", "big"]:
|
||||||
|
n = 50
|
||||||
|
if pat == "rand":
|
||||||
|
a = rng.normal(size=n).astype(t) * 10
|
||||||
|
elif pat == "const":
|
||||||
|
a = np.full(n, 2.25, t)
|
||||||
|
elif pat == "neg":
|
||||||
|
a = -np.abs(rng.normal(size=n)).astype(t) * 1000
|
||||||
|
elif pat == "big":
|
||||||
|
a = rng.normal(size=n).astype(t) * 1e6
|
||||||
|
else:
|
||||||
|
if fill is None:
|
||||||
|
continue
|
||||||
|
a = np.full(n, fill, t)
|
||||||
|
make(f"d{i}", a, so, fill, f"{bo}{dt} so={so} fill={fill} pat={pat}")
|
||||||
|
i += 1
|
||||||
|
# What libhdf5 decodes, stored uncompressed in the same datatype.
|
||||||
|
with h5py.File(out, "r") as f, h5py.File(expect, "w") as e:
|
||||||
|
for name in made:
|
||||||
|
e.create_dataset(name, data=f[name][()])
|
||||||
|
print(len(made))
|
||||||
|
"#;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn every_scale_offset_dataset_reads_as_h5py_reads_it() {
|
||||||
|
if !python_available() {
|
||||||
|
assert!(
|
||||||
|
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
|
||||||
|
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
|
||||||
|
);
|
||||||
|
eprintln!("SKIP: python3 with h5py not available");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let written = dir.path().join("scaleoffset.h5");
|
||||||
|
let expected = dir.path().join("expected.h5");
|
||||||
|
let out = Command::new(python())
|
||||||
|
.args(["-c", GENERATE])
|
||||||
|
.arg(&written)
|
||||||
|
.arg(&expected)
|
||||||
|
.output()
|
||||||
|
.expect("failed to run python");
|
||||||
|
assert!(
|
||||||
|
out.status.success(),
|
||||||
|
"python failed:\n{}",
|
||||||
|
String::from_utf8_lossy(&out.stderr)
|
||||||
|
);
|
||||||
|
let count: usize = String::from_utf8_lossy(&out.stdout).trim().parse().unwrap();
|
||||||
|
assert!(count > 1400, "only {count} datasets written");
|
||||||
|
|
||||||
|
let file = File::open(&written).unwrap();
|
||||||
|
let reference = File::open(&expected).unwrap();
|
||||||
|
let mut names = reference.root().datasets().unwrap();
|
||||||
|
names.sort();
|
||||||
|
assert_eq!(names.len(), count);
|
||||||
|
let mut wrong = Vec::new();
|
||||||
|
for name in &names {
|
||||||
|
let want = reference.read_multi(&[name]).unwrap().remove(0);
|
||||||
|
let got = file.read_multi(&[name]).map(|mut v| v.remove(0));
|
||||||
|
if got.as_ref().ok() != Some(&want) {
|
||||||
|
let desc = file.dataset(name).unwrap().attrs().unwrap().remove("desc");
|
||||||
|
wrong.push(format!("{name} {desc:?}: {:?}", got.map(|g| g.len())));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
assert!(
|
||||||
|
wrong.is_empty(),
|
||||||
|
"{} of {count} scale-offset datasets differ from h5py:\n{}",
|
||||||
|
wrong.len(),
|
||||||
|
wrong.join("\n")
|
||||||
|
);
|
||||||
|
}
|
||||||
+39
-1
@@ -63,6 +63,43 @@ tank with `concurrent_read` against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`,
|
|||||||
`CHANGELOG.md`). The chunked-read scaling item above is still open.
|
`CHANGELOG.md`). The chunked-read scaling item above is still open.
|
||||||
Values are correct; this is speed only.
|
Values are correct; this is speed only.
|
||||||
|
|
||||||
|
## Scale-offset data read back wrong values
|
||||||
|
|
||||||
|
**Status:** fixed 2026-09-26, after v2.7.0. **Every release that decoded
|
||||||
|
the scale-offset filter (v2.2.0 to v2.7.0) is affected**, on ordinary
|
||||||
|
files h5py writes, with no error.
|
||||||
|
|
||||||
|
Found by the review of the 2026-09-26 conformance work: h5py wrote 1480
|
||||||
|
scale-offset datasets (every integer type `i1` .. `u8`, `f4` and `f8`,
|
||||||
|
little- and big-endian, with no fill value and with the type's minimum,
|
||||||
|
maximum or another value as fill, random, constant, all-fill and extreme
|
||||||
|
data, `scaleoffset` 0, 1, 3, full width less one and full width for
|
||||||
|
integers, decimal scale factors 0, 1, 2, 4 and 7 for floats). v2.7.0's
|
||||||
|
decoder read 332 of them differently from h5py: **151 returned wrong
|
||||||
|
values with no error**, 181 failed to read.
|
||||||
|
|
||||||
|
| Case | Datasets | v2.7.0 |
|
||||||
|
|---|---|---|
|
||||||
|
| integer, `scaleoffset=0` (libhdf5 picks the bits), data spanning most of the type's range | 82 | wrong values |
|
||||||
|
| integer, `scaleoffset` = full width: `i4`/`u4` (10), `i8`/`u8` (41) | 51 | wrong values |
|
||||||
|
| integer, `scaleoffset` = full width, the other datasets (every `i1`..`u2` one, most `i4`/`u4`, some `i8`/`u8`) | 181 | "truncated minval" / "implausible minbits" |
|
||||||
|
| `f4` D-scale, factor 4 or 7, values up to about 10^6 | 18 | wrong values |
|
||||||
|
|
||||||
|
In every case libhdf5 stored a chunk at full width (`minbits` equal to the
|
||||||
|
type's width): the chunk then holds the elements as they are, and they
|
||||||
|
were decoded as offsets from `minval`. Two more differences were found on
|
||||||
|
crafted files and fixed with them: the packed codes start at byte 21
|
||||||
|
whatever size the chunk records for `minval` (`cve-2025-44905`
|
||||||
|
`/Scale_offset_short_data_be`), and a chunk with `minbits` 0 and a fill
|
||||||
|
value is all fill values (it read as `minval`).
|
||||||
|
|
||||||
|
**Fix:** `clawhdf5_format::filters` decodes a scale-offset chunk as
|
||||||
|
`H5Z__filter_scaleoffset` does. **Test:** the whole matrix is
|
||||||
|
`crates/clawhdf5/tests/scaleoffset_interop.rs`, generated by h5py at test
|
||||||
|
time and compared dataset by dataset; on v2.7.0's decoder it reports the
|
||||||
|
332. **Existing data:** the files were always right; only reads were
|
||||||
|
wrong, so re-reading with a fixed build gives the correct values.
|
||||||
|
|
||||||
## Silent wrong data found by the 2026-09-25 HDF5 audit
|
## Silent wrong data found by the 2026-09-25 HDF5 audit
|
||||||
|
|
||||||
**Status:** fixed after v2.7.0 (2026-09-25). **Every release up
|
**Status:** fixed after v2.7.0 (2026-09-25). **Every release up
|
||||||
@@ -230,7 +267,8 @@ fill-value item that did is fixed).
|
|||||||
under *Known not-our-bug*. Scale-offset did decode three cases
|
under *Known not-our-bug*. Scale-offset did decode three cases
|
||||||
differently from libhdf5 (codes after a `minval` of recorded size other
|
differently from libhdf5 (codes after a `minval` of recorded size other
|
||||||
than 8, `minbits` 0 with a fill value, full-width `minbits`): fixed
|
than 8, `minbits` 0 with a fill value, full-width `minbits`): fixed
|
||||||
2026-09-26.
|
2026-09-26, and the full-width case was silent wrong data on ordinary
|
||||||
|
h5py files (see *Scale-offset data read back wrong values* above).
|
||||||
- **Filters:** blosc, blosc2, bitshuffle, bzip2, LZF and zfp are not
|
- **Filters:** blosc, blosc2, bitshuffle, bzip2, LZF and zfp are not
|
||||||
implemented. **Fixed 2026-09-26** for LZF (default-on `lzf` feature),
|
implemented. **Fixed 2026-09-26** for LZF (default-on `lzf` feature),
|
||||||
bitshuffle, bzip2 and Blosc 1 (`bitshuffle`, `bzip2`, `blosc`, or
|
bitshuffle, bzip2 and Blosc 1 (`bitshuffle`, `bzip2`, `blosc`, or
|
||||||
|
|||||||
Reference in New Issue
Block a user