Merge branch 'fix/p2b-remaining-conformance' into feat/p2b-scale

# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-format/src/chunked_read.rs
This commit is contained in:
osobh
2026-09-26 11:57:35 -05:00
36 changed files with 3297 additions and 370 deletions
+114
View File
@@ -124,6 +124,98 @@
allocation for these frames and for 45,000 fuzzed ones: at most 6x the
chunk size, twice the input and 2 MiB of Zstandard state.
### Remaining conformance errors (2026-09-26)
Conformance on tank, `conformance/run.sh --no-fetch`: 598 of 697 files
ok (575 before). Of the 5 our-errors left, 3 are corrupt data HDF5 2.0
reads only through a bug (listed in `CONFORMANCE.md`), 2 are the Blosc2 and
ZFP filters; the 2 mismatches are the known h5py big-endian VL bug. The
five files whose cache image libhdf5 cannot load (`cve-2025-6269-*`,
`cve-2025-6516`) count as ok because the library, like libhdf5, opens them
and fails their objects (see below).
- **Metadata cache images are read.** A file written with a metadata cache
image keeps its metadata cache entries in an image block the superblock
extension points at, and libhdf5 reads them in place of the file's own
bytes; in `h5clear_mdc_image.h5` the root group exists only there, and
every reader failed with `InvalidObjectHeaderVersion(0)`. `File`,
`MmapFile` and `LazyFile` (and `h5rs`) now apply the image at open
(`clawhdf5_format::superblock_ext::CacheImage`), with libhdf5's checks.
The file is not copied to do it: a mapped file gets the image's entries
written into a private copy-on-write mapping
(`clawhdf5_io::HDF5Read::private_copy`, `MAP_PRIVATE`), so only the pages
they land on are copied, and a buffer the opener owns (`File::from_bytes`,
`open_buffered`) is patched in place; files without an image are read
from the mapping exactly as before. (An interim version copied the whole
file onto the heap: 2 GB of memory to open a 1 GiB sparse file with an
image, and an abort for an 8 GiB one; `tests/cache_image_memory.rs`
guards it.) An image entry that runs past the end of file is refused
(libhdf5 checks only its start; the images it writes never do this). A
file whose image libhdf5 cannot load (`cve-2025-6269-*`, `cve-2025-6516`)
opens, as in libhdf5, and every object lookup fails with the image's
error (`File`, `MmapFile`; `LazyFile` reads the root group at open, so
its open fails). libhdf5 fails only its first metadata read and then
reads the file's own, possibly stale, bytes; those are never read here.
An interim version refused such a file at `File::open` while the
conformance probe reported it as libhdf5 does, so the gate counted five
files as agreeing with h5py that the library did not open; probe and
library now take the decision from the same
`superblock_ext::cache_image_state`.
- **Every other opener applies the superblock extension and the cache
image too** (`superblock_ext::apply_cache_image_in_place`, writing into
the buffer each already owns): `clawhdf5_io`'s `NativeVol` (at `open`,
and on read for `from_bytes`), `AsyncHDF5File`, `MpiVol` (a minimal edit
through the same `vol::load_hdf5`; the `mpi-io` feature cannot be built
without an MPI installation, so it was not compiled), and the external
source files of a virtual dataset. They read a file with an image from
its own bytes — stale metadata, or none (`h5clear_mdc_image.h5` failed
with `InvalidObjectHeaderVersion(0)`) — and skipped the extension checks
`File::open` makes. These readers cannot open a file and fail each
object, so an image libhdf5 cannot load is refused with the image's
error.
- **The superblock extension is decoded at open, as libhdf5 does:** a File
Space Info or Metadata Cache Image message libhdf5 cannot decode makes the
open fail (`cve-2020-10810`, `cve-2020-10812` were opened).
`FormatError::InvalidSuperblockExtension`, `InvalidCacheImage`.
- **Dataset storage libhdf5 refuses at open is refused at open**
(`FormatError::InvalidDatasetStorage`, `data_read::check_dataset_storage`):
an element count times element size that overflows (`cve-2024-32624`
`/Dset_OBJREF` opened and reported its shape), contiguous storage past the
end of the file, compact data of the wrong size. An empty contiguous
dataset at a defined address, which clawhdf5 up to v2.7.0 wrote, still
opens.
- **Wrong or missing data fixed:**
- a simple dataspace of rank 0 holds one element (it held 0;
`cve-2020-18494`), and contiguous storage larger than the dataset reads
(`cve-2024-32623`, `cve-2025-2309`; libhdf5 ignores the excess);
- scale-offset returned wrong values for ordinary h5py files — see
*Correctness* below; E-scale is refused, as in libhdf5; codes past the
end of the chunk stay an error (`cve-2025-2308`, where HDF5 2.0 reads
past its buffer);
- shuffle uses its own parameter as the element size, as libhdf5 does
(`cve-2025-44905`);
- an unfiltered chunk the index records at other than the chunk's size is
refused (it read with zeros for the missing bytes; `cve-2025-44904`);
- a v1 B-tree chunk index is read as libhdf5 reads it: each chunk is
looked up the way `H5B_find` / `H5D__btree_cmp3` / `H5D__btree_found`
look it up, and a chunk that lookup does not find reads as fill values.
A key with a non-zero element-size coordinate is found in a 1-D dataset
and not in one of rank 2 or more (`cve-2025-44905`
`/Shuffle_float_data_le`, which read the chunk's data where h5py reads
fill values); an interim fix refused every such key, including 1-D
files libhdf5 reads correctly.
- **Refused as libhdf5 refuses them:** a v1 group with an empty link name
fails its listing (`FormatError::InvalidLinkName`; lookups still work,
`cve-2021-46244`); dataspaces with more than 32 dimensions, a rank on a
scalar or null dataspace, or a dimension over its maximum
(`FormatError::InvalidDataspace`).
- `ObjectHeader::object_class` classifies a header as libhdf5 does (a
dataset needs a datatype *and* a dataspace).
- Conformance harness: user-defined links were listed as objects by the
reference, unopenable objects were not deduplicated, nested array types
were hashed wrong (`tarray3.h5`), and the attributes of objects h5py
cannot open were compared; all fixed. `CONFORMANCE.md` lists the corrupt
objects HDF5 2.0 reads through a bug (`bad_nbit_parms_walk.h5` among
them: libhdf5's own test now requires that read to fail).
### Concurrent reads (2026-09-26)
- **Full reads of chunked datasets scale with threads again when rayon's
pool has one thread.** Each full read handed its chunks to rayon to
@@ -910,6 +1002,28 @@
- CI keeps zlib-ng building and tested; the arm64 job no longer needs cmake.
### Correctness
- **Scale-offset data read wrong values in every release that decoded it
(v2.2.0 to v2.7.0), silently, on ordinary h5py files** (fixed
2026-09-26). Of 1480 scale-offset datasets h5py writes across every
integer type (`i1` .. `u8`), `f4` and `f8`, both byte orders, with and
without a fill value, and `scaleoffset` from 0 to the full width, 332 did
not read as h5py reads them: **151 returned wrong values with no error**
and 181 failed to read. The common cause was a chunk libhdf5 stores at
full width (`minbits` equal to the type's width), which it does for any
full-width `scaleoffset` and on its own whenever a chunk's values span
most of the type's range: `scaleoffset=0` integer data with a wide range
(82 datasets, all wrong values), full-width `u4`/`i4`/`u8`/`i8` (51 wrong
values; the narrower types and the rest failed with "truncated minval" or
"implausible minbits"), and `f4` D-scale data with a large range (18,
wrong values). Such a chunk holds the elements as they are; they were
decoded as offsets from `minval`. Also fixed, found on crafted files: the
packed codes start at byte 21 whatever size the chunk records for
`minval` (`cve-2025-44905` `/Scale_offset_short_data_be`), and a chunk
with `minbits` 0 and a fill value is all fill values (it read as
`minval`). The whole matrix is now an interop test
(`crates/clawhdf5/tests/scaleoffset_interop.rs`, generated by h5py at
test time, every dataset compared); on v2.7.0's decoder it reports the
332. See `docs/known-issues.md`.
- **Corrupt files libhdf5 refuses are now refused instead of read.** On the
HDF Group's CVE reproducers, 18 objects that libhdf5 (HDF5 2.0, through
h5py) refuses to open were read by clawhdf5, some as wrong data (a chunk
+7 -1
View File
@@ -146,7 +146,13 @@ for rel in files:
if a.get("kind") != b.get("kind") and "error" not in b and "error" not in a:
issues.append(("mismatch", f"{p}: kind {a.get('kind')} vs ours {b.get('kind')}", "kind", b))
ok = False
# h5py could not open the object at all: it read none of its
# attributes or links, so there is nothing to compare ours with
# (the object's own error is compared above and below).
ref_unopened = a.get("kind") == "unknown" and "error" in a
for k in ("error", "list_error", "attrs_error"):
if ref_unopened and k != "error":
continue
if k in b and k not in a:
issues.append(("our-error", f"{p}: {k}: {b[k]}", b[k], b))
ok = False
@@ -164,7 +170,7 @@ for rel in files:
issues.append(("mismatch", f"{p}: values differ (h5py {a.get('dtype')} vs ours {b.get('dtype')})", "values", b | {"ref_head": a.get("head"), "ref_dtype": a.get("dtype")}))
ok = False
ra, oa = a.get("attrs") or {}, b.get("attrs") or {}
if "attrs_error" not in b and "attrs_error" not in a:
if "attrs_error" not in b and "attrs_error" not in a and not ref_unopened:
for an in sorted(set(ra) | set(oa)):
x, y = ra.get(an), oa.get(an)
if x is None:
+99 -24
View File
@@ -31,7 +31,7 @@ use clawhdf5_format::filter_pipeline::FilterPipeline;
use clawhdf5_format::group_v1::{self, GroupEntry};
use clawhdf5_format::group_v2;
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::object_header::{ObjectClass, ObjectHeader};
use clawhdf5_format::signature;
use clawhdf5_format::superblock::Superblock;
use clawhdf5_format::symbol_table::SymbolTableMessage;
@@ -308,18 +308,20 @@ impl<'a> Ctx<'a> {
)
.map_err(e)?;
}
let (shape, n) = Self::shape(&ds);
rec.insert("shape".into(), shape);
if n.saturating_mul(dt.type_size() as u64) > MAX_BYTES {
rec.insert("skipped".into(), Value::String("too large".into()));
return Ok(());
}
let lm = h
.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
.ok_or("MissingMessage(DataLayout)")?;
let dl = DataLayout::parse(&lm.data, self.os, self.ls).map_err(e)?;
// What libhdf5 checks when it opens the dataset (as File::dataset).
data_read::check_dataset_storage(&dl, &ds, &dt, self.data.len() as u64).map_err(e)?;
let (shape, n) = Self::shape(&ds);
rec.insert("shape".into(), shape);
if n.saturating_mul(dt.type_size() as u64) > MAX_BYTES {
rec.insert("skipped".into(), Value::String("too large".into()));
return Ok(());
}
rec.insert(
"layout".into(),
Value::String(
@@ -657,6 +659,20 @@ fn is_group(h: &ObjectHeader) -> bool {
})
}
/// The probe's kind for an object header: libhdf5's object class
/// ([`ObjectHeader::object_class`]: group, then dataset — a datatype *and* a
/// dataspace — then named datatype), which is what h5py opens the object as.
/// The root group, and a header with only link messages, count as groups.
fn kind_of(h: &ObjectHeader, is_root: bool) -> &'static str {
match h.object_class() {
Some(ObjectClass::Group) => "group",
Some(ObjectClass::Dataset) => "dataset",
_ if is_root || is_group(h) => "group",
Some(ObjectClass::NamedDatatype) => "datatype",
None => "unknown",
}
}
fn main() {
install_hook();
let path = std::env::args().nth(1).expect("usage: probe <file>");
@@ -696,6 +712,42 @@ fn main() {
return;
}
};
// libhdf5 decodes the superblock extension at open (an error refuses
// the file), and loads a metadata cache image over the file's own
// metadata. It loads the image only when it first reads metadata — the
// root group — so a file whose image it cannot load still opens and
// that read fails. The library decides all three cases with the same
// `cache_image_state`: `File` and `MmapFile` open such a file and fail
// every object lookup with the image's error, which is what the probe
// records here (on the root group, where libhdf5 reports it).
use clawhdf5_format::superblock_ext::{self, CacheImageState};
let state = match guarded(|| superblock_ext::cache_image_state(hdf5, &sb).map_err(e)) {
Ok(x) => x,
Err(msg) => {
top.insert("open_error".into(), Value::String(msg));
println!("{}", Value::Object(top));
return;
}
};
let mut image_error = None;
let view = match state {
CacheImageState::Absent => None,
CacheImageState::Unloadable(err) => {
image_error = Some(e(err));
None
}
CacheImageState::Loaded(image) => {
let mut v = hdf5.to_vec();
match image.block(hdf5).and_then(|b| image.apply(b, &mut v)) {
Ok(()) => Some(v),
Err(err) => {
image_error = Some(e(err));
None
}
}
}
};
let hdf5: &[u8] = view.as_deref().unwrap_or(hdf5);
top.insert("superblock_version".into(), json!(sb.version));
let ctx = Ctx {
data: hdf5,
@@ -723,6 +775,9 @@ fn main() {
let mut rec = Map::new();
rec.insert("path".into(), Value::String(p.clone()));
let r = guarded(|| {
if let Some(msg) = &image_error {
return Err(msg.clone());
}
let h = ctx.header(addr)?;
Ok(h)
});
@@ -735,23 +790,7 @@ fn main() {
continue;
}
};
let is_ds = h
.messages
.iter()
.any(|m| m.msg_type == MessageType::DataLayout);
let kind = if is_ds {
"dataset"
} else if is_group(&h) || addr == sb.root_group_address {
"group"
} else if h
.messages
.iter()
.any(|m| m.msg_type == MessageType::Datatype)
{
"datatype"
} else {
"unknown"
};
let kind = kind_of(&h, addr == sb.root_group_address);
rec.insert("kind".into(), Value::String(kind.into()));
if kind == "dataset"
&& let Err(msg) = guarded(|| ctx.read_dataset(&h, &mut rec))
@@ -867,6 +906,42 @@ mod tests {
assert!(ieee_layout(&f32le));
}
#[test]
fn kind_follows_libhdf5_object_class() {
use clawhdf5_format::object_header::HeaderMessage;
let header = |types: &[MessageType]| ObjectHeader {
version: 2,
messages: types
.iter()
.map(|&msg_type| HeaderMessage {
msg_type,
size: 0,
flags: 0,
creation_order: None,
data: Vec::new(),
})
.collect(),
reference_count: None,
flags: 0,
access_time: None,
modification_time: None,
change_time: None,
birth_time: None,
};
use MessageType::*;
// cve-2024-33874 `/Dset1`: a datatype and a layout but no dataspace
// is a named datatype to libhdf5 (h5py opens it as one).
assert_eq!(kind_of(&header(&[Datatype, DataLayout]), false), "datatype");
assert_eq!(
kind_of(&header(&[Datatype, Dataspace, DataLayout]), false),
"dataset"
);
assert_eq!(kind_of(&header(&[SymbolTable]), false), "group");
assert_eq!(kind_of(&header(&[Link]), false), "group");
assert_eq!(kind_of(&header(&[]), true), "group");
assert_eq!(kind_of(&header(&[]), false), "unknown");
}
#[test]
fn partial_precision_int_is_shifted_and_sign_extended() {
let dt = Datatype::FixedPoint {
+24 -10
View File
@@ -111,8 +111,11 @@ def note_conversion(tid, dt, rec):
def hash_values(arr, dt, rec):
if dt.subdtype is not None:
# h5py expands an HDF5 array element type into trailing array dims
# h5py expands an HDF5 array element type into trailing array dims, a
# nested array type (an array of arrays) into all of them. Converting the
# expanded array back to the inner subarray type would broadcast every
# element into a whole subarray, so strip every level.
while dt.subdtype is not None:
dt = dt.subdtype[0]
arr = np.asarray(arr, dtype=dt)
if simple(dt):
@@ -173,9 +176,13 @@ def main(path):
return
objects = []
seen = set()
stack = [("/", None)]
# Objects h5py cannot open have no ObjectID to deduplicate by; they are
# deduplicated by the address their hard link points at instead, as the
# probe deduplicates every object by header address.
seen_unopenable = set()
stack = [("/", None, None)]
while stack:
p, obj = stack.pop()
p, obj, link_addr = stack.pop()
if len(objects) >= MAX_OBJECTS:
top["truncated"] = True
break
@@ -185,6 +192,10 @@ def main(path):
obj = f[p]
key = hash(obj.id) # h5py ObjectID hash = (fileno, object address/token)
except Exception as e: # noqa: BLE001
if link_addr is not None:
if link_addr in seen_unopenable:
continue
seen_unopenable.add(link_addr)
rec["kind"] = "unknown"
rec["error"] = err(e)
objects.append(rec)
@@ -232,15 +243,18 @@ def main(path):
base = "" if p == "/" else p
kids = []
for n in names:
# The link's own type: `obj.get(n, getlink=True)` reports
# a user-defined link (type 64-255) as a HardLink.
try:
link = obj.get(n, getlink=True)
info = obj.id.links.get_info(n.encode("utf-8", "surrogateescape"))
except Exception: # noqa: BLE001
link = None
if link is not None and not isinstance(link, h5py.HardLink):
info = None
if info is not None and info.type != h5py.h5l.TYPE_HARD:
continue
kids.append(f"{base}/{n}")
for k in reversed(kids):
stack.append((k, None))
addr = info.u if info is not None else None
kids.append((f"{base}/{n}", addr))
for k, addr in reversed(kids):
stack.append((k, None, addr))
except Exception as e: # noqa: BLE001
rec["list_error"] = err(e)
objects.append(rec)
+39 -6
View File
@@ -98,13 +98,37 @@ def is_h5py_be_vlen(i):
and ">" in (i.get("ours_dtype") or ""))
# Objects the reference (h5py 3.16 / HDF5 2.0) reads only because of an
# HDF5 2.0 bug, and that clawhdf5 refuses: each one reads past a buffer or
# returns bytes the file does not hold, and libhdf5's develop branch refuses all
# three. (file, object) -> why. Checked 2026-09-26 against HDF5 2.0.0
# and HDFGroup/hdf5 develop sources; see docs/known-issues.md.
LIBHDF5_BUGS = {
("cve_hdf5/cvefiles/cve-2025-2308.h5", "/Scale_offset_long_long_data_le"):
"scale-offset codes run past the end of the chunk: HDF5 2.0 reads past its buffer; "
"libhdf5's develop branch refuses the chunk (\"Buffer too short\")",
("cve_hdf5/cvefiles/cve-2025-44904.h5", "/Scale_offset_float_data_le"):
"unfiltered chunks of 38 and 37 bytes for 48-byte chunks: HDF5 2.0 fills the rest with "
"whatever its buffer held; libhdf5's develop branch refuses them (\"incorrect chunk size returned "
"from index for unfiltered chunk\")",
("hdf5/test/testfiles/bad_nbit_parms_walk.h5", "/Nbit_int_data_le"):
"an N-Bit parameter list one value short: HDF5 2.0 reads past the list; libhdf5's own "
"test (`test_filter_bad_params`, test/dsets.c) now requires the read to fail",
}
def is_libhdf5_bug(rel, i):
return i["kind"] == "our-error" and any(
f == rel and i["detail"].startswith(obj + ":") for (f, obj) in LIBHDF5_BUGS)
known = collections.defaultdict(list)
for r in rows:
if r["class"] != "mismatch":
continue
iss = issues.get(r["file"], [])
if iss and all(is_h5py_be_vlen(i) for i in iss):
if r["class"] == "mismatch" and iss and all(is_h5py_be_vlen(i) for i in iss):
known["h5py-be-vlen"].append(r["file"])
if r["class"] == "our-error" and iss and all(is_libhdf5_bug(r["file"], i) for i in iss):
known["libhdf5-2.0"].append(r["file"])
# --- the CVE corpus: clawhdf5 vs h5dump vs h5py ------------------------------
@@ -230,9 +254,13 @@ for c in sorted(by_corpus):
w(f"| {c} | {sum(cnt.values())} | " + " | ".join(str(cnt.get(k, 0)) for k in CLASSES) + " |")
w(f"| **all** | **{len(rows)}** | " + " | ".join(f"**{total.get(k, 0)}**" for k in CLASSES) + " |")
w("")
n_known = sum(len(v) for v in known.values())
if n_known:
w(f"{n_known} of the {total.get('mismatch', 0)} mismatches are a known h5py bug, not ours (see *Known not-our-bug*).")
if known["h5py-be-vlen"]:
w(f"{len(known['h5py-be-vlen'])} of the {total.get('mismatch', 0)} mismatches are a known h5py bug, "
"not ours (see *Known not-our-bug*).")
w("")
if known["libhdf5-2.0"]:
w(f"{len(known['libhdf5-2.0'])} of the {total.get('our-error', 0)} our-errors are corrupt data that "
"HDF5 2.0 reads only through a bug and clawhdf5 refuses (see *Known not-our-bug*).")
w("")
w("Corpora (fetched by `conformance/fetch-corpus.sh` into the gitignored `conformance/.cache/`):")
w("")
@@ -311,6 +339,11 @@ if res["incomparable"]:
w(" (FP8 -> float16, bfloat16 -> float32, x87 long double -> float128) the values are not")
w(" compared (shape and presence still are): "
+ ", ".join(f"{k} ({n}x)" for k, n in res["incomparable"]) + ".")
w("- **Corrupt data HDF5 2.0 reads through a bug.** clawhdf5 refuses these objects; h5py 3.16 /")
w(" HDF5 2.0 returns values for them that the file does not hold:")
for (f, obj), why in sorted(LIBHDF5_BUGS.items()):
here = "" if f in known["libhdf5-2.0"] else " (not an our-error in this run)"
w(f" - `{f}` `{obj}`: {why}{here}.")
w("- **References** are compared by presence only (`R`), not by target.")
w("")
if res.get("ref_only_errors"):
+279 -75
View File
@@ -588,26 +588,41 @@ pub fn collect_chunk_info(
offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
collect_chunk_info_inner(
let _ = length_size;
let mut chunks = Vec::new();
parse_chunk_node(
file_data,
btree_address,
ndims,
None,
offset_size,
length_size,
0,
)
&mut chunks,
)?;
Ok(chunks)
}
/// [`collect_chunk_info`] for a layout with these `chunk_dimensions` (the
/// layout message's list, element size last), checking every key of the
/// B-tree as libhdf5 does (`H5D__btree_decode_key`): each coordinate offset
/// must be a multiple of its chunk dimension. That includes the keys that
/// only bound a node (internal-node keys and each node's final key), which
/// is where a corrupt chunk dimension shows when the chunks themselves all
/// start at offset 0 in that dimension (`cve-2018-11205`). A key that fails
/// ("bad coordinate offset") means a corrupt index or chunk dimension; the
/// chunks were read at the wrong place, or the dataset read as fill values.
/// The chunks of a v1 B-tree chunk index as libhdf5 reads them, for a
/// layout with these `chunk_dimensions` (the layout message's list, element
/// size last).
///
/// Every key of the B-tree is checked as libhdf5 checks it
/// (`H5D__btree_decode_key`): each coordinate offset must be a multiple of
/// its chunk dimension. That includes the keys that only bound a node
/// (internal-node keys and each node's final key), which is where a
/// corrupt chunk dimension shows when the chunks themselves all start at
/// offset 0 in that dimension (`cve-2018-11205`). A key that fails ("bad
/// coordinate offset") means a corrupt index or chunk dimension.
///
/// libhdf5 does not read a chunk by walking the tree: it looks each chunk
/// up (`H5B_find` with `H5D__btree_cmp3` and `H5D__btree_found`), comparing
/// the element-size coordinate too, which it asks for as 0. So a chunk is
/// returned only where that lookup finds it: a key whose element-size
/// coordinate is not 0 is found in a 1-D dataset (the comparison looks at
/// that coordinate only against the next key) but not in a dataset of rank
/// 2 or more (`cve-2025-44905` `/Shuffle_float_data_le`), which then reads
/// as fill values, and a tree whose keys are out of order loses the chunks
/// libhdf5's binary search misses.
pub fn collect_chunk_info_checked(
file_data: &[u8],
btree_address: u64,
@@ -615,15 +630,141 @@ pub fn collect_chunk_info_checked(
offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
collect_chunk_info_inner(
let _ = length_size;
let ndims = chunk_dimensions.len();
if ndims == 0 {
return Err(FormatError::ChunkedReadError(
"chunk layout has no dimensions".into(),
));
}
let mut stored = Vec::new();
let mut root = parse_chunk_node(
file_data,
btree_address,
chunk_dimensions.len(),
ndims,
Some(chunk_dimensions),
offset_size,
length_size,
0,
)
&mut stored,
)?;
// Keys were checked to be multiples of non-zero dimensions.
root.scale_keys(chunk_dimensions);
// Look every stored chunk's position up. A lookup finds a chunk only at
// that chunk's own position, so each is returned at most once.
let mut returned = vec![false; stored.len()];
let mut wanted = vec![0u64; ndims];
for chunk in &stored {
for (w, (&o, &d)) in wanted
.iter_mut()
.zip(chunk.offsets.iter().zip(chunk_dimensions))
{
*w = o / u64::from(d);
}
wanted[ndims - 1] = 0;
if let Some(i) = root.find(&wanted) {
returned[i] = true;
}
}
Ok(stored
.into_iter()
.zip(returned)
.filter_map(|(c, r)| r.then_some(c))
.collect())
}
/// A node of a v1 B-tree chunk index: its `n + 1` keys (`ndims`
/// coordinates each, flattened; byte offsets as stored, or scaled by the
/// chunk dimensions after [`ChunkNode::scale_keys`]) and its `n` children,
/// a leaf's as indices into the list of stored chunks.
struct ChunkNode {
ndims: usize,
keys: Vec<u64>,
children: ChunkChildren,
}
enum ChunkChildren {
Nodes(Vec<ChunkNode>),
Chunks(Vec<usize>),
}
impl ChunkNode {
fn key(&self, i: usize) -> &[u64] {
&self.keys[i * self.ndims..(i + 1) * self.ndims]
}
fn len(&self) -> usize {
match &self.children {
ChunkChildren::Nodes(n) => n.len(),
ChunkChildren::Chunks(c) => c.len(),
}
}
fn scale_keys(&mut self, dims: &[u32]) {
for (k, &d) in self.keys.iter_mut().zip(dims.iter().cycle()) {
*k /= u64::from(d);
}
if let ChunkChildren::Nodes(nodes) = &mut self.children {
for n in nodes {
n.scale_keys(dims);
}
}
}
/// `H5B_find_helper` over scaled keys: binary search for the child
/// whose keys bracket `scaled`, then `H5D__btree_found` at the leaf.
/// Returns the index of the chunk found.
fn find(&self, scaled: &[u64]) -> Option<usize> {
let (mut lt, mut rt) = (0, self.len());
let mut idx = 0;
let mut cmp = core::cmp::Ordering::Greater;
while lt < rt && cmp != core::cmp::Ordering::Equal {
idx = (lt + rt) / 2;
cmp = btree_cmp3(self.key(idx), scaled, self.key(idx + 1));
if cmp == core::cmp::Ordering::Less {
rt = idx;
} else {
lt = idx + 1;
}
}
if cmp != core::cmp::Ordering::Equal {
return None;
}
match &self.children {
ChunkChildren::Nodes(nodes) => nodes[idx].find(scaled),
ChunkChildren::Chunks(chunks) => {
// "Is this *really* the requested chunk?"
let lt_key = self.key(idx);
let found = scaled
.iter()
.zip(lt_key)
.all(|(&s, &k)| s < k.wrapping_add(1));
found.then_some(chunks[idx])
}
}
}
}
/// `H5D__btree_cmp3`: where `scaled` falls against a child's left and
/// right keys. `Less` is left of the child, `Greater` right of it. With a
/// rank-1 dataset (two coordinates, element size last) libhdf5 compares
/// only the first coordinate, and the second against the right key.
fn btree_cmp3(lt: &[u64], scaled: &[u64], rt: &[u64]) -> core::cmp::Ordering {
use core::cmp::Ordering;
if scaled.len() == 2 {
if scaled[0] > rt[0] || (scaled[0] == rt[0] && scaled[1] >= rt[1]) {
Ordering::Greater
} else if scaled[0] < lt[0] {
Ordering::Less
} else {
Ordering::Equal
}
} else if scaled >= rt {
Ordering::Greater
} else if scaled < lt {
Ordering::Less
} else {
Ordering::Equal
}
}
/// Check one v1 B-tree chunk key's offsets (see
@@ -640,24 +781,25 @@ fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u32]) -> Result<(), Fo
}
/// Read the `ndims` 8-byte offsets of the chunk key at `pos` (after its
/// chunk size and filter mask) and check them when `chunk_dimensions` is
/// given.
/// chunk size and filter mask) into `out`, checking them when
/// `chunk_dimensions` is given.
fn read_key_offsets(
file_data: &[u8],
pos: usize,
ndims: usize,
chunk_dimensions: Option<&[u32]>,
) -> Result<Vec<u64>, FormatError> {
let mut offsets = Vec::with_capacity(ndims);
out: &mut Vec<u64>,
) -> Result<(), FormatError> {
let start = out.len();
let mut kp = pos + 8;
for _ in 0..ndims {
offsets.push(read_offset(file_data, kp, CHUNK_KEY_OFFSET_SIZE)?);
out.push(read_offset(file_data, kp, CHUNK_KEY_OFFSET_SIZE)?);
kp += CHUNK_KEY_OFFSET_SIZE as usize;
}
if let Some(dims) = chunk_dimensions {
check_key_offsets(&offsets, dims)?;
check_key_offsets(&out[start..], dims)?;
}
Ok(offsets)
Ok(())
}
/// Width of each chunk offset in a v1 chunk B-tree key, independent of the
@@ -668,15 +810,17 @@ const CHUNK_KEY_OFFSET_SIZE: u8 = 8;
/// protection), matching `btree_v1.rs`'s `MAX_BTREE_DEPTH`.
const MAX_CHUNK_BTREE_DEPTH: usize = 64;
fn collect_chunk_info_inner(
/// Parse the v1 B-tree chunk index node at `btree_address` and its
/// subtree, appending its chunks to `stored` in tree order.
fn parse_chunk_node(
file_data: &[u8],
btree_address: u64,
ndims: usize,
chunk_dimensions: Option<&[u32]>,
offset_size: u8,
_length_size: u8,
depth: usize,
) -> Result<Vec<ChunkInfo>, FormatError> {
stored: &mut Vec<ChunkInfo>,
) -> Result<ChunkNode, FormatError> {
if depth > MAX_CHUNK_BTREE_DEPTH {
return Err(FormatError::NestingDepthExceeded);
}
@@ -711,15 +855,14 @@ fn collect_chunk_info_inner(
.and_then(|n| n.checked_add(8))
.ok_or_else(|| FormatError::ChunkedReadError("chunk key too large".into()))?;
if node_level == 0 {
// Leaf node: keys and children interleaved
// key[0], child[0], key[1], child[1], ..., key[N-1], child[N-1], key[N]
let needed = entries_used * (key_size + os) + key_size;
ensure_len(file_data, pos, needed)?;
let mut chunks = Vec::with_capacity(entries_used);
let mut keys = Vec::with_capacity((entries_used + 1) * ndims);
let mut chunks = Vec::new();
let mut child_addrs = Vec::new();
for _ in 0..entries_used {
// Parse key
let chunk_size = u32::from_le_bytes([
file_data[pos],
file_data[pos + 1],
@@ -732,53 +875,48 @@ fn collect_chunk_info_inner(
file_data[pos + 6],
file_data[pos + 7],
]);
let offsets = read_key_offsets(file_data, pos, ndims, chunk_dimensions)?;
let k = keys.len();
read_key_offsets(file_data, pos, ndims, chunk_dimensions, &mut keys)?;
pos += key_size;
// Parse child address
let address = read_offset(file_data, pos, offset_size)?;
pos += os;
chunks.push(ChunkInfo {
if node_level == 0 {
chunks.push(stored.len());
stored.push(ChunkInfo {
chunk_size,
filter_mask,
offsets,
offsets: keys[k..].to_vec(),
address,
});
} else {
child_addrs.push(address);
}
}
// The final key only bounds the node; libhdf5 still checks it.
read_key_offsets(file_data, pos, ndims, chunk_dimensions)?;
Ok(chunks)
read_key_offsets(file_data, pos, ndims, chunk_dimensions, &mut keys)?;
let children = if node_level == 0 {
ChunkChildren::Chunks(chunks)
} else {
// Internal node: recurse into children
let needed = entries_used * (key_size + os) + key_size;
ensure_len(file_data, pos, needed)?;
let mut child_addrs = Vec::with_capacity(entries_used);
for _ in 0..entries_used {
read_key_offsets(file_data, pos, ndims, chunk_dimensions)?;
pos += key_size;
let child_addr = read_offset(file_data, pos, offset_size)?;
child_addrs.push(child_addr);
pos += os;
}
read_key_offsets(file_data, pos, ndims, chunk_dimensions)?;
let mut all_chunks = Vec::new();
let mut nodes = Vec::with_capacity(child_addrs.len());
for child_addr in child_addrs {
let child_chunks = collect_chunk_info_inner(
nodes.push(parse_chunk_node(
file_data,
child_addr,
ndims,
chunk_dimensions,
offset_size,
_length_size,
depth + 1,
)?;
all_chunks.extend(child_chunks);
}
Ok(all_chunks)
stored,
)?);
}
ChunkChildren::Nodes(nodes)
};
Ok(ChunkNode {
ndims,
keys,
children,
})
}
/// Generate ChunkInfo entries for an implicit index (v4 index type 2).
@@ -1087,6 +1225,47 @@ pub fn list_chunks(
Ok((chunks, chunk_dims))
}
/// [`list_chunks`] for reading the chunks through `pipeline`: a dataset
/// without filters stores every chunk at the chunk's full size, and a chunk
/// the index records at another size is refused, as libhdf5 refuses it
/// ("incorrect chunk size returned from index for unfiltered chunk"). Such
/// a chunk was read at its recorded size, with the rest of the chunk left
/// as zeros or fill values: `cve-2025-44904`'s `Scale_offset_float_data_le`
/// has chunks of 38 and 37 bytes for 48-byte chunks, where HDF5 2.0 reads
/// whatever its buffer held for the missing bytes.
pub fn list_chunks_for_read(
file_data: &[u8],
layout: &DataLayout,
dataspace: &Dataspace,
elem_size: usize,
pipeline: Option<&FilterPipeline>,
offset_size: u8,
length_size: u8,
) -> Result<(Vec<ChunkInfo>, Vec<usize>), FormatError> {
let (chunks, chunk_dims) = list_chunks(
file_data,
layout,
dataspace,
elem_size,
offset_size,
length_size,
)?;
if pipeline.is_none_or(|p| p.filters.is_empty()) {
let chunk_bytes = checked_chunk_byte_len(&chunk_dims, elem_size)?;
if let Some(c) = chunks
.iter()
.find(|c| c.address != u64::MAX && c.chunk_size as usize != chunk_bytes)
{
return Err(FormatError::ChunkedReadError(format!(
"incorrect chunk size returned from index for unfiltered chunk at {:?}: \
{} bytes, expected {chunk_bytes}",
c.offsets, c.chunk_size
)));
}
}
Ok((chunks, chunk_dims))
}
/// The chunk cache a full read may use (`None` without `std`).
#[cfg(feature = "std")]
pub(crate) type CacheRef<'a> = Option<&'a ChunkCache>;
@@ -1114,11 +1293,12 @@ pub(crate) fn read_chunked_full<O>(
check_chunk_element_size(layout, datatype, offset_size)?;
let elem_size = datatype.type_size() as usize;
let list = || {
list_chunks(
list_chunks_for_read(
file_data,
layout,
dataspace,
elem_size,
pipeline,
offset_size,
length_size,
)
@@ -1429,11 +1609,12 @@ pub fn read_chunked_data_sweep(
// lookup is keyed by this dataset's chunk-index address, so another
// dataset's index or chunks are never used for this read.
let chunks = cache.chunks_for(addr, rank, || {
list_chunks(
list_chunks_for_read(
file_data,
layout,
dataspace,
elem_size,
pipeline,
offset_size,
length_size,
)
@@ -1570,11 +1751,12 @@ pub fn read_chunked_data_indexed(
addr,
rank,
|| {
list_chunks(
list_chunks_for_read(
file_data,
layout,
dataspace,
elem_size,
pipeline,
offset_size,
length_size,
)
@@ -1874,6 +2056,25 @@ mod tests {
/// Build a B-tree v1 type 1 leaf node with given chunk infos.
fn build_chunk_btree_leaf(chunks: &[ChunkInfo], ndims: usize, offset_size: u8) -> Vec<u8> {
build_chunk_btree_leaf_to(chunks, &vec![0; ndims], offset_size)
}
/// A leaf whose final key is the one libhdf5 writes past the last chunk:
/// each coordinate of the last chunk plus its chunk dimension (the
/// element size last).
fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u32], offset_size: u8) -> Vec<u8> {
let last = &chunks.last().expect("a chunk").offsets;
let end: Vec<u64> = dims
.iter()
.enumerate()
.map(|(d, &c)| last.get(d).copied().unwrap_or(0) + u64::from(c))
.collect();
build_chunk_btree_leaf_to(chunks, &end, offset_size)
}
/// A leaf holding `chunks`, with final key `end`.
fn build_chunk_btree_leaf_to(chunks: &[ChunkInfo], end: &[u64], offset_size: u8) -> Vec<u8> {
let ndims = end.len();
let _os = offset_size as usize;
let entries_used = chunks.len() as u16;
let mut buf = Vec::new();
@@ -1915,8 +2116,8 @@ mod tests {
// checks; 0 always is)
buf.extend_from_slice(&0u32.to_le_bytes()); // chunk_size
buf.extend_from_slice(&0u32.to_le_bytes()); // filter_mask
for _ in 0..ndims {
write_offset(&mut buf, 0, 8);
for &e in end {
write_offset(&mut buf, e, 8);
}
buf
@@ -1932,8 +2133,11 @@ mod tests {
offsets,
address,
};
let good =
build_chunk_btree_leaf(&[chunk(vec![0, 0], 0x100), chunk(vec![10, 0], 0x200)], 2, 8);
let good = build_chunk_btree_leaf_dims(
&[chunk(vec![0, 0], 0x100), chunk(vec![10, 0], 0x200)],
&[10, 8],
8,
);
assert_eq!(
collect_chunk_info_checked(&good, 0, &[10, 8], 8, 8)
.unwrap()
@@ -2164,7 +2368,6 @@ mod tests {
) -> (Vec<u8>, DataLayout, Dataspace) {
let os: u8 = 8;
let elem_size = 8usize;
let ndims = 2; // rank(1) + 1
let total = values.len();
// Place chunk data starting at offset 0x2000
@@ -2195,12 +2398,13 @@ mod tests {
}
// Build B-tree at offset 0x100
let btree = build_chunk_btree_leaf(&chunk_infos, ndims, os);
let dims = [chunk_size_elems as u32, elem_size as u32];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
let layout = DataLayout::Chunked {
chunk_dimensions: vec![chunk_size_elems as u32, elem_size as u32],
chunk_dimensions: dims.to_vec(),
btree_address: Some(btree_addr as u64),
version: 3,
chunk_index_type: None,
@@ -2530,7 +2734,6 @@ mod tests {
let os: u8 = 8;
let elem_size = 8usize;
let ndims = 2;
let chunk_elems = 10usize;
let total = 20usize;
@@ -2569,12 +2772,13 @@ mod tests {
data_offset += compressed.len() + 16; // some padding
}
let btree = build_chunk_btree_leaf(&chunk_infos, ndims, os);
let dims = [chunk_elems as u32, elem_size as u32];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
let layout = DataLayout::Chunked {
chunk_dimensions: vec![chunk_elems as u32, elem_size as u32],
chunk_dimensions: dims.to_vec(),
btree_address: Some(btree_addr as u64),
version: 3,
chunk_index_type: None,
@@ -2612,7 +2816,6 @@ mod tests {
// 4x6 dataset with chunk size 2x3 => 4 chunks
let os: u8 = 8;
let elem_size = 4usize; // f32
let ndims = 3; // rank(2) + 1
let ds_dims = [4usize, 6];
let chunk_dims = [2usize, 3];
@@ -2652,12 +2855,13 @@ mod tests {
}
}
let btree = build_chunk_btree_leaf(&chunk_infos, ndims, os);
let dims = [chunk_dims[0] as u32, chunk_dims[1] as u32, elem_size as u32];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
let layout = DataLayout::Chunked {
chunk_dimensions: vec![chunk_dims[0] as u32, chunk_dims[1] as u32, elem_size as u32],
chunk_dimensions: dims.to_vec(),
btree_address: Some(btree_addr as u64),
version: 3,
chunk_index_type: None,
+141 -14
View File
@@ -32,6 +32,80 @@ fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatErr
Ok(())
}
/// The storage checks libhdf5 makes when it opens a dataset, before any
/// data is read (`H5D__contig_check`, `H5D__compact_init`), so a dataset
/// they refuse fails to open, as in libhdf5, instead of opening and
/// reporting a shape nothing can be read from:
///
/// - the element count times the element size must not overflow 64 bits
/// ("size of dataset's storage overflowed" — `cve-2024-32624`
/// `/Dset_OBJREF`, 2^62 references of 8 bytes);
/// - contiguous storage at a defined address must end within the file's
/// `file_len` bytes (the HDF5 data up to the end of file the superblock
/// records);
/// - compact data must be exactly the dataset's size.
///
/// Deliberately not refused, unlike libhdf5: an empty contiguous dataset at
/// a defined address (libhdf5's overflow test `addr + 0 <= addr` refuses
/// it), which clawhdf5 up to v2.7.0 wrote. Chunked and virtual layouts are
/// checked when their data is read.
pub fn check_dataset_storage(
layout: &DataLayout,
dataspace: &Dataspace,
datatype: &Datatype,
file_len: u64,
) -> Result<(), FormatError> {
if !matches!(
layout,
DataLayout::Contiguous { .. } | DataLayout::Compact { .. }
) {
return Ok(());
}
const OVERFLOWED: &str = "size of dataset's storage overflowed";
let n = dataspace
.checked_num_elements()
.map_err(|_| FormatError::InvalidDatasetStorage(OVERFLOWED))?;
let data_size = n
.checked_mul(u64::from(datatype.type_size()))
.ok_or(FormatError::InvalidDatasetStorage(OVERFLOWED))?;
match layout {
DataLayout::Contiguous {
address: Some(address),
..
} if address
.checked_add(data_size)
.is_none_or(|end| end > file_len) =>
{
Err(FormatError::InvalidDatasetStorage(
"invalid dataset size, likely file corruption",
))
}
DataLayout::Compact { data } if data.len() as u64 != data_size => {
Err(FormatError::InvalidDatasetStorage(
"bad value from dataset header - size of compact dataset's data buffer \
doesn't match size of dataset data",
))
}
_ => Ok(()),
}
}
/// How many bytes to read from a contiguous dataset's storage of
/// `storage_size` bytes (the layout message's size) holding `needed` bytes
/// of elements. libhdf5 reads the elements' bytes from the start of the
/// storage and ignores storage past them (`H5D__contig_check` checks only
/// that the elements fit in the file), so a larger storage reads; one too
/// small to hold the elements is an error.
pub fn contiguous_read_len(storage_size: u64, needed: usize) -> Result<usize, FormatError> {
if storage_size < needed as u64 {
return Err(FormatError::DataSizeMismatch {
expected: needed,
actual: usize::try_from(storage_size).unwrap_or(usize::MAX),
});
}
Ok(needed)
}
/// Zero-copy read of contiguous raw data, returning a borrowed slice.
///
/// For contiguous layouts, returns a direct `&[u8]` slice into `file_data`.
@@ -55,13 +129,7 @@ pub fn read_raw_data_zerocopy<'a>(
DataLayout::Contiguous { address, size } => {
let addr = address.ok_or(FormatError::NoDataAllocated)?;
let addr = addr as usize;
let sz = *size as usize;
if sz != expected_size {
return Err(FormatError::DataSizeMismatch {
expected: expected_size,
actual: sz,
});
}
let sz = contiguous_read_len(*size, expected_size)?;
ensure_len(file_data, addr, sz)?;
Ok(Some(&file_data[addr..addr + sz]))
}
@@ -172,13 +240,7 @@ fn read_raw_data_full_impl(
DataLayout::Contiguous { address, size } => {
let addr = address.ok_or(FormatError::NoDataAllocated)?;
let addr = addr as usize;
let sz = *size as usize;
if sz != expected_size {
return Err(FormatError::DataSizeMismatch {
expected: expected_size,
actual: sz,
});
}
let sz = contiguous_read_len(*size, expected_size)?;
ensure_len(file_data, addr, sz)?;
let mut out = crate::bulk_alloc::vec_for_bulk(sz);
out.extend_from_slice(&file_data[addr..addr + sz]);
@@ -2746,6 +2808,71 @@ mod tests {
assert_eq!(result.unwrap(), &[1.5f32, 2.5, 3.5]);
}
/// libhdf5 reads a contiguous dataset's elements from the start of its
/// storage and ignores storage past them (cve-2024-32623's scalar
/// `/Dset1` has 240 bytes of storage for one 4-byte element). Storage too
/// small for the elements is still an error.
#[test]
fn contiguous_storage_larger_than_the_elements_reads() {
let dt = make_f64_le_type();
let ds = make_simple_dataspace(&[2]);
let mut file_data = vec![0u8; 64];
file_data[..8].copy_from_slice(&1.5f64.to_le_bytes());
file_data[8..16].copy_from_slice(&2.5f64.to_le_bytes());
file_data[16..24].copy_from_slice(&9.0f64.to_le_bytes());
let layout = DataLayout::Contiguous {
address: Some(0),
size: 40,
};
let raw = read_raw_data(&file_data, &layout, &ds, &dt).unwrap();
assert_eq!(raw, file_data[..16]);
let zc = read_raw_data_zerocopy(&file_data, &layout, &ds, &dt).unwrap();
assert_eq!(zc, Some(&file_data[..16]));
let small = DataLayout::Contiguous {
address: Some(0),
size: 8,
};
assert!(matches!(
read_raw_data(&file_data, &small, &ds, &dt),
Err(FormatError::DataSizeMismatch { .. })
));
}
/// `H5D__contig_check` / `H5D__compact_init`, run when a dataset opens.
#[test]
fn dataset_storage_checks_at_open() {
let dt = make_f64_le_type();
let contiguous = |address| DataLayout::Contiguous { address, size: 0 };
// cve-2024-32624 `/Dset_OBJREF`: 2^62 + 2 elements of 8 bytes.
let huge = make_simple_dataspace(&[(1 << 62) + 2]);
assert_eq!(
check_dataset_storage(&contiguous(None), &huge, &dt, 1 << 20),
Err(FormatError::InvalidDatasetStorage(
"size of dataset's storage overflowed"
))
);
let ds = make_simple_dataspace(&[4]);
assert!(check_dataset_storage(&contiguous(Some(100)), &ds, &dt, 132).is_ok());
assert!(matches!(
check_dataset_storage(&contiguous(Some(100)), &ds, &dt, 131),
Err(FormatError::InvalidDatasetStorage(_))
));
assert!(matches!(
check_dataset_storage(&contiguous(Some(u64::MAX - 8)), &ds, &dt, u64::MAX),
Err(FormatError::InvalidDatasetStorage(_))
));
// Not allocated, and (unlike libhdf5) empty at a defined address.
assert!(check_dataset_storage(&contiguous(None), &ds, &dt, 0).is_ok());
let empty = make_simple_dataspace(&[0]);
assert!(check_dataset_storage(&contiguous(Some(64)), &empty, &dt, 64).is_ok());
let compact = |n: usize| DataLayout::Compact { data: vec![0; n] };
assert!(check_dataset_storage(&compact(32), &ds, &dt, 0).is_ok());
assert!(matches!(
check_dataset_storage(&compact(24), &ds, &dt, 0),
Err(FormatError::InvalidDatasetStorage(_))
));
}
#[test]
fn zerocopy_size_mismatch() {
let dt = make_f64_le_type();
+67 -13
View File
@@ -7,6 +7,9 @@ use alloc::vec::Vec;
use crate::error::FormatError;
/// Most dimensions a dataspace can have (`H5S_MAX_RANK`).
pub const MAX_RANK: u8 = 32;
/// Type of dataspace.
#[derive(Debug, Clone, PartialEq)]
pub enum DataspaceType {
@@ -67,6 +70,12 @@ impl Dataspace {
let version = data[0];
let rank = data[1];
let flags = data[2];
// H5O__sdspace_decode's checks.
if rank > MAX_RANK {
return Err(FormatError::InvalidDataspace(
"simple dataspace dimensionality is too large",
));
}
let (space_type, header_size) = match version {
1 => {
@@ -88,6 +97,11 @@ impl Dataspace {
2 => DataspaceType::Null,
_ => return Err(FormatError::InvalidDataspaceType(type_byte)),
};
if st != DataspaceType::Simple && rank > 0 {
return Err(FormatError::InvalidDataspace(
"invalid rank for scalar or NULL dataspace",
));
}
(st, 4usize)
}
_ => return Err(FormatError::InvalidDataspaceVersion(version)),
@@ -107,8 +121,13 @@ impl Dataspace {
// Read max dimensions if flags bit 0 is set
let max_dimensions = if flags & 0x01 != 0 {
let mut max_dims = Vec::with_capacity(rank as usize);
for _ in 0..rank {
for &dim in &dimensions {
let val = read_length(data, pos, length_size)?;
if dim > val {
return Err(FormatError::InvalidDataspace(
"dataspace dimension size is greater than its maximum size",
));
}
max_dims.push(val);
pos += ls;
}
@@ -176,7 +195,6 @@ impl Dataspace {
match self.space_type {
DataspaceType::Null => Ok(0),
DataspaceType::Scalar => Ok(1),
DataspaceType::Simple if self.dimensions.is_empty() => Ok(0),
DataspaceType::Simple => self
.dimensions
.iter()
@@ -195,18 +213,14 @@ impl Dataspace {
match self.space_type {
DataspaceType::Null => 0,
DataspaceType::Scalar => 1,
DataspaceType::Simple => {
if self.dimensions.is_empty() {
0
} else {
// Saturate rather than wrap: a wrapped product could
// under-size a buffer. Size-critical callers use
// `checked_num_elements`.
self.dimensions
// A simple dataspace of rank 0 holds one element, as in libhdf5
// (the product of no dimensions). Saturate rather than wrap: a
// wrapped product could under-size a buffer. Size-critical
// callers use `checked_num_elements`.
DataspaceType::Simple => self
.dimensions
.iter()
.fold(1u64, |acc, &d| acc.saturating_mul(d))
}
}
.fold(1u64, |acc, &d| acc.saturating_mul(d)),
}
}
}
@@ -352,4 +366,44 @@ mod tests {
let ds = Dataspace::parse(&data, 8).unwrap();
assert_eq!(ds.max_dimensions, Some(vec![10]));
}
/// A simple dataspace of rank 0 (cve-2020-18494's `/dset1`) holds one
/// element in libhdf5, which h5py reads as shape `()`. It was 0.
#[test]
fn simple_rank_zero_holds_one_element() {
let data = build_v2_dataspace(0, 0, 1, &[], None);
let ds = Dataspace::parse(&data, 8).unwrap();
assert_eq!(ds.space_type, DataspaceType::Simple);
assert_eq!(ds.num_elements(), 1);
assert_eq!(ds.checked_num_elements().unwrap(), 1);
}
/// `H5O__sdspace_decode`'s checks.
#[test]
fn refuses_what_libhdf5_refuses() {
let too_many = build_v2_dataspace(33, 0, 1, &[1; 33], None);
assert!(matches!(
Dataspace::parse(&too_many, 8),
Err(FormatError::InvalidDataspace(_))
));
let scalar_with_rank = build_v2_dataspace(1, 0, 0, &[4], None);
assert!(matches!(
Dataspace::parse(&scalar_with_rank, 8),
Err(FormatError::InvalidDataspace(_))
));
let null_with_rank = build_v2_dataspace(1, 0, 2, &[4], None);
assert!(matches!(
Dataspace::parse(&null_with_rank, 8),
Err(FormatError::InvalidDataspace(_))
));
let over_max = build_v1_dataspace(2, 0x01, &[5, 20], Some(&[10, 10]));
assert!(matches!(
Dataspace::parse(&over_max, 8),
Err(FormatError::InvalidDataspace(_))
));
// 32 dimensions, and a size equal to the maximum or unlimited, are fine.
assert!(Dataspace::parse(&build_v2_dataspace(32, 0, 1, &[1; 32], None), 8).is_ok());
let at_max = build_v1_dataspace(2, 0x01, &[10, 20], Some(&[10, u64::MAX]));
assert!(Dataspace::parse(&at_max, 8).is_ok());
}
}
+35
View File
@@ -223,6 +223,26 @@ pub enum FormatError {
/// The file's actual length in bytes.
actual_len: u64,
},
/// A link libhdf5 refuses to list: a symbol-table entry with an empty
/// name ("invalid link name"). Listing the group fails, as in libhdf5.
InvalidLinkName,
/// A dataspace message libhdf5 refuses to decode (the reason is
/// libhdf5's own error text): more than 32 dimensions, a rank on a
/// scalar or null dataspace, a dimension larger than its maximum.
InvalidDataspace(&'static str),
/// A dataset whose storage libhdf5 refuses when it opens the dataset
/// (the reason is libhdf5's own error text): an element count times
/// element size that overflows, contiguous storage past the end of the
/// file, compact data of the wrong size.
InvalidDatasetStorage(&'static str),
/// A superblock extension message libhdf5 refuses to decode when it
/// opens the file (the reason is libhdf5's own error text): a File Space
/// Info message that runs off its end or has a bad page size, a metadata
/// cache image outside the file, …
InvalidSuperblockExtension(&'static str),
/// A metadata cache image block libhdf5 refuses to load (the reason is
/// libhdf5's own error text).
InvalidCacheImage(&'static str),
}
impl fmt::Display for FormatError {
@@ -494,6 +514,21 @@ impl fmt::Display for FormatError {
but the file is {actual_len} bytes"
)
}
FormatError::InvalidLinkName => {
write!(f, "invalid link name: a group entry has an empty name")
}
FormatError::InvalidDataspace(why) => {
write!(f, "invalid dataspace: {why}")
}
FormatError::InvalidDatasetStorage(why) => {
write!(f, "invalid dataset storage: {why}")
}
FormatError::InvalidSuperblockExtension(why) => {
write!(f, "invalid superblock extension: {why}")
}
FormatError::InvalidCacheImage(why) => {
write!(f, "invalid metadata cache image: {why}")
}
}
}
}
+235 -139
View File
@@ -360,7 +360,7 @@ pub(crate) static BUILTIN_FILTERS: &[BuiltinFilter] = &[
BuiltinFilter {
id: FILTER_SHUFFLE,
name: "shuffle",
decode: |d, c| shuffle_decompress(d, c.element_size),
decode: |d, c| shuffle_decompress(d, shuffle_type_size(c.client_data(), c.element_size)?),
encode: Some(|d, c| shuffle_compress(d, c.element_size)),
},
BuiltinFilter {
@@ -464,23 +464,6 @@ pub(crate) static BUILTIN_FILTERS: &[BuiltinFilter] = &[
},
];
/// Decode the HDF5 scale-offset filter (id 6).
///
/// Supports all three scale-offset variants:
/// - `H5Z_SO_FLOAT_DSCALE` (0): `value = minval + code / 10^D`
/// - `H5Z_SO_FLOAT_ESCALE` (1): `value = minval + code * 2^E`
/// - `H5Z_SO_INT` (2): `value = minval + code`
///
/// Compressed buffer layout: `minbits` (u32 LE) · `minval_width` (1 byte)
/// · `minval` (`minval_width` bytes) · 8 reserved bytes · MSB-first packed
/// codes (`nelmts * minbits` bits). The all-ones code is reserved for the
/// defined fill value.
///
/// `cd` is the `H5Zscaleoffset.c` parameter block: `[0]`=scale type,
/// `[1]`=scale factor (decimal digits D for D-scale, binary exponent E for
/// E-scale, interpreted as i32 for negative exponents), `[2]`=element count,
/// `[4]`=element size, `[5]`=signed flag, `[6]`=byte order (1 = big-endian),
/// `[7]`=fill defined, `[8..]`=fill value bits.
/// `f64::powi` equivalent that works under `no_std` (no libm/std available).
/// Exponentiation by squaring, matching `powi`'s semantics for negative
/// exponents via reciprocal.
@@ -502,6 +485,32 @@ fn powi_f64(base: f64, mut exp: i32) -> f64 {
if neg { 1.0 / result } else { result }
}
/// Decode the HDF5 scale-offset filter (id 6) as libhdf5 does
/// (`H5Z__filter_scaleoffset`, reverse direction).
///
/// - Integers (`H5Z_SO_INT`): `value = minval + code`.
/// - Floats, D-scale (`H5Z_SO_FLOAT_DSCALE`): `value = code / 10^D + min`.
/// - Floats, E-scale: refused, as libhdf5 refuses it ("E-scaling method not
/// supported"); no library writes it.
///
/// Compressed buffer layout: `minbits` (u32 LE) · the size of `minval` in
/// bytes (1 byte; libhdf5 uses at most 8 of them) · `minval` · packed codes
/// at byte 21, whatever the stored size of `minval` (`buf_offset` is fixed)
/// · MSB-first, `minbits` bits per element. With a fill value defined, the
/// all-ones code of `minbits` bits is the fill value — for `minbits == 0`
/// that is every element. `minbits` equal to the element's full width means
/// the elements are stored as they are (in little-endian order), and an
/// integer scale factor of the full width means the filter left the chunk
/// untouched.
///
/// `cd` is the `H5Zscaleoffset.c` parameter block: `[0]` scale type, `[1]`
/// scale factor, `[2]` element count, `[3]` class (0 integer, 1 float),
/// `[4]` element size, `[5]` signed, `[6]` byte order (1 = big-endian), `[7]`
/// fill defined, `[8..]` fill value bits.
///
/// Packed data too short for its codes is an error (libhdf5 2.0 read past
/// the end of the chunk buffer — `cve-2025-2308` — and later releases
/// refuse it: "Buffer too short").
fn scaleoffset_decompress(
data: &[u8],
cd: &[u32],
@@ -510,74 +519,101 @@ fn scaleoffset_decompress(
const H5Z_SO_FLOAT_DSCALE: u32 = 0;
const H5Z_SO_FLOAT_ESCALE: u32 = 1;
const H5Z_SO_INT: u32 = 2;
/// Where the packed codes start (`buf_offset` in `H5Zscaleoffset.c`).
const BUF_OFFSET: usize = 21;
let err = |why: &str| FormatError::ChunkedReadError(format!("scale-offset: {why}"));
if cd.len() < 8 {
return Err(FormatError::ChunkedReadError(
"scale-offset: missing filter client data".into(),
));
return Err(err("missing filter client data"));
}
let scale_type = cd[0];
let is_float = scale_type == H5Z_SO_FLOAT_DSCALE || scale_type == H5Z_SO_FLOAT_ESCALE;
if scale_type != H5Z_SO_INT && !is_float {
return Err(FormatError::UnsupportedFilter(FILTER_SCALEOFFSET));
let is_float = match cd[3] {
0 => false,
1 => true,
_ => return Err(err("cannot use C integer datatype for cast")),
};
if is_float && scale_type != H5Z_SO_FLOAT_DSCALE && scale_type != H5Z_SO_FLOAT_ESCALE
|| !is_float && scale_type != H5Z_SO_INT
{
return Err(err("invalid scale type"));
}
if scale_type == H5Z_SO_FLOAT_ESCALE {
return Err(err("E-scaling method not supported"));
}
let nelmts = cd[2] as usize;
let elem_size = cd[4] as usize;
if elem_size == 0 || elem_size > 8 || (is_float && elem_size != 4 && elem_size != 8) {
return Err(FormatError::ChunkedReadError(
"scale-offset: unsupported element size".into(),
));
let size_ok = if is_float {
matches!(elem_size, 4 | 8)
} else {
matches!(elem_size, 1 | 2 | 4 | 8)
};
if !size_ok {
return Err(err("cannot use C integer datatype for cast"));
}
let full_bits = elem_size * 8;
// An integer's scale factor is the number of bits kept; all of them
// means the filter stored the chunk as it was.
if !is_float && (cd[1] as i32).max(0) as usize > full_bits {
return Err(err("minimum number of bits exceeds maximum"));
}
if !is_float && cd[1] as i32 == full_bits as i32 {
return Ok(data.to_vec());
}
// The decoded output must match the chunk's uncompressed size; reject an
// element count that would over-allocate (e.g. minbits == 0 with a huge
// nelmts and no packed payload to bound it).
let out_bytes = nelmts
.checked_mul(elem_size)
.ok_or_else(|| FormatError::ChunkedReadError("scale-offset: size overflow".into()))?;
.ok_or_else(|| err("size overflow"))?;
if expected_bytes != 0 && out_bytes > expected_bytes {
return Err(FormatError::ChunkedReadError(
"scale-offset: element count exceeds chunk size".into(),
));
return Err(err("element count exceeds chunk size"));
}
let signed = cd[5] == 1;
let big_endian = cd[6] == 1;
let fill_defined = cd[7] == 1;
// --- header: minbits, then minval, then 8 reserved bytes ---
// --- header: minbits, then the size of minval and minval ---
if data.len() < 5 {
return Err(FormatError::ChunkedReadError(
"scale-offset: truncated header".into(),
));
return Err(err("buffer too short"));
}
let minbits = u32::from_le_bytes([data[0], data[1], data[2], data[3]]) as usize;
let minval_width = data[4] as usize;
let minval_end = 5 + minval_width;
if data.len() < minval_end {
return Err(FormatError::ChunkedReadError(
"scale-offset: truncated minval".into(),
));
if minbits > full_bits {
return Err(err("minimum number of bits exceeds size of type"));
}
let minval_bytes = &data[5..minval_end];
let minval_size = usize::from(data[4]).min(8);
let minval_bytes = data
.get(5..5 + minval_size)
.ok_or_else(|| err("buffer too short"))?;
let minval = minval_bytes
.iter()
.rev()
.fold(0u64, |acc, &b| (acc << 8) | u64::from(b));
// --- unpack the per-element codes (MSB-first), shared by both variants ---
if minbits > 64 {
return Err(FormatError::ChunkedReadError(
"scale-offset: implausible minbits".into(),
));
// Full precision: the elements follow as they were, little-endian.
if minbits == full_bits {
let raw = data
.get(BUF_OFFSET..)
.and_then(|d| d.get(..out_bytes))
.ok_or_else(|| err("buffer too short"))?;
let mut out = raw.to_vec();
if big_endian {
for e in out.chunks_exact_mut(elem_size) {
e.reverse();
}
}
return Ok(out);
}
// --- unpack the per-element codes (MSB-first) ---
let codes: Vec<u64> = if minbits == 0 {
// No packed payload: every element equals minval.
// No packed payload: every code is 0.
vec![0u64; nelmts]
} else {
let packed = data.get(minval_end + 8..).ok_or_else(|| {
FormatError::ChunkedReadError("scale-offset: truncated packed data".into())
})?;
let packed = data.get(BUF_OFFSET..).unwrap_or(&[]);
let need_bits = nelmts
.checked_mul(minbits)
.ok_or_else(|| FormatError::ChunkedReadError("scale-offset: size overflow".into()))?;
.ok_or_else(|| err("size overflow"))?;
if packed.len() * 8 < need_bits {
return Err(FormatError::ChunkedReadError(
"scale-offset: packed data too short".into(),
));
return Err(err("packed data too short"));
}
let mut out = Vec::with_capacity(nelmts);
let mut bitpos = 0usize;
@@ -592,35 +628,28 @@ fn scaleoffset_decompress(
}
out
};
// The fill code (all ones) only exists when there are bits to pack.
let has_fill_code = fill_defined && minbits > 0 && minbits < 64;
// Computed for all 1..=64 widths; `1 << 64` would overflow, so saturate.
let fill_code: u64 = if minbits == 0 {
0
} else if minbits >= 64 {
u64::MAX
} else {
(1u64 << minbits) - 1
// With a fill value defined, the all-ones code of `minbits` bits (0
// when minbits is 0) stands for it. minbits < 64 here.
let fill_code: u64 = (1u64 << minbits) - 1;
let fill_bits = || {
let lo = u64::from(*cd.get(8).unwrap_or(&0));
let hi = u64::from(*cd.get(9).unwrap_or(&0));
lo | (hi << 32)
};
if is_float {
let is_escale = scale_type == H5Z_SO_FLOAT_ESCALE;
let scale_factor = cd[1] as i32;
let minval = read_le_float(minval_bytes, elem_size);
let minval = bits_to_float(minval, elem_size);
let fill_value = if fill_defined {
let lo = *cd.get(8).unwrap_or(&0) as u64;
let hi = *cd.get(9).unwrap_or(&0) as u64;
bits_to_float(lo | (hi << 32), elem_size)
bits_to_float(fill_bits(), elem_size)
} else {
0.0
};
let values: Vec<f64> = codes
.iter()
.map(|&code| {
if has_fill_code && code == fill_code {
if fill_defined && code == fill_code {
fill_value
} else if is_escale {
minval + code as f64 * powi_f64(2.0, scale_factor)
} else if elem_size == 4 {
// H5Z_scaleoffset_modify_3/4 for `float`: the code is
// read as an `int` and everything is single precision,
@@ -640,21 +669,19 @@ fn scaleoffset_decompress(
.collect();
Ok(write_floats(&values, elem_size, big_endian))
} else {
let minval = read_le_int(minval_bytes, signed);
let fill_value: i64 = if fill_defined {
let lo = *cd.get(8).unwrap_or(&0) as u64;
let hi = *cd.get(9).unwrap_or(&0) as u64;
sign_extend(lo | (hi << 32), elem_size, signed)
sign_extend(fill_bits(), elem_size, signed)
} else {
0
};
let values: Vec<i64> = codes
.iter()
.map(|&code| {
if has_fill_code && code == fill_code {
if fill_defined && code == fill_code {
fill_value
} else {
minval.wrapping_add(code as i64)
// `(type)(buf[i] + minval)`: wraps at the element width.
(code.wrapping_add(minval)) as i64
}
})
.collect();
@@ -662,21 +689,6 @@ fn scaleoffset_decompress(
}
}
/// Read a little-endian float of `size` bytes (4 = f32, otherwise f64) as f64.
fn read_le_float(bytes: &[u8], size: usize) -> f64 {
if size == 4 {
let mut b = [0u8; 4];
let n = bytes.len().min(4);
b[..n].copy_from_slice(&bytes[..n]);
f32::from_le_bytes(b) as f64
} else {
let mut b = [0u8; 8];
let n = bytes.len().min(8);
b[..n].copy_from_slice(&bytes[..n]);
f64::from_le_bytes(b)
}
}
/// Interpret the low bits of `raw` as an IEEE float of `size` bytes.
fn bits_to_float(raw: u64, size: usize) -> f64 {
if size == 4 {
@@ -709,16 +721,6 @@ fn write_floats(values: &[f64], elem_size: usize, big_endian: bool) -> Vec<u8> {
out
}
/// Read a little-endian integer of `bytes.len()` bytes, sign-extending when
/// `signed`. Used for the scale-offset `minval` field.
fn read_le_int(bytes: &[u8], signed: bool) -> i64 {
let mut raw: u64 = 0;
for (i, &b) in bytes.iter().enumerate().take(8) {
raw |= (b as u64) << (i * 8);
}
sign_extend(raw, bytes.len().min(8), signed)
}
/// Interpret the low `size` bytes of `raw` as a (possibly signed) integer.
fn sign_extend(raw: u64, size: usize, signed: bool) -> i64 {
if size == 0 || size >= 8 {
@@ -1422,6 +1424,23 @@ fn zstd_compress(data: &[u8], level: u32) -> Result<Vec<u8>, FormatError> {
/// Unshuffle (decompress direction): reconstruct interleaved element bytes.
/// On disk: all byte-0s of each element together, then all byte-1s, etc.
/// Output: elements in natural order.
/// The element size the shuffle filter works with: its parameter, as
/// libhdf5 uses it (`H5Z__filter_shuffle`), not the dataset's element size.
/// They are the same in every file a library wrote; a corrupt parameter
/// larger than the chunk makes libhdf5 leave the chunk as it is, and so
/// does [`shuffle_decompress`] (`cve-2025-44905`'s `Shuffle_float_data_be`).
/// A zero parameter is an error ("invalid shuffle parameters"); a pipeline
/// without the parameter (never written by libhdf5) uses the element size.
fn shuffle_type_size(cd: &[u32], element_size: usize) -> Result<usize, FormatError> {
match cd {
[] => Ok(element_size),
[0] | [_, _, ..] => Err(FormatError::FilterError(
"invalid shuffle parameters".into(),
)),
[size] => Ok(*size as usize),
}
}
fn shuffle_decompress(data: &[u8], element_size: usize) -> Result<Vec<u8>, FormatError> {
let mut result = Vec::new();
shuffle_decompress_into(data, element_size, &mut result);
@@ -2765,48 +2784,125 @@ mod tests {
}
}
fn as_f64(bytes: &[u8]) -> Vec<f64> {
bytes
.as_chunks::<8>()
.0
.iter()
.map(|c| f64::from_le_bytes(*c))
.collect()
/// E-scale: libhdf5 refuses it on read and write ("E-scaling method not
/// supported"); it was decoded here, never checked against anything.
#[test]
fn scaleoffset_float_escale_is_refused_as_in_libhdf5() {
let cd = [1u32, 1, 4, 1, 8, 0, 0, 0];
let mut raw = vec![2, 0, 0, 0, 8];
raw.extend_from_slice(&[0; 16]);
raw.push(0x1B);
assert!(scaleoffset_decompress(&raw, &cd, 0).is_err());
}
/// A scale type that does not match the class is refused, as libhdf5
/// refuses it ("invalid scale type").
#[test]
fn scaleoffset_float_escale_e1() {
// f64 [0.0, 2.0, 4.0, 6.0], E=1 (×2^1=2), fill_defined=0.
// cd: scale_type=1, E=1, nelmts=4, elem_size=8.
let cd = [1u32, 1, 4, 0, 8, 0, 0, 0];
let raw: &[u8] = &[
2, 0, 0, 0, // minbits=2
8, // minval_width=8
0, 0, 0, 0, 0, 0, 0, 0, // minval=0.0f64
0, 0, 0, 0, 0, 0, 0, 0, // 8 reserved bytes
0x1B, // packed codes: 00 01 10 11 MSB-first
];
let got = as_f64(&scaleoffset_decompress(raw, &cd, 0).unwrap());
assert_eq!(got, vec![0.0, 2.0, 4.0, 6.0]);
fn scaleoffset_scale_type_must_match_the_class() {
let mut raw = vec![2, 0, 0, 0, 8];
raw.extend_from_slice(&[0; 16]);
raw.push(0x1B);
assert!(scaleoffset_decompress(&raw, &[0, 0, 4, 0, 4, 1, 0, 0], 0).is_err());
assert!(scaleoffset_decompress(&raw, &[2, 0, 4, 1, 4, 1, 0, 0], 0).is_err());
assert!(scaleoffset_decompress(&raw, &[2, 0, 4, 7, 4, 1, 0, 0], 0).is_err());
}
/// `cve-2025-44905` `/Scale_offset_short_data_be`, chunk (4, 0): the
/// stored size of `minval` is 0. libhdf5 reads a `minval` of 0 and the
/// packed codes from byte 21 regardless; we read them from byte 13
/// (5 + size + 8), so the values differed from h5py's.
#[test]
fn scaleoffset_float_escale_neg_exp() {
// f64 [0.0, 0.5, 1.0, 1.5], E=-1 (×2^-1=0.5), fill_defined=0.
// cd[1] = 0xFFFF_FFFF which casts to i32 = -1.
let cd = [1u32, 0xFFFF_FFFF, 4, 0, 8, 0, 0, 0];
let raw: &[u8] = &[
2, 0, 0, 0, // minbits=2
8, // minval_width=8
0, 0, 0, 0, 0, 0, 0, 0, // minval=0.0f64
0, 0, 0, 0, 0, 0, 0, 0, // 8 reserved bytes
0x1B, // packed codes: 00 01 10 11 MSB-first
];
let got = as_f64(&scaleoffset_decompress(raw, &cd, 0).unwrap());
let exp = [0.0f64, 0.5, 1.0, 1.5];
for (g, e) in got.iter().zip(exp.iter()) {
assert!((g - e).abs() < 1e-9, "got {g} expected {e}");
fn scaleoffset_codes_start_at_byte_21_whatever_the_minval_size() {
// big-endian i16, 12 elements, fill -2 (cd 65534), minbits 3.
let mut cd = vec![2u32, 0, 12, 0, 2, 1, 1, 1, 65534];
cd.resize(20, 0);
let raw = unhex("0300000000d20e00000000000034000000000000000400000000");
let got = scaleoffset_decompress(&raw, &cd, 24).unwrap();
// Codes of 3 bits from byte 21 (04 00 00 00 00): 0, 1, 0, ...; h5py
// reads the chunk's first row as 0, 1, 0.
let want: Vec<i16> = vec![0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0];
let want: Vec<u8> = want.iter().flat_map(|v| v.to_be_bytes()).collect();
assert_eq!(got, want);
}
/// With a fill value defined, libhdf5 compares each code with the
/// all-ones code of `minbits` bits — which for `minbits == 0` is 0, so a
/// chunk with no packed codes reads as all fill values (the compressor
/// writes that for a chunk of nothing but fill values). It read as
/// `minval` here.
#[test]
fn scaleoffset_minbits_zero_with_a_fill_value_is_all_fill() {
let mut cd = vec![2u32, 0, 3, 0, 4, 1, 0, 1, (-7i32) as u32];
cd.resize(20, 0);
let mut raw = vec![0, 0, 0, 0, 8];
raw.extend_from_slice(&5i64.to_le_bytes());
raw.extend_from_slice(&[0; 8]);
assert_eq!(
scaleoffset_decompress(&raw, &cd, 12).unwrap(),
i32_le(&[-7, -7, -7])
);
// Without a fill value every element is minval.
cd[7] = 0;
assert_eq!(
scaleoffset_decompress(&raw, &cd, 12).unwrap(),
i32_le(&[5, 5, 5])
);
}
/// `minbits` of the full width stores the elements as they are
/// (little-endian), without `minval`; a full-width integer scale factor
/// means the filter left the chunk untouched.
#[test]
fn scaleoffset_full_width_is_stored_as_is() {
let mut cd = vec![2u32, 0, 2, 0, 2, 1, 1, 0];
cd.resize(20, 0);
let mut raw = vec![16, 0, 0, 0, 8];
raw.extend_from_slice(&100i64.to_le_bytes());
raw.extend_from_slice(&[0; 8]);
raw.extend_from_slice(&[0x34, 0x12, 0xfe, 0xff]);
assert_eq!(
scaleoffset_decompress(&raw, &cd, 4).unwrap(),
[0x12, 0x34, 0xff, 0xfe]
);
cd[1] = 16;
assert_eq!(
scaleoffset_decompress(&[1, 2, 3, 4], &cd, 4).unwrap(),
[1, 2, 3, 4]
);
cd[1] = 17;
assert!(scaleoffset_decompress(&[1, 2, 3, 4], &cd, 4).is_err());
// minbits wider than the type.
cd[1] = 0;
raw[0] = 17;
assert!(scaleoffset_decompress(&raw, &cd, 4).is_err());
}
/// The shuffle filter uses its own parameter as the element size, as
/// libhdf5 does; a parameter larger than the chunk leaves the chunk as
/// it is (`cve-2025-44905` `/Shuffle_float_data_be`, whose parameter is
/// 4261347332: h5py and h5dump read the stored bytes unshuffled).
#[test]
fn shuffle_uses_its_parameter() {
let data: Vec<u8> = (0..16).collect();
let shuffled = shuffle_compress(&data, 4).unwrap();
let pipeline = |cd: Vec<u32>| FilterPipeline {
version: 2,
filters: vec![one_filter(FILTER_SHUFFLE, cd)],
};
// The dataset's element size says 2; the parameter says 4.
assert_eq!(
decompress_chunk(&shuffled, &pipeline(vec![4]), 16, 2).unwrap(),
data
);
assert_eq!(
decompress_chunk(&shuffled, &pipeline(vec![4_261_347_332]), 16, 4).unwrap(),
shuffled
);
assert!(decompress_chunk(&shuffled, &pipeline(vec![0]), 16, 4).is_err());
assert_eq!(
decompress_chunk(&shuffled, &pipeline(vec![]), 16, 4).unwrap(),
data
);
}
// --- N-Bit (filter id 5) --------------------------------------------------
+41 -3
View File
@@ -21,12 +21,35 @@ pub struct GroupEntry {
pub cache_type: u32,
}
/// Given a SymbolTableMessage, resolve all group children.
/// Given a SymbolTableMessage, resolve all group children: the group's
/// listing.
///
/// An entry with an empty name fails the listing with
/// [`FormatError::InvalidLinkName`], as it fails libhdf5's link iteration
/// (`H5G__ent_to_link`: "invalid link name"). Looking a name up
/// ([`resolve_path`], and the path resolution in
/// [`crate::group_v2::resolve_path_any`]) still works in such a group, as it
/// does in libhdf5.
pub fn resolve_v1_group_entries(
file_data: &[u8],
sym_table_msg: &SymbolTableMessage,
offset_size: u8,
length_size: u8,
) -> Result<Vec<GroupEntry>, FormatError> {
let entries = v1_group_entries(file_data, sym_table_msg, offset_size, length_size)?;
if entries.iter().any(|e| e.name.is_empty()) {
return Err(FormatError::InvalidLinkName);
}
Ok(entries)
}
/// Every entry of a v1 group, empty names included — for looking a name up,
/// which never matches an empty name.
pub(crate) fn v1_group_entries(
file_data: &[u8],
sym_table_msg: &SymbolTableMessage,
offset_size: u8,
length_size: u8,
) -> Result<Vec<GroupEntry>, FormatError> {
// Parse local heap
let heap = LocalHeap::parse(
@@ -207,8 +230,7 @@ pub fn resolve_path(
let mut current_sym_table = root_sym_table.clone();
for (i, component) in components.iter().enumerate() {
let entries =
resolve_v1_group_entries(file_data, &current_sym_table, offset_size, length_size)?;
let entries = v1_group_entries(file_data, &current_sym_table, offset_size, length_size)?;
let found = entries.iter().find(|e| e.name == *component);
match found {
@@ -425,6 +447,22 @@ mod tests {
assert_eq!(entries[1].object_header_address, 0x2000);
}
/// cve-2021-46244 `/BAG_root`: a symbol-table entry with an empty name.
/// libhdf5 fails the group's listing ("invalid link name"); a lookup of
/// the other names still works.
#[test]
fn empty_entry_name_fails_the_listing_not_a_lookup() {
let (file, msg) = build_synthetic_group(&[("", 0x1000, 0), ("elevation", 0x2000, 0)], 8, 8);
assert_eq!(
resolve_v1_group_entries(&file, &msg, 8, 8).unwrap_err(),
FormatError::InvalidLinkName
);
assert_eq!(
resolve_path(&file, &msg, "elevation", 8, 8).unwrap(),
0x2000
);
}
#[test]
fn resolve_path_single_level() {
let (file, msg) =
+3 -1
View File
@@ -450,7 +450,9 @@ fn resolve_group_entries(
.find(|m| m.msg_type == MessageType::SymbolTable)
.ok_or_else(|| FormatError::PathNotFound(String::from("no symbol table message")))?;
let stm = SymbolTableMessage::parse(&sym_msg.data, offset_size)?;
group_v1::resolve_v1_group_entries(file_data, &stm, offset_size, length_size)
// A lookup: an entry with an empty name (which fails a listing) is
// skipped by the name comparison, as in libhdf5.
group_v1::v1_group_entries(file_data, &stm, offset_size, length_size)
} else if is_v2_group(object_header) {
resolve_v2_group_entries(file_data, object_header, offset_size, length_size)
} else {
+1
View File
@@ -121,6 +121,7 @@ pub mod selection;
pub mod shared_message;
pub mod signature;
pub mod superblock;
pub mod superblock_ext;
pub mod symbol_table;
#[cfg(all(
test,
@@ -75,7 +75,40 @@ fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
})
}
/// The kind of object an object header describes.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ObjectClass {
/// A group: the header has a Symbol Table or a Link Info message.
Group,
/// A dataset: the header has a Datatype and a Dataspace message.
Dataset,
/// A committed (named) datatype: a Datatype message, no Dataspace.
NamedDatatype,
}
impl ObjectHeader {
/// The kind of object this header describes, decided as libhdf5 decides
/// it (`H5O__obj_class_real`): group first (a Symbol Table or Link Info
/// message), then dataset (a Datatype *and* a Dataspace message — not a
/// Data Layout message), then named datatype (a Datatype message).
/// `None` when none applies; libhdf5 then cannot open the object
/// ("unable to determine object type").
///
/// A header with a Datatype and a Data Layout message but no Dataspace
/// is a named datatype to libhdf5, not a dataset.
pub fn object_class(&self) -> Option<ObjectClass> {
let has = |t: MessageType| self.messages.iter().any(|m| m.msg_type == t);
if has(MessageType::SymbolTable) || has(MessageType::LinkInfo) {
Some(ObjectClass::Group)
} else if has(MessageType::Datatype) && has(MessageType::Dataspace) {
Some(ObjectClass::Dataset)
} else if has(MessageType::Datatype) {
Some(ObjectClass::NamedDatatype)
} else {
None
}
}
/// Parse an object header at the given offset in the data buffer.
///
/// `offset_size` and `length_size` come from the superblock.
@@ -662,6 +695,54 @@ fn check_message(
mod tests {
use super::*;
fn header_with(types: &[MessageType]) -> ObjectHeader {
ObjectHeader {
version: 2,
messages: types
.iter()
.map(|&msg_type| HeaderMessage {
msg_type,
size: 0,
flags: 0,
creation_order: None,
data: Vec::new(),
})
.collect(),
reference_count: None,
flags: 0,
access_time: None,
modification_time: None,
change_time: None,
birth_time: None,
}
}
#[test]
fn object_class_follows_libhdf5() {
use MessageType::*;
let class = |t: &[MessageType]| header_with(t).object_class();
assert_eq!(
class(&[Datatype, Dataspace, DataLayout]),
Some(ObjectClass::Dataset)
);
// A Data Layout message does not make a dataset without a dataspace
// (cve-2024-33874 `/Dset1`: h5py opens it as a named datatype).
assert_eq!(
class(&[Datatype, DataLayout]),
Some(ObjectClass::NamedDatatype)
);
assert_eq!(class(&[Datatype]), Some(ObjectClass::NamedDatatype));
// Group messages win over dataset messages.
assert_eq!(
class(&[Datatype, Dataspace, SymbolTable]),
Some(ObjectClass::Group)
);
assert_eq!(class(&[LinkInfo]), Some(ObjectClass::Group));
// Link messages alone are not a group; nothing is not an object.
assert_eq!(class(&[Link]), None);
assert_eq!(class(&[]), None);
}
// Helper: build a v1 object header with given messages
fn build_v1_header(
messages: &[(u16, &[u8], u8)], // (type, data, flags)
+3 -2
View File
@@ -18,7 +18,7 @@ use alloc::{format, vec, vec::Vec};
#[cfg(feature = "std")]
use std::string as alloc_or_std;
use crate::chunked_read::{alloc_output, checked_byte_len, list_chunks};
use crate::chunked_read::{alloc_output, checked_byte_len, list_chunks_for_read};
use crate::data_layout::DataLayout;
use crate::data_read::extract_selection_from_buffer;
use crate::dataspace::Dataspace;
@@ -294,11 +294,12 @@ pub fn read_selection(
btree_address: Some(_),
..
} => {
let (chunks, chunk_dims) = list_chunks(
let (chunks, chunk_dims) = list_chunks_for_read(
file_data,
layout,
dataspace,
elem_size,
pipeline,
offset_size,
length_size,
)?;
@@ -0,0 +1,812 @@
//! The superblock extension of a version 2 or 3 superblock, and the
//! metadata cache image it can point to.
//!
//! libhdf5 reads the extension when it opens a file (`H5F__super_read`) and
//! decodes the messages that configure the file: v1 B-tree "K" values, File
//! Space Info, and the Metadata Cache Image. A message that does not decode
//! makes the file fail to open, so [`read_superblock_extension`] decodes and
//! checks them the way libhdf5 does.
//!
//! A metadata cache image (written with `H5Pset_mdc_image_config`) is a
//! block holding serialized metadata cache entries — object headers, B-tree
//! nodes, heaps — each with its file address. libhdf5 loads it into its
//! cache before it reads any other metadata (`H5C__load_cache_image`,
//! `H5C__reconstruct_cache_contents`), and the entries take the place of
//! the file's bytes at their addresses: the file itself may hold stale or
//! no metadata there (in `h5clear_mdc_image.h5` the root group's header is
//! only in the image). [`CacheImage::apply`] does the same with bytes: it
//! writes every entry at its address, so every parser reads what libhdf5
//! reads. It writes into whatever the opener gives it — a private
//! copy-on-write mapping of the file, or a buffer the opener owns — so the
//! file is never copied whole.
#[cfg(not(feature = "std"))]
use alloc::{collections::BTreeSet, vec::Vec};
#[cfg(feature = "std")]
use std::collections::BTreeSet;
use crate::error::FormatError;
use crate::message_type::MessageType;
use crate::object_header::ObjectHeader;
use crate::superblock::Superblock;
/// Message type of the File Space Info message.
const MSG_FSINFO: u16 = 0x0017;
/// Message type of the Metadata Cache Image message.
const MSG_MDCI: u16 = 0x0018;
/// Header message flag: the library did not know the message when it wrote
/// it back (`H5O_MSG_FLAG_WAS_UNKNOWN`); libhdf5 then ignores its contents.
const MSG_FLAG_WAS_UNKNOWN: u8 = 0x20;
/// `H5F_FILE_SPACE_PAGE_SIZE_MIN` / `_MAX`.
const PAGE_SIZE_MIN: u64 = 512;
const PAGE_SIZE_MAX: u64 = 1024 * 1024 * 1024;
/// libhdf5's default file space page size, used for a version 0 message.
const PAGE_SIZE_DEFAULT: u64 = 4096;
/// Free-space managers whose addresses a persisting version 1 File Space
/// Info message lists (`H5F_MEM_PAGE_SUPER` .. `H5F_MEM_PAGE_NTYPES`), and
/// a version 0 one (`H5FD_MEM_SUPER` .. `H5FD_MEM_NTYPES`).
const FSM_ADDRS_V1: usize = 12;
const FSM_ADDRS_V0: usize = 6;
/// Metadata cache image block limits (`H5Cimage.c`, `H5ACprivate.h`).
const MDCI_SIGNATURE: &[u8; 4] = b"MDCI";
const MDCI_HAVE_RESIZE_STATUS: u8 = 0x01;
const MDCI_ENTRY_IS_FD_PARENT: u8 = 0x04;
const MDCI_ENTRY_IS_FD_CHILD: u8 = 0x08;
/// `H5AC_NTYPES`: entry type ids are below this.
const MDCI_NTYPES: u8 = 30;
/// `H5C_RING_NTYPES`.
const MDCI_RING_NTYPES: u8 = 6;
/// `H5AC__CACHE_IMAGE__ENTRY_AGEOUT__MAX`.
const MDCI_AGE_MAX: u8 = 100;
/// A decoded File Space Info message (0x0017), mapped to version 1 as
/// libhdf5 maps a version 0 one.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct FileSpaceInfo {
/// Message version as stored (0 or 1).
pub version: u8,
/// File space strategy (`H5F_fspace_strategy_t`).
pub strategy: u8,
/// Whether free space is persisted.
pub persist: bool,
/// Free-space section threshold.
pub threshold: u64,
/// File space page size.
pub page_size: u64,
}
/// Where a metadata cache image block is (Metadata Cache Image message,
/// 0x0018).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct CacheImageLocation {
/// Address of the image block.
pub address: u64,
/// Length of the image block in bytes.
pub length: u64,
}
/// The messages of a superblock extension that libhdf5 decodes at open.
#[derive(Debug, Clone, Default, PartialEq, Eq)]
pub struct SuperblockExtension {
/// v1 B-tree "K" values (chunk index, symbol table node, symbol table
/// leaf), when the extension overrides the defaults.
pub btree_k: Option<(u16, u16, u16)>,
/// The File Space Info message.
pub file_space_info: Option<FileSpaceInfo>,
/// The metadata cache image, when the file has one.
pub cache_image: Option<CacheImageLocation>,
}
fn ext_err(why: &'static str) -> FormatError {
FormatError::InvalidSuperblockExtension(why)
}
const RAN_OFF: &str = "ran off end of input buffer while decoding";
/// A little-endian cursor over one message or block, failing with
/// `overrun` when it runs off the end.
struct Cursor<'a> {
data: &'a [u8],
pos: usize,
overrun: FormatError,
}
impl<'a> Cursor<'a> {
fn new(data: &'a [u8], overrun: FormatError) -> Self {
Cursor {
data,
pos: 0,
overrun,
}
}
fn take(&mut self, n: usize) -> Result<&'a [u8], FormatError> {
let end = self
.pos
.checked_add(n)
.filter(|&e| e <= self.data.len())
.ok_or_else(|| self.overrun.clone())?;
let s = &self.data[self.pos..end];
self.pos = end;
Ok(s)
}
fn u8(&mut self) -> Result<u8, FormatError> {
Ok(self.take(1)?[0])
}
fn uint(&mut self, width: u8) -> Result<u64, FormatError> {
let b = self.take(width as usize)?;
Ok(b.iter()
.rev()
.fold(0u64, |acc, &x| (acc << 8) | u64::from(x)))
}
/// An address of `width` bytes; `None` when undefined (all ones).
fn addr(&mut self, width: u8) -> Result<Option<u64>, FormatError> {
let v = self.uint(width)?;
let undef = if width >= 8 {
u64::MAX
} else {
(1u64 << (8 * u32::from(width))) - 1
};
Ok((v != undef).then_some(v))
}
}
/// Decode and check the superblock extension of `sb`, as libhdf5 does when
/// it opens the file. `data` is the file from the superblock on, up to the
/// end of file the superblock records (its end is libhdf5's "eoa").
///
/// Returns `Ok(None)` for a superblock without an extension (versions 0
/// and 1 have none). A message libhdf5 fails to decode, or a cache image
/// that does not lie inside the file, is an error: libhdf5 refuses to open
/// such a file (`cve-2020-10810`: a File Space Info message too short for
/// the free-space manager addresses it announces; `cve-2020-10812`: a cache
/// image past the end of the file).
pub fn read_superblock_extension(
data: &[u8],
sb: &Superblock,
) -> Result<Option<SuperblockExtension>, FormatError> {
let os = sb.offset_size;
let ls = sb.length_size;
let undef = if os >= 8 {
u64::MAX
} else {
(1u64 << (8 * u32::from(os))) - 1
};
let Some(addr) = sb.superblock_extension_address.filter(|&a| a != undef) else {
return Ok(None);
};
let addr = usize::try_from(addr).map_err(|_| ext_err("address out of range"))?;
let header = ObjectHeader::parse(data, addr, os, ls)?;
let eoa = data.len() as u64;
let mut ext = SuperblockExtension::default();
for msg in &header.messages {
match msg.msg_type {
MessageType::BTreeKValues => {
let mut c = Cursor::new(&msg.data, ext_err(RAN_OFF));
if c.u8()? != 0 {
return Err(ext_err("bad version number for v1 B-tree 'K' message"));
}
let chunk = c.uint(2)? as u16;
let snode = c.uint(2)? as u16;
let leaf = c.uint(2)? as u16;
ext.btree_k = Some((chunk, snode, leaf));
}
MessageType::Unknown(MSG_FSINFO) if msg.flags & MSG_FLAG_WAS_UNKNOWN == 0 => {
ext.file_space_info = Some(decode_fsinfo(&msg.data, os, ls)?);
}
MessageType::Unknown(MSG_MDCI) => {
let mut c = Cursor::new(&msg.data, ext_err(RAN_OFF));
if c.u8()? != 0 {
return Err(ext_err(
"bad version number for metadata cache image message",
));
}
let address = c.addr(os)?;
let length = c.uint(ls)?;
let Some(address) = address else {
return Err(ext_err("metadata cache image address is undefined"));
};
if address.checked_add(length).is_none_or(|end| end > eoa) {
return Err(ext_err(
"metadata cache image: address plus size exceeds file eoa",
));
}
ext.cache_image = Some(CacheImageLocation { address, length });
}
_ => {}
}
}
Ok(Some(ext))
}
/// `H5O__fsinfo_decode` plus the checks `H5F__super_read` makes on it.
fn decode_fsinfo(data: &[u8], os: u8, ls: u8) -> Result<FileSpaceInfo, FormatError> {
let mut c = Cursor::new(data, ext_err(RAN_OFF));
let version = c.u8()?;
let info = if version == 0 {
let old_strategy = c.u8()?;
let threshold = c.uint(ls)?;
// H5F_file_space_type_t: 1 ALL_PERSIST, 2 ALL, 3 AGGR_VFD, 4 VFD.
let (strategy, persist) = match old_strategy {
1 => {
for _ in 0..FSM_ADDRS_V0 {
c.addr(os)?;
}
(0, true)
}
2 => (0, false),
3 => (2, false),
4 => (3, false),
_ => return Err(ext_err("invalid file space strategy")),
};
FileSpaceInfo {
version,
strategy,
persist,
threshold,
page_size: PAGE_SIZE_DEFAULT,
}
} else {
if version > 1 {
return Err(ext_err("File space info message's version out of bounds"));
}
let strategy = c.u8()?;
let persist = c.u8()? != 0;
let threshold = c.uint(ls)?;
let page_size = c.uint(ls)?;
if page_size == 0 || page_size > PAGE_SIZE_MAX {
return Err(ext_err("invalid page size in file space info"));
}
c.uint(2)?; // page end metadata threshold
c.addr(os)?; // EOA before the free-space managers
if persist {
for _ in 0..FSM_ADDRS_V1 {
c.addr(os)?;
}
}
FileSpaceInfo {
version,
strategy,
persist,
threshold,
page_size,
}
};
if info.page_size < PAGE_SIZE_MIN {
return Err(ext_err("file space page size too small"));
}
Ok(info)
}
/// One entry of a metadata cache image: `len` bytes at `image_offset` in
/// the image block, belonging at file address `address`.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
struct ImageEntry {
address: u64,
image_offset: usize,
len: usize,
}
/// A decoded metadata cache image: where its block is, and the entries it
/// holds. [`CacheImage::apply`] writes the entries over a file's bytes.
///
/// Only the entry list is kept, never a copy of the file: an opener that
/// maps the file applies the image to a private copy-on-write mapping, so
/// only the pages the entries land on are copied.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct CacheImage {
location: CacheImageLocation,
entries: Vec<ImageEntry>,
}
/// What an opener must do about a file's metadata cache image.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum CacheImageState {
/// The file has no image: its bytes are its metadata.
Absent,
/// The file has an image that loads: apply it with [`CacheImage::apply`].
Loaded(CacheImage),
/// The file has an image libhdf5 fails to load. libhdf5 still opens the
/// file (the image loads at the first metadata read), and that read
/// fails with this error.
Unloadable(FormatError),
}
impl CacheImage {
/// Decode the metadata cache image at `location` in `data` (the file
/// from the superblock on, up to its recorded end of file). The image is
/// checked as libhdf5 checks it (`H5C__decode_cache_image_header`,
/// `H5C__reconstruct_cache_entry`): signature and version, the image
/// length it records, entry types, rings and ages in range, entry
/// addresses inside the file and not repeated, flush-dependency parents
/// already in the cache.
///
/// One check is stricter than libhdf5's: an entry must end inside the
/// file. libhdf5 checks only that it starts there, and serves the rest
/// from the image; the images libhdf5 writes never do this (every entry
/// lies below the image block, which is written last), and the bytes an
/// entry would put past the end of file have nowhere to go in a view of
/// the file.
///
/// libhdf5 does not verify the block's trailing checksum when it loads
/// an image, so neither does this.
pub fn decode(
data: &[u8],
location: CacheImageLocation,
sb: &Superblock,
) -> Result<Self, FormatError> {
let (offset_size, length_size) = (sb.offset_size, sb.length_size);
let bad = FormatError::InvalidCacheImage;
let block = image_block(data, location)?;
let eoa = data.len() as u64;
let mut c = Cursor::new(block, bad(RAN_OFF));
// Header: signature, version, flags, image data length, entry count.
if c.take(4)? != MDCI_SIGNATURE {
return Err(bad("bad metadata cache image header signature"));
}
if c.u8()? != 0 {
return Err(bad("bad metadata cache image version"));
}
if c.u8()? & MDCI_HAVE_RESIZE_STATUS != 0 {
return Err(bad("MDC resize status not yet supported"));
}
if c.uint(length_size)? != location.length {
return Err(bad("bad metadata cache image data length"));
}
let n_entries = c.uint(4)?;
if n_entries == 0 {
return Err(bad("bad metadata cache entry count"));
}
let mut entries = Vec::new();
// What is in libhdf5's cache when it loads the image: the superblock
// and the superblock extension's object header (read to find the
// image). Each entry's flush-dependency parents are looked up in the
// cache as the entry is inserted (`H5C__reconstruct_cache_contents`
// searches the index inside the loop that inserts the entries, in
// HDF5 1.14.6 and 2.0.0 alike), so a parent must be one of those or
// an earlier entry.
let mut cached = BTreeSet::new();
cached.insert(0);
if let Some(ext) = sb.superblock_extension_address {
cached.insert(ext);
}
let mut seen = BTreeSet::new();
for _ in 0..n_entries {
let type_id = c.u8()?;
if type_id >= MDCI_NTYPES {
return Err(bad("type id is out of valid range"));
}
let flags = c.u8()?;
if c.u8()? >= MDCI_RING_NTYPES {
return Err(bad("ring is out of valid range"));
}
if c.u8()? > MDCI_AGE_MAX {
return Err(bad("entry age is out of policy range"));
}
let children = c.uint(2)?;
// libhdf5 checks the parent flag against the child count only in
// debug builds (release builds refuse any entry with children);
// the image format's own rule is checked here.
if (flags & MDCI_ENTRY_IS_FD_PARENT != 0) != (children > 0) {
return Err(bad("flush dependency parent flag and child count disagree"));
}
c.uint(2)?; // dirty dependency children: reset for a read-only open
let parents = c.uint(2)?;
if (flags & MDCI_ENTRY_IS_FD_CHILD != 0) != (parents > 0) {
return Err(bad("flush dependency child flag and parent count disagree"));
}
c.uint(4)?; // LRU rank
let address = c
.addr(offset_size)?
.filter(|&a| a < eoa)
.ok_or(bad("invalid entry address range"))?;
let size = c.uint(length_size)?;
if size == 0 {
return Err(bad("invalid entry size"));
}
for _ in 0..parents {
let parent = c
.addr(offset_size)?
.ok_or(bad("invalid flush dependency parent offset"))?;
if !seen.contains(&parent) && !cached.contains(&parent) {
return Err(bad("fd parent not in cache"));
}
}
let len = usize::try_from(size).map_err(|_| bad(RAN_OFF))?;
let image_offset = c.pos;
c.take(len)?;
if address.checked_add(size).is_none_or(|end| end > eoa) {
return Err(bad("entry extends past the end of file"));
}
if !seen.insert(address) {
return Err(bad("duplicate addresses in cache"));
}
entries.push(ImageEntry {
address,
image_offset,
len,
});
}
Ok(CacheImage { location, entries })
}
/// Where the image block is.
pub fn location(&self) -> CacheImageLocation {
self.location
}
/// The number of entries in the image.
pub fn len(&self) -> usize {
self.entries.len()
}
/// Whether the image has no entries (a decoded image always has some).
pub fn is_empty(&self) -> bool {
self.entries.is_empty()
}
/// The file ranges (address, length) the image's entries replace.
pub fn entry_ranges(&self) -> impl Iterator<Item = (u64, usize)> + '_ {
self.entries.iter().map(|e| (e.address, e.len))
}
/// The image block in `data`, the bytes [`Self::decode`] read it from.
pub fn block<'a>(&self, data: &'a [u8]) -> Result<&'a [u8], FormatError> {
image_block(data, self.location)
}
/// Write every entry over `dst`, the file's bytes from the superblock
/// on (as long as the `data` the image was decoded from), taking the
/// entries from `block` (the image block, see [`Self::block`]). `block`
/// must not alias `dst`: an entry may land on the block itself.
pub fn apply(&self, block: &[u8], dst: &mut [u8]) -> Result<(), FormatError> {
let short = || FormatError::InvalidCacheImage("image applied to the wrong file");
for e in &self.entries {
let src = block
.get(e.image_offset..e.image_offset + e.len)
.ok_or_else(short)?;
let at = usize::try_from(e.address).map_err(|_| short())?;
dst.get_mut(at..at + e.len)
.ok_or_else(short)?
.copy_from_slice(src);
}
Ok(())
}
}
fn image_block(data: &[u8], location: CacheImageLocation) -> Result<&[u8], FormatError> {
let bad = FormatError::InvalidCacheImage;
let start = usize::try_from(location.address).map_err(|_| bad("address out of range"))?;
let len = usize::try_from(location.length).map_err(|_| bad("length out of range"))?;
start
.checked_add(len)
.and_then(|end| data.get(start..end))
.ok_or(bad("image block extends past the end of the file"))
}
/// What an opener must do before reading a file's metadata: check the
/// superblock extension ([`read_superblock_extension`]; an error means
/// libhdf5 refuses to open the file) and decode any metadata cache image
/// ([`CacheImage::decode`]). `data` is the file from the superblock on, up
/// to its recorded end of file.
pub fn cache_image_state(data: &[u8], sb: &Superblock) -> Result<CacheImageState, FormatError> {
match read_superblock_extension(data, sb)? {
Some(SuperblockExtension {
cache_image: Some(location),
..
}) => Ok(match CacheImage::decode(data, location, sb) {
Ok(image) => CacheImageState::Loaded(image),
Err(e) => CacheImageState::Unloadable(e),
}),
_ => Ok(CacheImageState::Absent),
}
}
/// [`cache_image_state`] for a reader that holds the file's bytes in a
/// buffer of its own: check the superblock extension and write any cache
/// image over `data` in place (only the image block is copied). An image
/// libhdf5 cannot load is an error here: such a reader has no way to open
/// the file and fail each object instead.
pub fn apply_cache_image_in_place(data: &mut [u8], sb: &Superblock) -> Result<(), FormatError> {
match cache_image_state(data, sb)? {
CacheImageState::Absent => Ok(()),
CacheImageState::Unloadable(e) => Err(e),
CacheImageState::Loaded(image) => {
let block = image.block(data)?.to_vec();
image.apply(&block, data)
}
}
}
#[cfg(test)]
mod tests {
use super::*;
/// The file's bytes with the image at `loc` applied.
fn apply_cache_image(
data: &[u8],
loc: CacheImageLocation,
sb: &Superblock,
) -> Result<Vec<u8>, FormatError> {
let image = CacheImage::decode(data, loc, sb)?;
let mut out = data.to_vec();
image.apply(image.block(data)?, &mut out)?;
Ok(out)
}
fn sb_v2(ext: u64) -> Superblock {
Superblock {
version: 2,
offset_size: 8,
length_size: 8,
base_address: 0,
eof_address: 0,
root_group_address: 0,
group_leaf_node_k: None,
group_internal_node_k: None,
indexed_storage_internal_node_k: None,
free_space_address: None,
driver_info_address: None,
consistency_flags: 0,
superblock_extension_address: Some(ext),
checksum: None,
page_size: None,
}
}
/// A file whose superblock extension (a version 1 object header at 48)
/// holds the given messages, padded to `len` bytes.
fn file_with_ext(messages: &[(u16, &[u8])], len: usize) -> Vec<u8> {
let mut body = Vec::new();
for (t, d) in messages {
let padded = d.len().div_ceil(8) * 8;
body.extend_from_slice(&t.to_le_bytes());
body.extend_from_slice(&(padded as u16).to_le_bytes());
body.extend_from_slice(&[0x14, 0, 0, 0]);
body.extend_from_slice(d);
body.resize(body.len() + padded - d.len(), 0);
}
let mut f = vec![0u8; 48];
f.push(1);
f.push(0);
f.extend_from_slice(&(messages.len() as u16).to_le_bytes());
f.extend_from_slice(&1u32.to_le_bytes());
f.extend_from_slice(&(body.len() as u32).to_le_bytes());
f.extend_from_slice(&[0; 4]);
f.extend_from_slice(&body);
f.resize(len, 0);
f
}
fn fsinfo_v1(page_size: u64, persist: bool, n_addrs: usize) -> Vec<u8> {
let mut m = vec![1, 1, u8::from(persist)];
m.extend_from_slice(&1u64.to_le_bytes());
m.extend_from_slice(&page_size.to_le_bytes());
m.extend_from_slice(&0u16.to_le_bytes());
m.extend_from_slice(&u64::MAX.to_le_bytes());
for _ in 0..n_addrs {
m.extend_from_slice(&u64::MAX.to_le_bytes());
}
m
}
fn mdci(address: u64, length: u64) -> Vec<u8> {
let mut m = vec![0];
m.extend_from_slice(&address.to_le_bytes());
m.extend_from_slice(&length.to_le_bytes());
m
}
#[test]
fn no_extension() {
assert_eq!(
read_superblock_extension(&[0; 64], &sb_v2(u64::MAX)).unwrap(),
None
);
}
#[test]
fn file_space_info_as_libhdf5_decodes_it() {
// What FileWriter::with_page_size writes.
let f = file_with_ext(&[(MSG_FSINFO, &fsinfo_v1(4096, false, 0))], 256);
let ext = read_superblock_extension(&f, &sb_v2(48)).unwrap().unwrap();
assert_eq!(ext.file_space_info.unwrap().page_size, 4096);
let f = file_with_ext(&[(MSG_FSINFO, &fsinfo_v1(4096, true, 12))], 512);
assert!(read_superblock_extension(&f, &sb_v2(48)).is_ok());
let refused = |m: Vec<u8>| {
let f = file_with_ext(&[(MSG_FSINFO, &m)], 512);
read_superblock_extension(&f, &sb_v2(48)).unwrap_err()
};
// Persisting, but too short for the manager addresses.
let mut short = fsinfo_v1(4096, true, 12);
short.truncate(short.len() - 8);
assert_eq!(refused(short), ext_err(RAN_OFF));
assert!(matches!(
refused(fsinfo_v1(256, false, 0)),
FormatError::InvalidSuperblockExtension(_)
));
assert!(matches!(
refused(fsinfo_v1(0, false, 0)),
FormatError::InvalidSuperblockExtension(_)
));
let mut v2 = fsinfo_v1(4096, false, 0);
v2[0] = 2;
assert!(matches!(
refused(v2),
FormatError::InvalidSuperblockExtension(_)
));
// cve-2020-10810: version 0, strategy ALL_PERSIST, and a message of
// 32 bytes that cannot hold the six addresses that follow.
let mut v0 = vec![0u8, 1];
v0.extend_from_slice(&[0, 1, 0, 0, 0, 0, 0, 0]);
v0.resize(32, 0xff);
assert_eq!(refused(v0), ext_err(RAN_OFF));
// A version 0 message without persistence is fine.
let mut v0 = vec![0u8, 2];
v0.extend_from_slice(&[0; 8]);
let f = file_with_ext(&[(MSG_FSINFO, &v0)], 256);
assert!(read_superblock_extension(&f, &sb_v2(48)).is_ok());
}
#[test]
fn cache_image_location_must_be_inside_the_file() {
let f = file_with_ext(&[(MSG_MDCI, &mdci(128, 64))], 192);
let ext = read_superblock_extension(&f, &sb_v2(48)).unwrap().unwrap();
assert_eq!(
ext.cache_image,
Some(CacheImageLocation {
address: 128,
length: 64
})
);
// cve-2020-10812: 256 MiB at 0x10100 in a 2565-byte file.
let f = file_with_ext(&[(MSG_MDCI, &mdci(0x10100, 0x1000_0000))], 2565);
assert!(matches!(
read_superblock_extension(&f, &sb_v2(48)),
Err(FormatError::InvalidSuperblockExtension(_))
));
let f = file_with_ext(&[(MSG_MDCI, &mdci(u64::MAX, 8))], 256);
assert!(read_superblock_extension(&f, &sb_v2(48)).is_err());
}
/// A cache image block with `entries` of (address, bytes).
fn image(entries: &[(u64, &[u8])]) -> Vec<u8> {
let with_deps: Vec<_> = entries.iter().map(|&(a, b)| (a, b, 0, None)).collect();
image_with_deps(&with_deps)
}
/// A cache image block with `entries` of (address, bytes, flush
/// dependency children, flush dependency parent).
fn image_with_deps(entries: &[(u64, &[u8], u16, Option<u64>)]) -> Vec<u8> {
let mut b = Vec::new();
b.extend_from_slice(MDCI_SIGNATURE);
b.push(0);
b.push(0);
b.extend_from_slice(&0u64.to_le_bytes()); // length, patched below
b.extend_from_slice(&(entries.len() as u32).to_le_bytes());
for &(addr, bytes, children, parent) in entries {
let mut flags = 0x02; // in LRU
if children > 0 {
flags |= MDCI_ENTRY_IS_FD_PARENT;
}
if parent.is_some() {
flags |= MDCI_ENTRY_IS_FD_CHILD;
}
b.extend_from_slice(&[5, flags, 1, 0]); // type, flags, ring, age
b.extend_from_slice(&children.to_le_bytes());
b.extend_from_slice(&0u16.to_le_bytes()); // dirty children
b.extend_from_slice(&u16::from(parent.is_some()).to_le_bytes());
b.extend_from_slice(&0i32.to_le_bytes());
b.extend_from_slice(&addr.to_le_bytes());
b.extend_from_slice(&(bytes.len() as u64).to_le_bytes());
if let Some(p) = parent {
b.extend_from_slice(&p.to_le_bytes());
}
b.extend_from_slice(bytes);
}
b.extend_from_slice(&[0; 4]); // checksum (not verified, as in libhdf5)
let n = b.len() as u64;
b[6..14].copy_from_slice(&n.to_le_bytes());
b
}
#[test]
fn cache_image_entries_replace_the_file_bytes() {
let img = image(&[(16, b"HEADER"), (40, b"NODE")]);
let mut f = vec![0u8; 64];
let at = f.len() as u64;
f.extend_from_slice(&img);
let loc = CacheImageLocation {
address: at,
length: img.len() as u64,
};
let out = apply_cache_image(&f, loc, &sb_v2(u64::MAX)).unwrap();
assert_eq!(out.len(), f.len());
assert_eq!(&out[16..22], b"HEADER");
assert_eq!(&out[40..44], b"NODE");
assert_eq!(&out[..16], &f[..16]);
let bad = |img: Vec<u8>| {
let mut f = vec![0u8; 64];
f.extend_from_slice(&img);
let loc = CacheImageLocation {
address: 64,
length: img.len() as u64,
};
apply_cache_image(&f, loc, &sb_v2(u64::MAX)).unwrap_err()
};
let mut sig = image(&[(16, b"x")]);
sig[0] = b'X';
assert!(matches!(bad(sig), FormatError::InvalidCacheImage(_)));
assert!(matches!(
bad(image(&[(16, b"a"), (16, b"b")])),
FormatError::InvalidCacheImage("duplicate addresses in cache")
));
assert!(matches!(
bad(image(&[(1 << 20, b"far")])),
FormatError::InvalidCacheImage("invalid entry address range")
));
let mut len = image(&[(16, b"x")]);
len[6] ^= 1;
assert!(matches!(bad(len), FormatError::InvalidCacheImage(_)));
// An entry that starts inside the file (64 bytes, then a 60-byte
// image) but runs past its end.
assert!(matches!(
bad(image(&[(123, b"8 bytes!")])),
FormatError::InvalidCacheImage("entry extends past the end of file")
));
let mut cut = image(&[(16, b"abcdef")]);
let n = cut.len() as u64 - 8;
cut.truncate(cut.len() - 8);
cut[6..14].copy_from_slice(&n.to_le_bytes());
assert!(matches!(bad(cut), FormatError::InvalidCacheImage(_)));
}
/// libhdf5 resolves an entry's flush-dependency parents as it inserts
/// the entry (`H5C__reconstruct_cache_contents`): a parent must be an
/// earlier entry, or the superblock or its extension's object header,
/// which are cached before the image loads. A parent listed after its
/// child fails ("fd parent not in cache?!?").
#[test]
fn flush_dependency_parents_must_already_be_cached() {
let load = |img: Vec<u8>| {
let mut f = vec![0u8; 64];
f.extend_from_slice(&img);
let loc = CacheImageLocation {
address: 64,
length: img.len() as u64,
};
apply_cache_image(&f, loc, &sb_v2(48))
};
// Parent first, as libhdf5 writes images.
assert!(
load(image_with_deps(&[
(16, b"P", 1, None),
(40, b"C", 0, Some(16))
]))
.is_ok()
);
// Child first: libhdf5 does not find the parent.
assert_eq!(
load(image_with_deps(&[
(40, b"C", 0, Some(16)),
(16, b"P", 1, None)
]))
.unwrap_err(),
FormatError::InvalidCacheImage("fd parent not in cache")
);
// The superblock extension's header (at 48 here) is in the cache.
assert!(load(image_with_deps(&[(40, b"C", 0, Some(48))])).is_ok());
// An entry cannot be its own parent.
assert!(load(image_with_deps(&[(40, b"C", 1, Some(40))])).is_err());
}
}
+22 -1
View File
@@ -803,7 +803,11 @@ impl<'a, 'r> Sources<'a, 'r> {
let resolver = self.resolver.ok_or_else(|| {
vds_err("external-file virtual dataset sources require a file resolver")
})?;
self.cached_file = Some((String::from(name), resolver(name)?));
let mut bytes = resolver(name)?;
if let Some(b) = bytes.as_mut() {
load_source_file(b)?;
}
self.cached_file = Some((String::from(name), bytes));
}
// An external file is handed over whole; its addresses are relative
// to its superblock, so skip any user block.
@@ -851,6 +855,23 @@ impl<'a, 'r> Sources<'a, 'r> {
}
}
/// Check an external source file's superblock extension as libhdf5 does
/// when it opens the file, and write any metadata cache image over its
/// metadata in place: libhdf5 reads the image's entries instead of the
/// file's own, possibly stale, bytes (`crate::superblock_ext`). A source
/// file whose image cannot be loaded is an error, as other corrupt source
/// files are here.
fn load_source_file(whole: &mut [u8]) -> Result<(), FormatError> {
let base = crate::signature::find_signature(whole)?;
let sb = crate::superblock::Superblock::parse(&whole[base..], 0)?;
// The end of file the superblock records; a truncated source file is
// read as before, up to its length.
let end = sb
.data_end(base as u64, whole.len() as u64)
.map_or(whole.len(), |e| base + e as usize);
crate::superblock_ext::apply_cache_image_in_place(&mut whole[base..end], &sb)
}
/// Whether elements of `dt` contain addresses into their own file:
/// variable-length data (global-heap IDs) or references.
fn holds_file_addresses(dt: &Datatype) -> bool {
+43
View File
@@ -288,6 +288,12 @@ impl AsyncHDF5File {
// superblock records, as libhdf5 does.
let end = superblock.data_end(user_block as u64, whole_len)?;
data.truncate(end as usize);
// Check the superblock extension as libhdf5 does at open, and write
// any metadata cache image over the file's metadata (libhdf5 reads
// the image's entries instead of the file's own, possibly stale,
// bytes). An image libhdf5 cannot load is refused: this reader has
// no way to open the file and fail each object instead.
clawhdf5_format::superblock_ext::apply_cache_image_in_place(&mut data, &superblock)?;
Ok(Self { data, superblock })
}
@@ -430,6 +436,43 @@ mod tests {
fw.finish().unwrap()
}
/// libhdf5's `h5clear_mdc_image.h5` (see the clawhdf5 crate's
/// `tests/metadata_cache_image.rs`): the root group's header exists
/// only in the file's metadata cache image.
fn cache_image_fixture() -> Vec<u8> {
std::fs::read(concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5/tests/fixtures/h5clear_mdc_image.h5"
))
.unwrap()
}
#[tokio::test]
async fn reads_through_a_metadata_cache_image() {
let bytes = cache_image_fixture();
let file = AsyncHDF5File::from_bytes(bytes.clone()).unwrap();
let info = file.read_dataset_raw("DSET").await.unwrap();
assert_eq!(info.shape, [50, 100]);
let values: Vec<i32> = info
.raw
.as_chunks::<4>()
.0
.iter()
.map(|&b| i32::from_le_bytes(b))
.collect();
let expected: Vec<i32> = (0..50).flat_map(|i| (0..100).map(move |j| i * j)).collect();
assert_eq!(values, expected);
// An image libhdf5 cannot load is refused at open.
let mut bad = bytes;
let at = bad.windows(4).position(|w| w == b"MDCI").unwrap();
bad[at] = b'X';
assert!(matches!(
AsyncHDF5File::from_bytes(bad),
Err(AsyncHDF5Error::Format(FormatError::InvalidCacheImage(_)))
));
}
// --- AsyncMemoryReader tests ---
#[tokio::test]
+59
View File
@@ -41,6 +41,65 @@ pub trait HDF5Read {
fn is_empty(&self) -> bool {
self.as_bytes().is_empty()
}
/// A private, writable copy of [`Self::as_bytes`]: writes to it stay in
/// this process and never reach the underlying storage.
///
/// Readers use it to lay a file's metadata cache image over the file's
/// own metadata. The default copies the bytes; a memory-mapped reader
/// returns a copy-on-write mapping instead, so only the pages written to
/// are copied and the rest stay shared with the page cache.
fn private_copy(&self) -> io::Result<PrivateCopy> {
Ok(PrivateCopy::Owned(self.as_bytes().to_vec()))
}
}
/// A private, writable copy of a file's bytes (see
/// [`HDF5Read::private_copy`]).
pub enum PrivateCopy {
/// The bytes copied onto the heap.
Owned(Vec<u8>),
/// A copy-on-write mapping of the file: pages are copied only when
/// written to.
#[cfg(feature = "mmap")]
Mapped(memmap2::MmapMut),
}
impl PrivateCopy {
/// Whether this is a copy-on-write mapping rather than a heap copy.
pub fn is_mapped(&self) -> bool {
!matches!(self, PrivateCopy::Owned(_))
}
}
impl std::fmt::Debug for PrivateCopy {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("PrivateCopy")
.field("len", &self.len())
.field("mapped", &self.is_mapped())
.finish()
}
}
impl std::ops::Deref for PrivateCopy {
type Target = [u8];
fn deref(&self) -> &[u8] {
match self {
PrivateCopy::Owned(v) => v,
#[cfg(feature = "mmap")]
PrivateCopy::Mapped(m) => m,
}
}
}
impl std::ops::DerefMut for PrivateCopy {
fn deref_mut(&mut self) -> &mut [u8] {
match self {
PrivateCopy::Owned(v) => v,
#[cfg(feature = "mmap")]
PrivateCopy::Mapped(m) => m,
}
}
}
/// Read-write access to HDF5 data.
+29
View File
@@ -87,6 +87,19 @@ impl HDF5Read for MmapReader {
fn as_bytes(&self) -> &[u8] {
&self.mmap
}
/// A private copy-on-write mapping of the file (`MAP_PRIVATE`): only the
/// pages written to are copied.
fn private_copy(&self) -> io::Result<crate::PrivateCopy> {
if self.mmap.is_empty() {
return Ok(crate::PrivateCopy::Owned(Vec::new()));
}
// SAFETY: as for `open`: the caller keeps the file from being
// modified while the mapping is alive. Writes to a private mapping
// never reach the file.
let map = unsafe { memmap2::MmapOptions::new().map_copy(&self._file)? };
Ok(crate::PrivateCopy::Mapped(map))
}
}
/// Writable memory-mapped file for read-write HDF5 access.
@@ -218,6 +231,22 @@ mod tests {
fs::remove_file(&path).ok();
}
#[test]
fn private_copy_is_a_copy_on_write_mapping() {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("cow.bin");
fs::write(&path, [1u8, 2, 3, 4]).unwrap();
let reader = MmapReader::open(&path).unwrap();
let mut copy = reader.private_copy().unwrap();
assert!(copy.is_mapped());
copy[1] = 99;
assert_eq!(&copy[..], &[1, 99, 3, 4]);
// Neither the reader's mapping nor the file sees the write.
assert_eq!(reader.as_bytes(), &[1, 2, 3, 4]);
drop(copy);
assert_eq!(fs::read(&path).unwrap(), [1, 2, 3, 4]);
}
#[test]
fn mmap_reader_read_at() {
let dir = std::env::temp_dir();
+4 -1
View File
@@ -191,7 +191,10 @@ fn mpi_collective_read(vol: &MpiVol, location: &str, path: &str) -> Result<Vec<u
let mut len_buf = [0usize; 1];
if rank == 0 {
let file = std::fs::read(location).map_err(VolError::Io)?;
let mut file = std::fs::read(location).map_err(VolError::Io)?;
// Checked as libhdf5 checks a file at open, with any metadata cache
// image written over the metadata (see `vol::load_hdf5`).
crate::vol::load_hdf5(&mut file)?;
// From the superblock to the recorded end of file; truncated files
// are refused.
let (bytes, sb) = crate::vol::hdf5_view(&file)?;
+4
View File
@@ -195,6 +195,10 @@ impl<R: HDF5Read> HDF5Read for PrefetchReader<R> {
fn as_bytes(&self) -> &[u8] {
self.inner.as_bytes()
}
fn private_copy(&self) -> std::io::Result<crate::PrivateCopy> {
self.inner.private_copy()
}
}
// ---------------------------------------------------------------------------
+74 -4
View File
@@ -208,6 +208,9 @@ pub trait VirtualObjectLayer: Send + Sync {
pub struct NativeVol {
data: Option<Vec<u8>>,
location: Option<String>,
/// Why the bytes given to [`NativeVol::from_bytes`] cannot be read
/// (what `open` would have refused them for).
load_error: Option<String>,
}
/// The HDF5 bytes of a whole file and its superblock: from the superblock
@@ -228,12 +231,35 @@ pub(crate) fn hdf5_view(
Ok((&data[..end as usize], sb))
}
/// Check a whole file as libhdf5 does when it opens it ([`hdf5_view`], and
/// the superblock extension), and write any metadata cache image over the
/// file's metadata in place: libhdf5 reads the image's entries instead of
/// the file's own bytes at their addresses, which may be stale
/// (`clawhdf5_format::superblock_ext`). A file whose image libhdf5 cannot
/// load is refused: a connector that reads whole datasets has no way to
/// open the file and fail each object instead.
pub(crate) fn load_hdf5(whole: &mut [u8]) -> Result<(), VolError> {
use clawhdf5_format::superblock_ext::apply_cache_image_in_place;
let err = |e: clawhdf5_format::error::FormatError| VolError::DataError(e.to_string());
let (len, sb) = {
let (data, sb) = hdf5_view(whole)?;
(data.len(), sb)
};
// hdf5_view's bytes start at the superblock, after any user block.
let base = clawhdf5_format::signature::split_user_block(whole)
.map_err(err)?
.0
.len();
apply_cache_image_in_place(&mut whole[base..base + len], &sb).map_err(err)
}
impl NativeVol {
/// Create a new native VOL connector.
pub fn new() -> Self {
Self {
data: None,
location: None,
load_error: None,
}
}
@@ -245,10 +271,12 @@ impl NativeVol {
}
/// Create a native VOL connector from bytes already in memory.
pub fn from_bytes(data: Vec<u8>) -> Self {
pub fn from_bytes(mut data: Vec<u8>) -> Self {
let load_error = load_hdf5(&mut data).err().map(|e| e.to_string());
Self {
data: Some(data),
location: Some("<memory>".into()),
load_error,
}
}
@@ -281,10 +309,12 @@ impl VirtualObjectLayer for NativeVol {
}
fn open(&mut self, location: &str) -> Result<(), VolError> {
let data = std::fs::read(location)?;
// Refuse a truncated file at open, as libhdf5 does.
hdf5_view(&data)?;
let mut data = std::fs::read(location)?;
// Refuse a truncated file at open, as libhdf5 does, and load any
// metadata cache image.
load_hdf5(&mut data)?;
self.data = Some(data);
self.load_error = None;
self.location = Some(location.to_string());
Ok(())
}
@@ -299,6 +329,9 @@ impl VirtualObjectLayer for NativeVol {
let data = self.data.as_ref().ok_or_else(|| {
VolError::Io(io::Error::new(io::ErrorKind::NotConnected, "file not open"))
})?;
if let Some(e) = &self.load_error {
return Err(VolError::DataError(e.clone()));
}
use clawhdf5_format::{
data_layout::DataLayout, data_read::read_raw_data_full, dataspace::Dataspace,
@@ -434,6 +467,43 @@ mod tests {
assert_eq!(raw.len(), 24);
}
/// libhdf5's `h5clear_mdc_image.h5` (see the clawhdf5 crate's
/// `tests/metadata_cache_image.rs`): the root group's header exists
/// only in the file's metadata cache image, so reading the file's own
/// bytes finds zeros there.
#[test]
fn native_vol_reads_through_a_metadata_cache_image() {
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5/tests/fixtures/h5clear_mdc_image.h5"
);
let expected: Vec<u8> = (0..50)
.flat_map(|i| (0..100).map(move |j| i * j))
.flat_map(i32::to_le_bytes)
.collect();
let vol = NativeVol::open_path(path).unwrap();
assert_eq!(vol.read_dataset("DSET").unwrap(), expected);
let bytes = std::fs::read(path).unwrap();
let vol = NativeVol::from_bytes(bytes.clone());
assert_eq!(vol.read_dataset("DSET").unwrap(), expected);
// An image libhdf5 cannot load is refused, not read around.
let mut bad = bytes;
let at = bad.windows(4).position(|w| w == b"MDCI").unwrap();
bad[at] = b'X';
let err = NativeVol::from_bytes(bad.clone())
.read_dataset("DSET")
.unwrap_err();
assert!(err.to_string().contains("cache image"), "{err}");
let dir = tempfile::tempdir().unwrap();
let bad_path = dir.path().join("bad_image.h5");
std::fs::write(&bad_path, &bad).unwrap();
let err = NativeVol::open_path(bad_path.to_str().unwrap())
.err()
.unwrap();
assert!(err.to_string().contains("cache image"), "{err}");
}
#[test]
fn vol_error_display() {
let err = VolError::Unsupported("read_dataset".into());
+5
View File
@@ -234,6 +234,11 @@ impl H5 {
}
pub fn header(&self, addr: u64) -> Result<ObjectHeader> {
// A metadata cache image libhdf5 cannot load: the file opens, and
// every object fails (its bytes may hold stale metadata).
if let Some(e) = self.file.cache_image_error() {
return Err(Error::at(addr, format!("metadata cache image: {e}")));
}
let off = usize::try_from(addr).map_err(|_| Error::at(addr, "address out of range"))?;
ObjectHeader::parse(self.data(), off, self.os(), self.ls())
.map_err(|e| Error::at(addr, format!("object header: {e}")))
+85
View File
@@ -0,0 +1,85 @@
//! Metadata cache images, as the file openers apply them.
//!
//! A file written with a metadata cache image keeps metadata cache entries
//! (object headers, B-tree nodes, heaps) in an image block, and libhdf5
//! reads those entries in place of the file's own bytes at their addresses
//! (see `clawhdf5_format::superblock_ext`). The metadata parsers read one
//! contiguous byte slice, so the image has to be laid over the file's bytes
//! — without copying the file:
//!
//! - an opener that holds the file in a buffer it owns writes the entries
//! into that buffer (only the image block is copied, as libhdf5 copies
//! it);
//! - an opener that maps the file ([`File::open`](crate::File::open),
//! [`MmapFile`](crate::MmapFile), [`LazyFile::open_mmap`]
//! (crate::LazyFile::open_mmap)) writes them into a private copy-on-write
//! mapping of the file ([`clawhdf5_io::HDF5Read::private_copy`]): only the
//! pages the entries land on are copied, and the rest of the file stays
//! shared with the page cache;
//! - a file without an image is read from the original bytes, as before.
use clawhdf5_format::error::FormatError;
use clawhdf5_format::superblock::Superblock;
use clawhdf5_format::superblock_ext::{self, CacheImageState};
use clawhdf5_io::PrivateCopy;
use crate::error::Error;
/// A file's metadata, as an opener must read it.
pub(crate) enum ImageView {
/// Read the opener's bytes: the file has no cache image, or it was
/// written into a buffer the opener owns.
Plain,
/// The file with its cache image written in: a private copy of the
/// whole file (the HDF5 data at the same offsets as in the file).
Patched(PrivateCopy),
/// The file has a cache image libhdf5 cannot load.
Unloadable(FormatError),
}
/// Check the superblock extension of the file whose bytes are `whole` (the
/// HDF5 data in `base..end`) and lay any cache image over a private copy
/// of the file made by `copy`. An error means libhdf5 refuses the file.
pub(crate) fn private_view(
whole: &[u8],
base: usize,
end: usize,
sb: &Superblock,
copy: impl FnOnce() -> std::io::Result<PrivateCopy>,
) -> Result<ImageView, Error> {
let data = &whole[base..end];
Ok(match superblock_ext::cache_image_state(data, sb)? {
CacheImageState::Absent => ImageView::Plain,
CacheImageState::Unloadable(e) => ImageView::Unloadable(e),
CacheImageState::Loaded(image) => {
let mut view = copy().map_err(Error::Io)?;
let dst =
view.get_mut(base..end)
.ok_or(Error::Format(FormatError::InvalidCacheImage(
"the file changed while it was opened",
)))?;
image.apply(image.block(data)?, dst)?;
ImageView::Patched(view)
}
})
}
/// [`private_view`] for a file held in `whole`, a buffer the opener owns:
/// the image is written into it in place.
pub(crate) fn in_place(
whole: &mut [u8],
base: usize,
end: usize,
sb: &Superblock,
) -> Result<ImageView, Error> {
let data = &mut whole[base..end];
Ok(match superblock_ext::cache_image_state(data, sb)? {
CacheImageState::Absent => ImageView::Plain,
CacheImageState::Unloadable(e) => ImageView::Unloadable(e),
CacheImageState::Loaded(image) => {
let block = image.block(data)?.to_vec();
image.apply(&block, data)?;
ImageView::Plain
}
})
}
+54 -6
View File
@@ -46,6 +46,13 @@ pub struct LazyFile<R: HDF5Read> {
/// End of the HDF5 data (`Superblock::data_end`, absolute).
end: usize,
superblock: Superblock,
/// A file that holds a metadata cache image, with the image written in:
/// the reader's [`HDF5Read::private_copy`] of the whole file (a
/// copy-on-write mapping for [`clawhdf5_io::MmapReader`], so only the
/// pages the image's entries land on are copied; see
/// `crate::cache_image`). `None` for a file without an image, read
/// straight from the reader.
patched: Option<clawhdf5_io::PrivateCopy>,
root_header: ObjectHeader,
/// Cache of parsed object headers, keyed by address.
header_cache: RefCell<HashMap<u64, ObjectHeader>>,
@@ -82,7 +89,24 @@ impl<R: HDF5Read> LazyFile<R> {
let superblock = Superblock::parse(data, 0)?;
// Refuse a truncated file; read nothing past the recorded end of file.
let end = base + superblock.data_end(base as u64, whole_len)? as usize;
let data = &reader.as_bytes()[base..end];
// Decode the superblock extension as libhdf5 does at open, and load
// a metadata cache image over the file's metadata.
let view =
crate::cache_image::private_view(reader.as_bytes(), base, end, &superblock, || {
reader.private_copy()
})?;
let patched = match view {
crate::cache_image::ImageView::Plain => None,
crate::cache_image::ImageView::Patched(p) => Some(p),
// libhdf5 opens such a file and fails its first metadata read;
// a LazyFile reads the root group's header at open, so the open
// is that read.
crate::cache_image::ImageView::Unloadable(e) => return Err(e.into()),
};
let data = match &patched {
Some(p) => &p[base..end],
None => &reader.as_bytes()[base..end],
};
let root_header = ObjectHeader::parse(
data,
superblock.root_group_address as usize,
@@ -94,6 +118,7 @@ impl<R: HDF5Read> LazyFile<R> {
base,
end,
superblock,
patched,
root_header,
header_cache: RefCell::new(HashMap::new()),
})
@@ -111,7 +136,10 @@ impl<R: HDF5Read> LazyFile<R> {
}
fn hdf5_bytes(&self) -> &[u8] {
&self.reader.as_bytes()[self.base..self.end]
match &self.patched {
Some(p) => &p[self.base..self.end],
None => &self.reader.as_bytes()[self.base..self.end],
}
}
/// Returns a reference to the parsed superblock.
@@ -135,10 +163,11 @@ impl<R: HDF5Read> LazyFile<R> {
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(path.to_string()));
}
Ok(LazyDataset {
LazyDataset {
file: self,
header: hdr,
})
}
.check_open()
}
/// Resolve a path and return a `LazyGroup` handle.
@@ -284,10 +313,11 @@ impl<'f, R: HDF5Read> LazyGroup<'f, R> {
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(name.to_string()));
}
Ok(LazyDataset {
LazyDataset {
file: self.file,
header: hdr,
})
}
.check_open()
}
/// Get a subgroup within this group by name.
@@ -326,6 +356,24 @@ pub struct LazyDataset<'f, R: HDF5Read> {
}
impl<'f, R: HDF5Read> LazyDataset<'f, R> {
/// libhdf5's storage checks when it opens a dataset
/// ([`data_read::check_dataset_storage`]): a dataset whose element count
/// times element size overflows, or whose contiguous storage runs past
/// the end of the file, fails to open. A datatype, dataspace or layout
/// that does not decode is left for the read to report, as before (the
/// dataset still opens, and its attributes can be read).
fn check_open(self) -> Result<Self, Error> {
let decoded = (|| -> Result<_, Error> {
let data = self.required_payload(MessageType::Dataspace)?;
let ds = Dataspace::parse(&data, self.file.length_size())?;
Ok((self.datatype()?, ds, self.data_layout()?))
})();
if let Ok((dt, ds, dl)) = decoded {
data_read::check_dataset_storage(&dl, &ds, &dt, self.file.hdf5_bytes().len() as u64)?;
}
Ok(self)
}
/// Returns the shape (dimensions) of the dataset.
pub fn shape(&self) -> Result<Vec<u64>, Error> {
let ds = self.dataspace()?;
+1
View File
@@ -24,6 +24,7 @@
//! builder.write("output.h5").unwrap();
//! ```
mod cache_image;
pub mod error;
pub mod lazy;
#[cfg(feature = "mmap")]
+69 -16
View File
@@ -38,6 +38,14 @@ pub struct MmapFile {
/// End of the HDF5 data (`Superblock::data_end`, absolute).
end: usize,
superblock: Superblock,
/// A file that holds a metadata cache image, with the image written in:
/// a private copy-on-write mapping of the whole file, so only the pages
/// the image's entries land on are copied (see `crate::cache_image`).
/// `None` for a file without an image, read straight from the mapping.
patched: Option<clawhdf5_io::PrivateCopy>,
/// The file has a metadata cache image libhdf5 cannot load: every
/// object lookup fails with this error (see `File`).
image_error: Option<FormatError>,
}
impl MmapFile {
@@ -50,18 +58,34 @@ impl MmapFile {
let superblock = Superblock::parse(data, 0)?;
// Refuse a truncated file; read nothing past the recorded end of file.
let end = base + superblock.data_end(base as u64, whole_len)? as usize;
// Decode the superblock extension as libhdf5 does at open, and load
// a metadata cache image over the file's metadata.
let view =
crate::cache_image::private_view(reader.as_bytes(), base, end, &superblock, || {
clawhdf5_io::HDF5Read::private_copy(&reader)
})?;
let (patched, image_error) = match view {
crate::cache_image::ImageView::Plain => (None, None),
crate::cache_image::ImageView::Patched(p) => (Some(p), None),
crate::cache_image::ImageView::Unloadable(e) => (None, Some(e)),
};
Ok(Self {
reader,
base,
end,
superblock,
patched,
image_error,
})
}
/// The file's bytes from the superblock on — the space HDF5 addresses
/// index into.
fn hdf5_bytes(&self) -> &[u8] {
&self.reader.as_bytes()[self.base..self.end]
match &self.patched {
Some(p) => &p[self.base..self.end],
None => &self.reader.as_bytes()[self.base..self.end],
}
}
/// Size of the user block before the superblock (0 for most files).
@@ -79,21 +103,22 @@ impl MmapFile {
/// Resolve a path and return a `MmapDataset` handle.
pub fn dataset(&self, path: &str) -> Result<MmapDataset<'_>, Error> {
let data = self.hdf5_bytes();
let data = self.meta()?;
let addr = group_v2::resolve_path_any(data, &self.superblock, path)?;
let hdr = self.parse_header(addr)?;
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(path.to_string()));
}
Ok(MmapDataset {
MmapDataset {
file: self,
header: hdr,
})
}
.check_open()
}
/// Resolve a path and return a `MmapGroup` handle.
pub fn group(&self, path: &str) -> Result<MmapGroup<'_>, Error> {
let data = self.hdf5_bytes();
let data = self.meta()?;
let addr = group_v2::resolve_path_any(data, &self.superblock, path)?;
Ok(MmapGroup {
file: self,
@@ -108,14 +133,29 @@ impl MmapFile {
self.hdf5_bytes()
}
/// The error of a metadata cache image libhdf5 cannot load, when the
/// file has one (see [`crate::File::cache_image_error`]).
pub fn cache_image_error(&self) -> Option<&FormatError> {
self.image_error.as_ref()
}
/// Returns a reference to the parsed superblock.
pub fn superblock(&self) -> &Superblock {
&self.superblock
}
/// The bytes to read metadata from; fails for a file whose cache image
/// cannot be loaded.
fn meta(&self) -> Result<&[u8], FormatError> {
match &self.image_error {
Some(e) => Err(e.clone()),
None => Ok(self.hdf5_bytes()),
}
}
fn parse_header(&self, address: u64) -> Result<ObjectHeader, FormatError> {
ObjectHeader::parse(
self.hdf5_bytes(),
self.meta()?,
address as usize,
self.superblock.offset_size,
self.superblock.length_size,
@@ -211,10 +251,11 @@ impl<'f> MmapGroup<'f> {
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(name.to_string()));
}
Ok(MmapDataset {
MmapDataset {
file: self.file,
header: hdr,
})
}
.check_open()
}
/// Get a subgroup within this group by name.
@@ -235,7 +276,7 @@ impl<'f> MmapGroup<'f> {
/// [`group_v2::resolve_group_children`]); dangling, external and
/// user-defined links are left out.
fn children(&self) -> Result<Vec<GroupEntry>, Error> {
let data = self.file.hdf5_bytes();
let data = self.file.meta()?;
group_v2::resolve_group_children(data, &self.file.superblock, self.address)
.map_err(Error::Format)
}
@@ -253,6 +294,24 @@ pub struct MmapDataset<'f> {
}
impl<'f> MmapDataset<'f> {
/// libhdf5's storage checks when it opens a dataset
/// ([`data_read::check_dataset_storage`]): a dataset whose element count
/// times element size overflows, or whose contiguous storage runs past
/// the end of the file, fails to open. A datatype, dataspace or layout
/// that does not decode is left for the read to report, as before (the
/// dataset still opens, and its attributes can be read).
fn check_open(self) -> Result<Self, Error> {
let decoded = (|| -> Result<_, Error> {
let data = self.required_payload(MessageType::Dataspace)?;
let ds = Dataspace::parse(&data, self.file.length_size())?;
Ok((self.datatype()?, ds, self.data_layout()?))
})();
if let Ok((dt, ds, dl)) = decoded {
data_read::check_dataset_storage(&dl, &ds, &dt, self.file.hdf5_bytes().len() as u64)?;
}
Ok(self)
}
/// Returns the shape (dimensions) of the dataset.
pub fn shape(&self) -> Result<Vec<u64>, Error> {
let ds = self.dataspace()?;
@@ -405,13 +464,7 @@ impl<'f> MmapDataset<'f> {
match &dl {
DataLayout::Contiguous { address, size } => {
let addr = address.ok_or(Error::Format(FormatError::NoDataAllocated))?;
let sz = *size as usize;
if sz != expected {
return Err(Error::Format(FormatError::DataSizeMismatch {
expected,
actual: sz,
}));
}
let sz = clawhdf5_format::data_read::contiguous_read_len(*size, expected)?;
let data = self.file.hdf5_bytes();
let a = addr as usize;
if a + sz > data.len() {
+141 -16
View File
@@ -21,6 +21,7 @@ use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::signature;
use clawhdf5_format::superblock::Superblock;
use crate::cache_image::{self, ImageView};
use crate::error::Error;
use crate::types::{AttrValue, DType, classify_datatype, read_attrs};
@@ -55,12 +56,24 @@ struct FileData {
base: usize,
/// End of the HDF5 data in the file (`Superblock::data_end`, absolute).
end: usize,
/// A mapped file that holds a metadata cache image, with the image
/// written in: a private copy-on-write mapping of the whole file, so
/// only the pages the image's entries land on are copied (see
/// `crate::cache_image`). `None` for every other file: a file without
/// an image is read straight from the mapping, and an owned buffer has
/// the image written into it in place.
patched: Option<clawhdf5_io::PrivateCopy>,
/// The file has a metadata cache image libhdf5 cannot load. libhdf5
/// opens such a file and fails its first metadata read (the image loads
/// then); every object lookup here fails with this error, and no
/// metadata is read from the file's own, possibly stale, bytes.
image_error: Option<FormatError>,
}
impl FileData {
/// Locate the superblock and parse it. A truncated file is refused, and
/// bytes past the recorded end of file are not read, as in libhdf5.
fn new(backing: Backing) -> Result<(Self, Superblock), Error> {
fn new(mut backing: Backing) -> Result<(Self, Superblock), Error> {
let whole = backing.whole_file();
let (user_block, hdf5) = signature::split_user_block(whole)?;
let base = user_block.len();
@@ -68,16 +81,54 @@ impl FileData {
let end = superblock.data_end(base as u64, whole.len() as u64)?;
// data_end is at most the file length (less the user block).
let end = base + end as usize;
Ok((Self { backing, base, end }, superblock))
// libhdf5 decodes the superblock extension at open (a message it
// cannot decode fails the open) and loads a metadata cache image
// over the file's own metadata.
let view = match &mut backing {
Backing::Owned(v) => cache_image::in_place(v, base, end, &superblock)?,
#[cfg(feature = "mmap")]
Backing::Mmap(r) => {
cache_image::private_view(r.as_bytes(), base, end, &superblock, || {
clawhdf5_io::HDF5Read::private_copy(r)
})?
}
};
let (patched, image_error) = match view {
ImageView::Plain => (None, None),
ImageView::Patched(p) => (Some(p), None),
ImageView::Unloadable(e) => (None, Some(e)),
};
Ok((
Self {
backing,
base,
end,
patched,
image_error,
},
superblock,
))
}
fn as_bytes(&self) -> &[u8] {
&self.backing.whole_file()[self.base..self.end]
match &self.patched {
Some(p) => &p[self.base..self.end],
None => &self.backing.whole_file()[self.base..self.end],
}
}
fn len(&self) -> usize {
self.as_bytes().len()
}
/// The bytes to read metadata from; fails for a file whose cache image
/// cannot be loaded (see [`Self::image_error`]).
fn meta(&self) -> Result<&[u8], FormatError> {
match &self.image_error {
Some(e) => Err(e.clone()),
None => Ok(self.as_bytes()),
}
}
}
// ---------------------------------------------------------------------------
@@ -164,16 +215,17 @@ impl File {
///
/// The path uses `/` separators (e.g., `"group1/values"`).
pub fn dataset(&self, path: &str) -> Result<Dataset<'_>, Error> {
let data = self.data.as_bytes();
let data = self.data.meta()?;
let addr = group_v2::resolve_path_any(data, &self.superblock, path)?;
let hdr = self.parse_header(addr)?;
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(path.to_string()));
}
Ok(Dataset {
Dataset {
file: self,
header: hdr,
})
}
.check_open()
}
/// A `Dataset` handle for the object header at `address` (an address
@@ -186,10 +238,11 @@ impl File {
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(format!("object at address {address}")));
}
Ok(Dataset {
Dataset {
file: self,
header: hdr,
})
}
.check_open()
}
/// Resolve a path and return a `Group` handle.
@@ -197,7 +250,7 @@ impl File {
/// The path uses `/` separators (e.g., `"sensors"`).
/// Use `"/"` or `""` for the root group.
pub fn group(&self, path: &str) -> Result<Group<'_>, Error> {
let data = self.data.as_bytes();
let data = self.data.meta()?;
let addr = group_v2::resolve_path_any(data, &self.superblock, path)?;
Ok(Group {
file: self,
@@ -253,11 +306,23 @@ impl File {
/// Returns the file's bytes from the superblock on (after any user
/// block). Every HDF5 address in the file indexes this slice, so it is
/// what the `clawhdf5_format` parsers expect as `file_data`.
/// what the `clawhdf5_format` parsers expect as `file_data`. For a file
/// with a metadata cache image these are the bytes with the image
/// applied; when the image cannot be loaded they are the file's own
/// bytes, whose metadata may be stale (every object lookup fails then).
pub fn as_bytes(&self) -> &[u8] {
self.data.as_bytes()
}
/// The error of a metadata cache image libhdf5 cannot load, when the
/// file has one. Such a file opens, as in libhdf5, and every object
/// lookup fails with this error; code that parses [`Self::as_bytes`]
/// itself should check it first, since those bytes then hold the
/// file's own, possibly stale, metadata.
pub fn cache_image_error(&self) -> Option<&FormatError> {
self.data.image_error.as_ref()
}
/// Size of the user block before the superblock (0 for most files).
/// Matches h5py's `File.userblock_size`.
pub fn user_block_size(&self) -> u64 {
@@ -302,7 +367,7 @@ impl File {
raw: &[u8],
) -> Result<Vec<Vec<u8>>, Error> {
crate::vlen::decode_string_bytes(
self.as_bytes(),
self.data.meta()?,
datatype,
raw,
self.offset_size(),
@@ -320,7 +385,7 @@ impl File {
raw: &[u8],
) -> Result<Vec<Vec<T>>, Error> {
crate::vlen::decode_vlen(
self.as_bytes(),
self.data.meta()?,
datatype,
raw,
self.offset_size(),
@@ -330,7 +395,7 @@ impl File {
fn parse_header(&self, address: u64) -> Result<ObjectHeader, FormatError> {
ObjectHeader::parse(
self.data.as_bytes(),
self.data.meta()?,
address as usize,
self.superblock.offset_size,
self.superblock.length_size,
@@ -427,10 +492,11 @@ impl<'f> Group<'f> {
if !has_message(&hdr, MessageType::DataLayout) {
return Err(Error::NotADataset(name.to_string()));
}
Ok(Dataset {
Dataset {
file: self.file,
header: hdr,
})
}
.check_open()
}
/// Get a subgroup within this group by name.
@@ -451,7 +517,7 @@ impl<'f> Group<'f> {
/// [`group_v2::resolve_group_children`]); dangling, external and
/// user-defined links are left out.
fn children(&self) -> Result<Vec<GroupEntry>, Error> {
let data = self.file.data.as_bytes();
let data = self.file.data.meta()?;
group_v2::resolve_group_children(data, &self.file.superblock, self.address)
.map_err(Error::Format)
}
@@ -469,6 +535,24 @@ pub struct Dataset<'f> {
}
impl<'f> Dataset<'f> {
/// libhdf5's storage checks when it opens a dataset
/// ([`data_read::check_dataset_storage`]): a dataset whose element count
/// times element size overflows, or whose contiguous storage runs past
/// the end of the file, fails to open. A datatype, dataspace or layout
/// that does not decode is left for the read to report, as before (the
/// dataset still opens, and its attributes can be read).
fn check_open(self) -> Result<Self, Error> {
let decoded = (|| -> Result<_, Error> {
let data = self.required_payload(MessageType::Dataspace)?;
let ds = Dataspace::parse(&data, self.file.length_size())?;
Ok((self.datatype()?, ds, self.data_layout()?))
})();
if let Ok((dt, ds, dl)) = decoded {
data_read::check_dataset_storage(&dl, &ds, &dt, self.file.data.len() as u64)?;
}
Ok(self)
}
/// Returns the shape (dimensions) of the dataset.
pub fn shape(&self) -> Result<Vec<u64>, Error> {
let ds = self.dataspace()?;
@@ -1283,3 +1367,44 @@ mod sibling_file_name_tests {
}
}
}
#[cfg(all(test, feature = "mmap"))]
mod zero_copy_tests {
use super::*;
/// Where `File::open` reads metadata from: `Some(true)` for the file's
/// own mapping, `Some(false)` for a private copy-on-write mapping.
fn reads_from_the_mapping(f: &File) -> Option<bool> {
let Backing::Mmap(r) = &f.data.backing else {
return None;
};
let mapped = r.as_bytes()[f.data.base..].as_ptr();
match &f.data.patched {
None => Some(std::ptr::eq(f.as_bytes().as_ptr(), mapped)),
Some(p) => {
assert!(p.is_mapped(), "the image went into a heap copy of the file");
Some(false)
}
}
}
#[test]
fn a_file_without_a_cache_image_is_read_from_the_mapping() {
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("plain.h5");
let mut b = crate::FileBuilder::new();
b.create_dataset("d").with_f64_data(&[1.0, 2.0]);
b.write(&path).unwrap();
let f = File::open(&path).unwrap();
assert_eq!(reads_from_the_mapping(&f), Some(true));
assert_eq!(f.dataset("d").unwrap().read_f64().unwrap(), [1.0, 2.0]);
}
#[test]
fn a_cache_image_goes_into_a_copy_on_write_mapping() {
let path = std::path::Path::new(env!("CARGO_MANIFEST_DIR"))
.join("tests/fixtures/h5clear_mdc_image.h5");
let f = File::open(path).unwrap();
assert_eq!(reads_from_the_mapping(&f), Some(false));
}
}
+139
View File
@@ -0,0 +1,139 @@
//! Opening a file with a metadata cache image must not copy the file.
//!
//! The image's entries are laid over the file's bytes in a private
//! copy-on-write mapping (`File::open`, `MmapFile::open`,
//! `LazyFile::open_mmap`), so only the pages they land on are copied. The
//! first implementation copied the whole file onto the heap at open: a
//! 1 GiB file that takes a few KB on disk needed 2 GB of memory, and an
//! 8 GiB one aborted the process. Here libhdf5 itself (the library h5py
//! bundles, through ctypes: `H5Pset_mdc_image_config`) adds an image to a
//! 1 GiB sparse file, and the process's resident memory must stay far below
//! the file's size while each opener lists the file and reads its small
//! dataset.
//!
//! One test in its own binary, so no other test's allocations land in the
//! measurement. Linux only (it reads `VmRSS` from `/proc/self/status`).
//! Skipped when python3 with h5py is unavailable, unless
//! `CLAWHDF5_REQUIRE_INTEROP=1`.
#![cfg(target_os = "linux")]
use std::path::Path;
use std::process::Command;
use clawhdf5::{File, LazyFile, MmapFile};
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn python_available() -> bool {
Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
fn rss_bytes() -> u64 {
let status = std::fs::read_to_string("/proc/self/status").unwrap();
let line = status.lines().find(|l| l.starts_with("VmRSS:")).unwrap();
let kb: u64 = line.split_whitespace().nth(1).unwrap().parse().unwrap();
kb * 1024
}
/// A 1 GiB sparse file: `/big`, 2^27 `f8` with only its last element
/// written, `/small` = 0..10, with a metadata cache image added by libhdf5.
fn make_file(path: &Path) {
let script = format!(
r#"
import ctypes, glob, os, h5py, numpy as np
path = "{path}"
libs = glob.glob(os.path.join(os.path.dirname(h5py.__file__), os.pardir, "h5py.libs", "libhdf5-*.so*"))
assert libs, "no libhdf5 bundled with h5py"
lib = ctypes.CDLL(libs[0])
class Cfg(ctypes.Structure):
_fields_ = [("version", ctypes.c_int), ("generate_image", ctypes.c_bool),
("save_resize_status", ctypes.c_bool), ("entry_ageout", ctypes.c_int)]
with h5py.File(path, "w", libver="latest") as f:
d = f.create_dataset("big", shape=(2**27,), dtype="f8")
d[-1] = 7.5
f.create_dataset("small", data=np.arange(10, dtype="<i4"))
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
fapl.set_libver_bounds(h5py.h5f.LIBVER_LATEST, h5py.h5f.LIBVER_LATEST)
cfg = Cfg(1, True, False, -1)
assert lib.H5Pset_mdc_image_config(ctypes.c_int64(fapl.id), ctypes.byref(cfg)) >= 0
f = h5py.File(h5py.h5f.open(path.encode(), h5py.h5f.ACC_RDWR, fapl=fapl))
f["small"][()]; f["big"].shape
f.close()
assert os.path.getsize(path) >= 2**30
with open(path, "rb") as fh:
fh.seek(-(1 << 20), 2)
assert b"MDCI" in fh.read(), "libhdf5 wrote no cache image"
with h5py.File(path, "r") as f:
assert list(f["small"][()]) == list(range(10))
"#,
path = path.display()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.expect("failed to run python");
assert!(
out.status.success(),
"python failed:\n{}",
String::from_utf8_lossy(&out.stderr)
);
}
#[test]
fn a_cache_image_does_not_copy_the_file() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
);
eprintln!("SKIP: python3 with h5py not available");
return;
}
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("sparse_image.h5");
make_file(&path);
const LIMIT: u64 = 256 << 20;
let small: Vec<i32> = (0..10).collect();
let before = rss_bytes();
{
let f = File::open(&path).unwrap();
let mut names = f.root().datasets().unwrap();
names.sort();
assert_eq!(names, ["big", "small"]);
assert_eq!(f.dataset("small").unwrap().read_i32().unwrap(), small);
assert_eq!(f.dataset("big").unwrap().shape().unwrap(), [1 << 27]);
let grew = rss_bytes().saturating_sub(before);
assert!(
grew < LIMIT,
"File::open: resident memory grew {grew} bytes"
);
}
let before = rss_bytes();
{
let f = MmapFile::open(&path).unwrap();
assert_eq!(f.dataset("small").unwrap().read_i32().unwrap(), small);
let grew = rss_bytes().saturating_sub(before);
assert!(
grew < LIMIT,
"MmapFile::open: resident memory grew {grew} bytes"
);
}
let before = rss_bytes();
{
let f = LazyFile::open_mmap(&path).unwrap();
assert_eq!(f.dataset("small").unwrap().read_i32().unwrap(), small);
let grew = rss_bytes().saturating_sub(before);
assert!(
grew < LIMIT,
"LazyFile::open_mmap: resident memory grew {grew} bytes"
);
}
}
Binary file not shown.
@@ -42,8 +42,9 @@ macro_rules! skip_if_no_python {
/// Runs `body` (Python, with `h5py`, `numpy as np`, `struct` imported and
/// `d` the output directory) and then, for every `NAME.h5` it wrote,
/// prints `NAME ok` when h5py opens and reads dataset `d` and `NAME ERROR`
/// otherwise. Returns those lines, sorted.
/// prints `NAME ok` when h5py opens and reads dataset `d` (or the dataset
/// the body names in `DSET`) and `NAME ERROR` otherwise. Returns those
/// lines, sorted.
fn h5py_verdicts(dir: &Path, body: &str) -> Vec<String> {
let script = format!(
r#"
@@ -54,7 +55,7 @@ for path in sorted(glob.glob(os.path.join(d, "*.h5"))):
name = os.path.basename(path)[:-3]
try:
with h5py.File(path, "r") as f:
f["d"][()]
f[globals().get("DSET", "d")][()]
print(name, "ok")
except Exception:
print(name, "ERROR")
@@ -78,14 +79,14 @@ for path in sorted(glob.glob(os.path.join(d, "*.h5"))):
lines
}
/// Whether clawhdf5 opens and reads dataset `d` of `path` (as raw bytes of
/// whatever type it has).
fn clawhdf5_reads(path: &Path) -> Result<(), String> {
/// Whether clawhdf5 opens and reads dataset `dset` of `path` (as raw bytes
/// of whatever type it has).
fn clawhdf5_reads(path: &Path, dset: &str) -> Result<(), String> {
let file = File::open(path).map_err(|e| format!("open: {e}"))?;
let ds = file.dataset("d").map_err(|e| format!("dataset: {e}"))?;
let ds = file.dataset(dset).map_err(|e| format!("dataset: {e}"))?;
ds.dtype().map_err(|e| format!("dtype: {e}"))?;
ds.shape().map_err(|e| format!("shape: {e}"))?;
file.read_multi(&["d"])
file.read_multi(&[dset])
.map(|_| ())
.map_err(|e| format!("read: {e}"))
}
@@ -93,10 +94,15 @@ fn clawhdf5_reads(path: &Path) -> Result<(), String> {
/// h5py's verdict for each file must be `expected`, and clawhdf5 must read
/// exactly the files h5py reads.
fn assert_agrees_with_h5py(dir: &Path, verdicts: &[String], expected: &[&str]) {
assert_agrees_with_h5py_on(dir, verdicts, expected, "d");
}
/// [`assert_agrees_with_h5py`] reading dataset `dset`.
fn assert_agrees_with_h5py_on(dir: &Path, verdicts: &[String], expected: &[&str], dset: &str) {
assert_eq!(verdicts, expected, "h5py's view changed");
for line in verdicts {
let (name, verdict) = line.split_once(' ').unwrap();
let ours = clawhdf5_reads(&dir.join(format!("{name}.h5")));
let ours = clawhdf5_reads(&dir.join(format!("{name}.h5")), dset);
match verdict {
"ok" => assert!(ours.is_ok(), "{name}: h5py reads it, we fail: {ours:?}"),
_ => assert!(ours.is_err(), "{name}: h5py refuses it, we read it"),
@@ -244,7 +250,7 @@ for libver in ("earliest", "latest"):
],
);
for name in ["earliest_size2", "latest_size8"] {
let err = clawhdf5_reads(&dir.path().join(format!("{name}.h5"))).unwrap_err();
let err = clawhdf5_reads(&dir.path().join(format!("{name}.h5")), "d").unwrap_err();
assert!(
err.contains("stored datatype size in chunk layout"),
"{name}: {err}"
@@ -508,3 +514,205 @@ for libver in ("earliest", "latest"):
);
}
}
/// cve-2024-32624 `/Dset_OBJREF`: a dataspace whose element count times the
/// element size overflows 64 bits. libhdf5 refuses to open the dataset
/// ("size of dataset's storage overflowed"); `File::dataset` used to open it
/// and report its shape, and only reading failed. The same for storage that
/// runs past the end of the file ("invalid dataset size, likely file
/// corruption").
#[test]
fn dataset_storage_libhdf5_refuses_at_open_is_refused_at_open() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
// A contiguous int64 dataset of 3 elements, libver earliest (a version 1
// dataspace holding the size and the maximum size). "overflow" makes
// both 2^62 + 2 (x 8 bytes overflows); "pasteof" makes both 1000 (the
// storage runs past the end of the file).
let verdicts = h5py_verdicts(
dir.path(),
r#"
good = os.path.join(d, "good.h5")
with h5py.File(good, "w", libver="earliest") as f:
f.create_dataset("d", data=np.array([7, 8, 9], dtype="<i8"))
data = bytearray(open(good, "rb").read())
three = struct.pack("<QQ", 3, 3)
at = data.find(three)
assert at > 0 and data.find(three, at + 1) < 0
for name, n in (("overflow", (1 << 62) + 2), ("pasteof", 1000)):
bad = bytearray(data)
bad[at:at + 16] = struct.pack("<QQ", n, n)
open(os.path.join(d, f"{name}.h5"), "wb").write(bad)
"#,
);
assert_agrees_with_h5py(
dir.path(),
&verdicts,
&["good ok", "overflow ERROR", "pasteof ERROR"],
);
for name in ["overflow", "pasteof"] {
let path = dir.path().join(format!("{name}.h5"));
let file = File::open(&path).unwrap();
let err = file.dataset("d").unwrap_err();
assert!(
err.to_string().contains("invalid dataset storage"),
"{name}: {err}"
);
assert!(file.root().dataset("d").is_err(), "{name}");
let mm = clawhdf5::MmapFile::open(&path).unwrap();
assert!(mm.dataset("d").is_err(), "{name}: MmapFile");
}
}
/// libhdf5 decodes the superblock extension's messages when it opens a file
/// and refuses the file when one does not decode: cve-2020-10810 (a File
/// Space Info message too short for what it announces), cve-2020-10812 (a
/// metadata cache image past the end of the file). We did not look at those
/// messages and opened such files.
#[test]
fn superblock_extension_messages_libhdf5_refuses_are_refused() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let mdc = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures/h5clear_mdc_image.h5");
let body = format!(
r#"
{FIX_OHDR_PY}
def ext_addr(buf):
assert buf[8] in (2, 3)
return int.from_bytes(buf[20:28], "little")
def save(name, buf):
open(os.path.join(d, name + ".h5"), "wb").write(buf)
# A paged file: its superblock extension holds a File Space Info message
# (version 1, strategy page, not persisting, threshold 1, page size 4096).
DSET = "DSET"
good = os.path.join(d, "paged_good.h5")
with h5py.File(good, "w", libver="latest", fs_strategy="page", fs_page_size=4096) as f:
f.create_dataset("DSET", data=np.arange(10, dtype="<i4"))
data = bytearray(open(good, "rb").read())
ext = ext_addr(data)
at = data.find(struct.pack("<QQ", 1, 4096), ext)
assert at > ext
# Page size 256 (under libhdf5's minimum of 512).
bad = bytearray(data); bad[at + 8:at + 16] = struct.pack("<Q", 256); fix_ohdr(bad, ext)
save("paged_small_page", bad)
# Persisting free space, without the manager addresses that then follow.
bad = bytearray(data); bad[at - 1] = 1; fix_ohdr(bad, ext)
save("paged_persist_short", bad)
# h5clear_mdc_image.h5 (from libhdf5's tests): a metadata cache image, the
# root group's header only in the image. The Metadata Cache Image message
# (type 0x18) records the image's address and length.
data = bytearray(open(r"{mdc}", "rb").read())
save("mdc_good", data)
ext = ext_addr(data)
at = data.find(bytes([0x18, 17, 0]), ext)
assert at > ext
length = at + 5 + 8
bad = bytearray(data); bad[length:length + 8] = struct.pack("<Q", 1 << 28); fix_ohdr(bad, ext)
save("mdc_past_eof", bad)
"#,
mdc = mdc.display()
);
let verdicts = h5py_verdicts(dir.path(), &body);
assert_agrees_with_h5py_on(
dir.path(),
&verdicts,
&[
"mdc_good ok",
"mdc_past_eof ERROR",
"paged_good ok",
"paged_persist_short ERROR",
"paged_small_page ERROR",
],
"DSET",
);
}
/// An unfiltered chunk the index records at less than the chunk's size
/// (`cve-2025-44904`): HDF5 2.0 fills the rest of the chunk with whatever
/// its buffer held, and later libhdf5 releases refuse it ("incorrect chunk
/// size returned from index for unfiltered chunk"); we used to read the
/// rest as zeros, and now refuse it.
#[test]
fn unfiltered_chunk_of_the_wrong_size_is_refused() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
run_python(
dir.path(),
r#"
good = os.path.join(d, "good.h5")
with h5py.File(good, "w", libver="earliest") as f:
f.create_dataset("d", data=np.arange(100, dtype="<i4"), chunks=(37,), fillvalue=-1)
data = bytearray(open(good, "rb").read())
tree = data.find(b"TREE")
while data[tree + 4] != 1: # the chunk index, not the root group's B-tree
tree = data.find(b"TREE", tree + 1)
assert tree > 0 and data[tree + 5] == 0
key = tree + 8 + 16 # the first chunk's key: size, filter mask, offsets
second = key + 24 + 8 # a key (4 + 4 + 2 x 8 bytes), then a child address
assert struct.unpack_from("<IIQQ", data, second) == (148, 0, 37, 0)
bad = bytearray(data); struct.pack_into("<I", bad, second, 100)
open(os.path.join(d, "short_chunk.h5"), "wb").write(bad)
"#,
);
assert_eq!(
clawhdf5_reads(&dir.path().join("good.h5"), "d"),
Ok(()),
"good"
);
let err = clawhdf5_reads(&dir.path().join("short_chunk.h5"), "d").unwrap_err();
assert!(err.starts_with("read:"), "short_chunk: {err}");
}
/// A v1 B-tree chunk key whose element-size coordinate is not 0. libhdf5
/// looks each chunk up with that coordinate set to 0 (`H5B_find`,
/// `H5D__btree_cmp3`, `H5D__btree_found`): in a 1-D dataset it still finds
/// the chunk, in a 2-D one it does not and reads fill values
/// (`cve-2025-44905` `/Shuffle_float_data_le`). Both must read exactly what
/// h5py reads: the 1-D file was refused, and before that the 2-D chunk was
/// read as if its key were well formed.
#[test]
fn chunk_keys_with_an_element_offset_read_as_libhdf5_reads_them() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
run_python(
dir.path(),
r#"
def corrupt(name, shape, chunks, which, element_offset):
path = os.path.join(d, name + ".h5")
with h5py.File(path, "w", libver="earliest") as f:
f.create_dataset("d", data=np.arange(np.prod(shape), dtype="<i4").reshape(shape),
chunks=chunks, fillvalue=-1)
data = bytearray(open(path, "rb").read())
tree = data.find(b"TREE")
while data[tree + 4] != 1: # the chunk index, not the root group's B-tree
tree = data.find(b"TREE", tree + 1)
assert tree > 0 and data[tree + 5] == 0
ndims = len(shape) + 1
key_size = 8 + 8 * ndims
key = tree + 8 + 16 + which * (key_size + 8)
last = key + 8 + 8 * (ndims - 1)
assert struct.unpack_from("<Q", data, last) == (0,)
struct.pack_into("<Q", data, last, element_offset)
open(path, "wb").write(data)
with h5py.File(path, "r") as f:
open(path + ".txt", "w").write(" ".join(map(str, f["d"][()].ravel())))
corrupt("one_d", (100,), (37,), 1, 4096)
corrupt("two_d", (4, 6), (2, 3), 1, 4)
corrupt("two_d_first", (4, 6), (2, 3), 0, 8)
"#,
);
for name in ["one_d", "two_d", "two_d_first"] {
let path = dir.path().join(format!("{name}.h5"));
let expected: Vec<i32> = std::fs::read_to_string(dir.path().join(format!("{name}.h5.txt")))
.unwrap()
.split_whitespace()
.map(|v| v.parse().unwrap())
.collect();
let file = File::open(&path).unwrap();
let got = file.dataset("d").unwrap().read_i32();
assert_eq!(got.as_deref().ok(), Some(&expected[..]), "{name}: {got:?}");
}
}
@@ -0,0 +1,94 @@
//! Files with a metadata cache image read as libhdf5 reads them.
//!
//! `fixtures/h5clear_mdc_image.h5` comes from libhdf5's tool tests
//! (`tools/test/testfiles`): written with a metadata cache image, so its
//! superblock extension points at an image block holding the file's
//! metadata cache entries, and the root group's object header exists only
//! there — the header's own address in the file is zeros. It holds one
//! dataset, `/DSET`, 50 x 100 `int32` with `DSET[i][j] = i * j` (h5py
//! 3.16 / HDF5 2.0 reads that). Until 2026-09-26 every reader failed on the
//! root group: `InvalidObjectHeaderVersion(0)`.
use std::path::PathBuf;
use clawhdf5::{File, LazyFile, MmapFile};
fn fixture() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures/h5clear_mdc_image.h5")
}
fn expected() -> Vec<i32> {
(0..50).flat_map(|i| (0..100).map(move |j| i * j)).collect()
}
#[test]
fn file_reads_through_the_cache_image() {
let file = File::open(fixture()).unwrap();
assert_eq!(file.root().datasets().unwrap(), ["DSET"]);
let ds = file.dataset("DSET").unwrap();
assert_eq!(ds.shape().unwrap(), [50, 100]);
assert_eq!(ds.read_i32().unwrap(), expected());
let file = File::from_bytes(std::fs::read(fixture()).unwrap()).unwrap();
assert_eq!(
file.dataset("DSET").unwrap().read_i32().unwrap(),
expected()
);
}
#[test]
fn mmap_and_lazy_files_read_through_the_cache_image() {
let mm = MmapFile::open(fixture()).unwrap();
assert_eq!(mm.dataset("DSET").unwrap().read_i32().unwrap(), expected());
let lazy = LazyFile::from_bytes(std::fs::read(fixture()).unwrap()).unwrap();
assert_eq!(
lazy.dataset("DSET").unwrap().read_i32().unwrap(),
expected()
);
}
/// A file whose cache image libhdf5 cannot load: libhdf5 opens it, and the
/// first metadata read fails with the image's error (h5py:
/// `Unable to get group info (Bad metadata cache image header signature)`);
/// later reads get the file's own bytes, which here are stale (the root
/// group's header is zeros: "bad object header version number"). The
/// openers agree on the open and the failure: `File` and `MmapFile` open
/// and fail every object lookup with the image's error (never reading the
/// stale bytes), `LazyFile` reads the root group's header at open and so
/// fails there. `File::open` used to refuse the file, while the
/// conformance probe reported it as h5py does.
#[test]
fn an_image_libhdf5_cannot_load_fails_every_object() {
let mut bytes = std::fs::read(fixture()).unwrap();
let at = bytes.windows(4).position(|w| w == b"MDCI").unwrap();
bytes[at] = b'X';
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("bad_image.h5");
std::fs::write(&path, &bytes).unwrap();
let is_image_error = |e: clawhdf5::Error| {
matches!(
e,
clawhdf5::Error::Format(clawhdf5_format::error::FormatError::InvalidCacheImage(
"bad metadata cache image header signature"
))
)
};
for file in [
File::open(&path).unwrap(),
File::from_bytes(bytes.clone()).unwrap(),
] {
assert!(is_image_error(file.root().datasets().unwrap_err()));
assert!(is_image_error(file.root().attrs().unwrap_err()));
assert!(is_image_error(file.dataset("DSET").unwrap_err()));
assert!(is_image_error(file.group("/").map(|_| ()).unwrap_err()));
assert!(is_image_error(file.dataset_at(96).unwrap_err()));
}
let mm = MmapFile::open(&path).unwrap();
assert!(is_image_error(mm.root().datasets().unwrap_err()));
assert!(is_image_error(mm.dataset("DSET").unwrap_err()));
assert!(is_image_error(mm.group("/").map(|_| ()).unwrap_err()));
assert!(is_image_error(
LazyFile::from_bytes(bytes).map(|_| ()).unwrap_err()
));
}
@@ -0,0 +1,151 @@
//! Scale-offset datasets written by h5py (libhdf5) read exactly as h5py
//! reads them.
//!
//! h5py writes the whole matrix the filter has: every integer type (`i1` ..
//! `u8`) and `f4`/`f8`, both byte orders, with and without a fill value
//! (the type's minimum, maximum, or another value), with random, constant,
//! all-fill and extreme data, and every interesting `scaleoffset` setting
//! (0 = let libhdf5 choose the bits, a few bits, full width less one, full
//! width; decimal scale factors 0..7 for floats) — about 1480 datasets. For
//! each one h5py's decoded values are stored uncompressed next to it, and
//! the raw bytes clawhdf5 decodes must equal them.
//!
//! Until 2026-09-26 clawhdf5 silently returned wrong values for 332 of
//! these, in every release that decoded scale-offset: ordinary `u8`, `u4`,
//! `i8` and `f4` data with a wide range (where libhdf5 stores the elements
//! as they are, at full width), chunks whose `minval` field was recorded at
//! another size than 8 bytes, and chunks with `minbits` 0 and a fill value.
//!
//! Skipped when python3 with h5py is unavailable, unless
//! `CLAWHDF5_REQUIRE_INTEROP=1`.
use std::process::Command;
use clawhdf5::File;
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn python_available() -> bool {
Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
const GENERATE: &str = r#"
import sys, h5py, numpy as np
rng = np.random.default_rng(7)
out, expect = sys.argv[1], sys.argv[2]
made = []
with h5py.File(out, "w") as f:
def make(name, a, so, fv, desc):
try:
d = f.create_dataset(name, data=a, chunks=(16,), scaleoffset=so, fillvalue=fv)
except Exception:
return # a combination libhdf5 refuses to write
d.attrs["desc"] = desc
made.append(name)
i = 0
for dt in ["i1", "u1", "i2", "u2", "i4", "u4", "i8", "u8"]:
for bo in "<>":
t = np.dtype(bo + dt)
info = np.iinfo(t)
for so in [0, 1, 3, 8 * t.itemsize - 1, 8 * t.itemsize]:
for fill in [None, "min", "max", "mid"]:
for pat in ["rand", "const", "allfill", "extreme"]:
n = 64
fv = {None: None, "min": info.min, "max": info.max, "mid": t.type(7)}[fill]
if pat == "rand":
lo = max(info.min, -50) if so else info.min
hi = min(info.max, 50) if so else info.max
a = rng.integers(lo, hi, size=n, endpoint=True,
dtype=np.int64 if t.kind == "i" else np.uint64).astype(t)
elif pat == "const":
a = np.full(n, 3, t)
elif pat == "extreme":
a = np.array([info.min, info.max] * (n // 2), t)
else:
if fv is None:
continue
a = np.full(n, fv, t)
make(f"d{i}", a, so, fv, f"{bo}{dt} so={so} fill={fill} pat={pat}")
i += 1
for dt in ["f4", "f8"]:
for bo in "<>":
t = np.dtype(bo + dt)
for so in [0, 1, 2, 4, 7]:
for fill in [None, -1.5, 0.0]:
for pat in ["rand", "const", "allfill", "neg", "big"]:
n = 50
if pat == "rand":
a = rng.normal(size=n).astype(t) * 10
elif pat == "const":
a = np.full(n, 2.25, t)
elif pat == "neg":
a = -np.abs(rng.normal(size=n)).astype(t) * 1000
elif pat == "big":
a = rng.normal(size=n).astype(t) * 1e6
else:
if fill is None:
continue
a = np.full(n, fill, t)
make(f"d{i}", a, so, fill, f"{bo}{dt} so={so} fill={fill} pat={pat}")
i += 1
# What libhdf5 decodes, stored uncompressed in the same datatype.
with h5py.File(out, "r") as f, h5py.File(expect, "w") as e:
for name in made:
e.create_dataset(name, data=f[name][()])
print(len(made))
"#;
#[test]
fn every_scale_offset_dataset_reads_as_h5py_reads_it() {
if !python_available() {
assert!(
!std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
);
eprintln!("SKIP: python3 with h5py not available");
return;
}
let dir = tempfile::tempdir().unwrap();
let written = dir.path().join("scaleoffset.h5");
let expected = dir.path().join("expected.h5");
let out = Command::new(python())
.args(["-c", GENERATE])
.arg(&written)
.arg(&expected)
.output()
.expect("failed to run python");
assert!(
out.status.success(),
"python failed:\n{}",
String::from_utf8_lossy(&out.stderr)
);
let count: usize = String::from_utf8_lossy(&out.stdout).trim().parse().unwrap();
assert!(count > 1400, "only {count} datasets written");
let file = File::open(&written).unwrap();
let reference = File::open(&expected).unwrap();
let mut names = reference.root().datasets().unwrap();
names.sort();
assert_eq!(names.len(), count);
let mut wrong = Vec::new();
for name in &names {
let want = reference.read_multi(&[name]).unwrap().remove(0);
let got = file.read_multi(&[name]).map(|mut v| v.remove(0));
if got.as_ref().ok() != Some(&want) {
let desc = file.dataset(name).unwrap().attrs().unwrap().remove("desc");
wrong.push(format!("{name} {desc:?}: {:?}", got.map(|g| g.len())));
}
}
assert!(
wrong.is_empty(),
"{} of {count} scale-offset datasets differ from h5py:\n{}",
wrong.len(),
wrong.join("\n")
);
}
+23
View File
@@ -486,3 +486,26 @@ fn vds_libhdf5_test_files() {
vec![5, 10, 10]
);
}
/// A source file with a metadata cache image (libhdf5's
/// `h5clear_mdc_image.h5`, whose root group's header exists only in the
/// image) is read through the image, as libhdf5 opens it. Source files were
/// read from their own bytes, and this one failed on the root group.
#[test]
fn vds_source_file_with_a_metadata_cache_image() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let fixture = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures/h5clear_mdc_image.h5");
std::fs::copy(&fixture, dir.path().join("src.h5")).unwrap();
generate(
dir.path(),
r#"
with h5py.File("vds.h5", "w", libver="latest") as f:
lay = h5py.VirtualLayout(shape=(3, 100), dtype="i4")
lay[0:3, :] = h5py.VirtualSource("src.h5", "DSET", shape=(50, 100))[5:8, :]
f.create_virtual_dataset("v", lay)
expect("vds.h5", "v", "image_source")
"#,
);
assert_matches_libhdf5(dir.path(), "vds.h5", "v", "image_source");
}
+79 -6
View File
@@ -74,6 +74,43 @@ tank with `concurrent_read` against h5py 3.16 / HDF5 2.0 (`BENCHMARKS.md`,
`CHANGELOG.md`). The chunked-read scaling item above is still open.
Values are correct; this is speed only.
## Scale-offset data read back wrong values
**Status:** fixed 2026-09-26, after v2.7.0. **Every release that decoded
the scale-offset filter (v2.2.0 to v2.7.0) is affected**, on ordinary
files h5py writes, with no error.
Found by the review of the 2026-09-26 conformance work: h5py wrote 1480
scale-offset datasets (every integer type `i1` .. `u8`, `f4` and `f8`,
little- and big-endian, with no fill value and with the type's minimum,
maximum or another value as fill, random, constant, all-fill and extreme
data, `scaleoffset` 0, 1, 3, full width less one and full width for
integers, decimal scale factors 0, 1, 2, 4 and 7 for floats). v2.7.0's
decoder read 332 of them differently from h5py: **151 returned wrong
values with no error**, 181 failed to read.
| Case | Datasets | v2.7.0 |
|---|---|---|
| integer, `scaleoffset=0` (libhdf5 picks the bits), data spanning most of the type's range | 82 | wrong values |
| integer, `scaleoffset` = full width: `i4`/`u4` (10), `i8`/`u8` (41) | 51 | wrong values |
| integer, `scaleoffset` = full width, the other datasets (every `i1`..`u2` one, most `i4`/`u4`, some `i8`/`u8`) | 181 | "truncated minval" / "implausible minbits" |
| `f4` D-scale, factor 4 or 7, values up to about 10^6 | 18 | wrong values |
In every case libhdf5 stored a chunk at full width (`minbits` equal to the
type's width): the chunk then holds the elements as they are, and they
were decoded as offsets from `minval`. Two more differences were found on
crafted files and fixed with them: the packed codes start at byte 21
whatever size the chunk records for `minval` (`cve-2025-44905`
`/Scale_offset_short_data_be`), and a chunk with `minbits` 0 and a fill
value is all fill values (it read as `minval`).
**Fix:** `clawhdf5_format::filters` decodes a scale-offset chunk as
`H5Z__filter_scaleoffset` does. **Test:** the whole matrix is
`crates/clawhdf5/tests/scaleoffset_interop.rs`, generated by h5py at test
time and compared dataset by dataset; on v2.7.0's decoder it reports the
332. **Existing data:** the files were always right; only reads were
wrong, so re-reading with a fixed build gives the correct values.
## Silent wrong data found by the 2026-09-25 HDF5 audit
**Status:** fixed after v2.7.0 (2026-09-25). **Every release up
@@ -213,9 +250,36 @@ fill-value item that did is fixed).
its datatype message stores (12 with 4-byte offsets), and the global
heap is read with libhdf5's header padding
(`crates/clawhdf5/tests/vl_offset4_interop.rs`).
- Metadata cache images are not supported.
- Metadata cache images are not supported. **Fixed 2026-09-26:** the
image is applied at open, as libhdf5 loads it over the file's metadata
(`clawhdf5_format::superblock_ext`); `h5clear_mdc_image.h5` reads
(`crates/clawhdf5/tests/metadata_cache_image.rs`), without copying the
file (a private copy-on-write mapping takes the image's entries;
`tests/cache_image_memory.rs`). A file whose image libhdf5 cannot load
(`cve-2025-6269-*`, `cve-2025-6516`) opens, as in libhdf5, and every
object lookup fails with the image's error. Differences from libhdf5
that remain: libhdf5 fails only the first metadata read and then reads
the file's own (possibly stale) metadata, where we keep failing; an
image entry that runs past the end of file is refused (libhdf5 checks
only its start); a flush-dependency parent flag is checked against the
child count as libhdf5's debug build checks it (HDF5 2.0 release
builds refuse every entry that has children, even in images they
wrote); the superblock extension's driver-info and shared-message table
messages are not decoded at open.
- x87 long double and binary128 are refused.
- N-Bit on 64-bit scale-offset data and some N-Bit parameter layouts fail.
**Not our bug (checked 2026-09-26):** both are corrupt files HDF5 2.0
reads only by reading past a buffer. `cve-2025-2308`'s
`/Scale_offset_long_long_data_le` has scale-offset codes that run past
the end of the chunk (libhdf5's develop branch refuses it, "Buffer too
short"); `bad_nbit_parms_walk.h5` has an N-Bit parameter list one value
short (libhdf5's own `test_filter_bad_params` in `test/dsets.c` now
requires that read to fail). We refuse both; `CONFORMANCE.md` lists them
under *Known not-our-bug*. Scale-offset did decode three cases
differently from libhdf5 (codes after a `minval` of recorded size other
than 8, `minbits` 0 with a fill value, full-width `minbits`): fixed
2026-09-26, and the full-width case was silent wrong data on ordinary
h5py files (see *Scale-offset data read back wrong values* above).
- **Filters:** blosc, blosc2, bitshuffle, bzip2, LZF and zfp are not
implemented. **Fixed 2026-09-26** for LZF (default-on `lzf` feature),
bitshuffle, bzip2 and Blosc 1 (`bitshuffle`, `bzip2`, `blosc`, or
@@ -235,7 +299,11 @@ fill-value item that did is fixed).
read with zeros for the missing bytes** (any filter; found reviewing the plugin
filters). **Fixed 2026-09-26:** it is an error naming the chunk. A corrupt
chunk must never read as zeros. Unfiltered chunks are read at their stored
size and are not checked this way.
size and are not checked this way. **Fixed 2026-09-26** for unfiltered
chunks too: in a dataset without filters a chunk the index records at
other than the chunk's size is refused, as libhdf5's develop branch
refuses it (`cve-2025-44904`, where HDF5 2.0 fills the rest from its
buffer).
- **Crash:** a hostile Blosc chunk (frame size below its header) panicked in
builds with overflow checks. **Fixed 2026-09-26**; the new decoders are
fuzzed in the unit tests.
@@ -248,13 +316,18 @@ fill-value item that did is fixed).
refused. 17 of the 18 now
fail as in libhdf5 (conformance on tank, `conformance/run.sh --no-fetch`,
2026-09-26: 571 of 697 ok). Still read where libhdf5 refuses:
- `cve-2024-32624.h5` `/Dset_OBJREF`: a dataspace whose storage size
- ~~`cve-2024-32624.h5` `/Dset_OBJREF`: a dataspace whose storage size
overflows 64 bits. `File::dataset` and `shape()` succeed (libhdf5
refuses at open); reading the values fails.
- `cve-2020-10810.h5`, `cve-2020-10812.h5` (whole files libhdf5 cannot
refuses at open); reading the values fails.~~ **Fixed 2026-09-26:**
`File::dataset` (and `MmapFile`, `LazyFile`) refuse it at open
(`FormatError::InvalidDatasetStorage`), as they do contiguous storage
past the end of the file.
- ~~`cve-2020-10810.h5`, `cve-2020-10812.h5` (whole files libhdf5 cannot
open, not among the 18): libhdf5 decodes the superblock extension's File
Space Info and metadata-cache-image messages at open and refuses these
files; we do not decode those messages at open.
files; we do not decode those messages at open.~~ **Fixed 2026-09-26:**
the superblock extension is decoded at open with libhdf5's checks, and
both files are refused.
- Deliberately not refused, because clawhdf5 up to v2.7.0 wrote them: a
float sign bit position outside the type, and a size-0 string type.
- Not refused because current libhdf5 reads it though HDF5 2.0.0