Merge branch 'feat/p2-vl-strings' into feat/p2-perf-coverage

# Conflicts:
#	CHANGELOG.md
This commit is contained in:
osobh
2026-09-26 09:10:35 -05:00
24 changed files with 2405 additions and 300 deletions
+87 -2
View File
@@ -54,6 +54,91 @@
whole-row hyperslabs, points, empty selections; every type, both byte whole-row hyperslabs, points, empty selections; every type, both byte
orders, ranks 1-4). orders, ranks 1-4).
### Variable-length data (2026-09-26)
- **VL values in files with 4-byte offsets** (`sizeof_addr = 4`). A VL
string attribute came back as `AttrValue::Raw`, a VL member of a compound
failed with `GlobalHeapObjectNotFound`, and VL datasets failed with a
size mismatch. Two causes: `Datatype::type_size()` reported 16 for every
VL type (the element is 4 + offset size + 4 bytes: 12 here), and the
global heap was parsed without the padding libhdf5 puts after its
collection and object headers (`H5HG_SIZEOF_HDR`/`H5HG_SIZEOF_OBJHDR`
round up to 8), so with 4-byte lengths every object was looked up 4
bytes early. `Datatype::VariableLength` now carries the element `size`
stored in the datatype message (**breaking** for code that builds or
exhaustively destructures that variant; patterns with `..` are
unaffected), and it is written back as stored. Tested against h5py
(`crates/clawhdf5/tests/vl_offset4_interop.rs`).
- **Wrong data: VL strings with an embedded NUL, and VL elements whose heap
object has the wrong size.** libhdf5 hands VL strings over as C strings,
so h5py reads `"a\0b"` as `"a"`; `read_vl_strings` returned the NUL and
what followed. An element whose heap object is not exactly
`length × base size` bytes is refused by libhdf5 ("Expected global heap
object size does not match"); we returned the object cut or padded to
the length. Both now behave as libhdf5, and a heap address of 0 is a null
element (empty) whatever its length. The new
`clawhdf5_format::vl_data::VlResolver` does this and parses each global
heap collection once per read: `read_vl_strings` parsed the whole
collection again for every element. `vl_data::check_element_size` refuses
a VL datatype whose stored size is not 4 + offset size + 4 (libhdf5
ignores the stored size). The conformance probe resolves VL elements
with `VlResolver` too; conformance unchanged at 575 of 697.
- **VL data through the facade.** VL-string datasets (h5py's default `str`
dtype) failed `read_string` with "type mismatch: expected String, got
VariableLength". `Dataset::read_string` now reads fixed- and
variable-length strings; new `read_string_bytes` (a VL string's exact
bytes, as h5py's `Dataset[()]` returns them), `read_string_selection`,
`read_vlen::<T>()` / `read_vlen_selection::<T>()` for VL sequences of
numbers (`T` = `f64`, `f32`, `i64`, `i32`, `u64`; converted like the
other typed readers), and `File::decode_strings` / `decode_string_bytes`
/ `decode_vlen` for VL values in compound fields and `AttrValue::Raw`
attributes. `MmapDataset` and `LazyDataset` gain `read_string` for VL
strings, `read_string_bytes` and `read_vlen`. Checked against h5py with
8- and 4-byte offsets: scalar and 1-/2-D, ASCII and UTF-8, empty strings,
contiguous, compact, chunked with gzip/shuffle, never-written and
partly written chunks, hyperslab selections, VL members of compound
datasets and attributes (`crates/clawhdf5/tests/vl_data_interop.rs`).
NetCDF-4 `string` variables now read through
`clawhdf5_netcdf4::Variable::read_string` (checked against netCDF4-python
in `crates/clawhdf5-netcdf4/tests/interop_tests.rs`).
- **Crafted global heaps could exhaust memory.** `VlResolver` kept an owned
copy of every object of every heap collection it read, so collections
nested inside each other's object data made a 744 KB file take 1.58 GB
(and `read_vl_strings` before it did the same). The cache now records
where objects lie instead of copying them, is dropped past a 32 MiB
budget, and a collection overlapping one already read is an error
(libhdf5 never writes one). New `GlobalHeapCollection::parse_index`
locates a collection's objects without copying them; `parse` and
`parse_index` refuse a collection running past the end of the file or an
object running past its collection. Conformance unchanged at 575 of 697
(`crates/clawhdf5-format/tests/vl_heap_bounds.rs`).
- **Every reader resolves VL data the same way.** `h5rs` (`dump`, `ls`,
`diff`, `check --data`) had its own lenient VL decoder: a heap object
longer than the element's length was cut to it (h5py refuses it), a null
string printed `""` where h5dump prints `NULL`, the stored element size
was trusted, and each heap collection was kept as a copy for the whole
run. It now resolves through `VlResolver`, so `dump` matches h5dump byte
for byte on VL strings (`"a\0b"` as `"a"`, null as `NULL`), VL sequences
and 4-byte-offset files, `dump --json` gives h5py's values, and
`check --data` reports any heap object whose size is not exactly the
element's length × base size. `clawhdf5-wasm` already resolved VL strings
with `read_vl_strings`; it now uses `VlResolver` and refuses a VL type
whose stored element size disagrees with the file, as `File` does
(`crates/clawhdf5-tools/tests/h5rs_interop.rs`,
`crates/clawhdf5-wasm/tests/vl_strings.rs`). New
`VlResolver::element` / `string_element` resolve one element in place.
- **A VL element at the undefined heap address is an error**, as in
libhdf5 ("addr undefined"). One of length 0 read as `""` in every reader
(`File`, `h5rs`, `clawhdf5-wasm`, `read_vl_strings`, `read_vl_bytes`).
libhdf5 writes a null element with heap address 0, which still reads as
empty, and h5py writes `""` as a zero-size heap object at a real address,
so no file libhdf5 or h5py writes is affected
(`a_vl_element_at_the_undefined_heap_address_fails_like_h5py` in
`crates/clawhdf5/tests/vl_data_interop.rs`). `read_vl_bytes` now also
treats address 0 as null whatever the length, as `VlResolver` does.
### Plugin filters (2026-09-26) ### Plugin filters (2026-09-26)
- **LZF, bitshuffle, bzip2 and Blosc read and write, in pure Rust.** Files - **LZF, bitshuffle, bzip2 and Blosc read and write, in pure Rust.** Files
written by h5py with `compression="lzf"`, or with hdf5plugin's written by h5py with `compression="lzf"`, or with hdf5plugin's
@@ -257,10 +342,10 @@
printed with its address; exit 1 when there are any. libhdf5's h5check printed with its address; exit 1 when there are any. libhdf5's h5check
reads only the 1.8 format. On the conformance corpus it passes all 418 reads only the 1.8 format. On the conformance corpus it passes all 418
files that both clawhdf5 and h5py read in full, and `check --data` flags files that both clawhdf5 and h5py read in full, and `check --data` flags
134 of the 150 CVE and fuzzer files of the `cve_hdf5` corpus (tank, 135 of the 150 CVE and fuzzer files of the `cve_hdf5` corpus (tank,
2026-09-26). `--data` also follows variable-length data into its global 2026-09-26). `--data` also follows variable-length data into its global
heap collections and reports a damaged one at its address. It inherits heap collections and reports a damaged one at its address. It inherits
the library's tolerance, though: 9 of the 16 it passes are files h5dump the library's tolerance, though: 8 of the 15 it passes are files h5dump
1.14.6 rejects (see `docs/known-issues.md`, header checks). 1.14.6 rejects (see `docs/known-issues.md`, header checks).
- Values over `--max-bytes` (default 1 GiB) are reported instead of read; - Values over `--max-bytes` (default 1 GiB) are reported instead of read;
a panic is caught and reported as an internal error (exit 3). a panic is caught and reported as an internal error (exit 3).
+18 -57
View File
@@ -19,9 +19,8 @@
//! with its message, location and the clawhdf5 frames of its backtrace. //! with its message, location and the clawhdf5 frames of its backtrace.
use std::cell::RefCell; use std::cell::RefCell;
use std::collections::{HashMap, HashSet}; use std::collections::HashSet;
use std::panic::{self, AssertUnwindSafe}; use std::panic::{self, AssertUnwindSafe};
use std::rc::Rc;
use clawhdf5_format::attribute::extract_attributes_full; use clawhdf5_format::attribute::extract_attributes_full;
use clawhdf5_format::data_layout::DataLayout; use clawhdf5_format::data_layout::DataLayout;
@@ -29,7 +28,6 @@ use clawhdf5_format::data_read;
use clawhdf5_format::dataspace::{Dataspace, DataspaceType}; use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder}; use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
use clawhdf5_format::filter_pipeline::FilterPipeline; use clawhdf5_format::filter_pipeline::FilterPipeline;
use clawhdf5_format::global_heap::GlobalHeapCollection;
use clawhdf5_format::group_v1::{self, GroupEntry}; use clawhdf5_format::group_v1::{self, GroupEntry};
use clawhdf5_format::group_v2; use clawhdf5_format::group_v2;
use clawhdf5_format::message_type::MessageType; use clawhdf5_format::message_type::MessageType;
@@ -37,6 +35,7 @@ use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::signature; use clawhdf5_format::signature;
use clawhdf5_format::superblock::Superblock; use clawhdf5_format::superblock::Superblock;
use clawhdf5_format::symbol_table::SymbolTableMessage; use clawhdf5_format::symbol_table::SymbolTableMessage;
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
use serde_json::{Map, Value, json}; use serde_json::{Map, Value, json};
use sha2::{Digest, Sha256}; use sha2::{Digest, Sha256};
@@ -111,7 +110,10 @@ struct Ctx<'a> {
os: u8, os: u8,
ls: u8, ls: u8,
base_dir: std::path::PathBuf, base_dir: std::path::PathBuf,
heaps: RefCell<HashMap<u64, Result<Rc<GlobalHeapCollection>, String>>>, /// Resolves variable-length elements as the library does (null
/// elements, strings cut at a NUL, heap objects of the wrong size
/// refused), caching each heap collection.
vl: RefCell<VlResolver<'a>>,
} }
impl<'a> Ctx<'a> { impl<'a> Ctx<'a> {
@@ -130,33 +132,6 @@ impl<'a> Ctx<'a> {
} }
} }
fn heap_obj(&self, addr: u64, idx: u32) -> Result<Vec<u8>, String> {
let coll = {
let mut cache = self.heaps.borrow_mut();
cache
.entry(addr)
.or_insert_with(|| {
GlobalHeapCollection::parse(self.data, addr as usize, self.ls)
.map(Rc::new)
.map_err(e)
})
.clone()?
};
coll.get_object(idx as u16)
.map(|o| o.data.clone())
.ok_or_else(|| {
format!("GlobalHeapObjectNotFound {{ collection_address: {addr}, index: {idx} }}")
})
}
fn read_offset(&self, b: &[u8]) -> u64 {
let mut v = 0u64;
for (i, x) in b.iter().take(self.os as usize).enumerate() {
v |= (*x as u64) << (8 * i);
}
v
}
fn canon(&self, dt: &Datatype, b: &[u8], out: &mut Vec<u8>) -> Result<(), String> { fn canon(&self, dt: &Datatype, b: &[u8], out: &mut Vec<u8>) -> Result<(), String> {
let size = dt.type_size() as usize; let size = dt.type_size() as usize;
if b.len() < size { if b.len() < size {
@@ -204,41 +179,27 @@ impl<'a> Ctx<'a> {
} }
} }
Datatype::VariableLength { Datatype::VariableLength {
size: vl_size,
is_string, is_string,
base_type, base_type,
.. ..
} => { } => {
let len = u32::from_le_bytes([b[0], b[1], b[2], b[3]]) as usize; check_element_size(*vl_size, self.os).map_err(e)?;
let addr = self.read_offset(&b[4..]); let el = &b[..size];
let idx_off = 4 + self.os as usize;
let idx = u32::from_le_bytes([
b[idx_off],
b[idx_off + 1],
b[idx_off + 2],
b[idx_off + 3],
]);
let obj = if len == 0 || addr == 0 || addr == u64::MAX >> (64 - 8 * self.os as u32)
{
Vec::new()
} else {
self.heap_obj(addr, idx)?
};
if *is_string { if *is_string {
let l = len.min(obj.len()); let s = self.vl.borrow_mut().string_bytes(el).map_err(e)?;
canon_str(&obj[..l], out); canon_str(&s[0], out);
} else { } else {
let bs = base_type.type_size() as usize; let bs = base_type.type_size() as usize;
if bs == 0 { // The borrow ends here: the base type may itself be
return Err("canon: VL base size 0".into()); // variable-length.
} let seq = self.vl.borrow_mut().sequences(el, bs).map_err(e)?;
let need = len.checked_mul(bs).ok_or("canon: VL overflow")?; let seq = &seq[0];
if len > 0 && obj.len() < need { let len = seq.len() / bs;
return Err(format!("canon: VL object {} < {need}", obj.len()));
}
out.push(b'V'); out.push(b'V');
out.extend_from_slice(&(len as u32).to_le_bytes()); out.extend_from_slice(&(len as u32).to_le_bytes());
for i in 0..len { for i in 0..len {
self.canon(base_type, &obj[i * bs..], out)?; self.canon(base_type, &seq[i * bs..], out)?;
} }
} }
} }
@@ -744,7 +705,7 @@ fn main() {
.parent() .parent()
.map(|p| p.to_path_buf()) .map(|p| p.to_path_buf())
.unwrap_or_default(), .unwrap_or_default(),
heaps: RefCell::new(HashMap::new()), vl: RefCell::new(VlResolver::new(hdf5, sb.offset_size, sb.length_size)),
}; };
let mut objects: Vec<Value> = Vec::new(); let mut objects: Vec<Value> = Vec::new();
let mut visited = HashSet::new(); let mut visited = HashSet::new();
+26 -2
View File
@@ -125,6 +125,11 @@ pub enum Datatype {
}, },
/// Class 9: Variable-length type. /// Class 9: Variable-length type.
VariableLength { VariableLength {
/// Size of one element as stored in the file: a sequence length (4
/// bytes), a global heap collection address (the file's
/// `offset_size`) and an object index (4 bytes) — 16 in a file with
/// 8-byte offsets, 12 with 4-byte offsets.
size: u32,
is_string: bool, is_string: bool,
padding: Option<StringPadding>, padding: Option<StringPadding>,
charset: Option<CharacterSet>, charset: Option<CharacterSet>,
@@ -771,6 +776,7 @@ impl Datatype {
pos += consumed; pos += consumed;
Ok(( Ok((
Datatype::VariableLength { Datatype::VariableLength {
size,
is_string, is_string,
padding, padding,
charset, charset,
@@ -1017,6 +1023,7 @@ impl Datatype {
Self::build_header(3, 1, [bf0, 0, 0], *size) Self::build_header(3, 1, [bf0, 0, 0], *size)
} }
Datatype::VariableLength { Datatype::VariableLength {
size,
is_string, is_string,
padding, padding,
charset, charset,
@@ -1039,7 +1046,7 @@ impl Datatype {
} else { } else {
0 0
}; };
let mut buf = Self::build_header(9, 1, [bf0, bf1, 0], 16); let mut buf = Self::build_header(9, 1, [bf0, bf1, 0], *size);
buf.extend_from_slice(&base_type.serialize()); buf.extend_from_slice(&base_type.serialize());
buf buf
} }
@@ -1208,7 +1215,7 @@ impl Datatype {
Datatype::Compound { size, .. } => *size, Datatype::Compound { size, .. } => *size,
Datatype::Reference { size, .. } => *size, Datatype::Reference { size, .. } => *size,
Datatype::Enumeration { size, .. } => *size, Datatype::Enumeration { size, .. } => *size,
Datatype::VariableLength { .. } => 16, // typically pointer + length Datatype::VariableLength { size, .. } => *size,
Datatype::Array { Datatype::Array {
base_type, base_type,
dimensions, dimensions,
@@ -1889,11 +1896,13 @@ mod tests {
let (dt, _) = Datatype::parse(&buf).unwrap(); let (dt, _) = Datatype::parse(&buf).unwrap();
match dt { match dt {
Datatype::VariableLength { Datatype::VariableLength {
size,
is_string, is_string,
padding, padding,
charset, charset,
base_type, base_type,
} => { } => {
assert_eq!(size, 16);
assert!(is_string); assert!(is_string);
assert_eq!(padding, Some(StringPadding::NullTerminate)); assert_eq!(padding, Some(StringPadding::NullTerminate));
assert_eq!(charset, Some(CharacterSet::Utf8)); assert_eq!(charset, Some(CharacterSet::Utf8));
@@ -1914,11 +1923,13 @@ mod tests {
let (dt, _) = Datatype::parse(&buf).unwrap(); let (dt, _) = Datatype::parse(&buf).unwrap();
match dt { match dt {
Datatype::VariableLength { Datatype::VariableLength {
size,
is_string, is_string,
padding, padding,
charset, charset,
base_type, base_type,
} => { } => {
assert_eq!(size, 16);
assert!(!is_string); assert!(!is_string);
assert_eq!(padding, None); assert_eq!(padding, None);
assert_eq!(charset, None); assert_eq!(charset, None);
@@ -1928,6 +1939,19 @@ mod tests {
} }
} }
#[test]
fn variable_length_size_is_the_stored_size() {
// A file with 4-byte offsets stores 12-byte VL elements (length 4 +
// address 4 + index 4); the type used to report 16 regardless, so
// every read laid the elements out 16 bytes apart.
let mut buf = build_dt_header(9, 1, [0x01, 0x00, 0], 12);
buf.extend_from_slice(&build_fixed_point(1, false, false, 0, 8));
let (dt, _) = Datatype::parse(&buf).unwrap();
assert_eq!(dt.type_size(), 12);
// And it is written back as stored.
assert_eq!(dt.serialize()[4..8], 12u32.to_le_bytes());
}
#[test] #[test]
fn test_array_2d() { fn test_array_2d() {
// Array [3][4] of i32 LE, version 3 // Array [3][4] of i32 LE, version 3
+99 -24
View File
@@ -1,7 +1,7 @@
//! HDF5 Global Heap collection parsing. //! HDF5 Global Heap collection parsing.
#[cfg(not(feature = "std"))] #[cfg(not(feature = "std"))]
use alloc::vec::Vec; use alloc::{format, string::String, vec::Vec};
use crate::error::FormatError; use crate::error::FormatError;
@@ -52,11 +52,42 @@ fn read_length(data: &[u8], offset: usize, length_size: u8) -> Result<u64, Forma
}) })
} }
fn object_overrun_msg(index: u16, size: usize, collection_size: u64) -> String {
format!(
"global heap object {index} ({size} bytes) runs past the end of its \
{collection_size}-byte collection"
)
}
/// Round up to next multiple of 8. /// Round up to next multiple of 8.
fn pad8(x: usize) -> usize { fn pad8(x: usize) -> usize {
(x + 7) & !7 (x + 7) & !7
} }
/// Where one object of a global heap collection lies in the file, without
/// its data: see [`GlobalHeapCollection::parse_index`].
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct GlobalHeapObjectRef {
/// Object index (1-based; 0 is the free space marker).
pub index: u16,
/// Reference count.
pub reference_count: u16,
/// Offset of the object's data in the file data the collection was
/// parsed from.
pub offset: usize,
/// Size of the object's data in bytes.
pub size: usize,
}
/// A global heap collection's objects, located but not copied.
#[derive(Debug, Clone)]
pub struct GlobalHeapIndex {
/// Total size of this collection including header.
pub collection_size: u64,
/// The objects, in file order.
pub objects: Vec<GlobalHeapObjectRef>,
}
impl GlobalHeapCollection { impl GlobalHeapCollection {
/// Parse a global heap collection at the given offset in the file data. /// Parse a global heap collection at the given offset in the file data.
pub fn parse( pub fn parse(
@@ -64,8 +95,38 @@ impl GlobalHeapCollection {
offset: usize, offset: usize,
length_size: u8, length_size: u8,
) -> Result<GlobalHeapCollection, FormatError> { ) -> Result<GlobalHeapCollection, FormatError> {
// signature(4) + version(1) + reserved(3) + collection_size(length_size) let index = Self::parse_index(file_data, offset, length_size)?;
let header_size = 8 + length_size as usize; Ok(GlobalHeapCollection {
collection_size: index.collection_size,
objects: index
.objects
.iter()
.map(|o| GlobalHeapObject {
index: o.index,
reference_count: o.reference_count,
data: file_data[o.offset..o.offset + o.size].to_vec(),
})
.collect(),
})
}
/// Locate the objects of the global heap collection at `offset` without
/// copying their data, so a caller can keep many collections indexed
/// for the cost of their object headers.
///
/// The collection must lie inside `file_data`, and every object inside
/// the collection, as libhdf5 lays them out; an object that runs past
/// its collection is an error.
pub fn parse_index(
file_data: &[u8],
offset: usize,
length_size: u8,
) -> Result<GlobalHeapIndex, FormatError> {
// signature(4) + version(1) + reserved(3) + collection_size(length_size),
// padded to a multiple of 8 as libhdf5 lays it out (`H5HG_SIZEOF_HDR`).
// With 8-byte lengths the padding is 0; with 4-byte lengths it is 4,
// and reading without it put every object 4 bytes early.
let header_size = pad8(8 + length_size as usize);
ensure_len(file_data, offset, header_size)?; ensure_len(file_data, offset, header_size)?;
if file_data[offset..offset + 4] != GCOL_SIGNATURE { if file_data[offset..offset + 4] != GCOL_SIGNATURE {
@@ -78,25 +139,25 @@ impl GlobalHeapCollection {
} }
let collection_size = read_length(file_data, offset + 8, length_size)?; let collection_size = read_length(file_data, offset + 8, length_size)?;
let collection_size_usize = let collection_end = usize::try_from(collection_size)
usize::try_from(collection_size).map_err(|_| FormatError::UnexpectedEof { .ok()
expected: u64::MAX as usize, .and_then(|size| offset.checked_add(size))
available: file_data.len(),
})?;
let collection_end =
offset
.checked_add(collection_size_usize)
.ok_or(FormatError::UnexpectedEof { .ok_or(FormatError::UnexpectedEof {
expected: usize::MAX, expected: usize::MAX,
available: file_data.len(), available: file_data.len(),
})?; })?;
if collection_end > file_data.len() {
return Err(FormatError::UnexpectedEof {
expected: collection_end,
available: file_data.len(),
});
}
let mut pos = offset + header_size; let mut pos = offset + header_size;
let mut objects = Vec::new(); let mut objects = Vec::new();
// Parse objects until we hit index 0 (free space) or run out of space // Parse objects until we hit index 0 (free space) or run out of space
while pos + 2 <= collection_end { while pos + 2 <= collection_end {
ensure_len(file_data, pos, 2)?;
let object_index = u16::from_le_bytes([file_data[pos], file_data[pos + 1]]); let object_index = u16::from_le_bytes([file_data[pos], file_data[pos + 1]]);
if object_index == 0 { if object_index == 0 {
@@ -104,28 +165,39 @@ impl GlobalHeapCollection {
break; break;
} }
// object_index(2) + reference_count(2) + reserved(4) + object_size(length_size) // object_index(2) + reference_count(2) + reserved(4) +
let obj_header_size = 8 + length_size as usize; // object_size(length_size), padded to 8 (`H5HG_SIZEOF_OBJHDR`).
ensure_len(file_data, pos, obj_header_size)?; let obj_header_size = pad8(8 + length_size as usize);
ensure_len(&file_data[..collection_end], pos, obj_header_size)?;
let reference_count = u16::from_le_bytes([file_data[pos + 2], file_data[pos + 3]]); let reference_count = u16::from_le_bytes([file_data[pos + 2], file_data[pos + 3]]);
let object_size = read_length(file_data, pos + 8, length_size)? as usize; let object_size = usize::try_from(read_length(file_data, pos + 8, length_size)?)
.map_err(|_| FormatError::Overflow("global heap object size".into()))?;
pos += obj_header_size; pos += obj_header_size;
ensure_len(file_data, pos, object_size)?; if pos
let data = file_data[pos..pos + object_size].to_vec(); .checked_add(object_size)
.is_none_or(|end| end > collection_end)
{
return Err(FormatError::VlDataError(object_overrun_msg(
object_index,
object_size,
collection_size,
)));
}
objects.push(GlobalHeapObject { objects.push(GlobalHeapObjectRef {
index: object_index, index: object_index,
reference_count, reference_count,
data, offset: pos,
size: object_size,
}); });
// Advance past data + padding to 8-byte boundary // Advance past data + padding to 8-byte boundary
pos += pad8(object_size); pos = pos.saturating_add(pad8(object_size));
} }
Ok(GlobalHeapCollection { Ok(GlobalHeapIndex {
collection_size, collection_size,
objects, objects,
}) })
@@ -149,10 +221,11 @@ mod tests {
let ls = length_size as usize; let ls = length_size as usize;
// Calculate total size // Calculate total size
let header_size = 8 + ls; // libhdf5 pads both headers to a multiple of 8.
let header_size = pad8(8 + ls);
let mut obj_size_total = 0usize; let mut obj_size_total = 0usize;
for (_, _, data) in objects { for (_, _, data) in objects {
let obj_header = 8 + ls; let obj_header = pad8(8 + ls);
obj_size_total += obj_header + pad8(data.len()); obj_size_total += obj_header + pad8(data.len());
} }
// Free space marker (2 bytes for index 0) // Free space marker (2 bytes for index 0)
@@ -170,6 +243,7 @@ mod tests {
8 => buf.extend_from_slice(&(collection_size as u64).to_le_bytes()), 8 => buf.extend_from_slice(&(collection_size as u64).to_le_bytes()),
_ => panic!("unsupported length_size"), _ => panic!("unsupported length_size"),
} }
buf.resize(header_size, 0);
// Objects // Objects
for (index, ref_count, data) in objects { for (index, ref_count, data) in objects {
@@ -181,6 +255,7 @@ mod tests {
8 => buf.extend_from_slice(&(data.len() as u64).to_le_bytes()), 8 => buf.extend_from_slice(&(data.len() as u64).to_le_bytes()),
_ => panic!("unsupported"), _ => panic!("unsupported"),
} }
buf.resize(buf.len() + (pad8(8 + ls) - (8 + ls)), 0);
buf.extend_from_slice(data); buf.extend_from_slice(data);
// Pad to 8 bytes // Pad to 8 bytes
let padded = pad8(data.len()); let padded = pad8(data.len());
+452 -56
View File
@@ -5,10 +5,12 @@
//! `sequence_length(4 LE) + collection_address(offset_size LE) + object_index(4 LE)`. //! `sequence_length(4 LE) + collection_address(offset_size LE) + object_index(4 LE)`.
#[cfg(not(feature = "std"))] #[cfg(not(feature = "std"))]
use alloc::{string::String, vec::Vec}; use alloc::{collections::BTreeMap, format, string::String, vec, vec::Vec};
#[cfg(feature = "std")]
use std::collections::BTreeMap;
use crate::error::FormatError; use crate::error::FormatError;
use crate::global_heap::GlobalHeapCollection; use crate::global_heap::{GlobalHeapCollection, GlobalHeapIndex};
/// A parsed variable-length element reference (global heap ID). /// A parsed variable-length element reference (global heap ID).
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
@@ -109,7 +111,218 @@ fn is_undefined_address(addr: u64, offset_size: u8) -> bool {
} }
} }
/// The size of one variable-length element in a file with `offset_size`-byte
/// addresses: a sequence length (4), a global heap collection address and an
/// object index (4). libhdf5 computes it this way rather than trusting the
/// datatype message (`H5T_set_loc`).
pub fn element_size(offset_size: u8) -> usize {
4 + offset_size as usize + 4
}
/// Refuse a variable-length datatype whose stored element size is not the
/// one this file's offset size implies. Its elements would be laid out with
/// a stride libhdf5 does not use, so every value after the first would be
/// read from the wrong place.
pub fn check_element_size(stored_size: u32, offset_size: u8) -> Result<(), FormatError> {
let expected = element_size(offset_size);
if stored_size as usize != expected {
return Err(FormatError::VlDataError(format!(
"variable-length datatype stores {stored_size}-byte elements; a file with \
{offset_size}-byte offsets uses {expected}"
)));
}
Ok(())
}
/// A collection's objects, located in the file data but not copied:
/// `(index, offset, size)` of the first object with each index, sorted by
/// index.
struct CachedCollection {
objects: Vec<(u16, usize, usize)>,
}
impl CachedCollection {
fn new(index: GlobalHeapIndex) -> Self {
let mut objects: Vec<(u16, usize, usize)> = index
.objects
.iter()
.map(|o| (o.index, o.offset, o.size))
.collect();
// Stable, so the first object with a repeated index is kept.
objects.sort_by_key(|o| o.0);
objects.dedup_by_key(|o| o.0);
Self { objects }
}
/// What this entry costs to keep, in bytes (roughly).
fn cost(&self) -> usize {
64 + self.objects.len() * core::mem::size_of::<(u16, usize, usize)>()
}
fn get(&self, index: u32) -> Option<(usize, usize)> {
let index = u16::try_from(index).ok()?;
let i = self.objects.binary_search_by_key(&index, |o| o.0).ok()?;
Some((self.objects[i].1, self.objects[i].2))
}
}
/// How many bytes of collection indexes a [`VlResolver`] keeps before it
/// drops them and starts again. Values are never copied into the cache, so
/// this bounds what a read retains however many collections it visits.
const CACHE_BUDGET: usize = 32 << 20;
/// Resolves variable-length elements against a file's global heap, parsing
/// each heap collection once however many elements point into it.
///
/// Values follow libhdf5: an element whose heap address is 0 is null (an
/// empty string or sequence), and an element whose heap object is not
/// exactly `length × base size` bytes is an error ("Expected global heap
/// object size does not match"), not a truncated or padded value.
///
/// Memory stays bounded on hostile files: the cache holds where each
/// object lies, not a copy of it, up to a fixed budget; and collections
/// that overlap one another are refused (libhdf5 never writes them), so a
/// file cannot make the resolver parse the same bytes as the objects of
/// many collections.
pub struct VlResolver<'a> {
file_data: &'a [u8],
offset_size: u8,
length_size: u8,
cache: BTreeMap<u64, CachedCollection>,
cached_bytes: usize,
budget: usize,
/// Start → end of every collection parsed so far (kept when the cache
/// is dropped, to check overlaps).
extents: BTreeMap<usize, usize>,
}
impl<'a> VlResolver<'a> {
/// A resolver over `file_data` (the file from its superblock on), with
/// the superblock's offset and length sizes.
pub fn new(file_data: &'a [u8], offset_size: u8, length_size: u8) -> Self {
Self {
file_data,
offset_size,
length_size,
cache: BTreeMap::new(),
cached_bytes: 0,
budget: CACHE_BUDGET,
extents: BTreeMap::new(),
}
}
/// The size of one element in this file (see [`element_size`]).
pub fn element_size(&self) -> usize {
element_size(self.offset_size)
}
/// Split `raw` into elements; its length must be a whole number of them.
fn elements(&self, raw: &[u8]) -> Result<Vec<VlElement>, FormatError> {
let size = self.element_size();
if !raw.len().is_multiple_of(size) {
return Err(FormatError::VlDataError(format!(
"{} bytes is not a whole number of {size}-byte variable-length elements",
raw.len()
)));
}
parse_vl_references(raw, (raw.len() / size) as u64, self.offset_size)
}
/// The bytes of one element: `length × base_size` bytes from the heap,
/// or `None` for a null element.
fn resolve(
&mut self,
vl: &VlElement,
base_size: usize,
) -> Result<Option<&'a [u8]>, FormatError> {
let addr = vl.collection_address;
if addr == 0 {
return Ok(None);
}
let data = self.object(vl)?;
let expected = (vl.length as usize)
.checked_mul(base_size)
.ok_or_else(|| FormatError::Overflow("variable-length element size".into()))?;
if data.len() != expected {
return Err(FormatError::VlDataError(format!(
"global heap object {} in the collection at {addr} holds {} bytes; the element \
says {} × {base_size}",
vl.object_index,
data.len(),
vl.length
)));
}
Ok(Some(data))
}
/// One element (the first [`element_size`](Self::element_size) bytes of
/// `elem`) of a variable-length sequence whose base type is `base_size`
/// bytes: its `length × base_size` bytes, or `None` for a null element
/// (heap address 0).
pub fn element(
&mut self,
elem: &[u8],
base_size: usize,
) -> Result<Option<&'a [u8]>, FormatError> {
let vl = parse_vl_references(elem, 1, self.offset_size)?;
self.resolve(&vl[0], base_size)
}
/// One variable-length string element: its bytes up to the first NUL,
/// or `None` for a null element (h5dump prints it as `NULL`, h5py
/// returns it as empty).
pub fn string_element(&mut self, elem: &[u8]) -> Result<Option<&'a [u8]>, FormatError> {
Ok(self.element(elem, 1)?.map(cut_at_nul))
}
/// The strings of the variable-length string elements in `raw`, as
/// bytes. A string ends at its first NUL, as libhdf5 returns it (it
/// converts each to a C string); a null element is empty.
pub fn string_bytes(&mut self, raw: &[u8]) -> Result<Vec<Vec<u8>>, FormatError> {
self.elements(raw)?
.iter()
.map(|vl| Ok(self.resolve(vl, 1)?.map(cut_at_nul).unwrap_or(&[]).to_vec()))
.collect()
}
/// The strings of the variable-length string elements in `raw`, decoded
/// as UTF-8 with invalid sequences replaced by U+FFFD (see
/// [`string_bytes`](Self::string_bytes) for the exact bytes).
pub fn strings(&mut self, raw: &[u8]) -> Result<Vec<String>, FormatError> {
Ok(self
.string_bytes(raw)?
.into_iter()
.map(|b| match String::from_utf8(b) {
Ok(s) => s,
Err(e) => String::from_utf8_lossy(e.as_bytes()).into_owned(),
})
.collect())
}
/// The sequences of the variable-length sequence elements in `raw`, each
/// as its `length × base_size` bytes in the base type's encoding.
pub fn sequences(&mut self, raw: &[u8], base_size: usize) -> Result<Vec<Vec<u8>>, FormatError> {
if base_size == 0 {
return Err(FormatError::VlDataError(
"variable-length sequence of a zero-size base type".into(),
));
}
self.elements(raw)?
.iter()
.map(|vl| Ok(self.resolve(vl, base_size)?.unwrap_or(&[]).to_vec()))
.collect()
}
}
/// A string's bytes up to its first NUL.
fn cut_at_nul(s: &[u8]) -> &[u8] {
&s[..s.iter().position(|&b| b == 0).unwrap_or(s.len())]
}
/// Resolve VL strings from raw data by looking up each element in the global heap. /// Resolve VL strings from raw data by looking up each element in the global heap.
///
/// Reads the first `num_elements` elements of `raw`. Strings end at their
/// first NUL and invalid UTF-8 is replaced, as in [`VlResolver::strings`].
pub fn read_vl_strings( pub fn read_vl_strings(
file_data: &[u8], file_data: &[u8],
raw_data: &[u8], raw_data: &[u8],
@@ -117,35 +330,23 @@ pub fn read_vl_strings(
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Vec<String>, FormatError> { ) -> Result<Vec<String>, FormatError> {
let refs = parse_vl_references(raw_data, num_elements, offset_size)?; let raw = first_elements(raw_data, num_elements, offset_size)?;
let mut result = Vec::with_capacity(refs.len()); VlResolver::new(file_data, offset_size, length_size).strings(raw)
}
for vl in &refs { /// The first `num_elements` elements of `raw`, or an error if it is shorter.
if vl.length == 0 && is_undefined_address(vl.collection_address, offset_size) { fn first_elements(raw: &[u8], num_elements: u64, offset_size: u8) -> Result<&[u8], FormatError> {
result.push(String::new()); let total = usize::try_from(num_elements)
continue; .ok()
} .and_then(|n| n.checked_mul(element_size(offset_size)))
if vl.length == 0 && vl.collection_address == 0 { .ok_or(FormatError::UnexpectedEof {
result.push(String::new()); expected: usize::MAX,
continue; available: raw.len(),
} })?;
raw.get(..total).ok_or(FormatError::UnexpectedEof {
let coll = expected: total,
GlobalHeapCollection::parse(file_data, vl.collection_address as usize, length_size)?; available: raw.len(),
let obj = coll.get_object(vl.object_index as u16).ok_or( })
FormatError::GlobalHeapObjectNotFound {
collection_address: vl.collection_address,
index: vl.object_index as u16,
},
)?;
// The object data is the raw string bytes
let len = (vl.length as usize).min(obj.data.len());
let s = String::from_utf8_lossy(&obj.data[..len]).into_owned();
result.push(s);
}
Ok(result)
} }
/// Resolve VL sequences from raw data, returning each element's bytes. /// Resolve VL sequences from raw data, returning each element's bytes.
@@ -153,7 +354,9 @@ pub fn read_vl_strings(
/// Each element is the sequence's full encoding — element count × base type /// Each element is the sequence's full encoding — element count × base type
/// size bytes, in the base type's byte order — so a sequence of `i32` yields /// size bytes, in the base type's byte order — so a sequence of `i32` yields
/// four bytes per value. Decode it with the base type (e.g. /// four bytes per value. Decode it with the base type (e.g.
/// [`crate::data_read::read_as_i64`]). /// [`crate::data_read::read_as_i64`]). This does not know the base type, so
/// it returns each heap object whole; [`VlResolver::sequences`] also checks
/// the object's size against the element's length.
pub fn read_vl_bytes( pub fn read_vl_bytes(
file_data: &[u8], file_data: &[u8],
raw_data: &[u8], raw_data: &[u8],
@@ -162,35 +365,97 @@ pub fn read_vl_bytes(
length_size: u8, length_size: u8,
) -> Result<Vec<Vec<u8>>, FormatError> { ) -> Result<Vec<Vec<u8>>, FormatError> {
let refs = parse_vl_references(raw_data, num_elements, offset_size)?; let refs = parse_vl_references(raw_data, num_elements, offset_size)?;
let mut resolver = VlResolver::new(file_data, offset_size, length_size);
let mut result = Vec::with_capacity(refs.len()); let mut result = Vec::with_capacity(refs.len());
for vl in &refs { for vl in &refs {
if vl.length == 0 // A heap address of 0 is a null element, as in VlResolver.
&& (is_undefined_address(vl.collection_address, offset_size) if vl.collection_address == 0 {
|| vl.collection_address == 0)
{
result.push(Vec::new()); result.push(Vec::new());
continue; continue;
} }
let coll =
GlobalHeapCollection::parse(file_data, vl.collection_address as usize, length_size)?;
let obj = coll.get_object(vl.object_index as u16).ok_or(
FormatError::GlobalHeapObjectNotFound {
collection_address: vl.collection_address,
index: vl.object_index as u16,
},
)?;
// The heap object holds the whole sequence. `vl.length` counts // The heap object holds the whole sequence. `vl.length` counts
// elements, not bytes, so it is only the byte length when the base // elements, not bytes, so it is only the byte length when the base
// type is one byte wide. // type is one byte wide.
result.push(obj.data.clone()); let obj = resolver.object(vl)?;
result.push(obj.to_vec());
} }
Ok(result) Ok(result)
} }
impl<'a> VlResolver<'a> {
/// The heap object `vl` points to, whatever its size; its collection is
/// parsed on first use.
fn object(&mut self, vl: &VlElement) -> Result<&'a [u8], FormatError> {
let addr = vl.collection_address;
// libhdf5 writes a null element with address 0, never the undefined
// address, and fails to read one ("addr undefined") even when its
// length is 0; we returned an empty value.
if is_undefined_address(addr, self.offset_size) {
return Err(FormatError::VlDataError(format!(
"variable-length element (length {}) has the undefined global heap address",
vl.length
)));
}
if !self.cache.contains_key(&addr) {
let offset = usize::try_from(addr).map_err(|_| FormatError::UnexpectedEof {
expected: usize::MAX,
available: self.file_data.len(),
})?;
let index =
GlobalHeapCollection::parse_index(self.file_data, offset, self.length_size)?;
// parse_index checked that the collection lies in the file.
let end = offset + index.collection_size as usize;
self.check_overlap(offset, end)?;
let coll = CachedCollection::new(index);
if self.cached_bytes.saturating_add(coll.cost()) > self.budget {
self.cache.clear();
self.cached_bytes = 0;
}
self.cached_bytes += coll.cost();
self.cache.insert(addr, coll);
}
let (start, size) = self.cache[&addr].get(vl.object_index).ok_or(
FormatError::GlobalHeapObjectNotFound {
collection_address: addr,
index: vl.object_index as u16,
},
)?;
Ok(&self.file_data[start..start + size])
}
/// Record the collection at `start..end`, refusing one that overlaps a
/// collection already read. libhdf5 allocates each collection its own
/// block; overlapping ones only come from a crafted file, where they let
/// every byte be parsed again as the objects of each collection.
fn check_overlap(&mut self, start: usize, end: usize) -> Result<(), FormatError> {
if let Some(&known) = self.extents.get(&start) {
return if known == end {
Ok(())
} else {
Err(FormatError::VlDataError(format!(
"global heap collection at {start} changed size"
)))
};
}
let before = self.extents.range(..start).next_back();
let after = self.extents.range(start..).next();
let clash = match (before, after) {
(Some((&s, &e)), _) if e > start => Some(s),
(_, Some((&s, _))) if s < end => Some(s),
_ => None,
};
if let Some(other) = clash {
return Err(FormatError::VlDataError(format!(
"global heap collection at {start} overlaps the one at {other}"
)));
}
self.extents.insert(start, end);
Ok(())
}
}
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
@@ -285,16 +550,27 @@ mod tests {
} }
#[test] #[test]
fn null_vl_element_empty_string() { fn an_undefined_heap_address_is_an_error_even_at_length_0() {
// length=0, address=undefined // libhdf5 fails the read ("addr undefined"); h5py and libhdf5 write
let mut raw = Vec::new(); // a null element with address 0. We returned "".
raw.extend_from_slice(&0u32.to_le_bytes()); // length=0 let mut file_data = vec![0u8; 256];
raw.extend_from_slice(&u64::MAX.to_le_bytes()); // undefined address build_gcol_at(&mut file_data, 64, &[(1, b"x")]);
raw.extend_from_slice(&0u32.to_le_bytes()); // index for (os, undef) in [(8u8, u64::MAX), (4, 0xFFFF_FFFF)] {
for length in [0, 1] {
let file_data = vec![0u8; 16]; let mut raw = element(1, 64, 1, os);
let strings = read_vl_strings(&file_data, &raw, 1, 8, 8).unwrap(); raw.extend(element(length, undef, 1, os));
assert_eq!(strings, vec![""]); let mut r = VlResolver::new(&file_data, os, 8);
let e = r.string_bytes(&raw).unwrap_err().to_string();
assert!(e.contains("undefined"), "{e}");
assert!(r.sequences(&raw, 1).is_err());
assert!(r.string_element(&raw[raw.len() / 2..]).is_err());
let n = 2;
assert!(read_vl_strings(&file_data, &raw, n, os, 8).is_err());
assert!(read_vl_bytes(&file_data, &raw, n, os, 8).is_err());
// The defined element alone still reads.
assert_eq!(r.strings(&raw[..raw.len() / 2]).unwrap(), ["x"]);
}
}
} }
#[test] #[test]
@@ -333,6 +609,126 @@ mod tests {
assert_eq!(bytes, vec![vec![0xDE, 0xAD], vec![0xBE, 0xEF, 0xCA]]); assert_eq!(bytes, vec![vec![0xDE, 0xAD], vec![0xBE, 0xEF, 0xCA]]);
} }
fn element(length: u32, addr: u64, index: u32, offset_size: u8) -> Vec<u8> {
let mut raw = length.to_le_bytes().to_vec();
raw.extend_from_slice(&addr.to_le_bytes()[..offset_size as usize]);
raw.extend_from_slice(&index.to_le_bytes());
raw
}
#[test]
fn strings_end_at_the_first_nul() {
// libhdf5 hands each VL string over as a C string, so h5py sees
// "a\0b" as "a"; we used to return the NUL and what followed.
let mut file_data = vec![0u8; 512];
build_gcol_at(&mut file_data, 64, &[(1, b"a\0b"), (2, b"cd")]);
let mut raw = element(3, 64, 1, 8);
raw.extend(element(2, 64, 2, 8));
let mut r = VlResolver::new(&file_data, 8, 8);
assert_eq!(
r.string_bytes(&raw).unwrap(),
vec![b"a".to_vec(), b"cd".to_vec()]
);
assert_eq!(
read_vl_strings(&file_data, &raw, 2, 8, 8).unwrap(),
["a", "cd"]
);
}
#[test]
fn a_heap_object_of_the_wrong_size_is_an_error() {
// libhdf5: "Expected global heap object size does not match". We
// used to return the object cut to the element's length.
let mut file_data = vec![0u8; 512];
build_gcol_at(&mut file_data, 64, &[(1, b"cdefgh"), (2, &[1, 0, 0, 0])]);
let mut r = VlResolver::new(&file_data, 8, 8);
assert!(r.string_bytes(&element(3, 64, 1, 8)).is_err());
assert!(r.string_bytes(&element(9, 64, 1, 8)).is_err());
assert!(read_vl_strings(&file_data, &element(3, 64, 1, 8), 1, 8, 8).is_err());
// A sequence of one i32 is 4 bytes; of two, 8.
assert_eq!(
r.sequences(&element(1, 64, 2, 8), 4).unwrap(),
vec![vec![1, 0, 0, 0]]
);
assert!(r.sequences(&element(2, 64, 2, 8), 4).is_err());
assert!(r.sequences(&element(1, 64, 2, 8), 0).is_err());
}
#[test]
fn address_zero_is_null_whatever_the_length() {
// libhdf5 treats a heap address of 0 as a null element.
let file_data = vec![0u8; 64];
let mut r = VlResolver::new(&file_data, 8, 8);
assert_eq!(
r.string_bytes(&element(5, 0, 1, 8)).unwrap(),
vec![Vec::<u8>::new()]
);
assert_eq!(
r.sequences(&element(5, 0, 1, 8), 4).unwrap(),
vec![Vec::<u8>::new()]
);
}
#[test]
fn four_byte_offsets_use_twelve_byte_elements() {
let mut file_data = vec![0u8; 512];
build_gcol_at(&mut file_data, 64, &[(1, b"one"), (2, b""), (3, b"three")]);
let mut raw = element(3, 64, 1, 4);
raw.extend(element(0, 64, 2, 4));
raw.extend(element(5, 64, 3, 4));
assert_eq!(raw.len(), 36);
let mut r = VlResolver::new(&file_data, 4, 8);
assert_eq!(r.element_size(), 12);
assert_eq!(r.strings(&raw).unwrap(), ["one", "", "three"]);
// Not a whole number of elements.
assert!(r.strings(&raw[..30]).is_err());
}
#[test]
fn the_cache_stays_within_its_budget_and_rereads_what_it_dropped() {
// Twenty collections of three objects each; a budget that holds
// about two of them. Reading every element twice must still return
// the right strings after the cache is dropped.
let mut file_data = vec![0u8; 64];
let mut raw = Vec::new();
for c in 0..20u64 {
let at = file_data.len();
let names: Vec<String> = (0..3).map(|i| format!("c{c}o{i}")).collect();
let objs: Vec<(u16, &[u8])> = names
.iter()
.enumerate()
.map(|(i, n)| (i as u16 + 1, n.as_bytes()))
.collect();
build_gcol_at(&mut file_data, at, &objs);
for (i, n) in names.iter().enumerate() {
raw.extend(element(n.len() as u32, at as u64, i as u32 + 1, 8));
}
}
raw.extend(raw.clone());
let mut r = VlResolver::new(&file_data, 8, 8);
let one = CachedCollection {
objects: vec![(0, 0, 0); 3],
}
.cost();
r.budget = 2 * one + 1;
let want: Vec<String> = (0..2)
.flat_map(|_| (0..20).flat_map(|c| (0..3).map(move |i| format!("c{c}o{i}"))))
.collect();
for (k, chunk) in raw.chunks(16).enumerate() {
assert_eq!(r.strings(chunk).unwrap(), [want[k].clone()]);
assert!(r.cached_bytes <= r.budget);
assert!(r.cache.len() <= 2);
}
}
#[test]
fn element_size_is_checked_against_the_offset_size() {
assert!(check_element_size(16, 8).is_ok());
assert!(check_element_size(12, 4).is_ok());
assert!(check_element_size(16, 4).is_err());
assert!(check_element_size(524_304, 8).is_err());
}
#[test] #[test]
fn parse_vl_references_truncated_error() { fn parse_vl_references_truncated_error() {
let raw = vec![0u8; 10]; // too short for 1 element with offset_size=8 let raw = vec![0u8; 10]; // too short for 1 element with offset_size=8
@@ -0,0 +1,161 @@
//! Crafted files cannot make variable-length reads retain memory, or take
//! time, out of proportion to the file.
//!
//! `VlResolver` used to keep an owned copy of every object of every
//! collection it parsed, for the whole read. A file whose global heap
//! collections nest inside each other's object data — each element
//! pointing at a different one — then made retained memory O(K × file
//! size): a 744 KB file took 1.58 GB. The same nesting, with every
//! collection's object chain jumping to one shared run of tiny objects,
//! made the parse time O(K × M) as well. libhdf5 never writes overlapping
//! collections; they are now refused, and the cache holds only where
//! objects lie.
//!
//! Peak heap use is measured with a counting global allocator, so the
//! cases run one after another in a single test.
use std::alloc::{GlobalAlloc, Layout, System};
use std::sync::atomic::{AtomicUsize, Ordering};
use std::time::{Duration, Instant};
use clawhdf5_format::vl_data::VlResolver;
struct Counting;
static CURRENT: AtomicUsize = AtomicUsize::new(0);
static PEAK: AtomicUsize = AtomicUsize::new(0);
unsafe impl GlobalAlloc for Counting {
unsafe fn alloc(&self, layout: Layout) -> *mut u8 {
let p = unsafe { System.alloc(layout) };
if !p.is_null() {
let now = CURRENT.fetch_add(layout.size(), Ordering::Relaxed) + layout.size();
PEAK.fetch_max(now, Ordering::Relaxed);
}
p
}
unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) {
unsafe { System.dealloc(ptr, layout) };
CURRENT.fetch_sub(layout.size(), Ordering::Relaxed);
}
}
#[global_allocator]
static ALLOC: Counting = Counting;
/// Bytes allocated at the peak of `f`, above what was live when it started.
fn peak_during<T>(f: impl FnOnce() -> T) -> (T, usize) {
let base = CURRENT.load(Ordering::Relaxed);
PEAK.store(base, Ordering::Relaxed);
let out = f();
(out, PEAK.load(Ordering::Relaxed) - base)
}
fn put_header(file: &mut [u8], at: usize, size: u64) {
file[at..at + 4].copy_from_slice(b"GCOL");
file[at + 4] = 1;
file[at + 8..at + 16].copy_from_slice(&size.to_le_bytes());
}
fn put_object(file: &mut [u8], at: usize, index: u16, size: u64) {
file[at..at + 2].copy_from_slice(&index.to_le_bytes());
file[at + 2..at + 4].copy_from_slice(&1u16.to_le_bytes());
file[at + 8..at + 16].copy_from_slice(&size.to_le_bytes());
}
fn element(length: u32, addr: u64, index: u32) -> Vec<u8> {
let mut e = length.to_le_bytes().to_vec();
e.extend_from_slice(&addr.to_le_bytes());
e.extend_from_slice(&index.to_le_bytes());
e
}
/// K collections 32 bytes apart, each running to the end of the file with
/// one object covering the rest of it (and so every later collection).
/// Element i is that object of collection i.
fn nested(k: usize) -> (Vec<u8>, Vec<u8>) {
let base = 64;
let end = base + 32 * k + 64;
let mut file = vec![0u8; end];
let mut raw = Vec::new();
for i in 0..k {
let at = base + 32 * i;
put_header(&mut file, at, (end - at) as u64);
let obj = (end - at - 32) as u64;
put_object(&mut file, at + 16, 1, obj);
raw.extend(element(obj as u32, at as u64, 1));
}
(file, raw)
}
/// K collections 32 bytes apart, each with a first object that jumps over
/// the later collections to one shared run of M empty objects, so parsing
/// every collection walks all M.
fn shared_tail(k: usize, m: usize) -> (Vec<u8>, Vec<u8>) {
let base = 64;
let tail = base + 32 * k + 32;
let end = tail + 16 * m + 16;
let mut file = vec![0u8; end];
let mut raw = Vec::new();
for i in 0..k {
let at = base + 32 * i;
put_header(&mut file, at, (end - at) as u64);
let jump = (tail - at - 32) as u64;
put_object(&mut file, at + 16, 1, jump);
raw.extend(element(jump as u32, at as u64, 1));
}
for j in 0..m {
put_object(&mut file, tail + 16 * j, (j % 65_000 + 2) as u16, 0);
}
(file, raw)
}
#[test]
fn overlapping_collections_are_refused_in_bounded_memory_and_time() {
for (name, (file, raw)) in [
("nested", nested(2000)),
("shared tail", shared_tail(500, 10_000)),
] {
let start = Instant::now();
let (result, peak) = peak_during(|| {
let mut r = VlResolver::new(&file, 8, 8);
(r.string_bytes(&raw), r.sequences(&raw, 1).map(|s| s.len()))
});
let took = start.elapsed();
// libhdf5 never writes overlapping collections, and refuses these
// files; so do we, rather than returning what they claim.
let (strings, sequences) = result;
let e = strings.expect_err(name).to_string();
assert!(e.contains("overlaps"), "{name}: {e}");
assert!(sequences.is_err(), "{name}");
// Measured before the fix: 129 MB ("nested", 64 KB file) and 350 MB
// ("shared tail", 176 KB file) live at the peak; after, 97 KB and
// 0.9 MB.
assert!(
peak < 4 * file.len() + (1 << 20),
"{name}: peak {peak} bytes for a {}-byte file",
file.len()
);
assert!(took < Duration::from_secs(5), "{name}: took {took:?}");
}
}
/// Collections that do not overlap still read, however many elements point
/// into them, and the first object of a collection is returned for its
/// index (as before).
#[test]
fn separate_collections_still_read() {
let mut file = vec![0u8; 64 + 3 * 64];
let mut raw = Vec::new();
for i in 0..3usize {
let at = 64 + 64 * i;
put_header(&mut file, at, 64);
put_object(&mut file, at + 16, 1, 3);
file[at + 32..at + 35].copy_from_slice(format!("s{i}!").as_bytes());
raw.extend(element(3, at as u64, 1));
}
raw.extend(element(3, 64, 1));
let mut r = VlResolver::new(&file, 8, 8);
assert_eq!(r.strings(&raw).unwrap(), ["s0!", "s1!", "s2!", "s0!"]);
}
@@ -350,3 +350,30 @@ ds.close()
let press_vals = press_var.read_raw_f32().unwrap(); let press_vals = press_var.read_raw_f32().unwrap();
assert_eq!(press_vals, vec![1000.0f32, 850.0, 500.0, 200.0]); assert_eq!(press_vals, vec![1000.0f32, 850.0, 500.0, 200.0]);
} }
#[test]
fn netcdf4_python_string_variable_clawhdf5_reads() {
// NC_STRING variables are HDF5 variable-length strings, which
// `read_string` refused ("expected String, got VariableLength") until
// 2026-09-26.
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("strings.nc");
let path_str = path.display().to_string();
let script = format!(
r#"
import netCDF4 as nc
import numpy as np
ds = nc.Dataset("{path_str}", "w", format="NETCDF4")
ds.createDimension("station", 4)
v = ds.createVariable("name", str, ("station",))
v[:] = np.array(["Oslo", "", "São Paulo", "x"], dtype=object)
ds.close()
"#
);
run_python(&script);
let file = NetCDF4File::open(&path).unwrap();
let names = file.variable("name").unwrap().read_string().unwrap();
assert_eq!(names, vec!["Oslo", "", "São Paulo", "x"]);
}
+15 -7
View File
@@ -76,8 +76,13 @@ print the same bytes as h5dump 1.14.6 and as Debian's h5dump 1.14.5 (the
was run in that image on 2026-09-26) — `dump_matches_h5dump` in was run in that image on 2026-09-26) — `dump_matches_h5dump` in
`tests/h5rs_interop.rs` checks this, and `dump_shows_nul_padding_in_nested_strings` `tests/h5rs_interop.rs` checks this, and `dump_shows_nul_padding_in_nested_strings`
that null-padded strings show their NULs (`"a\000b"`) at any depth, as that null-padded strings show their NULs (`"a\000b"`) at any depth, as
h5dump's do. Not covered by those tests: references, opaque, bitfield, h5dump's do. `dump_prints_vl_data_like_h5dump` covers variable-length
variable-length sequences and virtual datasets. Known differences from strings (one with an embedded NUL, which prints up to the NUL; empty; null,
which prints `NULL`), variable-length sequences, a VL compound member and
a VL attribute, with 8- and 4-byte offsets. Not covered by those tests:
references, opaque, bitfield, non-ASCII UTF-8 (h5dump prints each byte
above 0x7f as a sign-extended octal escape, h5rs the character) and
virtual datasets. Known differences from
h5dump: h5dump:
- Floats print at their own precision (a `float32` 0.1 prints as `0.1`), - Floats print at their own precision (a `float32` 0.1 prints as `0.1`),
@@ -233,9 +238,11 @@ extension) and checks:
(which catches corrupt compressed data and Fletcher-32 mismatches), and (which catches corrupt compressed data and Fletcher-32 mismatches), and
follows every variable-length element (strings and sequences, also inside follows every variable-length element (strings and sequences, also inside
compounds and arrays) of every dataset and attribute into its global heap compounds and arrays) of every dataset and attribute into its global heap
collection: a collection that does not parse, a missing heap object, or a collection: a collection that does not parse or overlaps another, a missing
sequence longer than its heap object is a problem at the collection's heap object, or a heap object whose size is not exactly the element's
address. Data the length times its base size (libhdf5 refuses such an element) is a problem
at the collection's address. Variable-length elements are resolved by the
library's `VlResolver`, as `clawhdf5::File` resolves them. Data the
tool cannot decode (a filter it does not implement, such as szip, or a tool cannot decode (a filter it does not implement, such as szip, or a
dataset over `--max-bytes`) is a `note:`, not a problem. Every problem is dataset over `--max-bytes`) is a `note:`, not a problem. Every problem is
printed with the address of the structure involved; the exit status is 0 printed with the address of the structure involved; the exit status is 0
@@ -252,9 +259,10 @@ none at all without `--data`), and objects reachable only by external links. It
clawhdf5's parsers, so it accepts what they accept: some header damage that clawhdf5's parsers, so it accepts what they accept: some header damage that
libhdf5 refuses goes unreported. Of the 150 CVE and fuzzer files of the libhdf5 refuses goes unreported. Of the 150 CVE and fuzzer files of the
HDF Group's `cve_hdf5` corpus (`cvefiles/` and `fuzzerfiles/`), HDF Group's `cve_hdf5` corpus (`cvefiles/` and `fuzzerfiles/`),
`check --data` passes 16, and h5dump 1.14.6 rejects 9 of those (tank, `check --data` passes 15, and h5dump 1.14.6 rejects 8 of those (tank,
2026-09-26, `h5rs check --data F` and `h5dump F` per file; before the 2026-09-26, `h5rs check --data F` and `h5dump F` per file; before the
library's header checks it passed 28, of which h5dump rejects 21). library's header checks it passed 28, of which h5dump rejects 21, and 16
and 9 before a VL type's stored element size was checked).
## Robustness ## Robustness
+41 -41
View File
@@ -15,11 +15,13 @@ use clawhdf5_format::btree_v2::{BTreeV2Header, collect_btree_v2_records};
use clawhdf5_format::data_layout::DataLayout; use clawhdf5_format::data_layout::DataLayout;
use clawhdf5_format::dataspace::{Dataspace, DataspaceType}; use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
use clawhdf5_format::datatype::Datatype; use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::error::FormatError;
use clawhdf5_format::group_info::GroupInfoMessage; use clawhdf5_format::group_info::GroupInfoMessage;
use clawhdf5_format::link_info::LinkInfoMessage; use clawhdf5_format::link_info::LinkInfoMessage;
use clawhdf5_format::message_type::MessageType; use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader; use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::symbol_table::SymbolTableMessage; use clawhdf5_format::symbol_table::SymbolTableMessage;
use clawhdf5_format::vl_data::{VlResolver, check_element_size, parse_vl_references};
use crate::cli::{Args, Out}; use crate::cli::{Args, Out};
use crate::h5::{Error, ErrorKind, H5, Kind}; use crate::h5::{Error, ErrorKind, H5, Kind};
@@ -49,6 +51,15 @@ found, 3 internal error.";
const MAX_CHUNKS_CHECKED: usize = 10_000_000; const MAX_CHUNKS_CHECKED: usize = 10_000_000;
/// A variable-length element's problem, worded as `check` reports heap
/// problems ("global heap ...").
fn heap_problem(e: FormatError) -> String {
match e {
FormatError::VlDataError(m) if m.starts_with("global heap") => m,
e => format!("global heap: {e}"),
}
}
/// Whether values of `dt` hold variable-length data (in the global heap). /// Whether values of `dt` hold variable-length data (in the global heap).
fn has_vl(dt: &Datatype, depth: u32) -> bool { fn has_vl(dt: &Datatype, depth: u32) -> bool {
if depth > 32 { if depth > 32 {
@@ -104,6 +115,8 @@ struct Checker<'a> {
btrees_seen: HashSet<u64>, btrees_seen: HashSet<u64>,
/// Global heap collections already read (with --data). /// Global heap collections already read (with --data).
gcols_seen: HashSet<u64>, gcols_seen: HashSet<u64>,
/// Resolves variable-length elements (with --data), for the whole file.
vl: VlResolver<'a>,
panicked: bool, panicked: bool,
} }
@@ -157,6 +170,7 @@ pub fn run(args: &mut Args, out: &mut Out) -> std::io::Result<i32> {
heaps_seen: HashSet::new(), heaps_seen: HashSet::new(),
btrees_seen: HashSet::new(), btrees_seen: HashSet::new(),
gcols_seen: HashSet::new(), gcols_seen: HashSet::new(),
vl: VlResolver::new(h5.data(), h5.os(), h5.ls()),
panicked: false, panicked: false,
}; };
c.superblock(); c.superblock();
@@ -707,62 +721,48 @@ impl Checker<'_> {
} }
match dt { match dt {
Datatype::VariableLength { Datatype::VariableLength {
size,
is_string, is_string,
base_type, base_type,
.. ..
} => { } => {
let os = usize::from(self.h5.os()); // Resolved by the library's VlResolver, as every other
let (Some(lenb), Some(addrb), Some(idxb)) = // reader resolves them (and as libhdf5 does): a heap object
(b.get(..4), b.get(4..4 + os), b.get(4 + os..8 + os)) // whose size is not the element's length × base size, a
else { // collection that overlaps another, or a missing object is
// a problem at the collection's address.
let Ok(vl) = parse_vl_references(b, 1, self.h5.os()) else {
return; return;
}; };
let le = |x: &[u8]| { let gcol = vl[0].collection_address;
x.iter() if gcol == 0 || bad.contains_key(&gcol) {
.enumerate()
.fold(0u64, |a, (i, &v)| a | (u64::from(v) << (8 * i)))
};
let (len, gcol, idx) = (le(lenb), le(addrb), le(idxb));
let undef = if os >= 8 {
u64::MAX
} else {
(1u64 << (8 * os)) - 1
};
if len == 0 || gcol == 0 || gcol == undef || bad.contains_key(&gcol) {
return; return;
} }
let obj = match self.h5.heap_object(gcol, idx as u32) { if let Err(e) = check_element_size(*size, self.h5.os()) {
Ok(o) => o, bad.insert(gcol, e.to_string());
return;
}
let bs = if *is_string {
1
} else {
base_type.type_size() as usize
};
if bs == 0 {
return;
}
let obj = match self.vl.element(b, bs) {
Ok(o) => o.unwrap_or(&[]),
Err(e) => { Err(e) => {
bad.insert(e.addr.unwrap_or(gcol), e.msg); bad.insert(gcol, heap_problem(e));
return; return;
} }
}; };
if self.gcols_seen.insert(gcol) { if self.gcols_seen.insert(gcol) {
self.counts.global_heaps += 1; self.counts.global_heaps += 1;
} }
let bs = if *is_string { if !*is_string && has_vl(base_type, depth + 1) {
1 for eb in obj.chunks_exact(bs) {
} else { self.vl_element(base_type, eb, depth + 1, bad);
u64::from(base_type.type_size())
};
if len
.checked_mul(bs)
.is_none_or(|need| need > obj.len() as u64)
{
bad.insert(
gcol,
format!(
"global heap object {idx} holds {} bytes; the element needs {len} x {bs}",
obj.len()
),
);
return;
}
if !*is_string && bs > 0 && has_vl(base_type, depth + 1) {
let bs = bs as usize;
for k in 0..len as usize {
self.vl_element(base_type, &obj[k * bs..(k + 1) * bs], depth + 1, bad);
} }
} }
} }
+3
View File
@@ -699,6 +699,9 @@ impl Diff {
} }
match (x, y) { match (x, y) {
(Value::Str(p), Value::Str(q)) => p == q, (Value::Str(p), Value::Str(q)) => p == q,
// h5diff compares a null VL string equal to an empty one.
(Value::NullStr, Value::NullStr) => true,
(Value::NullStr, Value::Str(s)) | (Value::Str(s), Value::NullStr) => s.is_empty(),
(Value::Bytes(p), Value::Bytes(q)) | (Value::OtherRef(p), Value::OtherRef(q)) => p == q, (Value::Bytes(p), Value::Bytes(q)) | (Value::OtherRef(p), Value::OtherRef(q)) => p == q,
(Value::Compound(p), Value::Compound(q)) => { (Value::Compound(p), Value::Compound(q)) => {
p.len() == q.len() p.len() == q.len()
-27
View File
@@ -8,7 +8,6 @@
use std::cell::RefCell; use std::cell::RefCell;
use std::collections::HashMap; use std::collections::HashMap;
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};
use std::rc::Rc;
use clawhdf5::File; use clawhdf5::File;
use clawhdf5_format::attribute::{AttributeMessage, extract_attributes_tolerant}; use clawhdf5_format::attribute::{AttributeMessage, extract_attributes_tolerant};
@@ -20,7 +19,6 @@ use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::error::FormatError; use clawhdf5_format::error::FormatError;
use clawhdf5_format::filter_pipeline::FilterPipeline; use clawhdf5_format::filter_pipeline::FilterPipeline;
use clawhdf5_format::fractal_heap::FractalHeapHeader; use clawhdf5_format::fractal_heap::FractalHeapHeader;
use clawhdf5_format::global_heap::GlobalHeapCollection;
use clawhdf5_format::group_v1; use clawhdf5_format::group_v1;
use clawhdf5_format::link_info::LinkInfoMessage; use clawhdf5_format::link_info::LinkInfoMessage;
use clawhdf5_format::link_message::{LinkMessage, LinkTarget}; use clawhdf5_format::link_message::{LinkMessage, LinkTarget};
@@ -191,7 +189,6 @@ pub struct H5 {
pub path: PathBuf, pub path: PathBuf,
pub file: File, pub file: File,
pub max_bytes: u64, pub max_bytes: u64,
heaps: RefCell<HashMap<u64, std::result::Result<Rc<GlobalHeapCollection>, String>>>,
/// Fractal heaps whose blocks were verified: `None` = sound. /// Fractal heaps whose blocks were verified: `None` = sound.
verified_heaps: RefCell<HashMap<u64, Option<Error>>>, verified_heaps: RefCell<HashMap<u64, Option<Error>>>,
} }
@@ -211,7 +208,6 @@ impl H5 {
path: path.to_path_buf(), path: path.to_path_buf(),
file, file,
max_bytes: DEFAULT_MAX_BYTES, max_bytes: DEFAULT_MAX_BYTES,
heaps: RefCell::new(HashMap::new()),
verified_heaps: RefCell::new(HashMap::new()), verified_heaps: RefCell::new(HashMap::new()),
}) })
} }
@@ -433,29 +429,6 @@ impl H5 {
r.map_or(Ok(()), Err) r.map_or(Ok(()), Err)
} }
/// The global heap object `idx` of the collection at `addr` (cached per
/// collection).
pub fn heap_object(&self, addr: u64, idx: u32) -> Result<Vec<u8>> {
let coll = {
let mut cache = self.heaps.borrow_mut();
cache
.entry(addr)
.or_insert_with(|| match usize::try_from(addr) {
Ok(a) => GlobalHeapCollection::parse(self.data(), a, self.ls())
.map(Rc::new)
.map_err(|e| e.to_string()),
Err(_) => Err("address out of range".into()),
})
.clone()
.map_err(|e| Error::at(addr, format!("global heap: {e}")))?
};
let idx16 = u16::try_from(idx)
.map_err(|_| Error::at(addr, format!("global heap object index {idx} out of range")))?;
coll.get_object(idx16)
.map(|o| o.data.clone())
.ok_or_else(|| Error::at(addr, format!("global heap has no object {idx}")))
}
/// The dataspace of the dataset at `path` with a virtual dataset's /// The dataspace of the dataset at `path` with a virtual dataset's
/// extent resolved from its sources (as libhdf5 reports it) instead of /// extent resolved from its sources (as libhdf5 reports it) instead of
/// the stored one. /// the stored one.
+43 -55
View File
@@ -3,7 +3,10 @@
//! Decoding never panics: a short buffer, an unknown byte order or a //! Decoding never panics: a short buffer, an unknown byte order or a
//! dangling heap reference becomes [`Value::Error`]. //! dangling heap reference becomes [`Value::Error`].
use std::cell::RefCell;
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder, ReferenceType, StringPadding}; use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder, ReferenceType, StringPadding};
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
use serde_json::Value as J; use serde_json::Value as J;
use crate::dtype; use crate::dtype;
@@ -16,6 +19,9 @@ pub enum Value {
/// its own precision. /// its own precision.
Float(f64, u8), Float(f64, u8),
Str(String), Str(String),
/// A null variable-length string (heap address 0): h5dump prints it as
/// `NULL`, h5py reads it as empty.
NullStr,
/// Opaque, bitfield, time and oversized integers. /// Opaque, bitfield, time and oversized integers.
Bytes(Vec<u8>), Bytes(Vec<u8>),
/// An enum member (name, when the value matches one) and its value. /// An enum member (name, when the value matches one) and its value.
@@ -129,10 +135,10 @@ fn decode_float(dt: &Datatype, b: &[u8]) -> Value {
} }
} }
fn trim_string(b: &[u8], pad: Option<&StringPadding>) -> String { fn trim_string(b: &[u8], pad: &StringPadding) -> String {
let cut = b.iter().position(|&c| c == 0).unwrap_or(b.len()); let cut = b.iter().position(|&c| c == 0).unwrap_or(b.len());
let mut s = &b[..cut]; let mut s = &b[..cut];
if matches!(pad, Some(StringPadding::SpacePad)) { if matches!(pad, StringPadding::SpacePad) {
while let [rest @ .., b' '] = s { while let [rest @ .., b' '] = s {
s = rest; s = rest;
} }
@@ -140,22 +146,20 @@ fn trim_string(b: &[u8], pad: Option<&StringPadding>) -> String {
String::from_utf8_lossy(s).into_owned() String::from_utf8_lossy(s).into_owned()
} }
/// Little-endian unsigned integer of `b` (up to 8 bytes).
fn le(b: &[u8]) -> u64 {
b.iter()
.take(8)
.enumerate()
.fold(0u64, |a, (i, &x)| a | (u64::from(x) << (8 * i)))
}
/// Decodes elements of one file. /// Decodes elements of one file.
pub struct Decoder<'a> { pub struct Decoder<'a> {
pub h5: &'a H5, pub h5: &'a H5,
/// Variable-length elements are resolved as the library resolves them
/// (so as libhdf5 does), not by a decoder of our own.
vl: RefCell<VlResolver<'a>>,
} }
impl<'a> Decoder<'a> { impl<'a> Decoder<'a> {
pub fn new(h5: &'a H5) -> Self { pub fn new(h5: &'a H5) -> Self {
Self { h5 } Self {
h5,
vl: RefCell::new(VlResolver::new(h5.data(), h5.os(), h5.ls())),
}
} }
/// Decode element `i` of `raw`, an array of `dt` elements. /// Decode element `i` of `raw`, an array of `dt` elements.
@@ -187,7 +191,7 @@ impl<'a> Decoder<'a> {
Datatype::Time { .. } | Datatype::BitField { .. } | Datatype::Opaque { .. } => { Datatype::Time { .. } | Datatype::BitField { .. } | Datatype::Opaque { .. } => {
Value::Bytes(b.to_vec()) Value::Bytes(b.to_vec())
} }
Datatype::String { padding, .. } => Value::Str(trim_string(b, Some(padding))), Datatype::String { padding, .. } => Value::Str(trim_string(b, padding)),
Datatype::Compound { members, .. } => { Datatype::Compound { members, .. } => {
let mut out = Vec::with_capacity(members.len()); let mut out = Vec::with_capacity(members.len());
for m in members { for m in members {
@@ -246,61 +250,43 @@ impl<'a> Decoder<'a> {
Value::Array(out) Value::Array(out)
} }
Datatype::VariableLength { Datatype::VariableLength {
size,
is_string, is_string,
padding,
base_type, base_type,
.. ..
} => self.decode_vlen(*is_string, padding.as_ref(), base_type, b, depth), } => match check_element_size(*size, self.h5.os()) {
Ok(()) => self.decode_vlen(*is_string, base_type, b, depth),
Err(e) => Value::Error(e.to_string()),
},
} }
} }
fn decode_vlen( /// A variable-length element, resolved by the library's
&self, /// [`VlResolver`]: a string ends at its first NUL, a heap object whose
is_string: bool, /// size is not the element's length × base size is an error, and a
padding: Option<&StringPadding>, /// heap address of 0 is null — all as libhdf5 (and so h5dump and h5py)
base: &Datatype, /// has it.
b: &[u8], fn decode_vlen(&self, is_string: bool, base: &Datatype, b: &[u8], depth: u32) -> Value {
depth: u32,
) -> Value {
let os = usize::from(self.h5.os());
let (Some(lenb), Some(addrb), Some(idxb)) =
(b.get(..4), b.get(4..4 + os), b.get(4 + os..8 + os))
else {
return Value::Error("short VL element".into());
};
let len = le(lenb) as usize;
let addr = le(addrb);
let idx = le(idxb) as u32;
let undef = if os >= 8 {
u64::MAX
} else {
(1u64 << (8 * os)) - 1
};
let obj = if len == 0 || addr == 0 || addr == undef {
Vec::new()
} else {
match self.h5.heap_object(addr, idx) {
Ok(o) => o,
Err(e) => return Value::Error(e.to_string()),
}
};
if is_string { if is_string {
let l = len.min(obj.len()); return match self.vl.borrow_mut().string_element(b) {
return Value::Str(trim_string(&obj[..l], padding)); Ok(Some(s)) => Value::Str(String::from_utf8_lossy(s).into_owned()),
Ok(None) => Value::NullStr,
Err(e) => Value::Error(e.to_string()),
};
} }
let bs = base.type_size() as usize; let bs = base.type_size() as usize;
if bs == 0 { if bs == 0 {
return Value::Error("VL base type of size 0".into()); return Value::Error("VL base type of size 0".into());
} }
match len.checked_mul(bs) { let obj = match self.vl.borrow_mut().element(b, bs) {
Some(need) if need <= obj.len() => {} Ok(o) => o.unwrap_or(&[]),
_ => return Value::Error("VL sequence longer than its heap object".into()), Err(e) => return Value::Error(e.to_string()),
} };
let mut out = Vec::with_capacity(len); Value::Seq(
for k in 0..len { obj.chunks_exact(bs)
out.push(self.decode(base, &obj[k * bs..], depth + 1)); .map(|e| self.decode(base, e, depth + 1))
} .collect(),
Value::Seq(out) )
} }
} }
@@ -351,6 +337,7 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
Value::Int(i) => i.to_string(), Value::Int(i) => i.to_string(),
Value::Float(f, w) => fmt_float(*f, *w), Value::Float(f, w) => fmt_float(*f, *w),
Value::Str(s) => format!("\"{}\"", escape(s)), Value::Str(s) => format!("\"{}\"", escape(s)),
Value::NullStr => "NULL".into(),
Value::Bytes(b) => hex(b), Value::Bytes(b) => hex(b),
Value::Enum(Some(n), _) => n.clone(), Value::Enum(Some(n), _) => n.clone(),
Value::Enum(None, i) => i.to_string(), Value::Enum(None, i) => i.to_string(),
@@ -411,6 +398,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
} }
} }
Value::Str(s) => J::from(s.as_str()), Value::Str(s) => J::from(s.as_str()),
Value::NullStr => J::from(""),
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)), Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths), Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()), Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
+128
View File
@@ -0,0 +1,128 @@
"""Write the variable-length data files the h5rs VL tests run on.
usage: gen_vl_files.py OUTDIR
For 8-byte (`vl8`) and 4-byte (`vl4`) offsets, writes OUTDIR/vl8.h5 and
OUTDIR/vl4.h5, which libhdf5 reads in full, and OUTDIR/bad8.h5 and
OUTDIR/bad4.h5, whose `bad` and `badseq` elements 0 have a length that
disagrees with their global heap object (libhdf5: "Expected global heap
object size does not match"), and whose `undef` element 1 has length 0 and
the undefined heap address (libhdf5: "addr undefined"). h5py cannot write a VL string with a NUL in
it or a null element in a contiguous dataset, so those are patched in.
Prints one JSON object: for each file, each dataset's values as h5py reads
them one element at a time (strings as text, sequences as lists, a compound
as a list of its fields), with null for an element h5py cannot read; the
root attribute `va`; and the addresses of the `bad` elements' collections.
"""
import json
import os
import struct
import sys
import h5py
import numpy as np
out = sys.argv[1]
S = h5py.string_dtype("utf-8")
I4 = h5py.vlen_dtype(np.dtype("<i4"))
def create(path, sizes):
if sizes is None:
return h5py.File(path, "w")
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE)
fcpl.set_sizes(*sizes)
return h5py.File(h5py.h5f.create(path.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl))
def element(length, addr, index, os_):
return struct.pack("<I", length) + addr.to_bytes(os_, "little") + struct.pack("<I", index)
def good(path, sizes):
os_ = 8 if sizes is None else sizes[0]
with create(path, sizes) as f:
f.create_dataset(
"d", data=np.array(["aXb", "", "ok", "zz", "hello"], dtype=object), dtype=S
)
u = f.create_dataset("u", shape=(4,), dtype=S, chunks=(1,))
u[1] = "w"
s = f.create_dataset("seq", shape=(3,), dtype=I4)
s[0] = [1, 2, 3]
s[1] = []
s[2] = [-5]
s = f.create_dataset("sequ", shape=(3,), dtype=I4, chunks=(1,))
s[0] = [7, 8]
ct = np.dtype([("id", "<i4"), ("name", S)])
arr = np.zeros(3, dtype=ct)
arr["id"] = [1, 2, 3]
arr["name"] = ["one", "", "three"]
f.create_dataset("cmp", data=arr)
f.attrs.create("va", np.array(["p", "", "q"], dtype=object), dtype=S)
off = f["d"].id.get_offset()
b = bytearray(open(path, "rb").read())
i = b.index(b"aXb")
b[i + 1] = 0 # "a\0b"
es = 8 + os_
b[off + 2 * es : off + 3 * es] = element(2, 0, 1, os_) # "ok" -> null
open(path, "wb").write(bytes(b))
def bad(path, sizes):
os_ = 8 if sizes is None else sizes[0]
with create(path, sizes) as f:
f.create_dataset("bad", data=np.array(["cdefgh", "ok"], dtype=object), dtype=S)
s = f.create_dataset("badseq", shape=(2,), dtype=I4)
s[0] = [1, 2, 3]
s[1] = [4]
f.create_dataset("undef", data=np.array(["x", "", "yz"], dtype=object), dtype=S)
off, soff = f["bad"].id.get_offset(), f["badseq"].id.get_offset()
uoff = f["undef"].id.get_offset()
b = bytearray(open(path, "rb").read())
gcol = int.from_bytes(b[off + 4 : off + 4 + os_], "little")
struct.pack_into("<I", b, off, 3) # "cdefgh": length 6 -> 3
struct.pack_into("<I", b, soff, 2) # [1, 2, 3]: length 3 -> 2
# "": length 0 at the undefined address (all 0xff), which libhdf5 fails
# to read ("addr undefined"); it writes a null element as address 0.
es = 8 + os_
b[uoff + es : uoff + 2 * es] = element(0, (1 << (8 * os_)) - 1, 1, os_)
open(path, "wb").write(bytes(b))
return gcol
def value(v):
if isinstance(v, bytes):
return v.decode()
if isinstance(v, str):
return v
if isinstance(v, np.void):
return [value(x) for x in v]
if isinstance(v, np.ndarray):
return [value(x) for x in v]
return v.item() if hasattr(v, "item") else v
def read(ds):
got = []
for i in range(ds.shape[0]):
try:
got.append(value(ds[i]))
except OSError:
got.append(None)
return got
result = {}
for tag, sizes in (("8", None), ("4", (4, 4))):
g, x = os.path.join(out, f"vl{tag}.h5"), os.path.join(out, f"bad{tag}.h5")
good(g, sizes)
gcol = bad(x, sizes)
with h5py.File(g, "r") as f:
result[f"vl{tag}"] = {n: read(f[n]) for n in ("d", "u", "seq", "sequ", "cmp")}
result[f"vl{tag}"]["va"] = [value(s) for s in f.attrs["va"]]
with h5py.File(x, "r") as f:
result[f"bad{tag}"] = {n: read(f[n]) for n in ("bad", "badseq", "undef")}
result[f"bad{tag}"]["gcol"] = gcol
json.dump(result, sys.stdout)
+135
View File
@@ -773,3 +773,138 @@ fn every_subcommand_rejects_a_non_hdf5_file_cleanly() {
assert_eq!(code(&h5rs(&["ls"])), 2); assert_eq!(code(&h5rs(&["ls"])), 2);
assert_eq!(code(&h5rs(&["--help"])), 0); assert_eq!(code(&h5rs(&["--help"])), 0);
} }
// ---------------------------------------------------------------------------
// variable-length data
// ---------------------------------------------------------------------------
/// Runs `tests/gen_vl_files.py`: VL strings (with an embedded NUL, empty
/// and null elements), VL sequences, a VL compound member and a VL
/// attribute, with 8- and 4-byte offsets, plus files whose heap objects
/// disagree with their elements' lengths.
fn generate_vl() -> Option<Files> {
if missing(python_available(), "python3 with h5py") {
return None;
}
let dir = tempfile::tempdir().unwrap();
let script = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/gen_vl_files.py");
let out = Command::new(python())
.arg(&script)
.arg(dir.path())
.output()
.expect("run gen_vl_files.py");
assert!(
out.status.success(),
"gen_vl_files.py failed:\n{}",
String::from_utf8_lossy(&out.stderr)
);
let values = serde_json::from_slice(&out.stdout).expect("gen_vl_files.py output");
Some(Files { dir, values })
}
/// `dump` resolves VL elements through the library's `VlResolver`, as
/// libhdf5 does: "a\0b" prints as "a", a null string as NULL (it printed
/// ""), and with 4-byte offsets too; the output is h5dump's byte for byte.
#[test]
fn dump_prints_vl_data_like_h5dump() {
let Some(f) = generate_vl() else { return };
for name in ["vl8.h5", "vl4.h5"] {
let p = f.p(name);
let ours = stdout(&h5rs(&["dump", &p]));
assert!(
ours.contains(r#"(0): "a", "", NULL, "zz", "hello""#),
"{name}:\n{ours}"
);
assert!(ours.contains(r#"(0): NULL, "w", NULL, NULL"#), "{name}");
assert!(ours.contains("(0): (1, 2, 3), (), (-5)"), "{name}");
if missing(tool_available("h5dump"), "h5dump") {
continue;
}
let reference = run("h5dump", &[&p]);
assert!(reference.status.success(), "{name}: {reference:?}");
assert_eq!(ours, stdout(&reference).replacen(&p, name, 1), "{name}");
}
}
/// `dump --json` gives the values h5py reads, element by element; and an
/// element whose heap object is not its length × base size is an error, as
/// in h5py, not a truncated value (it printed "cde" and (1, 2)); so is a
/// length-0 element at the undefined heap address (it printed "").
#[test]
fn dump_json_vl_values_match_h5py() {
let Some(f) = generate_vl() else { return };
for tag in ["8", "4"] {
let (good, bad) = (format!("vl{tag}"), format!("bad{tag}"));
let o = h5rs(&["dump", "--json", &f.p(&format!("{good}.h5"))]);
assert!(o.status.success(), "{good}: {o:?}");
let doc: serde_json::Value = serde_json::from_slice(&o.stdout).unwrap();
let want = &f.values[&good];
for d in doc["datasets"].as_object().unwrap().values() {
let path = d["alias"][0].as_str().unwrap();
assert_eq!(d["value"], want[&path[1..]], "{good}: {path}");
}
let attrs = &doc["groups"][doc["root"].as_str().unwrap()]["attributes"];
assert_eq!(attrs[0]["name"], "va");
assert_eq!(attrs[0]["value"], want["va"], "{good}: va");
let o = h5rs(&["dump", "--json", &f.p(&format!("{bad}.h5"))]);
let doc: serde_json::Value = serde_json::from_slice(&o.stdout).unwrap();
let want = &f.values[&bad];
for d in doc["datasets"].as_object().unwrap().values() {
let path = d["alias"][0].as_str().unwrap();
let got = d["value"].as_array().unwrap();
let want = want[&path[1..]].as_array().unwrap();
assert_eq!(got.len(), want.len(), "{bad}: {path}");
for (g, w) in got.iter().zip(want) {
if w.is_null() {
// h5py cannot read it: neither can we.
let e = g["error"]
.as_str()
.unwrap_or_else(|| panic!("{bad}: {path}: {g}"));
let why = if path == "/undef" {
"undefined"
} else {
"holds"
};
assert!(e.contains(why), "{bad}: {path}: {e}");
} else {
assert_eq!(g, w, "{bad}: {path}");
}
}
}
}
}
/// `check --data` holds VL elements to libhdf5's rule: a heap object whose
/// size is not exactly the element's length × base size is a problem (it
/// only caught objects shorter than the element), and so is an element at
/// the undefined heap address.
#[test]
fn check_data_flags_mis_sized_vl_heap_objects() {
let Some(f) = generate_vl() else { return };
for tag in ["8", "4"] {
let o = h5rs(&["check", "--data", &f.p(&format!("vl{tag}.h5"))]);
let s = stdout(&o);
assert_eq!(code(&o), 0, "vl{tag}: {s}");
assert!(
s.contains("global heap collections read: 1"),
"vl{tag}: {s}"
);
let o = h5rs(&["check", "--data", &f.p(&format!("bad{tag}.h5"))]);
let s = stdout(&o);
assert_eq!(code(&o), 1, "bad{tag}: {s}");
let at = f.values[format!("bad{tag}")]["gcol"].as_u64().unwrap();
for (path, what) in [("/bad", "6 bytes"), ("/badseq", "12 bytes")] {
let want = format!("problem: {at:#x} {path}: variable-length data: global heap object");
assert!(s.contains(&want), "bad{tag}: no {want:?} in\n{s}");
assert!(s.contains(what), "bad{tag}: {s}");
}
// A length-0 element at the undefined heap address: libhdf5 fails
// to read it; check skipped it.
let undef: u64 = if tag == "8" { u64::MAX } else { 0xffff_ffff };
let want = format!("problem: {undef:#x} /undef: variable-length data: global heap:");
assert!(s.contains(&want), "bad{tag}: no {want:?} in\n{s}");
assert!(s.contains("undefined global heap address"), "bad{tag}: {s}");
}
}
+12 -14
View File
@@ -9,6 +9,7 @@
use clawhdf5::{AttrValue, File, Selection}; use clawhdf5::{AttrValue, File, Selection};
use clawhdf5_format::data_read; use clawhdf5_format::data_read;
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder}; use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
/// Errors are reported to JavaScript as messages. /// Errors are reported to JavaScript as messages.
pub type Result<T> = std::result::Result<T, String>; pub type Result<T> = std::result::Result<T, String>;
@@ -221,6 +222,12 @@ impl Reader {
None => (Selection::All, shape.clone()), None => (Selection::All, shape.clone()),
Some(h) => hyperslab_selection(h, &shape)?, Some(h) => hyperslab_selection(h, &shape)?,
}; };
// A VL type whose stored element size is not the one the file's
// offset size implies is refused before its data is read, as
// `File::read_string` refuses it.
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
}
let raw = ds.read_selection(&selection).map_err(err)?; let raw = ds.read_selection(&selection).map_err(err)?;
let data = self.decode(&raw, &dt)?; let data = self.decode(&raw, &dt)?;
out_shape.extend(element_shape(&dt)); out_shape.extend(element_shape(&dt));
@@ -270,22 +277,13 @@ impl Reader {
Datatype::VariableLength { Datatype::VariableLength {
is_string: true, .. is_string: true, ..
} if !is_array => { } if !is_array => {
let size = dt.type_size() as usize; // The library's resolver, as File::read_string uses: a
if size == 0 || !raw.len().is_multiple_of(size) { // string ends at its first NUL and a heap object of the
return Err(format!( // wrong size is an error, as in libhdf5 and h5py.
"{} bytes is not a whole number of {size}-byte string references",
raw.len()
));
}
let sb = self.file.superblock(); let sb = self.file.superblock();
Data::Strings( Data::Strings(
clawhdf5_format::vl_data::read_vl_strings( VlResolver::new(self.file.as_bytes(), sb.offset_size, sb.length_size)
self.file.as_bytes(), .strings(raw)
raw,
(raw.len() / size) as u64,
sb.offset_size,
sb.length_size,
)
.map_err(err)?, .map_err(err)?,
) )
} }
+169
View File
@@ -0,0 +1,169 @@
//! The wasm reader resolves VL strings with the library's `VlResolver`, so
//! it returns what `File::read_string` and h5py return: a string ends at
//! its first NUL, a null element is empty, a heap object of the wrong size
//! is an error, an element at the undefined heap address is an error, and a
//! VL datatype whose stored element size disagrees with the file's offset
//! size is refused. Checked with 8- and 4-byte offsets.
//!
//! Skipped when python3 with h5py is missing, unless
//! `CLAWHDF5_REQUIRE_INTEROP=1`. `CLAWHDF5_PYTHON` names the interpreter.
use std::process::Command;
use clawhdf5::File;
use clawhdf5_wasm::core::{Data, Reader};
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn h5py_available() -> bool {
let ok = Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.map(|o| o.status.success())
.unwrap_or(false);
if !ok {
assert!(
std::env::var("CLAWHDF5_REQUIRE_INTEROP").as_deref() != Ok("1"),
"CLAWHDF5_REQUIRE_INTEROP=1 but python with h5py is not available"
);
eprintln!("SKIP: python with h5py not available");
}
ok
}
/// For each offset size: `vl{8,4}.h5` with dataset `d` = "a\0b", "", null,
/// "zz" (patched: h5py writes neither a NUL nor a null element);
/// `bad{8,4}.h5` whose element 0 claims 3 bytes of a 6-byte heap object;
/// `size{8,4}.h5` whose VL datatype message stores a 24-byte element; and
/// `undef{8,4}.h5` whose element 1 has length 0 and the undefined heap
/// address.
/// Prints h5py's reading of each element as hex, or "error".
const SCRIPT: &str = r#"
import struct, sys, h5py, numpy as np
out = sys.argv[1]
S = h5py.string_dtype('utf-8')
def create(path, os_):
if os_ == 8:
return h5py.File(path, 'w', libver='earliest')
# The earliest format, so the patched object header has no checksum.
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE); fcpl.set_sizes(4, 4)
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
fapl.set_libver_bounds(h5py.h5f.LIBVER_EARLIEST, h5py.h5f.LIBVER_V18)
return h5py.File(h5py.h5f.create(path.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl, fapl=fapl))
def elem(length, addr, index, os_):
return struct.pack('<I', length) + addr.to_bytes(os_, 'little') + struct.pack('<I', index)
def make(path, os_, values):
with create(path, os_) as f:
f.create_dataset('d', data=np.array(values, dtype=object), dtype=S)
return f['d'].id.get_offset()
for os_ in (8, 4):
es = 8 + os_
p = '%s/vl%d.h5' % (out, os_)
off = make(p, os_, ['aXb', '', 'ok', 'zz'])
b = bytearray(open(p, 'rb').read())
b[b.index(b'aXb') + 1] = 0
b[off + 2 * es:off + 3 * es] = elem(2, 0, 1, os_)
open(p, 'wb').write(bytes(b))
p = '%s/bad%d.h5' % (out, os_)
off = make(p, os_, ['cdefgh', 'ok'])
b = bytearray(open(p, 'rb').read())
struct.pack_into('<I', b, off, 3)
open(p, 'wb').write(bytes(b))
p = '%s/size%d.h5' % (out, os_)
make(p, os_, ['x', 'yy'])
b = bytearray(open(p, 'rb').read())
# datatype message: version 1, class 9 (VL); string, null-terminated, UTF-8
pat = bytes([0x19, 0x01, 0x01, 0x00]) + struct.pack('<I', es)
assert b.count(pat) == 1, b.count(pat)
i = b.index(pat)
struct.pack_into('<I', b, i + 4, 24)
open(p, 'wb').write(bytes(b))
p = '%s/undef%d.h5' % (out, os_)
off = make(p, os_, ['x', '', 'yz'])
b = bytearray(open(p, 'rb').read())
b[off + es:off + 2 * es] = elem(0, (1 << (8 * os_)) - 1, 1, os_)
open(p, 'wb').write(bytes(b))
for name in ('vl', 'bad', 'size', 'undef'):
with h5py.File('%s/%s%d.h5' % (out, name, os_), 'r') as f:
got = []
for i in range(f['d'].shape[0]):
try:
got.append(f['d'][i].hex())
except OSError:
got.append('error')
print('%s%d\t%s' % (name, os_, ','.join(got)))
"#;
#[test]
fn vl_strings_read_like_file_and_h5py() {
if !h5py_available() {
return;
}
let dir = tempfile::tempdir().unwrap();
let out = Command::new(python())
.args(["-c", SCRIPT])
.arg(dir.path())
.output()
.expect("run python");
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let h5py: std::collections::HashMap<String, String> = String::from_utf8(out.stdout)
.unwrap()
.lines()
.map(|l| {
let (k, v) = l.split_once('\t').unwrap();
(k.to_string(), v.to_string())
})
.collect();
let read = |name: &str| {
let bytes = std::fs::read(dir.path().join(format!("{name}.h5"))).unwrap();
let wasm = Reader::open(bytes).unwrap().read("/d", None);
let file = File::open(dir.path().join(format!("{name}.h5")))
.unwrap()
.dataset("d")
.unwrap()
.read_string();
(wasm, file)
};
for os in [8, 4] {
// h5py: "a\0b" is "a"; the null element (address 0) is empty.
assert_eq!(h5py[&format!("vl{os}")], "61,,,7a7a");
let (wasm, file) = read(&format!("vl{os}"));
let Data::Strings(wasm) = wasm.unwrap().data else {
panic!("vl{os}: not strings")
};
assert_eq!(wasm, ["a", "", "", "zz"], "vl{os}");
assert_eq!(wasm, file.unwrap(), "vl{os}");
// h5py refuses the mis-sized element; so do both readers.
assert_eq!(h5py[&format!("bad{os}")], "error,6f6b");
let (wasm, file) = read(&format!("bad{os}"));
assert!(wasm.unwrap_err().contains("holds 6 bytes"), "bad{os}");
assert!(file.is_err(), "bad{os}");
// libhdf5 ignores the stored element size and reads the values;
// File refuses the datatype rather than guess its layout, and the
// wasm reader now does the same (it read with the stored size).
assert_eq!(h5py[&format!("size{os}")], "78,7979");
let (wasm, file) = read(&format!("size{os}"));
let e = wasm.unwrap_err();
assert!(e.contains("stores 24-byte elements"), "size{os}: {e}");
assert!(file.is_err(), "size{os}");
// Length 0 at the undefined heap address: libhdf5 fails the read
// ("addr undefined"); both readers returned "".
assert_eq!(h5py[&format!("undef{os}")], "78,error,797a");
let (wasm, file) = read(&format!("undef{os}"));
let e = wasm.unwrap_err();
assert!(
e.contains("undefined global heap address"),
"undef{os}: {e}"
);
assert!(file.is_err(), "undef{os}");
}
}
+38 -2
View File
@@ -422,11 +422,47 @@ impl<'f, R: HDF5Read> LazyDataset<'f, R> {
Ok(data_read::read_as_u64(&raw, &dt)?) Ok(data_read::read_as_u64(&raw, &dt)?)
} }
/// Read all data as `String` values. /// Read all data as `String` values: fixed- or variable-length strings
/// (see [`Dataset::read_string`](crate::Dataset::read_string)).
pub fn read_string(&self) -> Result<Vec<String>, Error> { pub fn read_string(&self) -> Result<Vec<String>, Error> {
let raw = self.read_raw()?; let raw = self.read_raw()?;
let dt = self.datatype()?; let dt = self.datatype()?;
Ok(data_read::read_as_strings(&raw, &dt)?) crate::vlen::decode_strings(
self.file.hdf5_bytes(),
&dt,
&raw,
self.file.offset_size(),
self.file.length_size(),
)
}
/// Read a variable-length string dataset as the exact bytes of each
/// string (see
/// [`Dataset::read_string_bytes`](crate::Dataset::read_string_bytes)).
pub fn read_string_bytes(&self) -> Result<Vec<Vec<u8>>, Error> {
let raw = self.read_raw()?;
let dt = self.datatype()?;
crate::vlen::decode_string_bytes(
self.file.hdf5_bytes(),
&dt,
&raw,
self.file.offset_size(),
self.file.length_size(),
)
}
/// Read a variable-length sequence dataset as one `Vec<T>` per element
/// (see [`Dataset::read_vlen`](crate::Dataset::read_vlen)).
pub fn read_vlen<T: crate::vlen::VlenValue>(&self) -> Result<Vec<Vec<T>>, Error> {
let raw = self.read_raw()?;
let dt = self.datatype()?;
crate::vlen::decode_vlen(
self.file.hdf5_bytes(),
&dt,
&raw,
self.file.offset_size(),
self.file.length_size(),
)
} }
/// Read all attributes of this dataset. /// Read all attributes of this dataset.
+2
View File
@@ -30,6 +30,7 @@ pub mod lazy;
pub mod mmap_file; pub mod mmap_file;
pub mod reader; pub mod reader;
pub mod types; pub mod types;
pub mod vlen;
pub mod writer; pub mod writer;
pub use error::Error; pub use error::Error;
@@ -38,6 +39,7 @@ pub use lazy::{LazyDataset, LazyFile, LazyGroup};
pub use mmap_file::{MmapDataset, MmapFile, MmapGroup}; pub use mmap_file::{MmapDataset, MmapFile, MmapGroup};
pub use reader::{Dataset, File, Group}; pub use reader::{Dataset, File, Group};
pub use types::{AttrValue, DType}; pub use types::{AttrValue, DType};
pub use vlen::VlenValue;
pub use writer::FileBuilder; pub use writer::FileBuilder;
#[cfg(feature = "parallel")] #[cfg(feature = "parallel")]
pub use writer::{DatasetSpec, create_datasets_parallel}; pub use writer::{DatasetSpec, create_datasets_parallel};
+38 -2
View File
@@ -336,11 +336,47 @@ impl<'f> MmapDataset<'f> {
Ok(data_read::read_as_u64(&raw, &dt)?) Ok(data_read::read_as_u64(&raw, &dt)?)
} }
/// Read all data as `String` values. /// Read all data as `String` values: fixed- or variable-length strings
/// (see [`Dataset::read_string`](crate::Dataset::read_string)).
pub fn read_string(&self) -> Result<Vec<String>, Error> { pub fn read_string(&self) -> Result<Vec<String>, Error> {
let raw = self.read_raw()?; let raw = self.read_raw()?;
let dt = self.datatype()?; let dt = self.datatype()?;
Ok(data_read::read_as_strings(&raw, &dt)?) crate::vlen::decode_strings(
self.file.hdf5_bytes(),
&dt,
&raw,
self.file.offset_size(),
self.file.length_size(),
)
}
/// Read a variable-length string dataset as the exact bytes of each
/// string (see
/// [`Dataset::read_string_bytes`](crate::Dataset::read_string_bytes)).
pub fn read_string_bytes(&self) -> Result<Vec<Vec<u8>>, Error> {
let raw = self.read_raw()?;
let dt = self.datatype()?;
crate::vlen::decode_string_bytes(
self.file.hdf5_bytes(),
&dt,
&raw,
self.file.offset_size(),
self.file.length_size(),
)
}
/// Read a variable-length sequence dataset as one `Vec<T>` per element
/// (see [`Dataset::read_vlen`](crate::Dataset::read_vlen)).
pub fn read_vlen<T: crate::vlen::VlenValue>(&self) -> Result<Vec<Vec<T>>, Error> {
let raw = self.read_raw()?;
let dt = self.datatype()?;
crate::vlen::decode_vlen(
self.file.hdf5_bytes(),
&dt,
&raw,
self.file.offset_size(),
self.file.length_size(),
)
} }
/// For contiguous datasets, return a zero-copy slice into the mmap. /// For contiguous datasets, return a zero-copy slice into the mmap.
+99 -2
View File
@@ -262,6 +262,56 @@ impl File {
} }
} }
/// Decode the strings in `raw`, a buffer of elements of `datatype` read
/// from this file — for instance a variable-length string field of a
/// compound ([`clawhdf5_format::data_read::read_compound_fields`]) or an
/// [`AttrValue::Raw`] attribute. Variable-length strings are resolved in
/// this file's global heap; see [`Dataset::read_string`] for the values.
pub fn decode_strings(&self, datatype: &Datatype, raw: &[u8]) -> Result<Vec<String>, Error> {
crate::vlen::decode_strings(
self.as_bytes(),
datatype,
raw,
self.offset_size(),
self.length_size(),
)
}
/// Like [`decode_strings`](Self::decode_strings) for variable-length
/// strings, returning each string's exact bytes (see
/// [`Dataset::read_string_bytes`]).
pub fn decode_string_bytes(
&self,
datatype: &Datatype,
raw: &[u8],
) -> Result<Vec<Vec<u8>>, Error> {
crate::vlen::decode_string_bytes(
self.as_bytes(),
datatype,
raw,
self.offset_size(),
self.length_size(),
)
}
/// Decode the variable-length sequences in `raw`, a buffer of elements
/// of the sequence type `datatype` read from this file (a compound
/// field, an [`AttrValue::Raw`] attribute, ...). See
/// [`Dataset::read_vlen`].
pub fn decode_vlen<T: crate::vlen::VlenValue>(
&self,
datatype: &Datatype,
raw: &[u8],
) -> Result<Vec<Vec<T>>, Error> {
crate::vlen::decode_vlen(
self.as_bytes(),
datatype,
raw,
self.offset_size(),
self.length_size(),
)
}
fn parse_header(&self, address: u64) -> Result<ObjectHeader, FormatError> { fn parse_header(&self, address: u64) -> Result<ObjectHeader, FormatError> {
ObjectHeader::parse( ObjectHeader::parse(
self.data.as_bytes(), self.data.as_bytes(),
@@ -498,11 +548,58 @@ impl<'f> Dataset<'f> {
Ok(data_read::read_as_u64(&raw, &dt)?) Ok(data_read::read_as_u64(&raw, &dt)?)
} }
/// Read all data as `String` values. /// Read all data as `String` values, in row-major order.
///
/// Works for fixed-length and variable-length string datasets (h5py's
/// default `str` dtype). A variable-length string ends at its first NUL
/// and a null element (e.g. never written) is `""`, as h5py returns
/// them; bytes that are not valid UTF-8 are replaced with U+FFFD — use
/// [`read_string_bytes`](Self::read_string_bytes) for the exact bytes.
pub fn read_string(&self) -> Result<Vec<String>, Error> { pub fn read_string(&self) -> Result<Vec<String>, Error> {
let raw = self.read_raw()?; let raw = self.read_raw()?;
let dt = self.datatype()?; let dt = self.datatype()?;
Ok(data_read::read_as_strings(&raw, &dt)?) self.file.decode_strings(&dt, &raw)
}
/// Read a variable-length string dataset as the exact bytes of each
/// string (what h5py's `Dataset[()]` returns), in row-major order.
pub fn read_string_bytes(&self) -> Result<Vec<Vec<u8>>, Error> {
let raw = self.read_raw()?;
let dt = self.datatype()?;
self.file.decode_string_bytes(&dt, &raw)
}
/// Read the selected elements of a fixed- or variable-length string
/// dataset (see [`read_string`](Self::read_string)).
pub fn read_string_selection(
&self,
selection: &clawhdf5_format::selection::Selection,
) -> Result<Vec<String>, Error> {
let raw = self.read_selection(selection)?;
let dt = self.datatype()?;
self.file.decode_strings(&dt, &raw)
}
/// Read a variable-length sequence dataset (h5py
/// `vlen_dtype(np.int32)`, ...) as one `Vec<T>` per element, in
/// row-major order. The base type must be an integer or float type; it
/// is converted to `T` as [`read_f64`](Self::read_f64) and the other
/// typed readers convert. A null element is an empty sequence.
pub fn read_vlen<T: crate::vlen::VlenValue>(&self) -> Result<Vec<Vec<T>>, Error> {
let raw = self.read_raw()?;
let dt = self.datatype()?;
self.file.decode_vlen(&dt, &raw)
}
/// Read the selected elements of a variable-length sequence dataset
/// (see [`read_vlen`](Self::read_vlen)).
pub fn read_vlen_selection<T: crate::vlen::VlenValue>(
&self,
selection: &clawhdf5_format::selection::Selection,
) -> Result<Vec<Vec<T>>, Error> {
let raw = self.read_selection(selection)?;
let dt = self.datatype()?;
self.file.decode_vlen(&dt, &raw)
} }
// ----- Selection-based read methods ----- // ----- Selection-based read methods -----
+143
View File
@@ -0,0 +1,143 @@
//! Variable-length data: VL strings and VL sequences of numbers.
//!
//! A variable-length element stores a reference into the file's global heap;
//! these helpers resolve the references in a buffer of raw elements (from a
//! dataset read, a selection, a compound field or an [`AttrValue::Raw`]
//! attribute) against the file they came from.
//!
//! Values match libhdf5 (and h5py): a string ends at its first NUL, a null
//! element is an empty string or sequence, and a heap object whose size
//! disagrees with its element is an error rather than a truncated value.
//!
//! [`AttrValue::Raw`]: crate::AttrValue::Raw
use clawhdf5_format::data_read;
use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::error::FormatError;
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
use crate::error::Error;
mod sealed {
pub trait Sealed {}
}
/// A number type that [`Dataset::read_vlen`](crate::Dataset::read_vlen) can
/// return: the sequence's base type is converted to it as libhdf5 converts
/// numbers (the same rules as `read_f64`, `read_i64`, ...).
pub trait VlenValue: sealed::Sealed + Sized {
#[doc(hidden)]
fn decode(raw: &[u8], base: &Datatype) -> Result<Vec<Self>, FormatError>;
}
macro_rules! vlen_value {
($t:ty, $f:path) => {
impl sealed::Sealed for $t {}
impl VlenValue for $t {
fn decode(raw: &[u8], base: &Datatype) -> Result<Vec<Self>, FormatError> {
$f(raw, base)
}
}
};
}
vlen_value!(f64, data_read::read_as_f64);
vlen_value!(f32, data_read::read_as_f32);
vlen_value!(i64, data_read::read_as_i64);
vlen_value!(i32, data_read::read_as_i32);
vlen_value!(u64, data_read::read_as_u64);
fn class_name(dt: &Datatype) -> &'static str {
match dt {
Datatype::FixedPoint { .. } => "integer",
Datatype::FloatingPoint { .. } => "float",
Datatype::Time { .. } => "time",
Datatype::String { .. } => "fixed-length string",
Datatype::BitField { .. } => "bitfield",
Datatype::Opaque { .. } => "opaque",
Datatype::Compound { .. } => "compound",
Datatype::Reference { .. } => "reference",
Datatype::Enumeration { .. } => "enum",
Datatype::VariableLength {
is_string: true, ..
} => "variable-length string",
Datatype::VariableLength { .. } => "variable-length sequence",
Datatype::Array { .. } => "array",
}
}
/// The strings in `raw`, elements of `dt`: fixed-length strings decoded as
/// `read_string` always has, variable-length strings resolved in the heap.
pub(crate) fn decode_strings(
file_data: &[u8],
dt: &Datatype,
raw: &[u8],
offset_size: u8,
length_size: u8,
) -> Result<Vec<String>, Error> {
match dt {
Datatype::VariableLength {
size,
is_string: true,
..
} => {
check_element_size(*size, offset_size)?;
Ok(VlResolver::new(file_data, offset_size, length_size).strings(raw)?)
}
_ => Ok(data_read::read_as_strings(raw, dt)?),
}
}
/// The exact bytes of the variable-length strings in `raw`.
pub(crate) fn decode_string_bytes(
file_data: &[u8],
dt: &Datatype,
raw: &[u8],
offset_size: u8,
length_size: u8,
) -> Result<Vec<Vec<u8>>, Error> {
match dt {
Datatype::VariableLength {
size,
is_string: true,
..
} => {
check_element_size(*size, offset_size)?;
Ok(VlResolver::new(file_data, offset_size, length_size).string_bytes(raw)?)
}
other => Err(Error::Format(FormatError::TypeMismatch {
expected: "variable-length string",
actual: class_name(other),
})),
}
}
/// The sequences in `raw`, elements of the variable-length sequence type
/// `dt`, converted to `T`.
pub(crate) fn decode_vlen<T: VlenValue>(
file_data: &[u8],
dt: &Datatype,
raw: &[u8],
offset_size: u8,
length_size: u8,
) -> Result<Vec<Vec<T>>, Error> {
let Datatype::VariableLength {
size,
is_string: false,
base_type,
..
} = dt
else {
return Err(Error::Format(FormatError::TypeMismatch {
expected: "variable-length sequence",
actual: class_name(dt),
}));
};
check_element_size(*size, offset_size)?;
let base_size = base_type.type_size() as usize;
VlResolver::new(file_data, offset_size, length_size)
.sequences(raw, base_size)?
.iter()
.map(|bytes| Ok(T::decode(bytes, base_type)?))
.collect()
}
+494
View File
@@ -0,0 +1,494 @@
//! Variable-length data (VL strings and VL sequences) read through the
//! facade, checked against h5py/libhdf5.
//!
//! h5py writes each file — once with the default 8-byte offsets and once
//! with 4-byte offsets and lengths (`sizeof_addr = 4`) — and prints what
//! libhdf5 reads back; `File`, `MmapFile` and `LazyFile` must return the same
//! values. Skipped when python3 with h5py is unavailable, unless
//! `CLAWHDF5_REQUIRE_INTEROP=1`.
// `Selection::slice(&[0..1])` is one range per dimension, not a Vec of a range.
#![allow(clippy::single_range_in_vec_init)]
use std::collections::HashMap;
use std::path::Path;
use std::process::Command;
use clawhdf5::{AttrValue, File, LazyFile, MmapFile, Selection};
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn interop_required() -> bool {
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
}
fn python_available() -> bool {
Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
macro_rules! skip_if_no_python {
() => {
if !python_available() {
assert!(
!interop_required(),
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
);
eprintln!("SKIP: python3 with h5py not available");
return;
}
};
}
/// Run `script` and return its stdout as `key -> value`, one
/// `key<TAB>value` line per key.
fn run_python(script: &str) -> HashMap<String, String> {
let output = Command::new(python())
.args(["-c", script])
.output()
.expect("failed to run python");
assert!(
output.status.success(),
"python failed:\n{}",
String::from_utf8_lossy(&output.stderr)
);
String::from_utf8_lossy(&output.stdout)
.lines()
.filter_map(|line| {
let (k, v) = line.split_once('\t')?;
Some((k.to_string(), v.to_string()))
})
.collect()
}
/// `hex,hex,...` -> the strings' bytes.
fn parse_strings(v: &str) -> Vec<Vec<u8>> {
v.split(',')
.map(|h| {
(0..h.len())
.step_by(2)
.map(|i| u8::from_str_radix(&h[i..i + 2], 16).unwrap())
.collect()
})
.collect()
}
/// `1 2 3|| -5` -> sequences.
fn parse_seqs(v: &str) -> Vec<Vec<f64>> {
v.split('|')
.map(|s| s.split_whitespace().map(|x| x.parse().unwrap()).collect())
.collect()
}
fn utf8(bytes: &[Vec<u8>]) -> Vec<String> {
bytes
.iter()
.map(|b| String::from_utf8(b.clone()).unwrap())
.collect()
}
/// Writes `vl8.h5` (8-byte offsets) and `vl4.h5` (4-byte offsets and
/// lengths) into `dir` and prints h5py's reading of both.
const SCRIPT: &str = r#"
import sys, h5py, numpy as np
d = sys.argv[1]
S = h5py.string_dtype('utf-8'); A = h5py.string_dtype('ascii')
def make(path, sizes):
if sizes:
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE); fcpl.set_sizes(*sizes)
f = h5py.File(h5py.h5f.create(path.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl))
else:
f = h5py.File(path, 'w')
f.create_dataset('scalar_utf8', data='héllo', dtype=S)
f.create_dataset('scalar_ascii', data=b'hello', dtype=A)
f.create_dataset('d1', data=np.array(['a', '', 'ccc', 'δδ'], dtype=object), dtype=S)
f.create_dataset('d2', data=np.array([['x', 'yy', 'zzz'], ['', 'w', 'vv']], dtype=object), dtype=S)
f.create_dataset('chunked', data=np.array(['s%d' % i * (i % 5) for i in range(100)], dtype=object),
dtype=S, chunks=(7,), compression='gzip')
f.create_dataset('chunked2d', data=np.array([['r%dc%d' % (r, c) for c in range(9)] for r in range(11)], dtype=object),
dtype=S, chunks=(4, 4), compression='gzip', shuffle=True)
f.create_dataset('unwritten', shape=(5,), dtype=S, chunks=(2,))
p = f.create_dataset('partial', shape=(6,), dtype=S, chunks=(2,)); p[0] = 'first'; p[5] = 'last'
f.create_dataset('contig_empty', shape=(3,), dtype=S)
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE); dcpl.set_layout(h5py.h5d.COMPACT)
f.create_dataset('compact', data=np.array(['c1', '', 'c3'], dtype=object), dtype=S, dcpl=dcpl)
assert f['compact'].id.get_create_plist().get_layout() == h5py.h5d.COMPACT
f.attrs['vlattr'] = 'attr-value'
f.attrs.create('vlattr_arr', np.array(['p', 'qq', ''], dtype=object), dtype=S)
ct = np.dtype([('id', '<i4'), ('name', S), ('v', '<f8')])
arr = np.zeros(3, dtype=ct); arr['id'] = [1, 2, 3]; arr['name'] = ['one', '', 'three']; arr['v'] = [.5, 1.5, 2.5]
f.create_dataset('compound', data=arr)
f.attrs.create('compound_attr', arr)
v = f.create_dataset('vlen_i4', shape=(3,), dtype=h5py.vlen_dtype(np.dtype('<i4')))
v[0] = [1, 2, 3]; v[1] = []; v[2] = [-5]
v = f.create_dataset('vlen_f8', shape=(2, 2), dtype=h5py.vlen_dtype(np.dtype('<f8')), chunks=(1, 2), compression='gzip')
v[0, 0] = [1.5]; v[0, 1] = [2.5, 3.5]; v[1, 1] = [9.0]
v = f.create_dataset('vlen_u2_be', shape=(2,), dtype=h5py.vlen_dtype(np.dtype('>u2')))
v[0] = [1, 65535]; v[1] = [300]
f.attrs.create('vlen_attr', np.array([np.array([1, 2], dtype='<i8'), np.array([3], dtype='<i8')], dtype=object),
dtype=h5py.vlen_dtype(np.dtype('<i8')))
f.close()
def hexes(a):
return ','.join(bytes(x).hex() for x in np.asarray(a, dtype=object).ravel())
def seqs(a):
return '|'.join(' '.join(repr(float(x)) for x in s) for s in np.asarray(a, dtype=object).ravel())
for tag, sizes in (('8', None), ('4', (4, 4))):
path = '%s/vl%s.h5' % (d, tag)
make(path, sizes)
with h5py.File(path, 'r') as f:
for name in ('compact', 'scalar_utf8', 'scalar_ascii', 'd1', 'd2', 'chunked', 'chunked2d', 'unwritten',
'partial', 'contig_empty'):
v = f[name][()]
print('%s:%s\t%s' % (tag, name, hexes([v] if np.ndim(v) == 0 else v)))
print('%s:d2[1,1:3]\t%s' % (tag, hexes(f['d2'][1, 1:3])))
print('%s:chunked[5:60:3]\t%s' % (tag, hexes(f['chunked'][5:60:3])))
print('%s:chunked2d[2:9:2,3:8]\t%s' % (tag, hexes(f['chunked2d'][2:9:2, 3:8])))
print('%s:compound.name\t%s' % (tag, hexes(f['compound']['name'])))
print('%s:compound_attr.name\t%s' % (tag, hexes(f.attrs['compound_attr']['name'])))
print('%s:vlattr\t%s' % (tag, hexes([f.attrs['vlattr'].encode()])))
print('%s:vlattr_arr\t%s' % (tag, hexes([s.encode() for s in f.attrs['vlattr_arr']])))
for name in ('vlen_i4', 'vlen_f8'):
print('%s:%s\t%s' % (tag, name, seqs(f[name][()])))
print('%s:vlen_f8[1,:]\t%s' % (tag, seqs(f['vlen_f8'][1, :])))
print('%s:vlen_attr\t%s' % (tag, seqs(f.attrs['vlen_attr'])))
"#;
fn make_files(dir: &Path) -> HashMap<String, String> {
let script = format!(
"import sys; sys.argv = ['x', {:?}]\n{SCRIPT}",
dir.display().to_string()
);
run_python(&script)
}
const STRING_DATASETS: [&str; 10] = [
"compact",
"scalar_utf8",
"scalar_ascii",
"d1",
"d2",
"chunked",
"chunked2d",
"unwritten",
"partial",
"contig_empty",
];
#[test]
fn vl_string_datasets_read_like_h5py() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let expected = make_files(dir.path());
for tag in ["8", "4"] {
let path = dir.path().join(format!("vl{tag}.h5"));
let file = File::open(&path).unwrap();
let mmap = MmapFile::open(&path).unwrap();
let lazy = LazyFile::open_mmap(&path).unwrap();
for name in STRING_DATASETS {
let want = parse_strings(&expected[&format!("{tag}:{name}")]);
let ctx = format!("vl{tag}.h5 {name}");
let ds = file.dataset(name).unwrap();
assert_eq!(ds.read_string_bytes().unwrap(), want, "{ctx}");
assert_eq!(ds.read_string().unwrap(), utf8(&want), "{ctx}");
let m = mmap.dataset(name).unwrap();
assert_eq!(m.read_string_bytes().unwrap(), want, "{ctx} (mmap)");
assert_eq!(m.read_string().unwrap(), utf8(&want), "{ctx} (mmap)");
let l = lazy.dataset(name).unwrap();
assert_eq!(l.read_string_bytes().unwrap(), want, "{ctx} (lazy)");
assert_eq!(l.read_string().unwrap(), utf8(&want), "{ctx} (lazy)");
}
}
}
#[test]
fn vl_string_selections_read_like_h5py() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let expected = make_files(dir.path());
let hyperslab = |start: &[u64], stride: &[u64], count: &[u64]| Selection::Hyperslab {
start: start.to_vec(),
stride: stride.to_vec(),
count: count.to_vec(),
block: vec![1; start.len()],
};
let cases = [
("d2", "d2[1,1:3]", Selection::slice(&[1..2, 1..3])),
("chunked", "chunked[5:60:3]", hyperslab(&[5], &[3], &[19])),
(
"chunked2d",
"chunked2d[2:9:2,3:8]",
hyperslab(&[2, 3], &[2, 1], &[4, 5]),
),
];
for tag in ["8", "4"] {
let file = File::open(dir.path().join(format!("vl{tag}.h5"))).unwrap();
for (name, key, sel) in &cases {
let want = utf8(&parse_strings(&expected[&format!("{tag}:{key}")]));
let got = file
.dataset(name)
.unwrap()
.read_string_selection(sel)
.unwrap();
assert_eq!(got, want, "vl{tag}.h5 {key}");
}
// A selection of VL integers is not strings.
assert!(
file.dataset("vlen_i4")
.unwrap()
.read_string_selection(&Selection::slice(&[0..1]))
.is_err()
);
}
}
#[test]
fn vl_values_in_compounds_and_attributes_read_like_h5py() {
// With 4-byte offsets these failed with GlobalHeapObjectNotFound or came
// back as `AttrValue::Raw`: the VL type claimed 16-byte elements and the
// global heap was read without the padding libhdf5 puts after its
// headers.
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let expected = make_files(dir.path());
for tag in ["8", "4"] {
let file = File::open(dir.path().join(format!("vl{tag}.h5"))).unwrap();
let want = |key: &str| utf8(&parse_strings(&expected[&format!("{tag}:{key}")]));
let attrs = file.root().attrs().unwrap();
match &attrs["vlattr"] {
AttrValue::String(s) => assert_eq!(*s, want("vlattr")[0], "vl{tag}.h5"),
other => panic!("vl{tag}.h5 vlattr: {other:?}"),
}
match &attrs["vlattr_arr"] {
AttrValue::StringArray(s) => assert_eq!(*s, want("vlattr_arr"), "vl{tag}.h5"),
other => panic!("vl{tag}.h5 vlattr_arr: {other:?}"),
}
// Compound with a VL string member: dataset and attribute.
let ds = file.dataset("compound").unwrap();
let dt = ds.raw_datatype().unwrap();
let raw = ds.read_selection(&Selection::All).unwrap();
let fields = clawhdf5_format::data_read::read_compound_fields(&raw, &dt).unwrap();
let name = fields.iter().find(|f| f.name == "name").unwrap();
assert_eq!(
file.decode_strings(&name.datatype, &name.raw_data).unwrap(),
want("compound.name"),
"vl{tag}.h5 compound"
);
let id = fields.iter().find(|f| f.name == "id").unwrap();
assert_eq!(
clawhdf5_format::data_read::read_as_i64(&id.raw_data, &id.datatype).unwrap(),
vec![1, 2, 3]
);
let v = fields.iter().find(|f| f.name == "v").unwrap();
assert_eq!(
clawhdf5_format::data_read::read_as_f64(&v.raw_data, &v.datatype).unwrap(),
vec![0.5, 1.5, 2.5]
);
let AttrValue::Raw { datatype, data, .. } = &attrs["compound_attr"] else {
panic!("compound attribute is Raw");
};
let fields = clawhdf5_format::data_read::read_compound_fields(data, datatype).unwrap();
let name = fields.iter().find(|f| f.name == "name").unwrap();
assert_eq!(
file.decode_strings(&name.datatype, &name.raw_data).unwrap(),
want("compound_attr.name"),
"vl{tag}.h5 compound attribute"
);
// A VL sequence attribute.
let AttrValue::Raw { datatype, data, .. } = &attrs["vlen_attr"] else {
panic!("vlen attribute is Raw");
};
let got: Vec<Vec<i64>> = file.decode_vlen(datatype, data).unwrap();
let want_seqs = parse_seqs(&expected[&format!("{tag}:vlen_attr")]);
assert_eq!(
got,
want_seqs
.iter()
.map(|s| s.iter().map(|&x| x as i64).collect::<Vec<_>>())
.collect::<Vec<_>>(),
"vl{tag}.h5 vlen_attr"
);
}
}
#[test]
fn vl_sequence_datasets_read_like_h5py() {
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let expected = make_files(dir.path());
for tag in ["8", "4"] {
let path = dir.path().join(format!("vl{tag}.h5"));
let file = File::open(&path).unwrap();
let seqs = |key: &str| parse_seqs(&expected[&format!("{tag}:{key}")]);
let i4 = file.dataset("vlen_i4").unwrap();
let want: Vec<Vec<i32>> = seqs("vlen_i4")
.iter()
.map(|s| s.iter().map(|&x| x as i32).collect())
.collect();
assert_eq!(i4.read_vlen::<i32>().unwrap(), want, "vl{tag}.h5 vlen_i4");
let as_f64: Vec<Vec<f64>> = i4.read_vlen().unwrap();
assert_eq!(as_f64, seqs("vlen_i4"));
let f8 = file.dataset("vlen_f8").unwrap();
assert_eq!(
f8.read_vlen::<f64>().unwrap(),
seqs("vlen_f8"),
"vl{tag}.h5"
);
assert_eq!(
f8.read_vlen_selection::<f64>(&Selection::slice(&[1..2, 0..2]))
.unwrap(),
seqs("vlen_f8[1,:]"),
"vl{tag}.h5 vlen_f8[1,:]"
);
// h5py returns big-endian VL elements byte-swapped (an h5py bug, see
// CONFORMANCE.md); the values written are [1, 65535] and [300].
assert_eq!(
file.dataset("vlen_u2_be")
.unwrap()
.read_vlen::<u64>()
.unwrap(),
vec![vec![1, 65535], vec![300]]
);
let mmap = MmapFile::open(&path).unwrap();
assert_eq!(
mmap.dataset("vlen_f8").unwrap().read_vlen::<f64>().unwrap(),
seqs("vlen_f8")
);
let lazy = LazyFile::open_mmap(&path).unwrap();
assert_eq!(
lazy.dataset("vlen_f8").unwrap().read_vlen::<f64>().unwrap(),
seqs("vlen_f8")
);
// Wrong kind of data is an error, not a value.
assert!(i4.read_string().is_err());
assert!(i4.read_string_bytes().is_err());
assert!(file.dataset("d1").unwrap().read_vlen::<f64>().is_err());
}
}
#[test]
fn vl_strings_end_at_nul_and_mis_sized_elements_fail_like_h5py() {
// h5py cannot write a VL string with a NUL in it, so the file is patched:
// one string gets an embedded NUL, and two elements get a length that
// disagrees with their heap object. libhdf5 returns the string up to the
// NUL and refuses the others ("Expected global heap object size does
// not match"); we used to return the NUL and a truncated string.
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("patched.h5");
let script = format!(
r#"
import struct, h5py, numpy as np
path = {path:?}
with h5py.File(path, 'w') as f:
f.create_dataset('d', data=np.array(['aXb', 'cdefgh', 'ij', 'ok'], dtype=object),
dtype=h5py.string_dtype())
s = f.create_dataset('seq', shape=(2,), dtype=h5py.vlen_dtype(np.dtype('<i4')))
s[0] = np.array([1, 2, 3], dtype='<i4'); s[1] = np.array([4], dtype='<i4')
off = f['d'].id.get_offset(); soff = f['seq'].id.get_offset()
b = bytearray(open(path, 'rb').read())
i = b.index(b'aXb'); b[i + 1] = 0
struct.pack_into('<I', b, off + 16, 3) # 'cdefgh': length 6 -> 3
struct.pack_into('<I', b, off + 32, 9) # 'ij': length 2 -> 9
struct.pack_into('<I', b, soff, 2) # [1, 2, 3]: length 3 -> 2
open(path, 'wb').write(bytes(b))
with h5py.File(path, 'r') as f:
for i in range(4):
try:
print('d%d\t%s' % (i, f['d'][i].hex()))
except OSError as e:
print('d%d\terror' % i)
for i in range(2):
try:
print('seq%d\t%s' % (i, ' '.join(str(x) for x in f['seq'][i])))
except OSError as e:
print('seq%d\terror' % i)
"#,
path = path.display().to_string()
);
let expected = run_python(&script);
assert_eq!(expected["d0"], "61", "h5py cuts 'a\\0b' at the NUL");
assert_eq!(expected["d1"], "error");
assert_eq!(expected["d2"], "error");
assert_eq!(expected["d3"], "6f6b");
assert_eq!(expected["seq0"], "error");
assert_eq!(expected["seq1"], "4");
let file = File::open(&path).unwrap();
let d = file.dataset("d").unwrap();
let one = |i: u64| d.read_string_selection(&Selection::slice(&[i..i + 1]));
assert_eq!(one(0).unwrap(), vec!["a"]);
assert!(one(1).is_err());
assert!(one(2).is_err());
assert_eq!(one(3).unwrap(), vec!["ok"]);
assert!(d.read_string().is_err());
let seq = file.dataset("seq").unwrap();
let one = |i: u64| seq.read_vlen_selection::<i32>(&Selection::slice(&[i..i + 1]));
assert!(one(0).is_err());
assert_eq!(one(1).unwrap(), vec![vec![4]]);
}
#[test]
fn a_vl_element_at_the_undefined_heap_address_fails_like_h5py() {
// libhdf5 writes a null element with heap address 0 (h5py reads it as
// b''), and an empty string as a real zero-size heap object; neither
// uses the undefined address. An element of length 0 at the undefined
// address fails in libhdf5 ("addr undefined"); we returned "".
skip_if_no_python!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("undef.h5");
let script = format!(
r#"
import struct, h5py, numpy as np
path = {path:?}
with h5py.File(path, 'w') as f:
f.create_dataset('d', data=np.array(['x', '', 'yz', ''], dtype=object), dtype=h5py.string_dtype())
off = f['d'].id.get_offset()
b = bytearray(open(path, 'rb').read())
# h5py's '' (element 3): length 0 at a real heap address, not 0 or all 0xff.
length, addr, _ = struct.unpack_from('<IQI', b, off + 48)
print('empty\t%d %d' % (length, addr not in (0, 2**64 - 1)))
struct.pack_into('<IQI', b, off + 16, 0, 2**64 - 1, 1)
open(path, 'wb').write(bytes(b))
with h5py.File(path, 'r') as f:
for i in range(4):
try:
print('d%d\t%s' % (i, f['d'][i].hex()))
except OSError as e:
print('d%d\terror %s' % (i, 'addr undefined' in str(e)))
"#,
path = path.display().to_string()
);
let expected = run_python(&script);
assert_eq!(expected["empty"], "0 1", "h5py writes '' at a real address");
assert_eq!(expected["d0"], "78");
assert_eq!(expected["d1"], "error True");
assert_eq!(expected["d2"], "797a");
assert_eq!(expected["d3"], "");
let file = File::open(&path).unwrap();
let d = file.dataset("d").unwrap();
let one = |i: u64| d.read_string_selection(&Selection::slice(&[i..i + 1]));
assert_eq!(one(0).unwrap(), vec!["x"]);
let e = one(1).unwrap_err().to_string();
assert!(e.contains("undefined global heap address"), "{e}");
assert_eq!(one(2).unwrap(), vec!["yz"]);
assert_eq!(one(3).unwrap(), vec![""]);
assert!(d.read_string().is_err());
assert!(d.read_string_bytes().is_err());
}
+126
View File
@@ -0,0 +1,126 @@
//! Variable-length values in files with 4-byte offsets and lengths
//! (`sizeof_addr = 4`), checked against h5py/libhdf5 through the
//! `clawhdf5_format` decoders.
//!
//! These failed with `GlobalHeapObjectNotFound` or came back as
//! `AttrValue::Raw`: the VL datatype claimed 16-byte elements whatever the
//! file's offset size, and the global heap was read without the padding
//! libhdf5 puts after its collection and object headers. Skipped when
//! python3 with h5py is unavailable, unless `CLAWHDF5_REQUIRE_INTEROP=1`.
use std::process::Command;
use clawhdf5::{AttrValue, File, Selection};
use clawhdf5_format::data_read::{read_as_i64, read_compound_fields};
use clawhdf5_format::vl_data::{read_vl_bytes, read_vl_strings};
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn interop_required() -> bool {
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
}
fn python_available() -> bool {
Command::new(python())
.args(["-c", "import h5py, numpy"])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
#[test]
fn vl_values_in_a_file_with_4_byte_offsets_read_like_h5py() {
if !python_available() {
assert!(
!interop_required(),
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
);
eprintln!("SKIP: python3 with h5py not available");
return;
}
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("offset4.h5");
let script = format!(
r#"
import h5py, numpy as np
S = h5py.string_dtype()
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE); fcpl.set_sizes(4, 4)
with h5py.File(h5py.h5f.create({path:?}.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl)) as f:
f.attrs['vlattr'] = 'attr-value'
f.attrs.create('vlattr_arr', np.array(['p', 'qq', ''], dtype=object), dtype=S)
ct = np.dtype([('id', '<i4'), ('name', S), ('v', '<f8')])
arr = np.zeros(3, dtype=ct); arr['id'] = [1, 2, 3]; arr['name'] = ['one', '', 'three']
f.create_dataset('compound', data=arr)
f.attrs.create('vlen_attr', np.array([np.array([1, 2], dtype='<i8'), np.array([3], dtype='<i8')],
dtype=object), dtype=h5py.vlen_dtype(np.dtype('<i8')))
with h5py.File({path:?}, 'r') as f:
assert f.id.get_create_plist().get_sizes() == (4, 4)
print(f.attrs['vlattr'])
print(','.join(f.attrs['vlattr_arr']))
print(','.join(s.decode() for s in f['compound']['name']))
print(';'.join(' '.join(str(x) for x in s) for s in f.attrs['vlen_attr']))
"#,
path = path.display().to_string()
);
let output = Command::new(python())
.args(["-c", &script])
.output()
.expect("failed to run python");
assert!(
output.status.success(),
"python failed:\n{}",
String::from_utf8_lossy(&output.stderr)
);
let stdout = String::from_utf8(output.stdout).unwrap();
let lines: Vec<&str> = stdout.lines().collect();
let (vlattr, vlattr_arr, names, seqs) = (lines[0], lines[1], lines[2], lines[3]);
let file = File::open(&path).unwrap();
let sb = file.superblock();
assert_eq!((sb.offset_size, sb.length_size), (4, 4));
let attrs = file.root().attrs().unwrap();
match &attrs["vlattr"] {
AttrValue::String(s) => assert_eq!(s, vlattr),
other => panic!("vlattr: {other:?}"),
}
match &attrs["vlattr_arr"] {
AttrValue::StringArray(s) => assert_eq!(s.join(","), vlattr_arr),
other => panic!("vlattr_arr: {other:?}"),
}
// The compound's VL string member.
let ds = file.dataset("compound").unwrap();
let dt = ds.raw_datatype().unwrap();
assert_eq!(dt.type_size(), 24, "4 + 12-byte VL element + 8");
let raw = ds.read_selection(&Selection::All).unwrap();
let fields = read_compound_fields(&raw, &dt).unwrap();
let name = fields.iter().find(|f| f.name == "name").unwrap();
assert_eq!(name.datatype.type_size(), 12);
let got = read_vl_strings(file.as_bytes(), &name.raw_data, 3, 4, 4).unwrap();
assert_eq!(got.join(","), names);
let id = fields.iter().find(|f| f.name == "id").unwrap();
assert_eq!(
read_as_i64(&id.raw_data, &id.datatype).unwrap(),
vec![1, 2, 3]
);
// A VL sequence attribute.
let AttrValue::Raw { datatype, data, .. } = &attrs["vlen_attr"] else {
panic!("vlen_attr is Raw");
};
let clawhdf5_format::datatype::Datatype::VariableLength { base_type, .. } = datatype else {
panic!("vlen_attr is VL");
};
let got: Vec<String> = read_vl_bytes(file.as_bytes(), data, 2, 4, 4)
.unwrap()
.iter()
.map(|b| {
let v = read_as_i64(b, base_type).unwrap();
v.iter().map(i64::to_string).collect::<Vec<_>>().join(" ")
})
.collect();
assert_eq!(got.join(";"), seqs);
}
+46 -6
View File
@@ -164,13 +164,24 @@ fill-value item that did is fixed).
is left out of `attrs()` (reported by `attrs_with_errors()`) instead of is left out of `attrs()` (reported by `attrs_with_errors()`) instead of
failing the others. failing the others.
- **Other readers:** - **Other readers:**
- VL-string datasets are not readable through `File`. - VL-string datasets are not readable through `File`. **Fixed
2026-09-26:** `read_string` reads them (also `read_string_bytes`,
`read_string_selection`, and on `MmapFile`/`LazyFile`), with h5py's
values: strings end at a NUL, null elements (heap address 0) are `""`,
and an element at the undefined heap address is an error as in libhdf5
(it read as `""` until 2026-09-26); VL sequences of
numbers read with `read_vlen::<T>()`, and VL values inside compounds or
`AttrValue::Raw` attributes decode with `File::decode_strings` /
`File::decode_vlen` (`crates/clawhdf5/tests/vl_data_interop.rs`).
- Variable-length values inside a compound (and VL-string attributes) in - Variable-length values inside a compound (and VL-string attributes) in
a file with 4-byte offsets (`sizeof_addr = 4`) fail with a file with 4-byte offsets (`sizeof_addr = 4`) fail with
`GlobalHeapObjectNotFound` or come back as `Raw`: these paths assume `GlobalHeapObjectNotFound` or come back as `Raw`: these paths assume
the 16-byte element of an 8-byte-offset file. The datatype itself reads the 16-byte element of an 8-byte-offset file. The datatype itself reads
(it was refused as "member overlaps with previous member" until (it was refused as "member overlaps with previous member" until
2026-09-26). 2026-09-26). **Fixed 2026-09-26:** a VL type's element size is the one
its datatype message stores (12 with 4-byte offsets), and the global
heap is read with libhdf5's header padding
(`crates/clawhdf5/tests/vl_offset4_interop.rs`).
- Metadata cache images are not supported. - Metadata cache images are not supported.
- x87 long double and binary128 are refused. - x87 long double and binary128 are refused.
- N-Bit on 64-bit scale-offset data and some N-Bit parameter layouts fail. - N-Bit on 64-bit scale-offset data and some N-Bit parameter layouts fail.
@@ -222,9 +233,10 @@ fill-value item that did is fixed).
- (`cve-2024-32616` `/group1/dset3` and `cve-2025-2309`'s `Comp_OBJREF` - (`cve-2024-32616` `/group1/dset3` and `cve-2025-2309`'s `Comp_OBJREF`
attribute are h5py/numpy type-mapping failures, not libhdf5 refusals.) attribute are h5py/numpy type-mapping failures, not libhdf5 refusals.)
- `h5rs check` validates with the library's parsers, so it inherits what - `h5rs check` validates with the library's parsers, so it inherits what
they accept: of the 150 CVE and fuzzer files, `check --data` passes 16, they accept: of the 150 CVE and fuzzer files, `check --data` passes 15,
and h5dump 1.14.6 rejects 9 of those (tank, 2026-09-26; 28 and 21 and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26; 28 and 21
before these checks). before these checks, 16 and 9 before a VL type's stored element size
was checked, which flags `cve-2024-32608`).
- **Writer:** - **Writer:**
- Nested groups beyond one level: path-like names are now refused, not - Nested groups beyond one level: path-like names are now refused, not
created. created.
@@ -424,6 +436,32 @@ has produced more records than the file could physically hold.
--- ---
## Crafted global heaps exhaust the variable-length reader's memory
**Status:** fixed on `feat/p2-vl-strings` (2026-09-26). Not a regression of
that branch: every earlier release is affected through `read_vl_strings`.
Reading variable-length values kept an owned copy of every object of every
global heap collection visited, for the whole read. A file whose collections
nest inside one another's object data (32 bytes apart, each element pointing
at a different one) made retained memory O(elements × file size): a 744 KB
file reached 1.58 GB. Letting every collection's object chain jump to one
shared run of tiny objects made the parse time O(elements × objects) too.
libhdf5 refuses such files.
Now `VlResolver` caches where each object lies instead of a copy, drops its
cache past a 32 MiB budget, and refuses a collection that overlaps one it
has already read (libhdf5 gives each collection its own block, so only a
crafted file has them). `GlobalHeapCollection::parse` (and the new
`parse_index`) also refuse a collection that runs past the end of the file,
or an object that runs past the end of its collection. Guarded by
`crates/clawhdf5-format/tests/vl_heap_bounds.rs`, which measures peak heap
use with a counting allocator. Still open: a file may point many elements
at one large heap object, and a VL-*sequence* read then returns that
object once per element, as h5py would.
---
## Extensible Array chunk indexes read back wrong data past the inline elements ## Extensible Array chunk indexes read back wrong data past the inline elements
**Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to **Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to
@@ -526,7 +564,9 @@ which is what libhdf5 itself writes.
followed (no file system). followed (no file system).
- Variable-length string datasets are read by decoding `read_selection`'s - Variable-length string datasets are read by decoding `read_selection`'s
bytes with `clawhdf5_format::vl_data` in the wasm crate; `File` itself still bytes with `clawhdf5_format::vl_data` in the wasm crate; `File` itself still
cannot (see the audit gaps above). cannot (see the audit gaps above). (`File` can since 2026-09-26. Since
2026-09-26 the wasm crate resolves them with the same `VlResolver` as
`File` and `h5rs`, so all three return h5py's values.)
## The Node.js package (`packages/clawhdf5-node`) does not work ## The Node.js package (`packages/clawhdf5-node`) does not work