Merge branch 'feat/p2-vl-strings' into feat/p2-perf-coverage
# Conflicts: # CHANGELOG.md
This commit is contained in:
+87
-2
@@ -54,6 +54,91 @@
|
|||||||
whole-row hyperslabs, points, empty selections; every type, both byte
|
whole-row hyperslabs, points, empty selections; every type, both byte
|
||||||
orders, ranks 1-4).
|
orders, ranks 1-4).
|
||||||
|
|
||||||
|
### Variable-length data (2026-09-26)
|
||||||
|
- **VL values in files with 4-byte offsets** (`sizeof_addr = 4`). A VL
|
||||||
|
string attribute came back as `AttrValue::Raw`, a VL member of a compound
|
||||||
|
failed with `GlobalHeapObjectNotFound`, and VL datasets failed with a
|
||||||
|
size mismatch. Two causes: `Datatype::type_size()` reported 16 for every
|
||||||
|
VL type (the element is 4 + offset size + 4 bytes: 12 here), and the
|
||||||
|
global heap was parsed without the padding libhdf5 puts after its
|
||||||
|
collection and object headers (`H5HG_SIZEOF_HDR`/`H5HG_SIZEOF_OBJHDR`
|
||||||
|
round up to 8), so with 4-byte lengths every object was looked up 4
|
||||||
|
bytes early. `Datatype::VariableLength` now carries the element `size`
|
||||||
|
stored in the datatype message (**breaking** for code that builds or
|
||||||
|
exhaustively destructures that variant; patterns with `..` are
|
||||||
|
unaffected), and it is written back as stored. Tested against h5py
|
||||||
|
(`crates/clawhdf5/tests/vl_offset4_interop.rs`).
|
||||||
|
- **Wrong data: VL strings with an embedded NUL, and VL elements whose heap
|
||||||
|
object has the wrong size.** libhdf5 hands VL strings over as C strings,
|
||||||
|
so h5py reads `"a\0b"` as `"a"`; `read_vl_strings` returned the NUL and
|
||||||
|
what followed. An element whose heap object is not exactly
|
||||||
|
`length × base size` bytes is refused by libhdf5 ("Expected global heap
|
||||||
|
object size does not match"); we returned the object cut or padded to
|
||||||
|
the length. Both now behave as libhdf5, and a heap address of 0 is a null
|
||||||
|
element (empty) whatever its length. The new
|
||||||
|
`clawhdf5_format::vl_data::VlResolver` does this and parses each global
|
||||||
|
heap collection once per read: `read_vl_strings` parsed the whole
|
||||||
|
collection again for every element. `vl_data::check_element_size` refuses
|
||||||
|
a VL datatype whose stored size is not 4 + offset size + 4 (libhdf5
|
||||||
|
ignores the stored size). The conformance probe resolves VL elements
|
||||||
|
with `VlResolver` too; conformance unchanged at 575 of 697.
|
||||||
|
- **VL data through the facade.** VL-string datasets (h5py's default `str`
|
||||||
|
dtype) failed `read_string` with "type mismatch: expected String, got
|
||||||
|
VariableLength". `Dataset::read_string` now reads fixed- and
|
||||||
|
variable-length strings; new `read_string_bytes` (a VL string's exact
|
||||||
|
bytes, as h5py's `Dataset[()]` returns them), `read_string_selection`,
|
||||||
|
`read_vlen::<T>()` / `read_vlen_selection::<T>()` for VL sequences of
|
||||||
|
numbers (`T` = `f64`, `f32`, `i64`, `i32`, `u64`; converted like the
|
||||||
|
other typed readers), and `File::decode_strings` / `decode_string_bytes`
|
||||||
|
/ `decode_vlen` for VL values in compound fields and `AttrValue::Raw`
|
||||||
|
attributes. `MmapDataset` and `LazyDataset` gain `read_string` for VL
|
||||||
|
strings, `read_string_bytes` and `read_vlen`. Checked against h5py with
|
||||||
|
8- and 4-byte offsets: scalar and 1-/2-D, ASCII and UTF-8, empty strings,
|
||||||
|
contiguous, compact, chunked with gzip/shuffle, never-written and
|
||||||
|
partly written chunks, hyperslab selections, VL members of compound
|
||||||
|
datasets and attributes (`crates/clawhdf5/tests/vl_data_interop.rs`).
|
||||||
|
NetCDF-4 `string` variables now read through
|
||||||
|
`clawhdf5_netcdf4::Variable::read_string` (checked against netCDF4-python
|
||||||
|
in `crates/clawhdf5-netcdf4/tests/interop_tests.rs`).
|
||||||
|
|
||||||
|
- **Crafted global heaps could exhaust memory.** `VlResolver` kept an owned
|
||||||
|
copy of every object of every heap collection it read, so collections
|
||||||
|
nested inside each other's object data made a 744 KB file take 1.58 GB
|
||||||
|
(and `read_vl_strings` before it did the same). The cache now records
|
||||||
|
where objects lie instead of copying them, is dropped past a 32 MiB
|
||||||
|
budget, and a collection overlapping one already read is an error
|
||||||
|
(libhdf5 never writes one). New `GlobalHeapCollection::parse_index`
|
||||||
|
locates a collection's objects without copying them; `parse` and
|
||||||
|
`parse_index` refuse a collection running past the end of the file or an
|
||||||
|
object running past its collection. Conformance unchanged at 575 of 697
|
||||||
|
(`crates/clawhdf5-format/tests/vl_heap_bounds.rs`).
|
||||||
|
|
||||||
|
- **Every reader resolves VL data the same way.** `h5rs` (`dump`, `ls`,
|
||||||
|
`diff`, `check --data`) had its own lenient VL decoder: a heap object
|
||||||
|
longer than the element's length was cut to it (h5py refuses it), a null
|
||||||
|
string printed `""` where h5dump prints `NULL`, the stored element size
|
||||||
|
was trusted, and each heap collection was kept as a copy for the whole
|
||||||
|
run. It now resolves through `VlResolver`, so `dump` matches h5dump byte
|
||||||
|
for byte on VL strings (`"a\0b"` as `"a"`, null as `NULL`), VL sequences
|
||||||
|
and 4-byte-offset files, `dump --json` gives h5py's values, and
|
||||||
|
`check --data` reports any heap object whose size is not exactly the
|
||||||
|
element's length × base size. `clawhdf5-wasm` already resolved VL strings
|
||||||
|
with `read_vl_strings`; it now uses `VlResolver` and refuses a VL type
|
||||||
|
whose stored element size disagrees with the file, as `File` does
|
||||||
|
(`crates/clawhdf5-tools/tests/h5rs_interop.rs`,
|
||||||
|
`crates/clawhdf5-wasm/tests/vl_strings.rs`). New
|
||||||
|
`VlResolver::element` / `string_element` resolve one element in place.
|
||||||
|
|
||||||
|
- **A VL element at the undefined heap address is an error**, as in
|
||||||
|
libhdf5 ("addr undefined"). One of length 0 read as `""` in every reader
|
||||||
|
(`File`, `h5rs`, `clawhdf5-wasm`, `read_vl_strings`, `read_vl_bytes`).
|
||||||
|
libhdf5 writes a null element with heap address 0, which still reads as
|
||||||
|
empty, and h5py writes `""` as a zero-size heap object at a real address,
|
||||||
|
so no file libhdf5 or h5py writes is affected
|
||||||
|
(`a_vl_element_at_the_undefined_heap_address_fails_like_h5py` in
|
||||||
|
`crates/clawhdf5/tests/vl_data_interop.rs`). `read_vl_bytes` now also
|
||||||
|
treats address 0 as null whatever the length, as `VlResolver` does.
|
||||||
|
|
||||||
### Plugin filters (2026-09-26)
|
### Plugin filters (2026-09-26)
|
||||||
- **LZF, bitshuffle, bzip2 and Blosc read and write, in pure Rust.** Files
|
- **LZF, bitshuffle, bzip2 and Blosc read and write, in pure Rust.** Files
|
||||||
written by h5py with `compression="lzf"`, or with hdf5plugin's
|
written by h5py with `compression="lzf"`, or with hdf5plugin's
|
||||||
@@ -257,10 +342,10 @@
|
|||||||
printed with its address; exit 1 when there are any. libhdf5's h5check
|
printed with its address; exit 1 when there are any. libhdf5's h5check
|
||||||
reads only the 1.8 format. On the conformance corpus it passes all 418
|
reads only the 1.8 format. On the conformance corpus it passes all 418
|
||||||
files that both clawhdf5 and h5py read in full, and `check --data` flags
|
files that both clawhdf5 and h5py read in full, and `check --data` flags
|
||||||
134 of the 150 CVE and fuzzer files of the `cve_hdf5` corpus (tank,
|
135 of the 150 CVE and fuzzer files of the `cve_hdf5` corpus (tank,
|
||||||
2026-09-26). `--data` also follows variable-length data into its global
|
2026-09-26). `--data` also follows variable-length data into its global
|
||||||
heap collections and reports a damaged one at its address. It inherits
|
heap collections and reports a damaged one at its address. It inherits
|
||||||
the library's tolerance, though: 9 of the 16 it passes are files h5dump
|
the library's tolerance, though: 8 of the 15 it passes are files h5dump
|
||||||
1.14.6 rejects (see `docs/known-issues.md`, header checks).
|
1.14.6 rejects (see `docs/known-issues.md`, header checks).
|
||||||
- Values over `--max-bytes` (default 1 GiB) are reported instead of read;
|
- Values over `--max-bytes` (default 1 GiB) are reported instead of read;
|
||||||
a panic is caught and reported as an internal error (exit 3).
|
a panic is caught and reported as an internal error (exit 3).
|
||||||
|
|||||||
@@ -19,9 +19,8 @@
|
|||||||
//! with its message, location and the clawhdf5 frames of its backtrace.
|
//! with its message, location and the clawhdf5 frames of its backtrace.
|
||||||
|
|
||||||
use std::cell::RefCell;
|
use std::cell::RefCell;
|
||||||
use std::collections::{HashMap, HashSet};
|
use std::collections::HashSet;
|
||||||
use std::panic::{self, AssertUnwindSafe};
|
use std::panic::{self, AssertUnwindSafe};
|
||||||
use std::rc::Rc;
|
|
||||||
|
|
||||||
use clawhdf5_format::attribute::extract_attributes_full;
|
use clawhdf5_format::attribute::extract_attributes_full;
|
||||||
use clawhdf5_format::data_layout::DataLayout;
|
use clawhdf5_format::data_layout::DataLayout;
|
||||||
@@ -29,7 +28,6 @@ use clawhdf5_format::data_read;
|
|||||||
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
|
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
|
||||||
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
||||||
use clawhdf5_format::filter_pipeline::FilterPipeline;
|
use clawhdf5_format::filter_pipeline::FilterPipeline;
|
||||||
use clawhdf5_format::global_heap::GlobalHeapCollection;
|
|
||||||
use clawhdf5_format::group_v1::{self, GroupEntry};
|
use clawhdf5_format::group_v1::{self, GroupEntry};
|
||||||
use clawhdf5_format::group_v2;
|
use clawhdf5_format::group_v2;
|
||||||
use clawhdf5_format::message_type::MessageType;
|
use clawhdf5_format::message_type::MessageType;
|
||||||
@@ -37,6 +35,7 @@ use clawhdf5_format::object_header::ObjectHeader;
|
|||||||
use clawhdf5_format::signature;
|
use clawhdf5_format::signature;
|
||||||
use clawhdf5_format::superblock::Superblock;
|
use clawhdf5_format::superblock::Superblock;
|
||||||
use clawhdf5_format::symbol_table::SymbolTableMessage;
|
use clawhdf5_format::symbol_table::SymbolTableMessage;
|
||||||
|
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
||||||
use serde_json::{Map, Value, json};
|
use serde_json::{Map, Value, json};
|
||||||
use sha2::{Digest, Sha256};
|
use sha2::{Digest, Sha256};
|
||||||
|
|
||||||
@@ -111,7 +110,10 @@ struct Ctx<'a> {
|
|||||||
os: u8,
|
os: u8,
|
||||||
ls: u8,
|
ls: u8,
|
||||||
base_dir: std::path::PathBuf,
|
base_dir: std::path::PathBuf,
|
||||||
heaps: RefCell<HashMap<u64, Result<Rc<GlobalHeapCollection>, String>>>,
|
/// Resolves variable-length elements as the library does (null
|
||||||
|
/// elements, strings cut at a NUL, heap objects of the wrong size
|
||||||
|
/// refused), caching each heap collection.
|
||||||
|
vl: RefCell<VlResolver<'a>>,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl<'a> Ctx<'a> {
|
impl<'a> Ctx<'a> {
|
||||||
@@ -130,33 +132,6 @@ impl<'a> Ctx<'a> {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn heap_obj(&self, addr: u64, idx: u32) -> Result<Vec<u8>, String> {
|
|
||||||
let coll = {
|
|
||||||
let mut cache = self.heaps.borrow_mut();
|
|
||||||
cache
|
|
||||||
.entry(addr)
|
|
||||||
.or_insert_with(|| {
|
|
||||||
GlobalHeapCollection::parse(self.data, addr as usize, self.ls)
|
|
||||||
.map(Rc::new)
|
|
||||||
.map_err(e)
|
|
||||||
})
|
|
||||||
.clone()?
|
|
||||||
};
|
|
||||||
coll.get_object(idx as u16)
|
|
||||||
.map(|o| o.data.clone())
|
|
||||||
.ok_or_else(|| {
|
|
||||||
format!("GlobalHeapObjectNotFound {{ collection_address: {addr}, index: {idx} }}")
|
|
||||||
})
|
|
||||||
}
|
|
||||||
|
|
||||||
fn read_offset(&self, b: &[u8]) -> u64 {
|
|
||||||
let mut v = 0u64;
|
|
||||||
for (i, x) in b.iter().take(self.os as usize).enumerate() {
|
|
||||||
v |= (*x as u64) << (8 * i);
|
|
||||||
}
|
|
||||||
v
|
|
||||||
}
|
|
||||||
|
|
||||||
fn canon(&self, dt: &Datatype, b: &[u8], out: &mut Vec<u8>) -> Result<(), String> {
|
fn canon(&self, dt: &Datatype, b: &[u8], out: &mut Vec<u8>) -> Result<(), String> {
|
||||||
let size = dt.type_size() as usize;
|
let size = dt.type_size() as usize;
|
||||||
if b.len() < size {
|
if b.len() < size {
|
||||||
@@ -204,41 +179,27 @@ impl<'a> Ctx<'a> {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
|
size: vl_size,
|
||||||
is_string,
|
is_string,
|
||||||
base_type,
|
base_type,
|
||||||
..
|
..
|
||||||
} => {
|
} => {
|
||||||
let len = u32::from_le_bytes([b[0], b[1], b[2], b[3]]) as usize;
|
check_element_size(*vl_size, self.os).map_err(e)?;
|
||||||
let addr = self.read_offset(&b[4..]);
|
let el = &b[..size];
|
||||||
let idx_off = 4 + self.os as usize;
|
|
||||||
let idx = u32::from_le_bytes([
|
|
||||||
b[idx_off],
|
|
||||||
b[idx_off + 1],
|
|
||||||
b[idx_off + 2],
|
|
||||||
b[idx_off + 3],
|
|
||||||
]);
|
|
||||||
let obj = if len == 0 || addr == 0 || addr == u64::MAX >> (64 - 8 * self.os as u32)
|
|
||||||
{
|
|
||||||
Vec::new()
|
|
||||||
} else {
|
|
||||||
self.heap_obj(addr, idx)?
|
|
||||||
};
|
|
||||||
if *is_string {
|
if *is_string {
|
||||||
let l = len.min(obj.len());
|
let s = self.vl.borrow_mut().string_bytes(el).map_err(e)?;
|
||||||
canon_str(&obj[..l], out);
|
canon_str(&s[0], out);
|
||||||
} else {
|
} else {
|
||||||
let bs = base_type.type_size() as usize;
|
let bs = base_type.type_size() as usize;
|
||||||
if bs == 0 {
|
// The borrow ends here: the base type may itself be
|
||||||
return Err("canon: VL base size 0".into());
|
// variable-length.
|
||||||
}
|
let seq = self.vl.borrow_mut().sequences(el, bs).map_err(e)?;
|
||||||
let need = len.checked_mul(bs).ok_or("canon: VL overflow")?;
|
let seq = &seq[0];
|
||||||
if len > 0 && obj.len() < need {
|
let len = seq.len() / bs;
|
||||||
return Err(format!("canon: VL object {} < {need}", obj.len()));
|
|
||||||
}
|
|
||||||
out.push(b'V');
|
out.push(b'V');
|
||||||
out.extend_from_slice(&(len as u32).to_le_bytes());
|
out.extend_from_slice(&(len as u32).to_le_bytes());
|
||||||
for i in 0..len {
|
for i in 0..len {
|
||||||
self.canon(base_type, &obj[i * bs..], out)?;
|
self.canon(base_type, &seq[i * bs..], out)?;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -744,7 +705,7 @@ fn main() {
|
|||||||
.parent()
|
.parent()
|
||||||
.map(|p| p.to_path_buf())
|
.map(|p| p.to_path_buf())
|
||||||
.unwrap_or_default(),
|
.unwrap_or_default(),
|
||||||
heaps: RefCell::new(HashMap::new()),
|
vl: RefCell::new(VlResolver::new(hdf5, sb.offset_size, sb.length_size)),
|
||||||
};
|
};
|
||||||
let mut objects: Vec<Value> = Vec::new();
|
let mut objects: Vec<Value> = Vec::new();
|
||||||
let mut visited = HashSet::new();
|
let mut visited = HashSet::new();
|
||||||
|
|||||||
@@ -125,6 +125,11 @@ pub enum Datatype {
|
|||||||
},
|
},
|
||||||
/// Class 9: Variable-length type.
|
/// Class 9: Variable-length type.
|
||||||
VariableLength {
|
VariableLength {
|
||||||
|
/// Size of one element as stored in the file: a sequence length (4
|
||||||
|
/// bytes), a global heap collection address (the file's
|
||||||
|
/// `offset_size`) and an object index (4 bytes) — 16 in a file with
|
||||||
|
/// 8-byte offsets, 12 with 4-byte offsets.
|
||||||
|
size: u32,
|
||||||
is_string: bool,
|
is_string: bool,
|
||||||
padding: Option<StringPadding>,
|
padding: Option<StringPadding>,
|
||||||
charset: Option<CharacterSet>,
|
charset: Option<CharacterSet>,
|
||||||
@@ -771,6 +776,7 @@ impl Datatype {
|
|||||||
pos += consumed;
|
pos += consumed;
|
||||||
Ok((
|
Ok((
|
||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
is_string,
|
is_string,
|
||||||
padding,
|
padding,
|
||||||
charset,
|
charset,
|
||||||
@@ -1017,6 +1023,7 @@ impl Datatype {
|
|||||||
Self::build_header(3, 1, [bf0, 0, 0], *size)
|
Self::build_header(3, 1, [bf0, 0, 0], *size)
|
||||||
}
|
}
|
||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
is_string,
|
is_string,
|
||||||
padding,
|
padding,
|
||||||
charset,
|
charset,
|
||||||
@@ -1039,7 +1046,7 @@ impl Datatype {
|
|||||||
} else {
|
} else {
|
||||||
0
|
0
|
||||||
};
|
};
|
||||||
let mut buf = Self::build_header(9, 1, [bf0, bf1, 0], 16);
|
let mut buf = Self::build_header(9, 1, [bf0, bf1, 0], *size);
|
||||||
buf.extend_from_slice(&base_type.serialize());
|
buf.extend_from_slice(&base_type.serialize());
|
||||||
buf
|
buf
|
||||||
}
|
}
|
||||||
@@ -1208,7 +1215,7 @@ impl Datatype {
|
|||||||
Datatype::Compound { size, .. } => *size,
|
Datatype::Compound { size, .. } => *size,
|
||||||
Datatype::Reference { size, .. } => *size,
|
Datatype::Reference { size, .. } => *size,
|
||||||
Datatype::Enumeration { size, .. } => *size,
|
Datatype::Enumeration { size, .. } => *size,
|
||||||
Datatype::VariableLength { .. } => 16, // typically pointer + length
|
Datatype::VariableLength { size, .. } => *size,
|
||||||
Datatype::Array {
|
Datatype::Array {
|
||||||
base_type,
|
base_type,
|
||||||
dimensions,
|
dimensions,
|
||||||
@@ -1889,11 +1896,13 @@ mod tests {
|
|||||||
let (dt, _) = Datatype::parse(&buf).unwrap();
|
let (dt, _) = Datatype::parse(&buf).unwrap();
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
is_string,
|
is_string,
|
||||||
padding,
|
padding,
|
||||||
charset,
|
charset,
|
||||||
base_type,
|
base_type,
|
||||||
} => {
|
} => {
|
||||||
|
assert_eq!(size, 16);
|
||||||
assert!(is_string);
|
assert!(is_string);
|
||||||
assert_eq!(padding, Some(StringPadding::NullTerminate));
|
assert_eq!(padding, Some(StringPadding::NullTerminate));
|
||||||
assert_eq!(charset, Some(CharacterSet::Utf8));
|
assert_eq!(charset, Some(CharacterSet::Utf8));
|
||||||
@@ -1914,11 +1923,13 @@ mod tests {
|
|||||||
let (dt, _) = Datatype::parse(&buf).unwrap();
|
let (dt, _) = Datatype::parse(&buf).unwrap();
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
is_string,
|
is_string,
|
||||||
padding,
|
padding,
|
||||||
charset,
|
charset,
|
||||||
base_type,
|
base_type,
|
||||||
} => {
|
} => {
|
||||||
|
assert_eq!(size, 16);
|
||||||
assert!(!is_string);
|
assert!(!is_string);
|
||||||
assert_eq!(padding, None);
|
assert_eq!(padding, None);
|
||||||
assert_eq!(charset, None);
|
assert_eq!(charset, None);
|
||||||
@@ -1928,6 +1939,19 @@ mod tests {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn variable_length_size_is_the_stored_size() {
|
||||||
|
// A file with 4-byte offsets stores 12-byte VL elements (length 4 +
|
||||||
|
// address 4 + index 4); the type used to report 16 regardless, so
|
||||||
|
// every read laid the elements out 16 bytes apart.
|
||||||
|
let mut buf = build_dt_header(9, 1, [0x01, 0x00, 0], 12);
|
||||||
|
buf.extend_from_slice(&build_fixed_point(1, false, false, 0, 8));
|
||||||
|
let (dt, _) = Datatype::parse(&buf).unwrap();
|
||||||
|
assert_eq!(dt.type_size(), 12);
|
||||||
|
// And it is written back as stored.
|
||||||
|
assert_eq!(dt.serialize()[4..8], 12u32.to_le_bytes());
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn test_array_2d() {
|
fn test_array_2d() {
|
||||||
// Array [3][4] of i32 LE, version 3
|
// Array [3][4] of i32 LE, version 3
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
//! HDF5 Global Heap collection parsing.
|
//! HDF5 Global Heap collection parsing.
|
||||||
|
|
||||||
#[cfg(not(feature = "std"))]
|
#[cfg(not(feature = "std"))]
|
||||||
use alloc::vec::Vec;
|
use alloc::{format, string::String, vec::Vec};
|
||||||
|
|
||||||
use crate::error::FormatError;
|
use crate::error::FormatError;
|
||||||
|
|
||||||
@@ -52,11 +52,42 @@ fn read_length(data: &[u8], offset: usize, length_size: u8) -> Result<u64, Forma
|
|||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn object_overrun_msg(index: u16, size: usize, collection_size: u64) -> String {
|
||||||
|
format!(
|
||||||
|
"global heap object {index} ({size} bytes) runs past the end of its \
|
||||||
|
{collection_size}-byte collection"
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
/// Round up to next multiple of 8.
|
/// Round up to next multiple of 8.
|
||||||
fn pad8(x: usize) -> usize {
|
fn pad8(x: usize) -> usize {
|
||||||
(x + 7) & !7
|
(x + 7) & !7
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Where one object of a global heap collection lies in the file, without
|
||||||
|
/// its data: see [`GlobalHeapCollection::parse_index`].
|
||||||
|
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||||
|
pub struct GlobalHeapObjectRef {
|
||||||
|
/// Object index (1-based; 0 is the free space marker).
|
||||||
|
pub index: u16,
|
||||||
|
/// Reference count.
|
||||||
|
pub reference_count: u16,
|
||||||
|
/// Offset of the object's data in the file data the collection was
|
||||||
|
/// parsed from.
|
||||||
|
pub offset: usize,
|
||||||
|
/// Size of the object's data in bytes.
|
||||||
|
pub size: usize,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A global heap collection's objects, located but not copied.
|
||||||
|
#[derive(Debug, Clone)]
|
||||||
|
pub struct GlobalHeapIndex {
|
||||||
|
/// Total size of this collection including header.
|
||||||
|
pub collection_size: u64,
|
||||||
|
/// The objects, in file order.
|
||||||
|
pub objects: Vec<GlobalHeapObjectRef>,
|
||||||
|
}
|
||||||
|
|
||||||
impl GlobalHeapCollection {
|
impl GlobalHeapCollection {
|
||||||
/// Parse a global heap collection at the given offset in the file data.
|
/// Parse a global heap collection at the given offset in the file data.
|
||||||
pub fn parse(
|
pub fn parse(
|
||||||
@@ -64,8 +95,38 @@ impl GlobalHeapCollection {
|
|||||||
offset: usize,
|
offset: usize,
|
||||||
length_size: u8,
|
length_size: u8,
|
||||||
) -> Result<GlobalHeapCollection, FormatError> {
|
) -> Result<GlobalHeapCollection, FormatError> {
|
||||||
// signature(4) + version(1) + reserved(3) + collection_size(length_size)
|
let index = Self::parse_index(file_data, offset, length_size)?;
|
||||||
let header_size = 8 + length_size as usize;
|
Ok(GlobalHeapCollection {
|
||||||
|
collection_size: index.collection_size,
|
||||||
|
objects: index
|
||||||
|
.objects
|
||||||
|
.iter()
|
||||||
|
.map(|o| GlobalHeapObject {
|
||||||
|
index: o.index,
|
||||||
|
reference_count: o.reference_count,
|
||||||
|
data: file_data[o.offset..o.offset + o.size].to_vec(),
|
||||||
|
})
|
||||||
|
.collect(),
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Locate the objects of the global heap collection at `offset` without
|
||||||
|
/// copying their data, so a caller can keep many collections indexed
|
||||||
|
/// for the cost of their object headers.
|
||||||
|
///
|
||||||
|
/// The collection must lie inside `file_data`, and every object inside
|
||||||
|
/// the collection, as libhdf5 lays them out; an object that runs past
|
||||||
|
/// its collection is an error.
|
||||||
|
pub fn parse_index(
|
||||||
|
file_data: &[u8],
|
||||||
|
offset: usize,
|
||||||
|
length_size: u8,
|
||||||
|
) -> Result<GlobalHeapIndex, FormatError> {
|
||||||
|
// signature(4) + version(1) + reserved(3) + collection_size(length_size),
|
||||||
|
// padded to a multiple of 8 as libhdf5 lays it out (`H5HG_SIZEOF_HDR`).
|
||||||
|
// With 8-byte lengths the padding is 0; with 4-byte lengths it is 4,
|
||||||
|
// and reading without it put every object 4 bytes early.
|
||||||
|
let header_size = pad8(8 + length_size as usize);
|
||||||
ensure_len(file_data, offset, header_size)?;
|
ensure_len(file_data, offset, header_size)?;
|
||||||
|
|
||||||
if file_data[offset..offset + 4] != GCOL_SIGNATURE {
|
if file_data[offset..offset + 4] != GCOL_SIGNATURE {
|
||||||
@@ -78,25 +139,25 @@ impl GlobalHeapCollection {
|
|||||||
}
|
}
|
||||||
|
|
||||||
let collection_size = read_length(file_data, offset + 8, length_size)?;
|
let collection_size = read_length(file_data, offset + 8, length_size)?;
|
||||||
let collection_size_usize =
|
let collection_end = usize::try_from(collection_size)
|
||||||
usize::try_from(collection_size).map_err(|_| FormatError::UnexpectedEof {
|
.ok()
|
||||||
expected: u64::MAX as usize,
|
.and_then(|size| offset.checked_add(size))
|
||||||
|
.ok_or(FormatError::UnexpectedEof {
|
||||||
|
expected: usize::MAX,
|
||||||
available: file_data.len(),
|
available: file_data.len(),
|
||||||
})?;
|
})?;
|
||||||
let collection_end =
|
if collection_end > file_data.len() {
|
||||||
offset
|
return Err(FormatError::UnexpectedEof {
|
||||||
.checked_add(collection_size_usize)
|
expected: collection_end,
|
||||||
.ok_or(FormatError::UnexpectedEof {
|
available: file_data.len(),
|
||||||
expected: usize::MAX,
|
});
|
||||||
available: file_data.len(),
|
}
|
||||||
})?;
|
|
||||||
|
|
||||||
let mut pos = offset + header_size;
|
let mut pos = offset + header_size;
|
||||||
let mut objects = Vec::new();
|
let mut objects = Vec::new();
|
||||||
|
|
||||||
// Parse objects until we hit index 0 (free space) or run out of space
|
// Parse objects until we hit index 0 (free space) or run out of space
|
||||||
while pos + 2 <= collection_end {
|
while pos + 2 <= collection_end {
|
||||||
ensure_len(file_data, pos, 2)?;
|
|
||||||
let object_index = u16::from_le_bytes([file_data[pos], file_data[pos + 1]]);
|
let object_index = u16::from_le_bytes([file_data[pos], file_data[pos + 1]]);
|
||||||
|
|
||||||
if object_index == 0 {
|
if object_index == 0 {
|
||||||
@@ -104,28 +165,39 @@ impl GlobalHeapCollection {
|
|||||||
break;
|
break;
|
||||||
}
|
}
|
||||||
|
|
||||||
// object_index(2) + reference_count(2) + reserved(4) + object_size(length_size)
|
// object_index(2) + reference_count(2) + reserved(4) +
|
||||||
let obj_header_size = 8 + length_size as usize;
|
// object_size(length_size), padded to 8 (`H5HG_SIZEOF_OBJHDR`).
|
||||||
ensure_len(file_data, pos, obj_header_size)?;
|
let obj_header_size = pad8(8 + length_size as usize);
|
||||||
|
ensure_len(&file_data[..collection_end], pos, obj_header_size)?;
|
||||||
|
|
||||||
let reference_count = u16::from_le_bytes([file_data[pos + 2], file_data[pos + 3]]);
|
let reference_count = u16::from_le_bytes([file_data[pos + 2], file_data[pos + 3]]);
|
||||||
let object_size = read_length(file_data, pos + 8, length_size)? as usize;
|
let object_size = usize::try_from(read_length(file_data, pos + 8, length_size)?)
|
||||||
|
.map_err(|_| FormatError::Overflow("global heap object size".into()))?;
|
||||||
|
|
||||||
pos += obj_header_size;
|
pos += obj_header_size;
|
||||||
ensure_len(file_data, pos, object_size)?;
|
if pos
|
||||||
let data = file_data[pos..pos + object_size].to_vec();
|
.checked_add(object_size)
|
||||||
|
.is_none_or(|end| end > collection_end)
|
||||||
|
{
|
||||||
|
return Err(FormatError::VlDataError(object_overrun_msg(
|
||||||
|
object_index,
|
||||||
|
object_size,
|
||||||
|
collection_size,
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
|
||||||
objects.push(GlobalHeapObject {
|
objects.push(GlobalHeapObjectRef {
|
||||||
index: object_index,
|
index: object_index,
|
||||||
reference_count,
|
reference_count,
|
||||||
data,
|
offset: pos,
|
||||||
|
size: object_size,
|
||||||
});
|
});
|
||||||
|
|
||||||
// Advance past data + padding to 8-byte boundary
|
// Advance past data + padding to 8-byte boundary
|
||||||
pos += pad8(object_size);
|
pos = pos.saturating_add(pad8(object_size));
|
||||||
}
|
}
|
||||||
|
|
||||||
Ok(GlobalHeapCollection {
|
Ok(GlobalHeapIndex {
|
||||||
collection_size,
|
collection_size,
|
||||||
objects,
|
objects,
|
||||||
})
|
})
|
||||||
@@ -149,10 +221,11 @@ mod tests {
|
|||||||
let ls = length_size as usize;
|
let ls = length_size as usize;
|
||||||
|
|
||||||
// Calculate total size
|
// Calculate total size
|
||||||
let header_size = 8 + ls;
|
// libhdf5 pads both headers to a multiple of 8.
|
||||||
|
let header_size = pad8(8 + ls);
|
||||||
let mut obj_size_total = 0usize;
|
let mut obj_size_total = 0usize;
|
||||||
for (_, _, data) in objects {
|
for (_, _, data) in objects {
|
||||||
let obj_header = 8 + ls;
|
let obj_header = pad8(8 + ls);
|
||||||
obj_size_total += obj_header + pad8(data.len());
|
obj_size_total += obj_header + pad8(data.len());
|
||||||
}
|
}
|
||||||
// Free space marker (2 bytes for index 0)
|
// Free space marker (2 bytes for index 0)
|
||||||
@@ -170,6 +243,7 @@ mod tests {
|
|||||||
8 => buf.extend_from_slice(&(collection_size as u64).to_le_bytes()),
|
8 => buf.extend_from_slice(&(collection_size as u64).to_le_bytes()),
|
||||||
_ => panic!("unsupported length_size"),
|
_ => panic!("unsupported length_size"),
|
||||||
}
|
}
|
||||||
|
buf.resize(header_size, 0);
|
||||||
|
|
||||||
// Objects
|
// Objects
|
||||||
for (index, ref_count, data) in objects {
|
for (index, ref_count, data) in objects {
|
||||||
@@ -181,6 +255,7 @@ mod tests {
|
|||||||
8 => buf.extend_from_slice(&(data.len() as u64).to_le_bytes()),
|
8 => buf.extend_from_slice(&(data.len() as u64).to_le_bytes()),
|
||||||
_ => panic!("unsupported"),
|
_ => panic!("unsupported"),
|
||||||
}
|
}
|
||||||
|
buf.resize(buf.len() + (pad8(8 + ls) - (8 + ls)), 0);
|
||||||
buf.extend_from_slice(data);
|
buf.extend_from_slice(data);
|
||||||
// Pad to 8 bytes
|
// Pad to 8 bytes
|
||||||
let padded = pad8(data.len());
|
let padded = pad8(data.len());
|
||||||
|
|||||||
@@ -5,10 +5,12 @@
|
|||||||
//! `sequence_length(4 LE) + collection_address(offset_size LE) + object_index(4 LE)`.
|
//! `sequence_length(4 LE) + collection_address(offset_size LE) + object_index(4 LE)`.
|
||||||
|
|
||||||
#[cfg(not(feature = "std"))]
|
#[cfg(not(feature = "std"))]
|
||||||
use alloc::{string::String, vec::Vec};
|
use alloc::{collections::BTreeMap, format, string::String, vec, vec::Vec};
|
||||||
|
#[cfg(feature = "std")]
|
||||||
|
use std::collections::BTreeMap;
|
||||||
|
|
||||||
use crate::error::FormatError;
|
use crate::error::FormatError;
|
||||||
use crate::global_heap::GlobalHeapCollection;
|
use crate::global_heap::{GlobalHeapCollection, GlobalHeapIndex};
|
||||||
|
|
||||||
/// A parsed variable-length element reference (global heap ID).
|
/// A parsed variable-length element reference (global heap ID).
|
||||||
#[derive(Debug, Clone)]
|
#[derive(Debug, Clone)]
|
||||||
@@ -109,7 +111,218 @@ fn is_undefined_address(addr: u64, offset_size: u8) -> bool {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The size of one variable-length element in a file with `offset_size`-byte
|
||||||
|
/// addresses: a sequence length (4), a global heap collection address and an
|
||||||
|
/// object index (4). libhdf5 computes it this way rather than trusting the
|
||||||
|
/// datatype message (`H5T_set_loc`).
|
||||||
|
pub fn element_size(offset_size: u8) -> usize {
|
||||||
|
4 + offset_size as usize + 4
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Refuse a variable-length datatype whose stored element size is not the
|
||||||
|
/// one this file's offset size implies. Its elements would be laid out with
|
||||||
|
/// a stride libhdf5 does not use, so every value after the first would be
|
||||||
|
/// read from the wrong place.
|
||||||
|
pub fn check_element_size(stored_size: u32, offset_size: u8) -> Result<(), FormatError> {
|
||||||
|
let expected = element_size(offset_size);
|
||||||
|
if stored_size as usize != expected {
|
||||||
|
return Err(FormatError::VlDataError(format!(
|
||||||
|
"variable-length datatype stores {stored_size}-byte elements; a file with \
|
||||||
|
{offset_size}-byte offsets uses {expected}"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A collection's objects, located in the file data but not copied:
|
||||||
|
/// `(index, offset, size)` of the first object with each index, sorted by
|
||||||
|
/// index.
|
||||||
|
struct CachedCollection {
|
||||||
|
objects: Vec<(u16, usize, usize)>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl CachedCollection {
|
||||||
|
fn new(index: GlobalHeapIndex) -> Self {
|
||||||
|
let mut objects: Vec<(u16, usize, usize)> = index
|
||||||
|
.objects
|
||||||
|
.iter()
|
||||||
|
.map(|o| (o.index, o.offset, o.size))
|
||||||
|
.collect();
|
||||||
|
// Stable, so the first object with a repeated index is kept.
|
||||||
|
objects.sort_by_key(|o| o.0);
|
||||||
|
objects.dedup_by_key(|o| o.0);
|
||||||
|
Self { objects }
|
||||||
|
}
|
||||||
|
|
||||||
|
/// What this entry costs to keep, in bytes (roughly).
|
||||||
|
fn cost(&self) -> usize {
|
||||||
|
64 + self.objects.len() * core::mem::size_of::<(u16, usize, usize)>()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn get(&self, index: u32) -> Option<(usize, usize)> {
|
||||||
|
let index = u16::try_from(index).ok()?;
|
||||||
|
let i = self.objects.binary_search_by_key(&index, |o| o.0).ok()?;
|
||||||
|
Some((self.objects[i].1, self.objects[i].2))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// How many bytes of collection indexes a [`VlResolver`] keeps before it
|
||||||
|
/// drops them and starts again. Values are never copied into the cache, so
|
||||||
|
/// this bounds what a read retains however many collections it visits.
|
||||||
|
const CACHE_BUDGET: usize = 32 << 20;
|
||||||
|
|
||||||
|
/// Resolves variable-length elements against a file's global heap, parsing
|
||||||
|
/// each heap collection once however many elements point into it.
|
||||||
|
///
|
||||||
|
/// Values follow libhdf5: an element whose heap address is 0 is null (an
|
||||||
|
/// empty string or sequence), and an element whose heap object is not
|
||||||
|
/// exactly `length × base size` bytes is an error ("Expected global heap
|
||||||
|
/// object size does not match"), not a truncated or padded value.
|
||||||
|
///
|
||||||
|
/// Memory stays bounded on hostile files: the cache holds where each
|
||||||
|
/// object lies, not a copy of it, up to a fixed budget; and collections
|
||||||
|
/// that overlap one another are refused (libhdf5 never writes them), so a
|
||||||
|
/// file cannot make the resolver parse the same bytes as the objects of
|
||||||
|
/// many collections.
|
||||||
|
pub struct VlResolver<'a> {
|
||||||
|
file_data: &'a [u8],
|
||||||
|
offset_size: u8,
|
||||||
|
length_size: u8,
|
||||||
|
cache: BTreeMap<u64, CachedCollection>,
|
||||||
|
cached_bytes: usize,
|
||||||
|
budget: usize,
|
||||||
|
/// Start → end of every collection parsed so far (kept when the cache
|
||||||
|
/// is dropped, to check overlaps).
|
||||||
|
extents: BTreeMap<usize, usize>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl<'a> VlResolver<'a> {
|
||||||
|
/// A resolver over `file_data` (the file from its superblock on), with
|
||||||
|
/// the superblock's offset and length sizes.
|
||||||
|
pub fn new(file_data: &'a [u8], offset_size: u8, length_size: u8) -> Self {
|
||||||
|
Self {
|
||||||
|
file_data,
|
||||||
|
offset_size,
|
||||||
|
length_size,
|
||||||
|
cache: BTreeMap::new(),
|
||||||
|
cached_bytes: 0,
|
||||||
|
budget: CACHE_BUDGET,
|
||||||
|
extents: BTreeMap::new(),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The size of one element in this file (see [`element_size`]).
|
||||||
|
pub fn element_size(&self) -> usize {
|
||||||
|
element_size(self.offset_size)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Split `raw` into elements; its length must be a whole number of them.
|
||||||
|
fn elements(&self, raw: &[u8]) -> Result<Vec<VlElement>, FormatError> {
|
||||||
|
let size = self.element_size();
|
||||||
|
if !raw.len().is_multiple_of(size) {
|
||||||
|
return Err(FormatError::VlDataError(format!(
|
||||||
|
"{} bytes is not a whole number of {size}-byte variable-length elements",
|
||||||
|
raw.len()
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
parse_vl_references(raw, (raw.len() / size) as u64, self.offset_size)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The bytes of one element: `length × base_size` bytes from the heap,
|
||||||
|
/// or `None` for a null element.
|
||||||
|
fn resolve(
|
||||||
|
&mut self,
|
||||||
|
vl: &VlElement,
|
||||||
|
base_size: usize,
|
||||||
|
) -> Result<Option<&'a [u8]>, FormatError> {
|
||||||
|
let addr = vl.collection_address;
|
||||||
|
if addr == 0 {
|
||||||
|
return Ok(None);
|
||||||
|
}
|
||||||
|
let data = self.object(vl)?;
|
||||||
|
let expected = (vl.length as usize)
|
||||||
|
.checked_mul(base_size)
|
||||||
|
.ok_or_else(|| FormatError::Overflow("variable-length element size".into()))?;
|
||||||
|
if data.len() != expected {
|
||||||
|
return Err(FormatError::VlDataError(format!(
|
||||||
|
"global heap object {} in the collection at {addr} holds {} bytes; the element \
|
||||||
|
says {} × {base_size}",
|
||||||
|
vl.object_index,
|
||||||
|
data.len(),
|
||||||
|
vl.length
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
Ok(Some(data))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// One element (the first [`element_size`](Self::element_size) bytes of
|
||||||
|
/// `elem`) of a variable-length sequence whose base type is `base_size`
|
||||||
|
/// bytes: its `length × base_size` bytes, or `None` for a null element
|
||||||
|
/// (heap address 0).
|
||||||
|
pub fn element(
|
||||||
|
&mut self,
|
||||||
|
elem: &[u8],
|
||||||
|
base_size: usize,
|
||||||
|
) -> Result<Option<&'a [u8]>, FormatError> {
|
||||||
|
let vl = parse_vl_references(elem, 1, self.offset_size)?;
|
||||||
|
self.resolve(&vl[0], base_size)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// One variable-length string element: its bytes up to the first NUL,
|
||||||
|
/// or `None` for a null element (h5dump prints it as `NULL`, h5py
|
||||||
|
/// returns it as empty).
|
||||||
|
pub fn string_element(&mut self, elem: &[u8]) -> Result<Option<&'a [u8]>, FormatError> {
|
||||||
|
Ok(self.element(elem, 1)?.map(cut_at_nul))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The strings of the variable-length string elements in `raw`, as
|
||||||
|
/// bytes. A string ends at its first NUL, as libhdf5 returns it (it
|
||||||
|
/// converts each to a C string); a null element is empty.
|
||||||
|
pub fn string_bytes(&mut self, raw: &[u8]) -> Result<Vec<Vec<u8>>, FormatError> {
|
||||||
|
self.elements(raw)?
|
||||||
|
.iter()
|
||||||
|
.map(|vl| Ok(self.resolve(vl, 1)?.map(cut_at_nul).unwrap_or(&[]).to_vec()))
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The strings of the variable-length string elements in `raw`, decoded
|
||||||
|
/// as UTF-8 with invalid sequences replaced by U+FFFD (see
|
||||||
|
/// [`string_bytes`](Self::string_bytes) for the exact bytes).
|
||||||
|
pub fn strings(&mut self, raw: &[u8]) -> Result<Vec<String>, FormatError> {
|
||||||
|
Ok(self
|
||||||
|
.string_bytes(raw)?
|
||||||
|
.into_iter()
|
||||||
|
.map(|b| match String::from_utf8(b) {
|
||||||
|
Ok(s) => s,
|
||||||
|
Err(e) => String::from_utf8_lossy(e.as_bytes()).into_owned(),
|
||||||
|
})
|
||||||
|
.collect())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The sequences of the variable-length sequence elements in `raw`, each
|
||||||
|
/// as its `length × base_size` bytes in the base type's encoding.
|
||||||
|
pub fn sequences(&mut self, raw: &[u8], base_size: usize) -> Result<Vec<Vec<u8>>, FormatError> {
|
||||||
|
if base_size == 0 {
|
||||||
|
return Err(FormatError::VlDataError(
|
||||||
|
"variable-length sequence of a zero-size base type".into(),
|
||||||
|
));
|
||||||
|
}
|
||||||
|
self.elements(raw)?
|
||||||
|
.iter()
|
||||||
|
.map(|vl| Ok(self.resolve(vl, base_size)?.unwrap_or(&[]).to_vec()))
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A string's bytes up to its first NUL.
|
||||||
|
fn cut_at_nul(s: &[u8]) -> &[u8] {
|
||||||
|
&s[..s.iter().position(|&b| b == 0).unwrap_or(s.len())]
|
||||||
|
}
|
||||||
|
|
||||||
/// Resolve VL strings from raw data by looking up each element in the global heap.
|
/// Resolve VL strings from raw data by looking up each element in the global heap.
|
||||||
|
///
|
||||||
|
/// Reads the first `num_elements` elements of `raw`. Strings end at their
|
||||||
|
/// first NUL and invalid UTF-8 is replaced, as in [`VlResolver::strings`].
|
||||||
pub fn read_vl_strings(
|
pub fn read_vl_strings(
|
||||||
file_data: &[u8],
|
file_data: &[u8],
|
||||||
raw_data: &[u8],
|
raw_data: &[u8],
|
||||||
@@ -117,35 +330,23 @@ pub fn read_vl_strings(
|
|||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
length_size: u8,
|
length_size: u8,
|
||||||
) -> Result<Vec<String>, FormatError> {
|
) -> Result<Vec<String>, FormatError> {
|
||||||
let refs = parse_vl_references(raw_data, num_elements, offset_size)?;
|
let raw = first_elements(raw_data, num_elements, offset_size)?;
|
||||||
let mut result = Vec::with_capacity(refs.len());
|
VlResolver::new(file_data, offset_size, length_size).strings(raw)
|
||||||
|
}
|
||||||
|
|
||||||
for vl in &refs {
|
/// The first `num_elements` elements of `raw`, or an error if it is shorter.
|
||||||
if vl.length == 0 && is_undefined_address(vl.collection_address, offset_size) {
|
fn first_elements(raw: &[u8], num_elements: u64, offset_size: u8) -> Result<&[u8], FormatError> {
|
||||||
result.push(String::new());
|
let total = usize::try_from(num_elements)
|
||||||
continue;
|
.ok()
|
||||||
}
|
.and_then(|n| n.checked_mul(element_size(offset_size)))
|
||||||
if vl.length == 0 && vl.collection_address == 0 {
|
.ok_or(FormatError::UnexpectedEof {
|
||||||
result.push(String::new());
|
expected: usize::MAX,
|
||||||
continue;
|
available: raw.len(),
|
||||||
}
|
})?;
|
||||||
|
raw.get(..total).ok_or(FormatError::UnexpectedEof {
|
||||||
let coll =
|
expected: total,
|
||||||
GlobalHeapCollection::parse(file_data, vl.collection_address as usize, length_size)?;
|
available: raw.len(),
|
||||||
let obj = coll.get_object(vl.object_index as u16).ok_or(
|
})
|
||||||
FormatError::GlobalHeapObjectNotFound {
|
|
||||||
collection_address: vl.collection_address,
|
|
||||||
index: vl.object_index as u16,
|
|
||||||
},
|
|
||||||
)?;
|
|
||||||
|
|
||||||
// The object data is the raw string bytes
|
|
||||||
let len = (vl.length as usize).min(obj.data.len());
|
|
||||||
let s = String::from_utf8_lossy(&obj.data[..len]).into_owned();
|
|
||||||
result.push(s);
|
|
||||||
}
|
|
||||||
|
|
||||||
Ok(result)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Resolve VL sequences from raw data, returning each element's bytes.
|
/// Resolve VL sequences from raw data, returning each element's bytes.
|
||||||
@@ -153,7 +354,9 @@ pub fn read_vl_strings(
|
|||||||
/// Each element is the sequence's full encoding — element count × base type
|
/// Each element is the sequence's full encoding — element count × base type
|
||||||
/// size bytes, in the base type's byte order — so a sequence of `i32` yields
|
/// size bytes, in the base type's byte order — so a sequence of `i32` yields
|
||||||
/// four bytes per value. Decode it with the base type (e.g.
|
/// four bytes per value. Decode it with the base type (e.g.
|
||||||
/// [`crate::data_read::read_as_i64`]).
|
/// [`crate::data_read::read_as_i64`]). This does not know the base type, so
|
||||||
|
/// it returns each heap object whole; [`VlResolver::sequences`] also checks
|
||||||
|
/// the object's size against the element's length.
|
||||||
pub fn read_vl_bytes(
|
pub fn read_vl_bytes(
|
||||||
file_data: &[u8],
|
file_data: &[u8],
|
||||||
raw_data: &[u8],
|
raw_data: &[u8],
|
||||||
@@ -162,35 +365,97 @@ pub fn read_vl_bytes(
|
|||||||
length_size: u8,
|
length_size: u8,
|
||||||
) -> Result<Vec<Vec<u8>>, FormatError> {
|
) -> Result<Vec<Vec<u8>>, FormatError> {
|
||||||
let refs = parse_vl_references(raw_data, num_elements, offset_size)?;
|
let refs = parse_vl_references(raw_data, num_elements, offset_size)?;
|
||||||
|
let mut resolver = VlResolver::new(file_data, offset_size, length_size);
|
||||||
let mut result = Vec::with_capacity(refs.len());
|
let mut result = Vec::with_capacity(refs.len());
|
||||||
|
|
||||||
for vl in &refs {
|
for vl in &refs {
|
||||||
if vl.length == 0
|
// A heap address of 0 is a null element, as in VlResolver.
|
||||||
&& (is_undefined_address(vl.collection_address, offset_size)
|
if vl.collection_address == 0 {
|
||||||
|| vl.collection_address == 0)
|
|
||||||
{
|
|
||||||
result.push(Vec::new());
|
result.push(Vec::new());
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
|
|
||||||
let coll =
|
|
||||||
GlobalHeapCollection::parse(file_data, vl.collection_address as usize, length_size)?;
|
|
||||||
let obj = coll.get_object(vl.object_index as u16).ok_or(
|
|
||||||
FormatError::GlobalHeapObjectNotFound {
|
|
||||||
collection_address: vl.collection_address,
|
|
||||||
index: vl.object_index as u16,
|
|
||||||
},
|
|
||||||
)?;
|
|
||||||
|
|
||||||
// The heap object holds the whole sequence. `vl.length` counts
|
// The heap object holds the whole sequence. `vl.length` counts
|
||||||
// elements, not bytes, so it is only the byte length when the base
|
// elements, not bytes, so it is only the byte length when the base
|
||||||
// type is one byte wide.
|
// type is one byte wide.
|
||||||
result.push(obj.data.clone());
|
let obj = resolver.object(vl)?;
|
||||||
|
result.push(obj.to_vec());
|
||||||
}
|
}
|
||||||
|
|
||||||
Ok(result)
|
Ok(result)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
impl<'a> VlResolver<'a> {
|
||||||
|
/// The heap object `vl` points to, whatever its size; its collection is
|
||||||
|
/// parsed on first use.
|
||||||
|
fn object(&mut self, vl: &VlElement) -> Result<&'a [u8], FormatError> {
|
||||||
|
let addr = vl.collection_address;
|
||||||
|
// libhdf5 writes a null element with address 0, never the undefined
|
||||||
|
// address, and fails to read one ("addr undefined") even when its
|
||||||
|
// length is 0; we returned an empty value.
|
||||||
|
if is_undefined_address(addr, self.offset_size) {
|
||||||
|
return Err(FormatError::VlDataError(format!(
|
||||||
|
"variable-length element (length {}) has the undefined global heap address",
|
||||||
|
vl.length
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
if !self.cache.contains_key(&addr) {
|
||||||
|
let offset = usize::try_from(addr).map_err(|_| FormatError::UnexpectedEof {
|
||||||
|
expected: usize::MAX,
|
||||||
|
available: self.file_data.len(),
|
||||||
|
})?;
|
||||||
|
let index =
|
||||||
|
GlobalHeapCollection::parse_index(self.file_data, offset, self.length_size)?;
|
||||||
|
// parse_index checked that the collection lies in the file.
|
||||||
|
let end = offset + index.collection_size as usize;
|
||||||
|
self.check_overlap(offset, end)?;
|
||||||
|
let coll = CachedCollection::new(index);
|
||||||
|
if self.cached_bytes.saturating_add(coll.cost()) > self.budget {
|
||||||
|
self.cache.clear();
|
||||||
|
self.cached_bytes = 0;
|
||||||
|
}
|
||||||
|
self.cached_bytes += coll.cost();
|
||||||
|
self.cache.insert(addr, coll);
|
||||||
|
}
|
||||||
|
let (start, size) = self.cache[&addr].get(vl.object_index).ok_or(
|
||||||
|
FormatError::GlobalHeapObjectNotFound {
|
||||||
|
collection_address: addr,
|
||||||
|
index: vl.object_index as u16,
|
||||||
|
},
|
||||||
|
)?;
|
||||||
|
Ok(&self.file_data[start..start + size])
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Record the collection at `start..end`, refusing one that overlaps a
|
||||||
|
/// collection already read. libhdf5 allocates each collection its own
|
||||||
|
/// block; overlapping ones only come from a crafted file, where they let
|
||||||
|
/// every byte be parsed again as the objects of each collection.
|
||||||
|
fn check_overlap(&mut self, start: usize, end: usize) -> Result<(), FormatError> {
|
||||||
|
if let Some(&known) = self.extents.get(&start) {
|
||||||
|
return if known == end {
|
||||||
|
Ok(())
|
||||||
|
} else {
|
||||||
|
Err(FormatError::VlDataError(format!(
|
||||||
|
"global heap collection at {start} changed size"
|
||||||
|
)))
|
||||||
|
};
|
||||||
|
}
|
||||||
|
let before = self.extents.range(..start).next_back();
|
||||||
|
let after = self.extents.range(start..).next();
|
||||||
|
let clash = match (before, after) {
|
||||||
|
(Some((&s, &e)), _) if e > start => Some(s),
|
||||||
|
(_, Some((&s, _))) if s < end => Some(s),
|
||||||
|
_ => None,
|
||||||
|
};
|
||||||
|
if let Some(other) = clash {
|
||||||
|
return Err(FormatError::VlDataError(format!(
|
||||||
|
"global heap collection at {start} overlaps the one at {other}"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
self.extents.insert(start, end);
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
#[cfg(test)]
|
#[cfg(test)]
|
||||||
mod tests {
|
mod tests {
|
||||||
use super::*;
|
use super::*;
|
||||||
@@ -285,16 +550,27 @@ mod tests {
|
|||||||
}
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn null_vl_element_empty_string() {
|
fn an_undefined_heap_address_is_an_error_even_at_length_0() {
|
||||||
// length=0, address=undefined
|
// libhdf5 fails the read ("addr undefined"); h5py and libhdf5 write
|
||||||
let mut raw = Vec::new();
|
// a null element with address 0. We returned "".
|
||||||
raw.extend_from_slice(&0u32.to_le_bytes()); // length=0
|
let mut file_data = vec![0u8; 256];
|
||||||
raw.extend_from_slice(&u64::MAX.to_le_bytes()); // undefined address
|
build_gcol_at(&mut file_data, 64, &[(1, b"x")]);
|
||||||
raw.extend_from_slice(&0u32.to_le_bytes()); // index
|
for (os, undef) in [(8u8, u64::MAX), (4, 0xFFFF_FFFF)] {
|
||||||
|
for length in [0, 1] {
|
||||||
let file_data = vec![0u8; 16];
|
let mut raw = element(1, 64, 1, os);
|
||||||
let strings = read_vl_strings(&file_data, &raw, 1, 8, 8).unwrap();
|
raw.extend(element(length, undef, 1, os));
|
||||||
assert_eq!(strings, vec![""]);
|
let mut r = VlResolver::new(&file_data, os, 8);
|
||||||
|
let e = r.string_bytes(&raw).unwrap_err().to_string();
|
||||||
|
assert!(e.contains("undefined"), "{e}");
|
||||||
|
assert!(r.sequences(&raw, 1).is_err());
|
||||||
|
assert!(r.string_element(&raw[raw.len() / 2..]).is_err());
|
||||||
|
let n = 2;
|
||||||
|
assert!(read_vl_strings(&file_data, &raw, n, os, 8).is_err());
|
||||||
|
assert!(read_vl_bytes(&file_data, &raw, n, os, 8).is_err());
|
||||||
|
// The defined element alone still reads.
|
||||||
|
assert_eq!(r.strings(&raw[..raw.len() / 2]).unwrap(), ["x"]);
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
@@ -333,6 +609,126 @@ mod tests {
|
|||||||
assert_eq!(bytes, vec![vec![0xDE, 0xAD], vec![0xBE, 0xEF, 0xCA]]);
|
assert_eq!(bytes, vec![vec![0xDE, 0xAD], vec![0xBE, 0xEF, 0xCA]]);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
fn element(length: u32, addr: u64, index: u32, offset_size: u8) -> Vec<u8> {
|
||||||
|
let mut raw = length.to_le_bytes().to_vec();
|
||||||
|
raw.extend_from_slice(&addr.to_le_bytes()[..offset_size as usize]);
|
||||||
|
raw.extend_from_slice(&index.to_le_bytes());
|
||||||
|
raw
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn strings_end_at_the_first_nul() {
|
||||||
|
// libhdf5 hands each VL string over as a C string, so h5py sees
|
||||||
|
// "a\0b" as "a"; we used to return the NUL and what followed.
|
||||||
|
let mut file_data = vec![0u8; 512];
|
||||||
|
build_gcol_at(&mut file_data, 64, &[(1, b"a\0b"), (2, b"cd")]);
|
||||||
|
let mut raw = element(3, 64, 1, 8);
|
||||||
|
raw.extend(element(2, 64, 2, 8));
|
||||||
|
let mut r = VlResolver::new(&file_data, 8, 8);
|
||||||
|
assert_eq!(
|
||||||
|
r.string_bytes(&raw).unwrap(),
|
||||||
|
vec![b"a".to_vec(), b"cd".to_vec()]
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
read_vl_strings(&file_data, &raw, 2, 8, 8).unwrap(),
|
||||||
|
["a", "cd"]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_heap_object_of_the_wrong_size_is_an_error() {
|
||||||
|
// libhdf5: "Expected global heap object size does not match". We
|
||||||
|
// used to return the object cut to the element's length.
|
||||||
|
let mut file_data = vec![0u8; 512];
|
||||||
|
build_gcol_at(&mut file_data, 64, &[(1, b"cdefgh"), (2, &[1, 0, 0, 0])]);
|
||||||
|
let mut r = VlResolver::new(&file_data, 8, 8);
|
||||||
|
assert!(r.string_bytes(&element(3, 64, 1, 8)).is_err());
|
||||||
|
assert!(r.string_bytes(&element(9, 64, 1, 8)).is_err());
|
||||||
|
assert!(read_vl_strings(&file_data, &element(3, 64, 1, 8), 1, 8, 8).is_err());
|
||||||
|
// A sequence of one i32 is 4 bytes; of two, 8.
|
||||||
|
assert_eq!(
|
||||||
|
r.sequences(&element(1, 64, 2, 8), 4).unwrap(),
|
||||||
|
vec![vec![1, 0, 0, 0]]
|
||||||
|
);
|
||||||
|
assert!(r.sequences(&element(2, 64, 2, 8), 4).is_err());
|
||||||
|
assert!(r.sequences(&element(1, 64, 2, 8), 0).is_err());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn address_zero_is_null_whatever_the_length() {
|
||||||
|
// libhdf5 treats a heap address of 0 as a null element.
|
||||||
|
let file_data = vec![0u8; 64];
|
||||||
|
let mut r = VlResolver::new(&file_data, 8, 8);
|
||||||
|
assert_eq!(
|
||||||
|
r.string_bytes(&element(5, 0, 1, 8)).unwrap(),
|
||||||
|
vec![Vec::<u8>::new()]
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
r.sequences(&element(5, 0, 1, 8), 4).unwrap(),
|
||||||
|
vec![Vec::<u8>::new()]
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn four_byte_offsets_use_twelve_byte_elements() {
|
||||||
|
let mut file_data = vec![0u8; 512];
|
||||||
|
build_gcol_at(&mut file_data, 64, &[(1, b"one"), (2, b""), (3, b"three")]);
|
||||||
|
let mut raw = element(3, 64, 1, 4);
|
||||||
|
raw.extend(element(0, 64, 2, 4));
|
||||||
|
raw.extend(element(5, 64, 3, 4));
|
||||||
|
assert_eq!(raw.len(), 36);
|
||||||
|
let mut r = VlResolver::new(&file_data, 4, 8);
|
||||||
|
assert_eq!(r.element_size(), 12);
|
||||||
|
assert_eq!(r.strings(&raw).unwrap(), ["one", "", "three"]);
|
||||||
|
// Not a whole number of elements.
|
||||||
|
assert!(r.strings(&raw[..30]).is_err());
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn the_cache_stays_within_its_budget_and_rereads_what_it_dropped() {
|
||||||
|
// Twenty collections of three objects each; a budget that holds
|
||||||
|
// about two of them. Reading every element twice must still return
|
||||||
|
// the right strings after the cache is dropped.
|
||||||
|
let mut file_data = vec![0u8; 64];
|
||||||
|
let mut raw = Vec::new();
|
||||||
|
for c in 0..20u64 {
|
||||||
|
let at = file_data.len();
|
||||||
|
let names: Vec<String> = (0..3).map(|i| format!("c{c}o{i}")).collect();
|
||||||
|
let objs: Vec<(u16, &[u8])> = names
|
||||||
|
.iter()
|
||||||
|
.enumerate()
|
||||||
|
.map(|(i, n)| (i as u16 + 1, n.as_bytes()))
|
||||||
|
.collect();
|
||||||
|
build_gcol_at(&mut file_data, at, &objs);
|
||||||
|
for (i, n) in names.iter().enumerate() {
|
||||||
|
raw.extend(element(n.len() as u32, at as u64, i as u32 + 1, 8));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
raw.extend(raw.clone());
|
||||||
|
let mut r = VlResolver::new(&file_data, 8, 8);
|
||||||
|
let one = CachedCollection {
|
||||||
|
objects: vec![(0, 0, 0); 3],
|
||||||
|
}
|
||||||
|
.cost();
|
||||||
|
r.budget = 2 * one + 1;
|
||||||
|
let want: Vec<String> = (0..2)
|
||||||
|
.flat_map(|_| (0..20).flat_map(|c| (0..3).map(move |i| format!("c{c}o{i}"))))
|
||||||
|
.collect();
|
||||||
|
for (k, chunk) in raw.chunks(16).enumerate() {
|
||||||
|
assert_eq!(r.strings(chunk).unwrap(), [want[k].clone()]);
|
||||||
|
assert!(r.cached_bytes <= r.budget);
|
||||||
|
assert!(r.cache.len() <= 2);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn element_size_is_checked_against_the_offset_size() {
|
||||||
|
assert!(check_element_size(16, 8).is_ok());
|
||||||
|
assert!(check_element_size(12, 4).is_ok());
|
||||||
|
assert!(check_element_size(16, 4).is_err());
|
||||||
|
assert!(check_element_size(524_304, 8).is_err());
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn parse_vl_references_truncated_error() {
|
fn parse_vl_references_truncated_error() {
|
||||||
let raw = vec![0u8; 10]; // too short for 1 element with offset_size=8
|
let raw = vec![0u8; 10]; // too short for 1 element with offset_size=8
|
||||||
|
|||||||
@@ -0,0 +1,161 @@
|
|||||||
|
//! Crafted files cannot make variable-length reads retain memory, or take
|
||||||
|
//! time, out of proportion to the file.
|
||||||
|
//!
|
||||||
|
//! `VlResolver` used to keep an owned copy of every object of every
|
||||||
|
//! collection it parsed, for the whole read. A file whose global heap
|
||||||
|
//! collections nest inside each other's object data — each element
|
||||||
|
//! pointing at a different one — then made retained memory O(K × file
|
||||||
|
//! size): a 744 KB file took 1.58 GB. The same nesting, with every
|
||||||
|
//! collection's object chain jumping to one shared run of tiny objects,
|
||||||
|
//! made the parse time O(K × M) as well. libhdf5 never writes overlapping
|
||||||
|
//! collections; they are now refused, and the cache holds only where
|
||||||
|
//! objects lie.
|
||||||
|
//!
|
||||||
|
//! Peak heap use is measured with a counting global allocator, so the
|
||||||
|
//! cases run one after another in a single test.
|
||||||
|
|
||||||
|
use std::alloc::{GlobalAlloc, Layout, System};
|
||||||
|
use std::sync::atomic::{AtomicUsize, Ordering};
|
||||||
|
use std::time::{Duration, Instant};
|
||||||
|
|
||||||
|
use clawhdf5_format::vl_data::VlResolver;
|
||||||
|
|
||||||
|
struct Counting;
|
||||||
|
|
||||||
|
static CURRENT: AtomicUsize = AtomicUsize::new(0);
|
||||||
|
static PEAK: AtomicUsize = AtomicUsize::new(0);
|
||||||
|
|
||||||
|
unsafe impl GlobalAlloc for Counting {
|
||||||
|
unsafe fn alloc(&self, layout: Layout) -> *mut u8 {
|
||||||
|
let p = unsafe { System.alloc(layout) };
|
||||||
|
if !p.is_null() {
|
||||||
|
let now = CURRENT.fetch_add(layout.size(), Ordering::Relaxed) + layout.size();
|
||||||
|
PEAK.fetch_max(now, Ordering::Relaxed);
|
||||||
|
}
|
||||||
|
p
|
||||||
|
}
|
||||||
|
|
||||||
|
unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) {
|
||||||
|
unsafe { System.dealloc(ptr, layout) };
|
||||||
|
CURRENT.fetch_sub(layout.size(), Ordering::Relaxed);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[global_allocator]
|
||||||
|
static ALLOC: Counting = Counting;
|
||||||
|
|
||||||
|
/// Bytes allocated at the peak of `f`, above what was live when it started.
|
||||||
|
fn peak_during<T>(f: impl FnOnce() -> T) -> (T, usize) {
|
||||||
|
let base = CURRENT.load(Ordering::Relaxed);
|
||||||
|
PEAK.store(base, Ordering::Relaxed);
|
||||||
|
let out = f();
|
||||||
|
(out, PEAK.load(Ordering::Relaxed) - base)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn put_header(file: &mut [u8], at: usize, size: u64) {
|
||||||
|
file[at..at + 4].copy_from_slice(b"GCOL");
|
||||||
|
file[at + 4] = 1;
|
||||||
|
file[at + 8..at + 16].copy_from_slice(&size.to_le_bytes());
|
||||||
|
}
|
||||||
|
|
||||||
|
fn put_object(file: &mut [u8], at: usize, index: u16, size: u64) {
|
||||||
|
file[at..at + 2].copy_from_slice(&index.to_le_bytes());
|
||||||
|
file[at + 2..at + 4].copy_from_slice(&1u16.to_le_bytes());
|
||||||
|
file[at + 8..at + 16].copy_from_slice(&size.to_le_bytes());
|
||||||
|
}
|
||||||
|
|
||||||
|
fn element(length: u32, addr: u64, index: u32) -> Vec<u8> {
|
||||||
|
let mut e = length.to_le_bytes().to_vec();
|
||||||
|
e.extend_from_slice(&addr.to_le_bytes());
|
||||||
|
e.extend_from_slice(&index.to_le_bytes());
|
||||||
|
e
|
||||||
|
}
|
||||||
|
|
||||||
|
/// K collections 32 bytes apart, each running to the end of the file with
|
||||||
|
/// one object covering the rest of it (and so every later collection).
|
||||||
|
/// Element i is that object of collection i.
|
||||||
|
fn nested(k: usize) -> (Vec<u8>, Vec<u8>) {
|
||||||
|
let base = 64;
|
||||||
|
let end = base + 32 * k + 64;
|
||||||
|
let mut file = vec![0u8; end];
|
||||||
|
let mut raw = Vec::new();
|
||||||
|
for i in 0..k {
|
||||||
|
let at = base + 32 * i;
|
||||||
|
put_header(&mut file, at, (end - at) as u64);
|
||||||
|
let obj = (end - at - 32) as u64;
|
||||||
|
put_object(&mut file, at + 16, 1, obj);
|
||||||
|
raw.extend(element(obj as u32, at as u64, 1));
|
||||||
|
}
|
||||||
|
(file, raw)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// K collections 32 bytes apart, each with a first object that jumps over
|
||||||
|
/// the later collections to one shared run of M empty objects, so parsing
|
||||||
|
/// every collection walks all M.
|
||||||
|
fn shared_tail(k: usize, m: usize) -> (Vec<u8>, Vec<u8>) {
|
||||||
|
let base = 64;
|
||||||
|
let tail = base + 32 * k + 32;
|
||||||
|
let end = tail + 16 * m + 16;
|
||||||
|
let mut file = vec![0u8; end];
|
||||||
|
let mut raw = Vec::new();
|
||||||
|
for i in 0..k {
|
||||||
|
let at = base + 32 * i;
|
||||||
|
put_header(&mut file, at, (end - at) as u64);
|
||||||
|
let jump = (tail - at - 32) as u64;
|
||||||
|
put_object(&mut file, at + 16, 1, jump);
|
||||||
|
raw.extend(element(jump as u32, at as u64, 1));
|
||||||
|
}
|
||||||
|
for j in 0..m {
|
||||||
|
put_object(&mut file, tail + 16 * j, (j % 65_000 + 2) as u16, 0);
|
||||||
|
}
|
||||||
|
(file, raw)
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn overlapping_collections_are_refused_in_bounded_memory_and_time() {
|
||||||
|
for (name, (file, raw)) in [
|
||||||
|
("nested", nested(2000)),
|
||||||
|
("shared tail", shared_tail(500, 10_000)),
|
||||||
|
] {
|
||||||
|
let start = Instant::now();
|
||||||
|
let (result, peak) = peak_during(|| {
|
||||||
|
let mut r = VlResolver::new(&file, 8, 8);
|
||||||
|
(r.string_bytes(&raw), r.sequences(&raw, 1).map(|s| s.len()))
|
||||||
|
});
|
||||||
|
let took = start.elapsed();
|
||||||
|
// libhdf5 never writes overlapping collections, and refuses these
|
||||||
|
// files; so do we, rather than returning what they claim.
|
||||||
|
let (strings, sequences) = result;
|
||||||
|
let e = strings.expect_err(name).to_string();
|
||||||
|
assert!(e.contains("overlaps"), "{name}: {e}");
|
||||||
|
assert!(sequences.is_err(), "{name}");
|
||||||
|
// Measured before the fix: 129 MB ("nested", 64 KB file) and 350 MB
|
||||||
|
// ("shared tail", 176 KB file) live at the peak; after, 97 KB and
|
||||||
|
// 0.9 MB.
|
||||||
|
assert!(
|
||||||
|
peak < 4 * file.len() + (1 << 20),
|
||||||
|
"{name}: peak {peak} bytes for a {}-byte file",
|
||||||
|
file.len()
|
||||||
|
);
|
||||||
|
assert!(took < Duration::from_secs(5), "{name}: took {took:?}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Collections that do not overlap still read, however many elements point
|
||||||
|
/// into them, and the first object of a collection is returned for its
|
||||||
|
/// index (as before).
|
||||||
|
#[test]
|
||||||
|
fn separate_collections_still_read() {
|
||||||
|
let mut file = vec![0u8; 64 + 3 * 64];
|
||||||
|
let mut raw = Vec::new();
|
||||||
|
for i in 0..3usize {
|
||||||
|
let at = 64 + 64 * i;
|
||||||
|
put_header(&mut file, at, 64);
|
||||||
|
put_object(&mut file, at + 16, 1, 3);
|
||||||
|
file[at + 32..at + 35].copy_from_slice(format!("s{i}!").as_bytes());
|
||||||
|
raw.extend(element(3, at as u64, 1));
|
||||||
|
}
|
||||||
|
raw.extend(element(3, 64, 1));
|
||||||
|
let mut r = VlResolver::new(&file, 8, 8);
|
||||||
|
assert_eq!(r.strings(&raw).unwrap(), ["s0!", "s1!", "s2!", "s0!"]);
|
||||||
|
}
|
||||||
@@ -350,3 +350,30 @@ ds.close()
|
|||||||
let press_vals = press_var.read_raw_f32().unwrap();
|
let press_vals = press_var.read_raw_f32().unwrap();
|
||||||
assert_eq!(press_vals, vec![1000.0f32, 850.0, 500.0, 200.0]);
|
assert_eq!(press_vals, vec![1000.0f32, 850.0, 500.0, 200.0]);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn netcdf4_python_string_variable_clawhdf5_reads() {
|
||||||
|
// NC_STRING variables are HDF5 variable-length strings, which
|
||||||
|
// `read_string` refused ("expected String, got VariableLength") until
|
||||||
|
// 2026-09-26.
|
||||||
|
skip_if_no_netcdf4!();
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = dir.path().join("strings.nc");
|
||||||
|
let path_str = path.display().to_string();
|
||||||
|
let script = format!(
|
||||||
|
r#"
|
||||||
|
import netCDF4 as nc
|
||||||
|
import numpy as np
|
||||||
|
ds = nc.Dataset("{path_str}", "w", format="NETCDF4")
|
||||||
|
ds.createDimension("station", 4)
|
||||||
|
v = ds.createVariable("name", str, ("station",))
|
||||||
|
v[:] = np.array(["Oslo", "", "São Paulo", "x"], dtype=object)
|
||||||
|
ds.close()
|
||||||
|
"#
|
||||||
|
);
|
||||||
|
run_python(&script);
|
||||||
|
|
||||||
|
let file = NetCDF4File::open(&path).unwrap();
|
||||||
|
let names = file.variable("name").unwrap().read_string().unwrap();
|
||||||
|
assert_eq!(names, vec!["Oslo", "", "São Paulo", "x"]);
|
||||||
|
}
|
||||||
|
|||||||
@@ -76,8 +76,13 @@ print the same bytes as h5dump 1.14.6 and as Debian's h5dump 1.14.5 (the
|
|||||||
was run in that image on 2026-09-26) — `dump_matches_h5dump` in
|
was run in that image on 2026-09-26) — `dump_matches_h5dump` in
|
||||||
`tests/h5rs_interop.rs` checks this, and `dump_shows_nul_padding_in_nested_strings`
|
`tests/h5rs_interop.rs` checks this, and `dump_shows_nul_padding_in_nested_strings`
|
||||||
that null-padded strings show their NULs (`"a\000b"`) at any depth, as
|
that null-padded strings show their NULs (`"a\000b"`) at any depth, as
|
||||||
h5dump's do. Not covered by those tests: references, opaque, bitfield,
|
h5dump's do. `dump_prints_vl_data_like_h5dump` covers variable-length
|
||||||
variable-length sequences and virtual datasets. Known differences from
|
strings (one with an embedded NUL, which prints up to the NUL; empty; null,
|
||||||
|
which prints `NULL`), variable-length sequences, a VL compound member and
|
||||||
|
a VL attribute, with 8- and 4-byte offsets. Not covered by those tests:
|
||||||
|
references, opaque, bitfield, non-ASCII UTF-8 (h5dump prints each byte
|
||||||
|
above 0x7f as a sign-extended octal escape, h5rs the character) and
|
||||||
|
virtual datasets. Known differences from
|
||||||
h5dump:
|
h5dump:
|
||||||
|
|
||||||
- Floats print at their own precision (a `float32` 0.1 prints as `0.1`),
|
- Floats print at their own precision (a `float32` 0.1 prints as `0.1`),
|
||||||
@@ -233,9 +238,11 @@ extension) and checks:
|
|||||||
(which catches corrupt compressed data and Fletcher-32 mismatches), and
|
(which catches corrupt compressed data and Fletcher-32 mismatches), and
|
||||||
follows every variable-length element (strings and sequences, also inside
|
follows every variable-length element (strings and sequences, also inside
|
||||||
compounds and arrays) of every dataset and attribute into its global heap
|
compounds and arrays) of every dataset and attribute into its global heap
|
||||||
collection: a collection that does not parse, a missing heap object, or a
|
collection: a collection that does not parse or overlaps another, a missing
|
||||||
sequence longer than its heap object is a problem at the collection's
|
heap object, or a heap object whose size is not exactly the element's
|
||||||
address. Data the
|
length times its base size (libhdf5 refuses such an element) is a problem
|
||||||
|
at the collection's address. Variable-length elements are resolved by the
|
||||||
|
library's `VlResolver`, as `clawhdf5::File` resolves them. Data the
|
||||||
tool cannot decode (a filter it does not implement, such as szip, or a
|
tool cannot decode (a filter it does not implement, such as szip, or a
|
||||||
dataset over `--max-bytes`) is a `note:`, not a problem. Every problem is
|
dataset over `--max-bytes`) is a `note:`, not a problem. Every problem is
|
||||||
printed with the address of the structure involved; the exit status is 0
|
printed with the address of the structure involved; the exit status is 0
|
||||||
@@ -252,9 +259,10 @@ none at all without `--data`), and objects reachable only by external links. It
|
|||||||
clawhdf5's parsers, so it accepts what they accept: some header damage that
|
clawhdf5's parsers, so it accepts what they accept: some header damage that
|
||||||
libhdf5 refuses goes unreported. Of the 150 CVE and fuzzer files of the
|
libhdf5 refuses goes unreported. Of the 150 CVE and fuzzer files of the
|
||||||
HDF Group's `cve_hdf5` corpus (`cvefiles/` and `fuzzerfiles/`),
|
HDF Group's `cve_hdf5` corpus (`cvefiles/` and `fuzzerfiles/`),
|
||||||
`check --data` passes 16, and h5dump 1.14.6 rejects 9 of those (tank,
|
`check --data` passes 15, and h5dump 1.14.6 rejects 8 of those (tank,
|
||||||
2026-09-26, `h5rs check --data F` and `h5dump F` per file; before the
|
2026-09-26, `h5rs check --data F` and `h5dump F` per file; before the
|
||||||
library's header checks it passed 28, of which h5dump rejects 21).
|
library's header checks it passed 28, of which h5dump rejects 21, and 16
|
||||||
|
and 9 before a VL type's stored element size was checked).
|
||||||
|
|
||||||
## Robustness
|
## Robustness
|
||||||
|
|
||||||
|
|||||||
@@ -15,11 +15,13 @@ use clawhdf5_format::btree_v2::{BTreeV2Header, collect_btree_v2_records};
|
|||||||
use clawhdf5_format::data_layout::DataLayout;
|
use clawhdf5_format::data_layout::DataLayout;
|
||||||
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
|
use clawhdf5_format::dataspace::{Dataspace, DataspaceType};
|
||||||
use clawhdf5_format::datatype::Datatype;
|
use clawhdf5_format::datatype::Datatype;
|
||||||
|
use clawhdf5_format::error::FormatError;
|
||||||
use clawhdf5_format::group_info::GroupInfoMessage;
|
use clawhdf5_format::group_info::GroupInfoMessage;
|
||||||
use clawhdf5_format::link_info::LinkInfoMessage;
|
use clawhdf5_format::link_info::LinkInfoMessage;
|
||||||
use clawhdf5_format::message_type::MessageType;
|
use clawhdf5_format::message_type::MessageType;
|
||||||
use clawhdf5_format::object_header::ObjectHeader;
|
use clawhdf5_format::object_header::ObjectHeader;
|
||||||
use clawhdf5_format::symbol_table::SymbolTableMessage;
|
use clawhdf5_format::symbol_table::SymbolTableMessage;
|
||||||
|
use clawhdf5_format::vl_data::{VlResolver, check_element_size, parse_vl_references};
|
||||||
|
|
||||||
use crate::cli::{Args, Out};
|
use crate::cli::{Args, Out};
|
||||||
use crate::h5::{Error, ErrorKind, H5, Kind};
|
use crate::h5::{Error, ErrorKind, H5, Kind};
|
||||||
@@ -49,6 +51,15 @@ found, 3 internal error.";
|
|||||||
|
|
||||||
const MAX_CHUNKS_CHECKED: usize = 10_000_000;
|
const MAX_CHUNKS_CHECKED: usize = 10_000_000;
|
||||||
|
|
||||||
|
/// A variable-length element's problem, worded as `check` reports heap
|
||||||
|
/// problems ("global heap ...").
|
||||||
|
fn heap_problem(e: FormatError) -> String {
|
||||||
|
match e {
|
||||||
|
FormatError::VlDataError(m) if m.starts_with("global heap") => m,
|
||||||
|
e => format!("global heap: {e}"),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// Whether values of `dt` hold variable-length data (in the global heap).
|
/// Whether values of `dt` hold variable-length data (in the global heap).
|
||||||
fn has_vl(dt: &Datatype, depth: u32) -> bool {
|
fn has_vl(dt: &Datatype, depth: u32) -> bool {
|
||||||
if depth > 32 {
|
if depth > 32 {
|
||||||
@@ -104,6 +115,8 @@ struct Checker<'a> {
|
|||||||
btrees_seen: HashSet<u64>,
|
btrees_seen: HashSet<u64>,
|
||||||
/// Global heap collections already read (with --data).
|
/// Global heap collections already read (with --data).
|
||||||
gcols_seen: HashSet<u64>,
|
gcols_seen: HashSet<u64>,
|
||||||
|
/// Resolves variable-length elements (with --data), for the whole file.
|
||||||
|
vl: VlResolver<'a>,
|
||||||
panicked: bool,
|
panicked: bool,
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -157,6 +170,7 @@ pub fn run(args: &mut Args, out: &mut Out) -> std::io::Result<i32> {
|
|||||||
heaps_seen: HashSet::new(),
|
heaps_seen: HashSet::new(),
|
||||||
btrees_seen: HashSet::new(),
|
btrees_seen: HashSet::new(),
|
||||||
gcols_seen: HashSet::new(),
|
gcols_seen: HashSet::new(),
|
||||||
|
vl: VlResolver::new(h5.data(), h5.os(), h5.ls()),
|
||||||
panicked: false,
|
panicked: false,
|
||||||
};
|
};
|
||||||
c.superblock();
|
c.superblock();
|
||||||
@@ -707,62 +721,48 @@ impl Checker<'_> {
|
|||||||
}
|
}
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
is_string,
|
is_string,
|
||||||
base_type,
|
base_type,
|
||||||
..
|
..
|
||||||
} => {
|
} => {
|
||||||
let os = usize::from(self.h5.os());
|
// Resolved by the library's VlResolver, as every other
|
||||||
let (Some(lenb), Some(addrb), Some(idxb)) =
|
// reader resolves them (and as libhdf5 does): a heap object
|
||||||
(b.get(..4), b.get(4..4 + os), b.get(4 + os..8 + os))
|
// whose size is not the element's length × base size, a
|
||||||
else {
|
// collection that overlaps another, or a missing object is
|
||||||
|
// a problem at the collection's address.
|
||||||
|
let Ok(vl) = parse_vl_references(b, 1, self.h5.os()) else {
|
||||||
return;
|
return;
|
||||||
};
|
};
|
||||||
let le = |x: &[u8]| {
|
let gcol = vl[0].collection_address;
|
||||||
x.iter()
|
if gcol == 0 || bad.contains_key(&gcol) {
|
||||||
.enumerate()
|
|
||||||
.fold(0u64, |a, (i, &v)| a | (u64::from(v) << (8 * i)))
|
|
||||||
};
|
|
||||||
let (len, gcol, idx) = (le(lenb), le(addrb), le(idxb));
|
|
||||||
let undef = if os >= 8 {
|
|
||||||
u64::MAX
|
|
||||||
} else {
|
|
||||||
(1u64 << (8 * os)) - 1
|
|
||||||
};
|
|
||||||
if len == 0 || gcol == 0 || gcol == undef || bad.contains_key(&gcol) {
|
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
let obj = match self.h5.heap_object(gcol, idx as u32) {
|
if let Err(e) = check_element_size(*size, self.h5.os()) {
|
||||||
Ok(o) => o,
|
bad.insert(gcol, e.to_string());
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let bs = if *is_string {
|
||||||
|
1
|
||||||
|
} else {
|
||||||
|
base_type.type_size() as usize
|
||||||
|
};
|
||||||
|
if bs == 0 {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let obj = match self.vl.element(b, bs) {
|
||||||
|
Ok(o) => o.unwrap_or(&[]),
|
||||||
Err(e) => {
|
Err(e) => {
|
||||||
bad.insert(e.addr.unwrap_or(gcol), e.msg);
|
bad.insert(gcol, heap_problem(e));
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
};
|
};
|
||||||
if self.gcols_seen.insert(gcol) {
|
if self.gcols_seen.insert(gcol) {
|
||||||
self.counts.global_heaps += 1;
|
self.counts.global_heaps += 1;
|
||||||
}
|
}
|
||||||
let bs = if *is_string {
|
if !*is_string && has_vl(base_type, depth + 1) {
|
||||||
1
|
for eb in obj.chunks_exact(bs) {
|
||||||
} else {
|
self.vl_element(base_type, eb, depth + 1, bad);
|
||||||
u64::from(base_type.type_size())
|
|
||||||
};
|
|
||||||
if len
|
|
||||||
.checked_mul(bs)
|
|
||||||
.is_none_or(|need| need > obj.len() as u64)
|
|
||||||
{
|
|
||||||
bad.insert(
|
|
||||||
gcol,
|
|
||||||
format!(
|
|
||||||
"global heap object {idx} holds {} bytes; the element needs {len} x {bs}",
|
|
||||||
obj.len()
|
|
||||||
),
|
|
||||||
);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
if !*is_string && bs > 0 && has_vl(base_type, depth + 1) {
|
|
||||||
let bs = bs as usize;
|
|
||||||
for k in 0..len as usize {
|
|
||||||
self.vl_element(base_type, &obj[k * bs..(k + 1) * bs], depth + 1, bad);
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -699,6 +699,9 @@ impl Diff {
|
|||||||
}
|
}
|
||||||
match (x, y) {
|
match (x, y) {
|
||||||
(Value::Str(p), Value::Str(q)) => p == q,
|
(Value::Str(p), Value::Str(q)) => p == q,
|
||||||
|
// h5diff compares a null VL string equal to an empty one.
|
||||||
|
(Value::NullStr, Value::NullStr) => true,
|
||||||
|
(Value::NullStr, Value::Str(s)) | (Value::Str(s), Value::NullStr) => s.is_empty(),
|
||||||
(Value::Bytes(p), Value::Bytes(q)) | (Value::OtherRef(p), Value::OtherRef(q)) => p == q,
|
(Value::Bytes(p), Value::Bytes(q)) | (Value::OtherRef(p), Value::OtherRef(q)) => p == q,
|
||||||
(Value::Compound(p), Value::Compound(q)) => {
|
(Value::Compound(p), Value::Compound(q)) => {
|
||||||
p.len() == q.len()
|
p.len() == q.len()
|
||||||
|
|||||||
@@ -8,7 +8,6 @@
|
|||||||
use std::cell::RefCell;
|
use std::cell::RefCell;
|
||||||
use std::collections::HashMap;
|
use std::collections::HashMap;
|
||||||
use std::path::{Path, PathBuf};
|
use std::path::{Path, PathBuf};
|
||||||
use std::rc::Rc;
|
|
||||||
|
|
||||||
use clawhdf5::File;
|
use clawhdf5::File;
|
||||||
use clawhdf5_format::attribute::{AttributeMessage, extract_attributes_tolerant};
|
use clawhdf5_format::attribute::{AttributeMessage, extract_attributes_tolerant};
|
||||||
@@ -20,7 +19,6 @@ use clawhdf5_format::datatype::Datatype;
|
|||||||
use clawhdf5_format::error::FormatError;
|
use clawhdf5_format::error::FormatError;
|
||||||
use clawhdf5_format::filter_pipeline::FilterPipeline;
|
use clawhdf5_format::filter_pipeline::FilterPipeline;
|
||||||
use clawhdf5_format::fractal_heap::FractalHeapHeader;
|
use clawhdf5_format::fractal_heap::FractalHeapHeader;
|
||||||
use clawhdf5_format::global_heap::GlobalHeapCollection;
|
|
||||||
use clawhdf5_format::group_v1;
|
use clawhdf5_format::group_v1;
|
||||||
use clawhdf5_format::link_info::LinkInfoMessage;
|
use clawhdf5_format::link_info::LinkInfoMessage;
|
||||||
use clawhdf5_format::link_message::{LinkMessage, LinkTarget};
|
use clawhdf5_format::link_message::{LinkMessage, LinkTarget};
|
||||||
@@ -191,7 +189,6 @@ pub struct H5 {
|
|||||||
pub path: PathBuf,
|
pub path: PathBuf,
|
||||||
pub file: File,
|
pub file: File,
|
||||||
pub max_bytes: u64,
|
pub max_bytes: u64,
|
||||||
heaps: RefCell<HashMap<u64, std::result::Result<Rc<GlobalHeapCollection>, String>>>,
|
|
||||||
/// Fractal heaps whose blocks were verified: `None` = sound.
|
/// Fractal heaps whose blocks were verified: `None` = sound.
|
||||||
verified_heaps: RefCell<HashMap<u64, Option<Error>>>,
|
verified_heaps: RefCell<HashMap<u64, Option<Error>>>,
|
||||||
}
|
}
|
||||||
@@ -211,7 +208,6 @@ impl H5 {
|
|||||||
path: path.to_path_buf(),
|
path: path.to_path_buf(),
|
||||||
file,
|
file,
|
||||||
max_bytes: DEFAULT_MAX_BYTES,
|
max_bytes: DEFAULT_MAX_BYTES,
|
||||||
heaps: RefCell::new(HashMap::new()),
|
|
||||||
verified_heaps: RefCell::new(HashMap::new()),
|
verified_heaps: RefCell::new(HashMap::new()),
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
@@ -433,29 +429,6 @@ impl H5 {
|
|||||||
r.map_or(Ok(()), Err)
|
r.map_or(Ok(()), Err)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The global heap object `idx` of the collection at `addr` (cached per
|
|
||||||
/// collection).
|
|
||||||
pub fn heap_object(&self, addr: u64, idx: u32) -> Result<Vec<u8>> {
|
|
||||||
let coll = {
|
|
||||||
let mut cache = self.heaps.borrow_mut();
|
|
||||||
cache
|
|
||||||
.entry(addr)
|
|
||||||
.or_insert_with(|| match usize::try_from(addr) {
|
|
||||||
Ok(a) => GlobalHeapCollection::parse(self.data(), a, self.ls())
|
|
||||||
.map(Rc::new)
|
|
||||||
.map_err(|e| e.to_string()),
|
|
||||||
Err(_) => Err("address out of range".into()),
|
|
||||||
})
|
|
||||||
.clone()
|
|
||||||
.map_err(|e| Error::at(addr, format!("global heap: {e}")))?
|
|
||||||
};
|
|
||||||
let idx16 = u16::try_from(idx)
|
|
||||||
.map_err(|_| Error::at(addr, format!("global heap object index {idx} out of range")))?;
|
|
||||||
coll.get_object(idx16)
|
|
||||||
.map(|o| o.data.clone())
|
|
||||||
.ok_or_else(|| Error::at(addr, format!("global heap has no object {idx}")))
|
|
||||||
}
|
|
||||||
|
|
||||||
/// The dataspace of the dataset at `path` with a virtual dataset's
|
/// The dataspace of the dataset at `path` with a virtual dataset's
|
||||||
/// extent resolved from its sources (as libhdf5 reports it) instead of
|
/// extent resolved from its sources (as libhdf5 reports it) instead of
|
||||||
/// the stored one.
|
/// the stored one.
|
||||||
|
|||||||
@@ -3,7 +3,10 @@
|
|||||||
//! Decoding never panics: a short buffer, an unknown byte order or a
|
//! Decoding never panics: a short buffer, an unknown byte order or a
|
||||||
//! dangling heap reference becomes [`Value::Error`].
|
//! dangling heap reference becomes [`Value::Error`].
|
||||||
|
|
||||||
|
use std::cell::RefCell;
|
||||||
|
|
||||||
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder, ReferenceType, StringPadding};
|
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder, ReferenceType, StringPadding};
|
||||||
|
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
||||||
use serde_json::Value as J;
|
use serde_json::Value as J;
|
||||||
|
|
||||||
use crate::dtype;
|
use crate::dtype;
|
||||||
@@ -16,6 +19,9 @@ pub enum Value {
|
|||||||
/// its own precision.
|
/// its own precision.
|
||||||
Float(f64, u8),
|
Float(f64, u8),
|
||||||
Str(String),
|
Str(String),
|
||||||
|
/// A null variable-length string (heap address 0): h5dump prints it as
|
||||||
|
/// `NULL`, h5py reads it as empty.
|
||||||
|
NullStr,
|
||||||
/// Opaque, bitfield, time and oversized integers.
|
/// Opaque, bitfield, time and oversized integers.
|
||||||
Bytes(Vec<u8>),
|
Bytes(Vec<u8>),
|
||||||
/// An enum member (name, when the value matches one) and its value.
|
/// An enum member (name, when the value matches one) and its value.
|
||||||
@@ -129,10 +135,10 @@ fn decode_float(dt: &Datatype, b: &[u8]) -> Value {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn trim_string(b: &[u8], pad: Option<&StringPadding>) -> String {
|
fn trim_string(b: &[u8], pad: &StringPadding) -> String {
|
||||||
let cut = b.iter().position(|&c| c == 0).unwrap_or(b.len());
|
let cut = b.iter().position(|&c| c == 0).unwrap_or(b.len());
|
||||||
let mut s = &b[..cut];
|
let mut s = &b[..cut];
|
||||||
if matches!(pad, Some(StringPadding::SpacePad)) {
|
if matches!(pad, StringPadding::SpacePad) {
|
||||||
while let [rest @ .., b' '] = s {
|
while let [rest @ .., b' '] = s {
|
||||||
s = rest;
|
s = rest;
|
||||||
}
|
}
|
||||||
@@ -140,22 +146,20 @@ fn trim_string(b: &[u8], pad: Option<&StringPadding>) -> String {
|
|||||||
String::from_utf8_lossy(s).into_owned()
|
String::from_utf8_lossy(s).into_owned()
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Little-endian unsigned integer of `b` (up to 8 bytes).
|
|
||||||
fn le(b: &[u8]) -> u64 {
|
|
||||||
b.iter()
|
|
||||||
.take(8)
|
|
||||||
.enumerate()
|
|
||||||
.fold(0u64, |a, (i, &x)| a | (u64::from(x) << (8 * i)))
|
|
||||||
}
|
|
||||||
|
|
||||||
/// Decodes elements of one file.
|
/// Decodes elements of one file.
|
||||||
pub struct Decoder<'a> {
|
pub struct Decoder<'a> {
|
||||||
pub h5: &'a H5,
|
pub h5: &'a H5,
|
||||||
|
/// Variable-length elements are resolved as the library resolves them
|
||||||
|
/// (so as libhdf5 does), not by a decoder of our own.
|
||||||
|
vl: RefCell<VlResolver<'a>>,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl<'a> Decoder<'a> {
|
impl<'a> Decoder<'a> {
|
||||||
pub fn new(h5: &'a H5) -> Self {
|
pub fn new(h5: &'a H5) -> Self {
|
||||||
Self { h5 }
|
Self {
|
||||||
|
h5,
|
||||||
|
vl: RefCell::new(VlResolver::new(h5.data(), h5.os(), h5.ls())),
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Decode element `i` of `raw`, an array of `dt` elements.
|
/// Decode element `i` of `raw`, an array of `dt` elements.
|
||||||
@@ -187,7 +191,7 @@ impl<'a> Decoder<'a> {
|
|||||||
Datatype::Time { .. } | Datatype::BitField { .. } | Datatype::Opaque { .. } => {
|
Datatype::Time { .. } | Datatype::BitField { .. } | Datatype::Opaque { .. } => {
|
||||||
Value::Bytes(b.to_vec())
|
Value::Bytes(b.to_vec())
|
||||||
}
|
}
|
||||||
Datatype::String { padding, .. } => Value::Str(trim_string(b, Some(padding))),
|
Datatype::String { padding, .. } => Value::Str(trim_string(b, padding)),
|
||||||
Datatype::Compound { members, .. } => {
|
Datatype::Compound { members, .. } => {
|
||||||
let mut out = Vec::with_capacity(members.len());
|
let mut out = Vec::with_capacity(members.len());
|
||||||
for m in members {
|
for m in members {
|
||||||
@@ -246,61 +250,43 @@ impl<'a> Decoder<'a> {
|
|||||||
Value::Array(out)
|
Value::Array(out)
|
||||||
}
|
}
|
||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
is_string,
|
is_string,
|
||||||
padding,
|
|
||||||
base_type,
|
base_type,
|
||||||
..
|
..
|
||||||
} => self.decode_vlen(*is_string, padding.as_ref(), base_type, b, depth),
|
} => match check_element_size(*size, self.h5.os()) {
|
||||||
|
Ok(()) => self.decode_vlen(*is_string, base_type, b, depth),
|
||||||
|
Err(e) => Value::Error(e.to_string()),
|
||||||
|
},
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn decode_vlen(
|
/// A variable-length element, resolved by the library's
|
||||||
&self,
|
/// [`VlResolver`]: a string ends at its first NUL, a heap object whose
|
||||||
is_string: bool,
|
/// size is not the element's length × base size is an error, and a
|
||||||
padding: Option<&StringPadding>,
|
/// heap address of 0 is null — all as libhdf5 (and so h5dump and h5py)
|
||||||
base: &Datatype,
|
/// has it.
|
||||||
b: &[u8],
|
fn decode_vlen(&self, is_string: bool, base: &Datatype, b: &[u8], depth: u32) -> Value {
|
||||||
depth: u32,
|
|
||||||
) -> Value {
|
|
||||||
let os = usize::from(self.h5.os());
|
|
||||||
let (Some(lenb), Some(addrb), Some(idxb)) =
|
|
||||||
(b.get(..4), b.get(4..4 + os), b.get(4 + os..8 + os))
|
|
||||||
else {
|
|
||||||
return Value::Error("short VL element".into());
|
|
||||||
};
|
|
||||||
let len = le(lenb) as usize;
|
|
||||||
let addr = le(addrb);
|
|
||||||
let idx = le(idxb) as u32;
|
|
||||||
let undef = if os >= 8 {
|
|
||||||
u64::MAX
|
|
||||||
} else {
|
|
||||||
(1u64 << (8 * os)) - 1
|
|
||||||
};
|
|
||||||
let obj = if len == 0 || addr == 0 || addr == undef {
|
|
||||||
Vec::new()
|
|
||||||
} else {
|
|
||||||
match self.h5.heap_object(addr, idx) {
|
|
||||||
Ok(o) => o,
|
|
||||||
Err(e) => return Value::Error(e.to_string()),
|
|
||||||
}
|
|
||||||
};
|
|
||||||
if is_string {
|
if is_string {
|
||||||
let l = len.min(obj.len());
|
return match self.vl.borrow_mut().string_element(b) {
|
||||||
return Value::Str(trim_string(&obj[..l], padding));
|
Ok(Some(s)) => Value::Str(String::from_utf8_lossy(s).into_owned()),
|
||||||
|
Ok(None) => Value::NullStr,
|
||||||
|
Err(e) => Value::Error(e.to_string()),
|
||||||
|
};
|
||||||
}
|
}
|
||||||
let bs = base.type_size() as usize;
|
let bs = base.type_size() as usize;
|
||||||
if bs == 0 {
|
if bs == 0 {
|
||||||
return Value::Error("VL base type of size 0".into());
|
return Value::Error("VL base type of size 0".into());
|
||||||
}
|
}
|
||||||
match len.checked_mul(bs) {
|
let obj = match self.vl.borrow_mut().element(b, bs) {
|
||||||
Some(need) if need <= obj.len() => {}
|
Ok(o) => o.unwrap_or(&[]),
|
||||||
_ => return Value::Error("VL sequence longer than its heap object".into()),
|
Err(e) => return Value::Error(e.to_string()),
|
||||||
}
|
};
|
||||||
let mut out = Vec::with_capacity(len);
|
Value::Seq(
|
||||||
for k in 0..len {
|
obj.chunks_exact(bs)
|
||||||
out.push(self.decode(base, &obj[k * bs..], depth + 1));
|
.map(|e| self.decode(base, e, depth + 1))
|
||||||
}
|
.collect(),
|
||||||
Value::Seq(out)
|
)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -351,6 +337,7 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
|
|||||||
Value::Int(i) => i.to_string(),
|
Value::Int(i) => i.to_string(),
|
||||||
Value::Float(f, w) => fmt_float(*f, *w),
|
Value::Float(f, w) => fmt_float(*f, *w),
|
||||||
Value::Str(s) => format!("\"{}\"", escape(s)),
|
Value::Str(s) => format!("\"{}\"", escape(s)),
|
||||||
|
Value::NullStr => "NULL".into(),
|
||||||
Value::Bytes(b) => hex(b),
|
Value::Bytes(b) => hex(b),
|
||||||
Value::Enum(Some(n), _) => n.clone(),
|
Value::Enum(Some(n), _) => n.clone(),
|
||||||
Value::Enum(None, i) => i.to_string(),
|
Value::Enum(None, i) => i.to_string(),
|
||||||
@@ -411,6 +398,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
Value::Str(s) => J::from(s.as_str()),
|
Value::Str(s) => J::from(s.as_str()),
|
||||||
|
Value::NullStr => J::from(""),
|
||||||
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
|
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
|
||||||
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
|
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
|
||||||
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
|
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
|
||||||
|
|||||||
@@ -0,0 +1,128 @@
|
|||||||
|
"""Write the variable-length data files the h5rs VL tests run on.
|
||||||
|
|
||||||
|
usage: gen_vl_files.py OUTDIR
|
||||||
|
|
||||||
|
For 8-byte (`vl8`) and 4-byte (`vl4`) offsets, writes OUTDIR/vl8.h5 and
|
||||||
|
OUTDIR/vl4.h5, which libhdf5 reads in full, and OUTDIR/bad8.h5 and
|
||||||
|
OUTDIR/bad4.h5, whose `bad` and `badseq` elements 0 have a length that
|
||||||
|
disagrees with their global heap object (libhdf5: "Expected global heap
|
||||||
|
object size does not match"), and whose `undef` element 1 has length 0 and
|
||||||
|
the undefined heap address (libhdf5: "addr undefined"). h5py cannot write a VL string with a NUL in
|
||||||
|
it or a null element in a contiguous dataset, so those are patched in.
|
||||||
|
|
||||||
|
Prints one JSON object: for each file, each dataset's values as h5py reads
|
||||||
|
them one element at a time (strings as text, sequences as lists, a compound
|
||||||
|
as a list of its fields), with null for an element h5py cannot read; the
|
||||||
|
root attribute `va`; and the addresses of the `bad` elements' collections.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import struct
|
||||||
|
import sys
|
||||||
|
|
||||||
|
import h5py
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
out = sys.argv[1]
|
||||||
|
S = h5py.string_dtype("utf-8")
|
||||||
|
I4 = h5py.vlen_dtype(np.dtype("<i4"))
|
||||||
|
|
||||||
|
|
||||||
|
def create(path, sizes):
|
||||||
|
if sizes is None:
|
||||||
|
return h5py.File(path, "w")
|
||||||
|
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE)
|
||||||
|
fcpl.set_sizes(*sizes)
|
||||||
|
return h5py.File(h5py.h5f.create(path.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl))
|
||||||
|
|
||||||
|
|
||||||
|
def element(length, addr, index, os_):
|
||||||
|
return struct.pack("<I", length) + addr.to_bytes(os_, "little") + struct.pack("<I", index)
|
||||||
|
|
||||||
|
|
||||||
|
def good(path, sizes):
|
||||||
|
os_ = 8 if sizes is None else sizes[0]
|
||||||
|
with create(path, sizes) as f:
|
||||||
|
f.create_dataset(
|
||||||
|
"d", data=np.array(["aXb", "", "ok", "zz", "hello"], dtype=object), dtype=S
|
||||||
|
)
|
||||||
|
u = f.create_dataset("u", shape=(4,), dtype=S, chunks=(1,))
|
||||||
|
u[1] = "w"
|
||||||
|
s = f.create_dataset("seq", shape=(3,), dtype=I4)
|
||||||
|
s[0] = [1, 2, 3]
|
||||||
|
s[1] = []
|
||||||
|
s[2] = [-5]
|
||||||
|
s = f.create_dataset("sequ", shape=(3,), dtype=I4, chunks=(1,))
|
||||||
|
s[0] = [7, 8]
|
||||||
|
ct = np.dtype([("id", "<i4"), ("name", S)])
|
||||||
|
arr = np.zeros(3, dtype=ct)
|
||||||
|
arr["id"] = [1, 2, 3]
|
||||||
|
arr["name"] = ["one", "", "three"]
|
||||||
|
f.create_dataset("cmp", data=arr)
|
||||||
|
f.attrs.create("va", np.array(["p", "", "q"], dtype=object), dtype=S)
|
||||||
|
off = f["d"].id.get_offset()
|
||||||
|
b = bytearray(open(path, "rb").read())
|
||||||
|
i = b.index(b"aXb")
|
||||||
|
b[i + 1] = 0 # "a\0b"
|
||||||
|
es = 8 + os_
|
||||||
|
b[off + 2 * es : off + 3 * es] = element(2, 0, 1, os_) # "ok" -> null
|
||||||
|
open(path, "wb").write(bytes(b))
|
||||||
|
|
||||||
|
|
||||||
|
def bad(path, sizes):
|
||||||
|
os_ = 8 if sizes is None else sizes[0]
|
||||||
|
with create(path, sizes) as f:
|
||||||
|
f.create_dataset("bad", data=np.array(["cdefgh", "ok"], dtype=object), dtype=S)
|
||||||
|
s = f.create_dataset("badseq", shape=(2,), dtype=I4)
|
||||||
|
s[0] = [1, 2, 3]
|
||||||
|
s[1] = [4]
|
||||||
|
f.create_dataset("undef", data=np.array(["x", "", "yz"], dtype=object), dtype=S)
|
||||||
|
off, soff = f["bad"].id.get_offset(), f["badseq"].id.get_offset()
|
||||||
|
uoff = f["undef"].id.get_offset()
|
||||||
|
b = bytearray(open(path, "rb").read())
|
||||||
|
gcol = int.from_bytes(b[off + 4 : off + 4 + os_], "little")
|
||||||
|
struct.pack_into("<I", b, off, 3) # "cdefgh": length 6 -> 3
|
||||||
|
struct.pack_into("<I", b, soff, 2) # [1, 2, 3]: length 3 -> 2
|
||||||
|
# "": length 0 at the undefined address (all 0xff), which libhdf5 fails
|
||||||
|
# to read ("addr undefined"); it writes a null element as address 0.
|
||||||
|
es = 8 + os_
|
||||||
|
b[uoff + es : uoff + 2 * es] = element(0, (1 << (8 * os_)) - 1, 1, os_)
|
||||||
|
open(path, "wb").write(bytes(b))
|
||||||
|
return gcol
|
||||||
|
|
||||||
|
|
||||||
|
def value(v):
|
||||||
|
if isinstance(v, bytes):
|
||||||
|
return v.decode()
|
||||||
|
if isinstance(v, str):
|
||||||
|
return v
|
||||||
|
if isinstance(v, np.void):
|
||||||
|
return [value(x) for x in v]
|
||||||
|
if isinstance(v, np.ndarray):
|
||||||
|
return [value(x) for x in v]
|
||||||
|
return v.item() if hasattr(v, "item") else v
|
||||||
|
|
||||||
|
|
||||||
|
def read(ds):
|
||||||
|
got = []
|
||||||
|
for i in range(ds.shape[0]):
|
||||||
|
try:
|
||||||
|
got.append(value(ds[i]))
|
||||||
|
except OSError:
|
||||||
|
got.append(None)
|
||||||
|
return got
|
||||||
|
|
||||||
|
|
||||||
|
result = {}
|
||||||
|
for tag, sizes in (("8", None), ("4", (4, 4))):
|
||||||
|
g, x = os.path.join(out, f"vl{tag}.h5"), os.path.join(out, f"bad{tag}.h5")
|
||||||
|
good(g, sizes)
|
||||||
|
gcol = bad(x, sizes)
|
||||||
|
with h5py.File(g, "r") as f:
|
||||||
|
result[f"vl{tag}"] = {n: read(f[n]) for n in ("d", "u", "seq", "sequ", "cmp")}
|
||||||
|
result[f"vl{tag}"]["va"] = [value(s) for s in f.attrs["va"]]
|
||||||
|
with h5py.File(x, "r") as f:
|
||||||
|
result[f"bad{tag}"] = {n: read(f[n]) for n in ("bad", "badseq", "undef")}
|
||||||
|
result[f"bad{tag}"]["gcol"] = gcol
|
||||||
|
json.dump(result, sys.stdout)
|
||||||
@@ -773,3 +773,138 @@ fn every_subcommand_rejects_a_non_hdf5_file_cleanly() {
|
|||||||
assert_eq!(code(&h5rs(&["ls"])), 2);
|
assert_eq!(code(&h5rs(&["ls"])), 2);
|
||||||
assert_eq!(code(&h5rs(&["--help"])), 0);
|
assert_eq!(code(&h5rs(&["--help"])), 0);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ---------------------------------------------------------------------------
|
||||||
|
// variable-length data
|
||||||
|
// ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
/// Runs `tests/gen_vl_files.py`: VL strings (with an embedded NUL, empty
|
||||||
|
/// and null elements), VL sequences, a VL compound member and a VL
|
||||||
|
/// attribute, with 8- and 4-byte offsets, plus files whose heap objects
|
||||||
|
/// disagree with their elements' lengths.
|
||||||
|
fn generate_vl() -> Option<Files> {
|
||||||
|
if missing(python_available(), "python3 with h5py") {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let script = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/gen_vl_files.py");
|
||||||
|
let out = Command::new(python())
|
||||||
|
.arg(&script)
|
||||||
|
.arg(dir.path())
|
||||||
|
.output()
|
||||||
|
.expect("run gen_vl_files.py");
|
||||||
|
assert!(
|
||||||
|
out.status.success(),
|
||||||
|
"gen_vl_files.py failed:\n{}",
|
||||||
|
String::from_utf8_lossy(&out.stderr)
|
||||||
|
);
|
||||||
|
let values = serde_json::from_slice(&out.stdout).expect("gen_vl_files.py output");
|
||||||
|
Some(Files { dir, values })
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `dump` resolves VL elements through the library's `VlResolver`, as
|
||||||
|
/// libhdf5 does: "a\0b" prints as "a", a null string as NULL (it printed
|
||||||
|
/// ""), and with 4-byte offsets too; the output is h5dump's byte for byte.
|
||||||
|
#[test]
|
||||||
|
fn dump_prints_vl_data_like_h5dump() {
|
||||||
|
let Some(f) = generate_vl() else { return };
|
||||||
|
for name in ["vl8.h5", "vl4.h5"] {
|
||||||
|
let p = f.p(name);
|
||||||
|
let ours = stdout(&h5rs(&["dump", &p]));
|
||||||
|
assert!(
|
||||||
|
ours.contains(r#"(0): "a", "", NULL, "zz", "hello""#),
|
||||||
|
"{name}:\n{ours}"
|
||||||
|
);
|
||||||
|
assert!(ours.contains(r#"(0): NULL, "w", NULL, NULL"#), "{name}");
|
||||||
|
assert!(ours.contains("(0): (1, 2, 3), (), (-5)"), "{name}");
|
||||||
|
if missing(tool_available("h5dump"), "h5dump") {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
let reference = run("h5dump", &[&p]);
|
||||||
|
assert!(reference.status.success(), "{name}: {reference:?}");
|
||||||
|
assert_eq!(ours, stdout(&reference).replacen(&p, name, 1), "{name}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `dump --json` gives the values h5py reads, element by element; and an
|
||||||
|
/// element whose heap object is not its length × base size is an error, as
|
||||||
|
/// in h5py, not a truncated value (it printed "cde" and (1, 2)); so is a
|
||||||
|
/// length-0 element at the undefined heap address (it printed "").
|
||||||
|
#[test]
|
||||||
|
fn dump_json_vl_values_match_h5py() {
|
||||||
|
let Some(f) = generate_vl() else { return };
|
||||||
|
for tag in ["8", "4"] {
|
||||||
|
let (good, bad) = (format!("vl{tag}"), format!("bad{tag}"));
|
||||||
|
let o = h5rs(&["dump", "--json", &f.p(&format!("{good}.h5"))]);
|
||||||
|
assert!(o.status.success(), "{good}: {o:?}");
|
||||||
|
let doc: serde_json::Value = serde_json::from_slice(&o.stdout).unwrap();
|
||||||
|
let want = &f.values[&good];
|
||||||
|
for d in doc["datasets"].as_object().unwrap().values() {
|
||||||
|
let path = d["alias"][0].as_str().unwrap();
|
||||||
|
assert_eq!(d["value"], want[&path[1..]], "{good}: {path}");
|
||||||
|
}
|
||||||
|
let attrs = &doc["groups"][doc["root"].as_str().unwrap()]["attributes"];
|
||||||
|
assert_eq!(attrs[0]["name"], "va");
|
||||||
|
assert_eq!(attrs[0]["value"], want["va"], "{good}: va");
|
||||||
|
|
||||||
|
let o = h5rs(&["dump", "--json", &f.p(&format!("{bad}.h5"))]);
|
||||||
|
let doc: serde_json::Value = serde_json::from_slice(&o.stdout).unwrap();
|
||||||
|
let want = &f.values[&bad];
|
||||||
|
for d in doc["datasets"].as_object().unwrap().values() {
|
||||||
|
let path = d["alias"][0].as_str().unwrap();
|
||||||
|
let got = d["value"].as_array().unwrap();
|
||||||
|
let want = want[&path[1..]].as_array().unwrap();
|
||||||
|
assert_eq!(got.len(), want.len(), "{bad}: {path}");
|
||||||
|
for (g, w) in got.iter().zip(want) {
|
||||||
|
if w.is_null() {
|
||||||
|
// h5py cannot read it: neither can we.
|
||||||
|
let e = g["error"]
|
||||||
|
.as_str()
|
||||||
|
.unwrap_or_else(|| panic!("{bad}: {path}: {g}"));
|
||||||
|
let why = if path == "/undef" {
|
||||||
|
"undefined"
|
||||||
|
} else {
|
||||||
|
"holds"
|
||||||
|
};
|
||||||
|
assert!(e.contains(why), "{bad}: {path}: {e}");
|
||||||
|
} else {
|
||||||
|
assert_eq!(g, w, "{bad}: {path}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `check --data` holds VL elements to libhdf5's rule: a heap object whose
|
||||||
|
/// size is not exactly the element's length × base size is a problem (it
|
||||||
|
/// only caught objects shorter than the element), and so is an element at
|
||||||
|
/// the undefined heap address.
|
||||||
|
#[test]
|
||||||
|
fn check_data_flags_mis_sized_vl_heap_objects() {
|
||||||
|
let Some(f) = generate_vl() else { return };
|
||||||
|
for tag in ["8", "4"] {
|
||||||
|
let o = h5rs(&["check", "--data", &f.p(&format!("vl{tag}.h5"))]);
|
||||||
|
let s = stdout(&o);
|
||||||
|
assert_eq!(code(&o), 0, "vl{tag}: {s}");
|
||||||
|
assert!(
|
||||||
|
s.contains("global heap collections read: 1"),
|
||||||
|
"vl{tag}: {s}"
|
||||||
|
);
|
||||||
|
|
||||||
|
let o = h5rs(&["check", "--data", &f.p(&format!("bad{tag}.h5"))]);
|
||||||
|
let s = stdout(&o);
|
||||||
|
assert_eq!(code(&o), 1, "bad{tag}: {s}");
|
||||||
|
let at = f.values[format!("bad{tag}")]["gcol"].as_u64().unwrap();
|
||||||
|
for (path, what) in [("/bad", "6 bytes"), ("/badseq", "12 bytes")] {
|
||||||
|
let want = format!("problem: {at:#x} {path}: variable-length data: global heap object");
|
||||||
|
assert!(s.contains(&want), "bad{tag}: no {want:?} in\n{s}");
|
||||||
|
assert!(s.contains(what), "bad{tag}: {s}");
|
||||||
|
}
|
||||||
|
// A length-0 element at the undefined heap address: libhdf5 fails
|
||||||
|
// to read it; check skipped it.
|
||||||
|
let undef: u64 = if tag == "8" { u64::MAX } else { 0xffff_ffff };
|
||||||
|
let want = format!("problem: {undef:#x} /undef: variable-length data: global heap:");
|
||||||
|
assert!(s.contains(&want), "bad{tag}: no {want:?} in\n{s}");
|
||||||
|
assert!(s.contains("undefined global heap address"), "bad{tag}: {s}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
@@ -9,6 +9,7 @@
|
|||||||
use clawhdf5::{AttrValue, File, Selection};
|
use clawhdf5::{AttrValue, File, Selection};
|
||||||
use clawhdf5_format::data_read;
|
use clawhdf5_format::data_read;
|
||||||
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
|
||||||
|
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
||||||
|
|
||||||
/// Errors are reported to JavaScript as messages.
|
/// Errors are reported to JavaScript as messages.
|
||||||
pub type Result<T> = std::result::Result<T, String>;
|
pub type Result<T> = std::result::Result<T, String>;
|
||||||
@@ -221,6 +222,12 @@ impl Reader {
|
|||||||
None => (Selection::All, shape.clone()),
|
None => (Selection::All, shape.clone()),
|
||||||
Some(h) => hyperslab_selection(h, &shape)?,
|
Some(h) => hyperslab_selection(h, &shape)?,
|
||||||
};
|
};
|
||||||
|
// A VL type whose stored element size is not the one the file's
|
||||||
|
// offset size implies is refused before its data is read, as
|
||||||
|
// `File::read_string` refuses it.
|
||||||
|
if let Datatype::VariableLength { size, .. } = array_base(&dt) {
|
||||||
|
check_element_size(*size, self.file.superblock().offset_size).map_err(err)?;
|
||||||
|
}
|
||||||
let raw = ds.read_selection(&selection).map_err(err)?;
|
let raw = ds.read_selection(&selection).map_err(err)?;
|
||||||
let data = self.decode(&raw, &dt)?;
|
let data = self.decode(&raw, &dt)?;
|
||||||
out_shape.extend(element_shape(&dt));
|
out_shape.extend(element_shape(&dt));
|
||||||
@@ -270,23 +277,14 @@ impl Reader {
|
|||||||
Datatype::VariableLength {
|
Datatype::VariableLength {
|
||||||
is_string: true, ..
|
is_string: true, ..
|
||||||
} if !is_array => {
|
} if !is_array => {
|
||||||
let size = dt.type_size() as usize;
|
// The library's resolver, as File::read_string uses: a
|
||||||
if size == 0 || !raw.len().is_multiple_of(size) {
|
// string ends at its first NUL and a heap object of the
|
||||||
return Err(format!(
|
// wrong size is an error, as in libhdf5 and h5py.
|
||||||
"{} bytes is not a whole number of {size}-byte string references",
|
|
||||||
raw.len()
|
|
||||||
));
|
|
||||||
}
|
|
||||||
let sb = self.file.superblock();
|
let sb = self.file.superblock();
|
||||||
Data::Strings(
|
Data::Strings(
|
||||||
clawhdf5_format::vl_data::read_vl_strings(
|
VlResolver::new(self.file.as_bytes(), sb.offset_size, sb.length_size)
|
||||||
self.file.as_bytes(),
|
.strings(raw)
|
||||||
raw,
|
.map_err(err)?,
|
||||||
(raw.len() / size) as u64,
|
|
||||||
sb.offset_size,
|
|
||||||
sb.length_size,
|
|
||||||
)
|
|
||||||
.map_err(err)?,
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
Datatype::Enumeration { .. } if !is_array => {
|
Datatype::Enumeration { .. } if !is_array => {
|
||||||
|
|||||||
@@ -0,0 +1,169 @@
|
|||||||
|
//! The wasm reader resolves VL strings with the library's `VlResolver`, so
|
||||||
|
//! it returns what `File::read_string` and h5py return: a string ends at
|
||||||
|
//! its first NUL, a null element is empty, a heap object of the wrong size
|
||||||
|
//! is an error, an element at the undefined heap address is an error, and a
|
||||||
|
//! VL datatype whose stored element size disagrees with the file's offset
|
||||||
|
//! size is refused. Checked with 8- and 4-byte offsets.
|
||||||
|
//!
|
||||||
|
//! Skipped when python3 with h5py is missing, unless
|
||||||
|
//! `CLAWHDF5_REQUIRE_INTEROP=1`. `CLAWHDF5_PYTHON` names the interpreter.
|
||||||
|
|
||||||
|
use std::process::Command;
|
||||||
|
|
||||||
|
use clawhdf5::File;
|
||||||
|
use clawhdf5_wasm::core::{Data, Reader};
|
||||||
|
|
||||||
|
fn python() -> String {
|
||||||
|
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn h5py_available() -> bool {
|
||||||
|
let ok = Command::new(python())
|
||||||
|
.args(["-c", "import h5py, numpy"])
|
||||||
|
.output()
|
||||||
|
.map(|o| o.status.success())
|
||||||
|
.unwrap_or(false);
|
||||||
|
if !ok {
|
||||||
|
assert!(
|
||||||
|
std::env::var("CLAWHDF5_REQUIRE_INTEROP").as_deref() != Ok("1"),
|
||||||
|
"CLAWHDF5_REQUIRE_INTEROP=1 but python with h5py is not available"
|
||||||
|
);
|
||||||
|
eprintln!("SKIP: python with h5py not available");
|
||||||
|
}
|
||||||
|
ok
|
||||||
|
}
|
||||||
|
|
||||||
|
/// For each offset size: `vl{8,4}.h5` with dataset `d` = "a\0b", "", null,
|
||||||
|
/// "zz" (patched: h5py writes neither a NUL nor a null element);
|
||||||
|
/// `bad{8,4}.h5` whose element 0 claims 3 bytes of a 6-byte heap object;
|
||||||
|
/// `size{8,4}.h5` whose VL datatype message stores a 24-byte element; and
|
||||||
|
/// `undef{8,4}.h5` whose element 1 has length 0 and the undefined heap
|
||||||
|
/// address.
|
||||||
|
/// Prints h5py's reading of each element as hex, or "error".
|
||||||
|
const SCRIPT: &str = r#"
|
||||||
|
import struct, sys, h5py, numpy as np
|
||||||
|
out = sys.argv[1]
|
||||||
|
S = h5py.string_dtype('utf-8')
|
||||||
|
def create(path, os_):
|
||||||
|
if os_ == 8:
|
||||||
|
return h5py.File(path, 'w', libver='earliest')
|
||||||
|
# The earliest format, so the patched object header has no checksum.
|
||||||
|
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE); fcpl.set_sizes(4, 4)
|
||||||
|
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
|
||||||
|
fapl.set_libver_bounds(h5py.h5f.LIBVER_EARLIEST, h5py.h5f.LIBVER_V18)
|
||||||
|
return h5py.File(h5py.h5f.create(path.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl, fapl=fapl))
|
||||||
|
def elem(length, addr, index, os_):
|
||||||
|
return struct.pack('<I', length) + addr.to_bytes(os_, 'little') + struct.pack('<I', index)
|
||||||
|
def make(path, os_, values):
|
||||||
|
with create(path, os_) as f:
|
||||||
|
f.create_dataset('d', data=np.array(values, dtype=object), dtype=S)
|
||||||
|
return f['d'].id.get_offset()
|
||||||
|
for os_ in (8, 4):
|
||||||
|
es = 8 + os_
|
||||||
|
p = '%s/vl%d.h5' % (out, os_)
|
||||||
|
off = make(p, os_, ['aXb', '', 'ok', 'zz'])
|
||||||
|
b = bytearray(open(p, 'rb').read())
|
||||||
|
b[b.index(b'aXb') + 1] = 0
|
||||||
|
b[off + 2 * es:off + 3 * es] = elem(2, 0, 1, os_)
|
||||||
|
open(p, 'wb').write(bytes(b))
|
||||||
|
p = '%s/bad%d.h5' % (out, os_)
|
||||||
|
off = make(p, os_, ['cdefgh', 'ok'])
|
||||||
|
b = bytearray(open(p, 'rb').read())
|
||||||
|
struct.pack_into('<I', b, off, 3)
|
||||||
|
open(p, 'wb').write(bytes(b))
|
||||||
|
p = '%s/size%d.h5' % (out, os_)
|
||||||
|
make(p, os_, ['x', 'yy'])
|
||||||
|
b = bytearray(open(p, 'rb').read())
|
||||||
|
# datatype message: version 1, class 9 (VL); string, null-terminated, UTF-8
|
||||||
|
pat = bytes([0x19, 0x01, 0x01, 0x00]) + struct.pack('<I', es)
|
||||||
|
assert b.count(pat) == 1, b.count(pat)
|
||||||
|
i = b.index(pat)
|
||||||
|
struct.pack_into('<I', b, i + 4, 24)
|
||||||
|
open(p, 'wb').write(bytes(b))
|
||||||
|
p = '%s/undef%d.h5' % (out, os_)
|
||||||
|
off = make(p, os_, ['x', '', 'yz'])
|
||||||
|
b = bytearray(open(p, 'rb').read())
|
||||||
|
b[off + es:off + 2 * es] = elem(0, (1 << (8 * os_)) - 1, 1, os_)
|
||||||
|
open(p, 'wb').write(bytes(b))
|
||||||
|
for name in ('vl', 'bad', 'size', 'undef'):
|
||||||
|
with h5py.File('%s/%s%d.h5' % (out, name, os_), 'r') as f:
|
||||||
|
got = []
|
||||||
|
for i in range(f['d'].shape[0]):
|
||||||
|
try:
|
||||||
|
got.append(f['d'][i].hex())
|
||||||
|
except OSError:
|
||||||
|
got.append('error')
|
||||||
|
print('%s%d\t%s' % (name, os_, ','.join(got)))
|
||||||
|
"#;
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn vl_strings_read_like_file_and_h5py() {
|
||||||
|
if !h5py_available() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let out = Command::new(python())
|
||||||
|
.args(["-c", SCRIPT])
|
||||||
|
.arg(dir.path())
|
||||||
|
.output()
|
||||||
|
.expect("run python");
|
||||||
|
assert!(
|
||||||
|
out.status.success(),
|
||||||
|
"{}",
|
||||||
|
String::from_utf8_lossy(&out.stderr)
|
||||||
|
);
|
||||||
|
let h5py: std::collections::HashMap<String, String> = String::from_utf8(out.stdout)
|
||||||
|
.unwrap()
|
||||||
|
.lines()
|
||||||
|
.map(|l| {
|
||||||
|
let (k, v) = l.split_once('\t').unwrap();
|
||||||
|
(k.to_string(), v.to_string())
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
let read = |name: &str| {
|
||||||
|
let bytes = std::fs::read(dir.path().join(format!("{name}.h5"))).unwrap();
|
||||||
|
let wasm = Reader::open(bytes).unwrap().read("/d", None);
|
||||||
|
let file = File::open(dir.path().join(format!("{name}.h5")))
|
||||||
|
.unwrap()
|
||||||
|
.dataset("d")
|
||||||
|
.unwrap()
|
||||||
|
.read_string();
|
||||||
|
(wasm, file)
|
||||||
|
};
|
||||||
|
for os in [8, 4] {
|
||||||
|
// h5py: "a\0b" is "a"; the null element (address 0) is empty.
|
||||||
|
assert_eq!(h5py[&format!("vl{os}")], "61,,,7a7a");
|
||||||
|
let (wasm, file) = read(&format!("vl{os}"));
|
||||||
|
let Data::Strings(wasm) = wasm.unwrap().data else {
|
||||||
|
panic!("vl{os}: not strings")
|
||||||
|
};
|
||||||
|
assert_eq!(wasm, ["a", "", "", "zz"], "vl{os}");
|
||||||
|
assert_eq!(wasm, file.unwrap(), "vl{os}");
|
||||||
|
|
||||||
|
// h5py refuses the mis-sized element; so do both readers.
|
||||||
|
assert_eq!(h5py[&format!("bad{os}")], "error,6f6b");
|
||||||
|
let (wasm, file) = read(&format!("bad{os}"));
|
||||||
|
assert!(wasm.unwrap_err().contains("holds 6 bytes"), "bad{os}");
|
||||||
|
assert!(file.is_err(), "bad{os}");
|
||||||
|
|
||||||
|
// libhdf5 ignores the stored element size and reads the values;
|
||||||
|
// File refuses the datatype rather than guess its layout, and the
|
||||||
|
// wasm reader now does the same (it read with the stored size).
|
||||||
|
assert_eq!(h5py[&format!("size{os}")], "78,7979");
|
||||||
|
let (wasm, file) = read(&format!("size{os}"));
|
||||||
|
let e = wasm.unwrap_err();
|
||||||
|
assert!(e.contains("stores 24-byte elements"), "size{os}: {e}");
|
||||||
|
assert!(file.is_err(), "size{os}");
|
||||||
|
|
||||||
|
// Length 0 at the undefined heap address: libhdf5 fails the read
|
||||||
|
// ("addr undefined"); both readers returned "".
|
||||||
|
assert_eq!(h5py[&format!("undef{os}")], "78,error,797a");
|
||||||
|
let (wasm, file) = read(&format!("undef{os}"));
|
||||||
|
let e = wasm.unwrap_err();
|
||||||
|
assert!(
|
||||||
|
e.contains("undefined global heap address"),
|
||||||
|
"undef{os}: {e}"
|
||||||
|
);
|
||||||
|
assert!(file.is_err(), "undef{os}");
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -422,11 +422,47 @@ impl<'f, R: HDF5Read> LazyDataset<'f, R> {
|
|||||||
Ok(data_read::read_as_u64(&raw, &dt)?)
|
Ok(data_read::read_as_u64(&raw, &dt)?)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Read all data as `String` values.
|
/// Read all data as `String` values: fixed- or variable-length strings
|
||||||
|
/// (see [`Dataset::read_string`](crate::Dataset::read_string)).
|
||||||
pub fn read_string(&self) -> Result<Vec<String>, Error> {
|
pub fn read_string(&self) -> Result<Vec<String>, Error> {
|
||||||
let raw = self.read_raw()?;
|
let raw = self.read_raw()?;
|
||||||
let dt = self.datatype()?;
|
let dt = self.datatype()?;
|
||||||
Ok(data_read::read_as_strings(&raw, &dt)?)
|
crate::vlen::decode_strings(
|
||||||
|
self.file.hdf5_bytes(),
|
||||||
|
&dt,
|
||||||
|
&raw,
|
||||||
|
self.file.offset_size(),
|
||||||
|
self.file.length_size(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read a variable-length string dataset as the exact bytes of each
|
||||||
|
/// string (see
|
||||||
|
/// [`Dataset::read_string_bytes`](crate::Dataset::read_string_bytes)).
|
||||||
|
pub fn read_string_bytes(&self) -> Result<Vec<Vec<u8>>, Error> {
|
||||||
|
let raw = self.read_raw()?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
crate::vlen::decode_string_bytes(
|
||||||
|
self.file.hdf5_bytes(),
|
||||||
|
&dt,
|
||||||
|
&raw,
|
||||||
|
self.file.offset_size(),
|
||||||
|
self.file.length_size(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read a variable-length sequence dataset as one `Vec<T>` per element
|
||||||
|
/// (see [`Dataset::read_vlen`](crate::Dataset::read_vlen)).
|
||||||
|
pub fn read_vlen<T: crate::vlen::VlenValue>(&self) -> Result<Vec<Vec<T>>, Error> {
|
||||||
|
let raw = self.read_raw()?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
crate::vlen::decode_vlen(
|
||||||
|
self.file.hdf5_bytes(),
|
||||||
|
&dt,
|
||||||
|
&raw,
|
||||||
|
self.file.offset_size(),
|
||||||
|
self.file.length_size(),
|
||||||
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Read all attributes of this dataset.
|
/// Read all attributes of this dataset.
|
||||||
|
|||||||
@@ -30,6 +30,7 @@ pub mod lazy;
|
|||||||
pub mod mmap_file;
|
pub mod mmap_file;
|
||||||
pub mod reader;
|
pub mod reader;
|
||||||
pub mod types;
|
pub mod types;
|
||||||
|
pub mod vlen;
|
||||||
pub mod writer;
|
pub mod writer;
|
||||||
|
|
||||||
pub use error::Error;
|
pub use error::Error;
|
||||||
@@ -38,6 +39,7 @@ pub use lazy::{LazyDataset, LazyFile, LazyGroup};
|
|||||||
pub use mmap_file::{MmapDataset, MmapFile, MmapGroup};
|
pub use mmap_file::{MmapDataset, MmapFile, MmapGroup};
|
||||||
pub use reader::{Dataset, File, Group};
|
pub use reader::{Dataset, File, Group};
|
||||||
pub use types::{AttrValue, DType};
|
pub use types::{AttrValue, DType};
|
||||||
|
pub use vlen::VlenValue;
|
||||||
pub use writer::FileBuilder;
|
pub use writer::FileBuilder;
|
||||||
#[cfg(feature = "parallel")]
|
#[cfg(feature = "parallel")]
|
||||||
pub use writer::{DatasetSpec, create_datasets_parallel};
|
pub use writer::{DatasetSpec, create_datasets_parallel};
|
||||||
|
|||||||
@@ -336,11 +336,47 @@ impl<'f> MmapDataset<'f> {
|
|||||||
Ok(data_read::read_as_u64(&raw, &dt)?)
|
Ok(data_read::read_as_u64(&raw, &dt)?)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Read all data as `String` values.
|
/// Read all data as `String` values: fixed- or variable-length strings
|
||||||
|
/// (see [`Dataset::read_string`](crate::Dataset::read_string)).
|
||||||
pub fn read_string(&self) -> Result<Vec<String>, Error> {
|
pub fn read_string(&self) -> Result<Vec<String>, Error> {
|
||||||
let raw = self.read_raw()?;
|
let raw = self.read_raw()?;
|
||||||
let dt = self.datatype()?;
|
let dt = self.datatype()?;
|
||||||
Ok(data_read::read_as_strings(&raw, &dt)?)
|
crate::vlen::decode_strings(
|
||||||
|
self.file.hdf5_bytes(),
|
||||||
|
&dt,
|
||||||
|
&raw,
|
||||||
|
self.file.offset_size(),
|
||||||
|
self.file.length_size(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read a variable-length string dataset as the exact bytes of each
|
||||||
|
/// string (see
|
||||||
|
/// [`Dataset::read_string_bytes`](crate::Dataset::read_string_bytes)).
|
||||||
|
pub fn read_string_bytes(&self) -> Result<Vec<Vec<u8>>, Error> {
|
||||||
|
let raw = self.read_raw()?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
crate::vlen::decode_string_bytes(
|
||||||
|
self.file.hdf5_bytes(),
|
||||||
|
&dt,
|
||||||
|
&raw,
|
||||||
|
self.file.offset_size(),
|
||||||
|
self.file.length_size(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read a variable-length sequence dataset as one `Vec<T>` per element
|
||||||
|
/// (see [`Dataset::read_vlen`](crate::Dataset::read_vlen)).
|
||||||
|
pub fn read_vlen<T: crate::vlen::VlenValue>(&self) -> Result<Vec<Vec<T>>, Error> {
|
||||||
|
let raw = self.read_raw()?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
crate::vlen::decode_vlen(
|
||||||
|
self.file.hdf5_bytes(),
|
||||||
|
&dt,
|
||||||
|
&raw,
|
||||||
|
self.file.offset_size(),
|
||||||
|
self.file.length_size(),
|
||||||
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// For contiguous datasets, return a zero-copy slice into the mmap.
|
/// For contiguous datasets, return a zero-copy slice into the mmap.
|
||||||
|
|||||||
@@ -262,6 +262,56 @@ impl File {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// Decode the strings in `raw`, a buffer of elements of `datatype` read
|
||||||
|
/// from this file — for instance a variable-length string field of a
|
||||||
|
/// compound ([`clawhdf5_format::data_read::read_compound_fields`]) or an
|
||||||
|
/// [`AttrValue::Raw`] attribute. Variable-length strings are resolved in
|
||||||
|
/// this file's global heap; see [`Dataset::read_string`] for the values.
|
||||||
|
pub fn decode_strings(&self, datatype: &Datatype, raw: &[u8]) -> Result<Vec<String>, Error> {
|
||||||
|
crate::vlen::decode_strings(
|
||||||
|
self.as_bytes(),
|
||||||
|
datatype,
|
||||||
|
raw,
|
||||||
|
self.offset_size(),
|
||||||
|
self.length_size(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Like [`decode_strings`](Self::decode_strings) for variable-length
|
||||||
|
/// strings, returning each string's exact bytes (see
|
||||||
|
/// [`Dataset::read_string_bytes`]).
|
||||||
|
pub fn decode_string_bytes(
|
||||||
|
&self,
|
||||||
|
datatype: &Datatype,
|
||||||
|
raw: &[u8],
|
||||||
|
) -> Result<Vec<Vec<u8>>, Error> {
|
||||||
|
crate::vlen::decode_string_bytes(
|
||||||
|
self.as_bytes(),
|
||||||
|
datatype,
|
||||||
|
raw,
|
||||||
|
self.offset_size(),
|
||||||
|
self.length_size(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Decode the variable-length sequences in `raw`, a buffer of elements
|
||||||
|
/// of the sequence type `datatype` read from this file (a compound
|
||||||
|
/// field, an [`AttrValue::Raw`] attribute, ...). See
|
||||||
|
/// [`Dataset::read_vlen`].
|
||||||
|
pub fn decode_vlen<T: crate::vlen::VlenValue>(
|
||||||
|
&self,
|
||||||
|
datatype: &Datatype,
|
||||||
|
raw: &[u8],
|
||||||
|
) -> Result<Vec<Vec<T>>, Error> {
|
||||||
|
crate::vlen::decode_vlen(
|
||||||
|
self.as_bytes(),
|
||||||
|
datatype,
|
||||||
|
raw,
|
||||||
|
self.offset_size(),
|
||||||
|
self.length_size(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
fn parse_header(&self, address: u64) -> Result<ObjectHeader, FormatError> {
|
fn parse_header(&self, address: u64) -> Result<ObjectHeader, FormatError> {
|
||||||
ObjectHeader::parse(
|
ObjectHeader::parse(
|
||||||
self.data.as_bytes(),
|
self.data.as_bytes(),
|
||||||
@@ -498,11 +548,58 @@ impl<'f> Dataset<'f> {
|
|||||||
Ok(data_read::read_as_u64(&raw, &dt)?)
|
Ok(data_read::read_as_u64(&raw, &dt)?)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Read all data as `String` values.
|
/// Read all data as `String` values, in row-major order.
|
||||||
|
///
|
||||||
|
/// Works for fixed-length and variable-length string datasets (h5py's
|
||||||
|
/// default `str` dtype). A variable-length string ends at its first NUL
|
||||||
|
/// and a null element (e.g. never written) is `""`, as h5py returns
|
||||||
|
/// them; bytes that are not valid UTF-8 are replaced with U+FFFD — use
|
||||||
|
/// [`read_string_bytes`](Self::read_string_bytes) for the exact bytes.
|
||||||
pub fn read_string(&self) -> Result<Vec<String>, Error> {
|
pub fn read_string(&self) -> Result<Vec<String>, Error> {
|
||||||
let raw = self.read_raw()?;
|
let raw = self.read_raw()?;
|
||||||
let dt = self.datatype()?;
|
let dt = self.datatype()?;
|
||||||
Ok(data_read::read_as_strings(&raw, &dt)?)
|
self.file.decode_strings(&dt, &raw)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read a variable-length string dataset as the exact bytes of each
|
||||||
|
/// string (what h5py's `Dataset[()]` returns), in row-major order.
|
||||||
|
pub fn read_string_bytes(&self) -> Result<Vec<Vec<u8>>, Error> {
|
||||||
|
let raw = self.read_raw()?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
self.file.decode_string_bytes(&dt, &raw)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read the selected elements of a fixed- or variable-length string
|
||||||
|
/// dataset (see [`read_string`](Self::read_string)).
|
||||||
|
pub fn read_string_selection(
|
||||||
|
&self,
|
||||||
|
selection: &clawhdf5_format::selection::Selection,
|
||||||
|
) -> Result<Vec<String>, Error> {
|
||||||
|
let raw = self.read_selection(selection)?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
self.file.decode_strings(&dt, &raw)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read a variable-length sequence dataset (h5py
|
||||||
|
/// `vlen_dtype(np.int32)`, ...) as one `Vec<T>` per element, in
|
||||||
|
/// row-major order. The base type must be an integer or float type; it
|
||||||
|
/// is converted to `T` as [`read_f64`](Self::read_f64) and the other
|
||||||
|
/// typed readers convert. A null element is an empty sequence.
|
||||||
|
pub fn read_vlen<T: crate::vlen::VlenValue>(&self) -> Result<Vec<Vec<T>>, Error> {
|
||||||
|
let raw = self.read_raw()?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
self.file.decode_vlen(&dt, &raw)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Read the selected elements of a variable-length sequence dataset
|
||||||
|
/// (see [`read_vlen`](Self::read_vlen)).
|
||||||
|
pub fn read_vlen_selection<T: crate::vlen::VlenValue>(
|
||||||
|
&self,
|
||||||
|
selection: &clawhdf5_format::selection::Selection,
|
||||||
|
) -> Result<Vec<Vec<T>>, Error> {
|
||||||
|
let raw = self.read_selection(selection)?;
|
||||||
|
let dt = self.datatype()?;
|
||||||
|
self.file.decode_vlen(&dt, &raw)
|
||||||
}
|
}
|
||||||
|
|
||||||
// ----- Selection-based read methods -----
|
// ----- Selection-based read methods -----
|
||||||
|
|||||||
@@ -0,0 +1,143 @@
|
|||||||
|
//! Variable-length data: VL strings and VL sequences of numbers.
|
||||||
|
//!
|
||||||
|
//! A variable-length element stores a reference into the file's global heap;
|
||||||
|
//! these helpers resolve the references in a buffer of raw elements (from a
|
||||||
|
//! dataset read, a selection, a compound field or an [`AttrValue::Raw`]
|
||||||
|
//! attribute) against the file they came from.
|
||||||
|
//!
|
||||||
|
//! Values match libhdf5 (and h5py): a string ends at its first NUL, a null
|
||||||
|
//! element is an empty string or sequence, and a heap object whose size
|
||||||
|
//! disagrees with its element is an error rather than a truncated value.
|
||||||
|
//!
|
||||||
|
//! [`AttrValue::Raw`]: crate::AttrValue::Raw
|
||||||
|
|
||||||
|
use clawhdf5_format::data_read;
|
||||||
|
use clawhdf5_format::datatype::Datatype;
|
||||||
|
use clawhdf5_format::error::FormatError;
|
||||||
|
use clawhdf5_format::vl_data::{VlResolver, check_element_size};
|
||||||
|
|
||||||
|
use crate::error::Error;
|
||||||
|
|
||||||
|
mod sealed {
|
||||||
|
pub trait Sealed {}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A number type that [`Dataset::read_vlen`](crate::Dataset::read_vlen) can
|
||||||
|
/// return: the sequence's base type is converted to it as libhdf5 converts
|
||||||
|
/// numbers (the same rules as `read_f64`, `read_i64`, ...).
|
||||||
|
pub trait VlenValue: sealed::Sealed + Sized {
|
||||||
|
#[doc(hidden)]
|
||||||
|
fn decode(raw: &[u8], base: &Datatype) -> Result<Vec<Self>, FormatError>;
|
||||||
|
}
|
||||||
|
|
||||||
|
macro_rules! vlen_value {
|
||||||
|
($t:ty, $f:path) => {
|
||||||
|
impl sealed::Sealed for $t {}
|
||||||
|
impl VlenValue for $t {
|
||||||
|
fn decode(raw: &[u8], base: &Datatype) -> Result<Vec<Self>, FormatError> {
|
||||||
|
$f(raw, base)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
vlen_value!(f64, data_read::read_as_f64);
|
||||||
|
vlen_value!(f32, data_read::read_as_f32);
|
||||||
|
vlen_value!(i64, data_read::read_as_i64);
|
||||||
|
vlen_value!(i32, data_read::read_as_i32);
|
||||||
|
vlen_value!(u64, data_read::read_as_u64);
|
||||||
|
|
||||||
|
fn class_name(dt: &Datatype) -> &'static str {
|
||||||
|
match dt {
|
||||||
|
Datatype::FixedPoint { .. } => "integer",
|
||||||
|
Datatype::FloatingPoint { .. } => "float",
|
||||||
|
Datatype::Time { .. } => "time",
|
||||||
|
Datatype::String { .. } => "fixed-length string",
|
||||||
|
Datatype::BitField { .. } => "bitfield",
|
||||||
|
Datatype::Opaque { .. } => "opaque",
|
||||||
|
Datatype::Compound { .. } => "compound",
|
||||||
|
Datatype::Reference { .. } => "reference",
|
||||||
|
Datatype::Enumeration { .. } => "enum",
|
||||||
|
Datatype::VariableLength {
|
||||||
|
is_string: true, ..
|
||||||
|
} => "variable-length string",
|
||||||
|
Datatype::VariableLength { .. } => "variable-length sequence",
|
||||||
|
Datatype::Array { .. } => "array",
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The strings in `raw`, elements of `dt`: fixed-length strings decoded as
|
||||||
|
/// `read_string` always has, variable-length strings resolved in the heap.
|
||||||
|
pub(crate) fn decode_strings(
|
||||||
|
file_data: &[u8],
|
||||||
|
dt: &Datatype,
|
||||||
|
raw: &[u8],
|
||||||
|
offset_size: u8,
|
||||||
|
length_size: u8,
|
||||||
|
) -> Result<Vec<String>, Error> {
|
||||||
|
match dt {
|
||||||
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
|
is_string: true,
|
||||||
|
..
|
||||||
|
} => {
|
||||||
|
check_element_size(*size, offset_size)?;
|
||||||
|
Ok(VlResolver::new(file_data, offset_size, length_size).strings(raw)?)
|
||||||
|
}
|
||||||
|
_ => Ok(data_read::read_as_strings(raw, dt)?),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The exact bytes of the variable-length strings in `raw`.
|
||||||
|
pub(crate) fn decode_string_bytes(
|
||||||
|
file_data: &[u8],
|
||||||
|
dt: &Datatype,
|
||||||
|
raw: &[u8],
|
||||||
|
offset_size: u8,
|
||||||
|
length_size: u8,
|
||||||
|
) -> Result<Vec<Vec<u8>>, Error> {
|
||||||
|
match dt {
|
||||||
|
Datatype::VariableLength {
|
||||||
|
size,
|
||||||
|
is_string: true,
|
||||||
|
..
|
||||||
|
} => {
|
||||||
|
check_element_size(*size, offset_size)?;
|
||||||
|
Ok(VlResolver::new(file_data, offset_size, length_size).string_bytes(raw)?)
|
||||||
|
}
|
||||||
|
other => Err(Error::Format(FormatError::TypeMismatch {
|
||||||
|
expected: "variable-length string",
|
||||||
|
actual: class_name(other),
|
||||||
|
})),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The sequences in `raw`, elements of the variable-length sequence type
|
||||||
|
/// `dt`, converted to `T`.
|
||||||
|
pub(crate) fn decode_vlen<T: VlenValue>(
|
||||||
|
file_data: &[u8],
|
||||||
|
dt: &Datatype,
|
||||||
|
raw: &[u8],
|
||||||
|
offset_size: u8,
|
||||||
|
length_size: u8,
|
||||||
|
) -> Result<Vec<Vec<T>>, Error> {
|
||||||
|
let Datatype::VariableLength {
|
||||||
|
size,
|
||||||
|
is_string: false,
|
||||||
|
base_type,
|
||||||
|
..
|
||||||
|
} = dt
|
||||||
|
else {
|
||||||
|
return Err(Error::Format(FormatError::TypeMismatch {
|
||||||
|
expected: "variable-length sequence",
|
||||||
|
actual: class_name(dt),
|
||||||
|
}));
|
||||||
|
};
|
||||||
|
check_element_size(*size, offset_size)?;
|
||||||
|
let base_size = base_type.type_size() as usize;
|
||||||
|
VlResolver::new(file_data, offset_size, length_size)
|
||||||
|
.sequences(raw, base_size)?
|
||||||
|
.iter()
|
||||||
|
.map(|bytes| Ok(T::decode(bytes, base_type)?))
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
@@ -0,0 +1,494 @@
|
|||||||
|
//! Variable-length data (VL strings and VL sequences) read through the
|
||||||
|
//! facade, checked against h5py/libhdf5.
|
||||||
|
//!
|
||||||
|
//! h5py writes each file — once with the default 8-byte offsets and once
|
||||||
|
//! with 4-byte offsets and lengths (`sizeof_addr = 4`) — and prints what
|
||||||
|
//! libhdf5 reads back; `File`, `MmapFile` and `LazyFile` must return the same
|
||||||
|
//! values. Skipped when python3 with h5py is unavailable, unless
|
||||||
|
//! `CLAWHDF5_REQUIRE_INTEROP=1`.
|
||||||
|
|
||||||
|
// `Selection::slice(&[0..1])` is one range per dimension, not a Vec of a range.
|
||||||
|
#![allow(clippy::single_range_in_vec_init)]
|
||||||
|
|
||||||
|
use std::collections::HashMap;
|
||||||
|
use std::path::Path;
|
||||||
|
use std::process::Command;
|
||||||
|
|
||||||
|
use clawhdf5::{AttrValue, File, LazyFile, MmapFile, Selection};
|
||||||
|
|
||||||
|
fn python() -> String {
|
||||||
|
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn interop_required() -> bool {
|
||||||
|
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
|
||||||
|
}
|
||||||
|
|
||||||
|
fn python_available() -> bool {
|
||||||
|
Command::new(python())
|
||||||
|
.args(["-c", "import h5py, numpy"])
|
||||||
|
.output()
|
||||||
|
.map(|o| o.status.success())
|
||||||
|
.unwrap_or(false)
|
||||||
|
}
|
||||||
|
|
||||||
|
macro_rules! skip_if_no_python {
|
||||||
|
() => {
|
||||||
|
if !python_available() {
|
||||||
|
assert!(
|
||||||
|
!interop_required(),
|
||||||
|
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
|
||||||
|
);
|
||||||
|
eprintln!("SKIP: python3 with h5py not available");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Run `script` and return its stdout as `key -> value`, one
|
||||||
|
/// `key<TAB>value` line per key.
|
||||||
|
fn run_python(script: &str) -> HashMap<String, String> {
|
||||||
|
let output = Command::new(python())
|
||||||
|
.args(["-c", script])
|
||||||
|
.output()
|
||||||
|
.expect("failed to run python");
|
||||||
|
assert!(
|
||||||
|
output.status.success(),
|
||||||
|
"python failed:\n{}",
|
||||||
|
String::from_utf8_lossy(&output.stderr)
|
||||||
|
);
|
||||||
|
String::from_utf8_lossy(&output.stdout)
|
||||||
|
.lines()
|
||||||
|
.filter_map(|line| {
|
||||||
|
let (k, v) = line.split_once('\t')?;
|
||||||
|
Some((k.to_string(), v.to_string()))
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `hex,hex,...` -> the strings' bytes.
|
||||||
|
fn parse_strings(v: &str) -> Vec<Vec<u8>> {
|
||||||
|
v.split(',')
|
||||||
|
.map(|h| {
|
||||||
|
(0..h.len())
|
||||||
|
.step_by(2)
|
||||||
|
.map(|i| u8::from_str_radix(&h[i..i + 2], 16).unwrap())
|
||||||
|
.collect()
|
||||||
|
})
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `1 2 3|| -5` -> sequences.
|
||||||
|
fn parse_seqs(v: &str) -> Vec<Vec<f64>> {
|
||||||
|
v.split('|')
|
||||||
|
.map(|s| s.split_whitespace().map(|x| x.parse().unwrap()).collect())
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn utf8(bytes: &[Vec<u8>]) -> Vec<String> {
|
||||||
|
bytes
|
||||||
|
.iter()
|
||||||
|
.map(|b| String::from_utf8(b.clone()).unwrap())
|
||||||
|
.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Writes `vl8.h5` (8-byte offsets) and `vl4.h5` (4-byte offsets and
|
||||||
|
/// lengths) into `dir` and prints h5py's reading of both.
|
||||||
|
const SCRIPT: &str = r#"
|
||||||
|
import sys, h5py, numpy as np
|
||||||
|
d = sys.argv[1]
|
||||||
|
S = h5py.string_dtype('utf-8'); A = h5py.string_dtype('ascii')
|
||||||
|
def make(path, sizes):
|
||||||
|
if sizes:
|
||||||
|
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE); fcpl.set_sizes(*sizes)
|
||||||
|
f = h5py.File(h5py.h5f.create(path.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl))
|
||||||
|
else:
|
||||||
|
f = h5py.File(path, 'w')
|
||||||
|
f.create_dataset('scalar_utf8', data='héllo', dtype=S)
|
||||||
|
f.create_dataset('scalar_ascii', data=b'hello', dtype=A)
|
||||||
|
f.create_dataset('d1', data=np.array(['a', '', 'ccc', 'δδ'], dtype=object), dtype=S)
|
||||||
|
f.create_dataset('d2', data=np.array([['x', 'yy', 'zzz'], ['', 'w', 'vv']], dtype=object), dtype=S)
|
||||||
|
f.create_dataset('chunked', data=np.array(['s%d' % i * (i % 5) for i in range(100)], dtype=object),
|
||||||
|
dtype=S, chunks=(7,), compression='gzip')
|
||||||
|
f.create_dataset('chunked2d', data=np.array([['r%dc%d' % (r, c) for c in range(9)] for r in range(11)], dtype=object),
|
||||||
|
dtype=S, chunks=(4, 4), compression='gzip', shuffle=True)
|
||||||
|
f.create_dataset('unwritten', shape=(5,), dtype=S, chunks=(2,))
|
||||||
|
p = f.create_dataset('partial', shape=(6,), dtype=S, chunks=(2,)); p[0] = 'first'; p[5] = 'last'
|
||||||
|
f.create_dataset('contig_empty', shape=(3,), dtype=S)
|
||||||
|
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE); dcpl.set_layout(h5py.h5d.COMPACT)
|
||||||
|
f.create_dataset('compact', data=np.array(['c1', '', 'c3'], dtype=object), dtype=S, dcpl=dcpl)
|
||||||
|
assert f['compact'].id.get_create_plist().get_layout() == h5py.h5d.COMPACT
|
||||||
|
f.attrs['vlattr'] = 'attr-value'
|
||||||
|
f.attrs.create('vlattr_arr', np.array(['p', 'qq', ''], dtype=object), dtype=S)
|
||||||
|
ct = np.dtype([('id', '<i4'), ('name', S), ('v', '<f8')])
|
||||||
|
arr = np.zeros(3, dtype=ct); arr['id'] = [1, 2, 3]; arr['name'] = ['one', '', 'three']; arr['v'] = [.5, 1.5, 2.5]
|
||||||
|
f.create_dataset('compound', data=arr)
|
||||||
|
f.attrs.create('compound_attr', arr)
|
||||||
|
v = f.create_dataset('vlen_i4', shape=(3,), dtype=h5py.vlen_dtype(np.dtype('<i4')))
|
||||||
|
v[0] = [1, 2, 3]; v[1] = []; v[2] = [-5]
|
||||||
|
v = f.create_dataset('vlen_f8', shape=(2, 2), dtype=h5py.vlen_dtype(np.dtype('<f8')), chunks=(1, 2), compression='gzip')
|
||||||
|
v[0, 0] = [1.5]; v[0, 1] = [2.5, 3.5]; v[1, 1] = [9.0]
|
||||||
|
v = f.create_dataset('vlen_u2_be', shape=(2,), dtype=h5py.vlen_dtype(np.dtype('>u2')))
|
||||||
|
v[0] = [1, 65535]; v[1] = [300]
|
||||||
|
f.attrs.create('vlen_attr', np.array([np.array([1, 2], dtype='<i8'), np.array([3], dtype='<i8')], dtype=object),
|
||||||
|
dtype=h5py.vlen_dtype(np.dtype('<i8')))
|
||||||
|
f.close()
|
||||||
|
|
||||||
|
def hexes(a):
|
||||||
|
return ','.join(bytes(x).hex() for x in np.asarray(a, dtype=object).ravel())
|
||||||
|
def seqs(a):
|
||||||
|
return '|'.join(' '.join(repr(float(x)) for x in s) for s in np.asarray(a, dtype=object).ravel())
|
||||||
|
|
||||||
|
for tag, sizes in (('8', None), ('4', (4, 4))):
|
||||||
|
path = '%s/vl%s.h5' % (d, tag)
|
||||||
|
make(path, sizes)
|
||||||
|
with h5py.File(path, 'r') as f:
|
||||||
|
for name in ('compact', 'scalar_utf8', 'scalar_ascii', 'd1', 'd2', 'chunked', 'chunked2d', 'unwritten',
|
||||||
|
'partial', 'contig_empty'):
|
||||||
|
v = f[name][()]
|
||||||
|
print('%s:%s\t%s' % (tag, name, hexes([v] if np.ndim(v) == 0 else v)))
|
||||||
|
print('%s:d2[1,1:3]\t%s' % (tag, hexes(f['d2'][1, 1:3])))
|
||||||
|
print('%s:chunked[5:60:3]\t%s' % (tag, hexes(f['chunked'][5:60:3])))
|
||||||
|
print('%s:chunked2d[2:9:2,3:8]\t%s' % (tag, hexes(f['chunked2d'][2:9:2, 3:8])))
|
||||||
|
print('%s:compound.name\t%s' % (tag, hexes(f['compound']['name'])))
|
||||||
|
print('%s:compound_attr.name\t%s' % (tag, hexes(f.attrs['compound_attr']['name'])))
|
||||||
|
print('%s:vlattr\t%s' % (tag, hexes([f.attrs['vlattr'].encode()])))
|
||||||
|
print('%s:vlattr_arr\t%s' % (tag, hexes([s.encode() for s in f.attrs['vlattr_arr']])))
|
||||||
|
for name in ('vlen_i4', 'vlen_f8'):
|
||||||
|
print('%s:%s\t%s' % (tag, name, seqs(f[name][()])))
|
||||||
|
print('%s:vlen_f8[1,:]\t%s' % (tag, seqs(f['vlen_f8'][1, :])))
|
||||||
|
print('%s:vlen_attr\t%s' % (tag, seqs(f.attrs['vlen_attr'])))
|
||||||
|
"#;
|
||||||
|
|
||||||
|
fn make_files(dir: &Path) -> HashMap<String, String> {
|
||||||
|
let script = format!(
|
||||||
|
"import sys; sys.argv = ['x', {:?}]\n{SCRIPT}",
|
||||||
|
dir.display().to_string()
|
||||||
|
);
|
||||||
|
run_python(&script)
|
||||||
|
}
|
||||||
|
|
||||||
|
const STRING_DATASETS: [&str; 10] = [
|
||||||
|
"compact",
|
||||||
|
"scalar_utf8",
|
||||||
|
"scalar_ascii",
|
||||||
|
"d1",
|
||||||
|
"d2",
|
||||||
|
"chunked",
|
||||||
|
"chunked2d",
|
||||||
|
"unwritten",
|
||||||
|
"partial",
|
||||||
|
"contig_empty",
|
||||||
|
];
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn vl_string_datasets_read_like_h5py() {
|
||||||
|
skip_if_no_python!();
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let expected = make_files(dir.path());
|
||||||
|
for tag in ["8", "4"] {
|
||||||
|
let path = dir.path().join(format!("vl{tag}.h5"));
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
let mmap = MmapFile::open(&path).unwrap();
|
||||||
|
let lazy = LazyFile::open_mmap(&path).unwrap();
|
||||||
|
for name in STRING_DATASETS {
|
||||||
|
let want = parse_strings(&expected[&format!("{tag}:{name}")]);
|
||||||
|
let ctx = format!("vl{tag}.h5 {name}");
|
||||||
|
let ds = file.dataset(name).unwrap();
|
||||||
|
assert_eq!(ds.read_string_bytes().unwrap(), want, "{ctx}");
|
||||||
|
assert_eq!(ds.read_string().unwrap(), utf8(&want), "{ctx}");
|
||||||
|
let m = mmap.dataset(name).unwrap();
|
||||||
|
assert_eq!(m.read_string_bytes().unwrap(), want, "{ctx} (mmap)");
|
||||||
|
assert_eq!(m.read_string().unwrap(), utf8(&want), "{ctx} (mmap)");
|
||||||
|
let l = lazy.dataset(name).unwrap();
|
||||||
|
assert_eq!(l.read_string_bytes().unwrap(), want, "{ctx} (lazy)");
|
||||||
|
assert_eq!(l.read_string().unwrap(), utf8(&want), "{ctx} (lazy)");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn vl_string_selections_read_like_h5py() {
|
||||||
|
skip_if_no_python!();
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let expected = make_files(dir.path());
|
||||||
|
let hyperslab = |start: &[u64], stride: &[u64], count: &[u64]| Selection::Hyperslab {
|
||||||
|
start: start.to_vec(),
|
||||||
|
stride: stride.to_vec(),
|
||||||
|
count: count.to_vec(),
|
||||||
|
block: vec![1; start.len()],
|
||||||
|
};
|
||||||
|
let cases = [
|
||||||
|
("d2", "d2[1,1:3]", Selection::slice(&[1..2, 1..3])),
|
||||||
|
("chunked", "chunked[5:60:3]", hyperslab(&[5], &[3], &[19])),
|
||||||
|
(
|
||||||
|
"chunked2d",
|
||||||
|
"chunked2d[2:9:2,3:8]",
|
||||||
|
hyperslab(&[2, 3], &[2, 1], &[4, 5]),
|
||||||
|
),
|
||||||
|
];
|
||||||
|
for tag in ["8", "4"] {
|
||||||
|
let file = File::open(dir.path().join(format!("vl{tag}.h5"))).unwrap();
|
||||||
|
for (name, key, sel) in &cases {
|
||||||
|
let want = utf8(&parse_strings(&expected[&format!("{tag}:{key}")]));
|
||||||
|
let got = file
|
||||||
|
.dataset(name)
|
||||||
|
.unwrap()
|
||||||
|
.read_string_selection(sel)
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(got, want, "vl{tag}.h5 {key}");
|
||||||
|
}
|
||||||
|
// A selection of VL integers is not strings.
|
||||||
|
assert!(
|
||||||
|
file.dataset("vlen_i4")
|
||||||
|
.unwrap()
|
||||||
|
.read_string_selection(&Selection::slice(&[0..1]))
|
||||||
|
.is_err()
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn vl_values_in_compounds_and_attributes_read_like_h5py() {
|
||||||
|
// With 4-byte offsets these failed with GlobalHeapObjectNotFound or came
|
||||||
|
// back as `AttrValue::Raw`: the VL type claimed 16-byte elements and the
|
||||||
|
// global heap was read without the padding libhdf5 puts after its
|
||||||
|
// headers.
|
||||||
|
skip_if_no_python!();
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let expected = make_files(dir.path());
|
||||||
|
for tag in ["8", "4"] {
|
||||||
|
let file = File::open(dir.path().join(format!("vl{tag}.h5"))).unwrap();
|
||||||
|
let want = |key: &str| utf8(&parse_strings(&expected[&format!("{tag}:{key}")]));
|
||||||
|
|
||||||
|
let attrs = file.root().attrs().unwrap();
|
||||||
|
match &attrs["vlattr"] {
|
||||||
|
AttrValue::String(s) => assert_eq!(*s, want("vlattr")[0], "vl{tag}.h5"),
|
||||||
|
other => panic!("vl{tag}.h5 vlattr: {other:?}"),
|
||||||
|
}
|
||||||
|
match &attrs["vlattr_arr"] {
|
||||||
|
AttrValue::StringArray(s) => assert_eq!(*s, want("vlattr_arr"), "vl{tag}.h5"),
|
||||||
|
other => panic!("vl{tag}.h5 vlattr_arr: {other:?}"),
|
||||||
|
}
|
||||||
|
|
||||||
|
// Compound with a VL string member: dataset and attribute.
|
||||||
|
let ds = file.dataset("compound").unwrap();
|
||||||
|
let dt = ds.raw_datatype().unwrap();
|
||||||
|
let raw = ds.read_selection(&Selection::All).unwrap();
|
||||||
|
let fields = clawhdf5_format::data_read::read_compound_fields(&raw, &dt).unwrap();
|
||||||
|
let name = fields.iter().find(|f| f.name == "name").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
file.decode_strings(&name.datatype, &name.raw_data).unwrap(),
|
||||||
|
want("compound.name"),
|
||||||
|
"vl{tag}.h5 compound"
|
||||||
|
);
|
||||||
|
let id = fields.iter().find(|f| f.name == "id").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
clawhdf5_format::data_read::read_as_i64(&id.raw_data, &id.datatype).unwrap(),
|
||||||
|
vec![1, 2, 3]
|
||||||
|
);
|
||||||
|
let v = fields.iter().find(|f| f.name == "v").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
clawhdf5_format::data_read::read_as_f64(&v.raw_data, &v.datatype).unwrap(),
|
||||||
|
vec![0.5, 1.5, 2.5]
|
||||||
|
);
|
||||||
|
|
||||||
|
let AttrValue::Raw { datatype, data, .. } = &attrs["compound_attr"] else {
|
||||||
|
panic!("compound attribute is Raw");
|
||||||
|
};
|
||||||
|
let fields = clawhdf5_format::data_read::read_compound_fields(data, datatype).unwrap();
|
||||||
|
let name = fields.iter().find(|f| f.name == "name").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
file.decode_strings(&name.datatype, &name.raw_data).unwrap(),
|
||||||
|
want("compound_attr.name"),
|
||||||
|
"vl{tag}.h5 compound attribute"
|
||||||
|
);
|
||||||
|
|
||||||
|
// A VL sequence attribute.
|
||||||
|
let AttrValue::Raw { datatype, data, .. } = &attrs["vlen_attr"] else {
|
||||||
|
panic!("vlen attribute is Raw");
|
||||||
|
};
|
||||||
|
let got: Vec<Vec<i64>> = file.decode_vlen(datatype, data).unwrap();
|
||||||
|
let want_seqs = parse_seqs(&expected[&format!("{tag}:vlen_attr")]);
|
||||||
|
assert_eq!(
|
||||||
|
got,
|
||||||
|
want_seqs
|
||||||
|
.iter()
|
||||||
|
.map(|s| s.iter().map(|&x| x as i64).collect::<Vec<_>>())
|
||||||
|
.collect::<Vec<_>>(),
|
||||||
|
"vl{tag}.h5 vlen_attr"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn vl_sequence_datasets_read_like_h5py() {
|
||||||
|
skip_if_no_python!();
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let expected = make_files(dir.path());
|
||||||
|
for tag in ["8", "4"] {
|
||||||
|
let path = dir.path().join(format!("vl{tag}.h5"));
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
let seqs = |key: &str| parse_seqs(&expected[&format!("{tag}:{key}")]);
|
||||||
|
|
||||||
|
let i4 = file.dataset("vlen_i4").unwrap();
|
||||||
|
let want: Vec<Vec<i32>> = seqs("vlen_i4")
|
||||||
|
.iter()
|
||||||
|
.map(|s| s.iter().map(|&x| x as i32).collect())
|
||||||
|
.collect();
|
||||||
|
assert_eq!(i4.read_vlen::<i32>().unwrap(), want, "vl{tag}.h5 vlen_i4");
|
||||||
|
let as_f64: Vec<Vec<f64>> = i4.read_vlen().unwrap();
|
||||||
|
assert_eq!(as_f64, seqs("vlen_i4"));
|
||||||
|
|
||||||
|
let f8 = file.dataset("vlen_f8").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
f8.read_vlen::<f64>().unwrap(),
|
||||||
|
seqs("vlen_f8"),
|
||||||
|
"vl{tag}.h5"
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
f8.read_vlen_selection::<f64>(&Selection::slice(&[1..2, 0..2]))
|
||||||
|
.unwrap(),
|
||||||
|
seqs("vlen_f8[1,:]"),
|
||||||
|
"vl{tag}.h5 vlen_f8[1,:]"
|
||||||
|
);
|
||||||
|
|
||||||
|
// h5py returns big-endian VL elements byte-swapped (an h5py bug, see
|
||||||
|
// CONFORMANCE.md); the values written are [1, 65535] and [300].
|
||||||
|
assert_eq!(
|
||||||
|
file.dataset("vlen_u2_be")
|
||||||
|
.unwrap()
|
||||||
|
.read_vlen::<u64>()
|
||||||
|
.unwrap(),
|
||||||
|
vec![vec![1, 65535], vec![300]]
|
||||||
|
);
|
||||||
|
|
||||||
|
let mmap = MmapFile::open(&path).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
mmap.dataset("vlen_f8").unwrap().read_vlen::<f64>().unwrap(),
|
||||||
|
seqs("vlen_f8")
|
||||||
|
);
|
||||||
|
let lazy = LazyFile::open_mmap(&path).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
lazy.dataset("vlen_f8").unwrap().read_vlen::<f64>().unwrap(),
|
||||||
|
seqs("vlen_f8")
|
||||||
|
);
|
||||||
|
|
||||||
|
// Wrong kind of data is an error, not a value.
|
||||||
|
assert!(i4.read_string().is_err());
|
||||||
|
assert!(i4.read_string_bytes().is_err());
|
||||||
|
assert!(file.dataset("d1").unwrap().read_vlen::<f64>().is_err());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn vl_strings_end_at_nul_and_mis_sized_elements_fail_like_h5py() {
|
||||||
|
// h5py cannot write a VL string with a NUL in it, so the file is patched:
|
||||||
|
// one string gets an embedded NUL, and two elements get a length that
|
||||||
|
// disagrees with their heap object. libhdf5 returns the string up to the
|
||||||
|
// NUL and refuses the others ("Expected global heap object size does
|
||||||
|
// not match"); we used to return the NUL and a truncated string.
|
||||||
|
skip_if_no_python!();
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = dir.path().join("patched.h5");
|
||||||
|
let script = format!(
|
||||||
|
r#"
|
||||||
|
import struct, h5py, numpy as np
|
||||||
|
path = {path:?}
|
||||||
|
with h5py.File(path, 'w') as f:
|
||||||
|
f.create_dataset('d', data=np.array(['aXb', 'cdefgh', 'ij', 'ok'], dtype=object),
|
||||||
|
dtype=h5py.string_dtype())
|
||||||
|
s = f.create_dataset('seq', shape=(2,), dtype=h5py.vlen_dtype(np.dtype('<i4')))
|
||||||
|
s[0] = np.array([1, 2, 3], dtype='<i4'); s[1] = np.array([4], dtype='<i4')
|
||||||
|
off = f['d'].id.get_offset(); soff = f['seq'].id.get_offset()
|
||||||
|
b = bytearray(open(path, 'rb').read())
|
||||||
|
i = b.index(b'aXb'); b[i + 1] = 0
|
||||||
|
struct.pack_into('<I', b, off + 16, 3) # 'cdefgh': length 6 -> 3
|
||||||
|
struct.pack_into('<I', b, off + 32, 9) # 'ij': length 2 -> 9
|
||||||
|
struct.pack_into('<I', b, soff, 2) # [1, 2, 3]: length 3 -> 2
|
||||||
|
open(path, 'wb').write(bytes(b))
|
||||||
|
with h5py.File(path, 'r') as f:
|
||||||
|
for i in range(4):
|
||||||
|
try:
|
||||||
|
print('d%d\t%s' % (i, f['d'][i].hex()))
|
||||||
|
except OSError as e:
|
||||||
|
print('d%d\terror' % i)
|
||||||
|
for i in range(2):
|
||||||
|
try:
|
||||||
|
print('seq%d\t%s' % (i, ' '.join(str(x) for x in f['seq'][i])))
|
||||||
|
except OSError as e:
|
||||||
|
print('seq%d\terror' % i)
|
||||||
|
"#,
|
||||||
|
path = path.display().to_string()
|
||||||
|
);
|
||||||
|
let expected = run_python(&script);
|
||||||
|
assert_eq!(expected["d0"], "61", "h5py cuts 'a\\0b' at the NUL");
|
||||||
|
assert_eq!(expected["d1"], "error");
|
||||||
|
assert_eq!(expected["d2"], "error");
|
||||||
|
assert_eq!(expected["d3"], "6f6b");
|
||||||
|
assert_eq!(expected["seq0"], "error");
|
||||||
|
assert_eq!(expected["seq1"], "4");
|
||||||
|
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
let d = file.dataset("d").unwrap();
|
||||||
|
let one = |i: u64| d.read_string_selection(&Selection::slice(&[i..i + 1]));
|
||||||
|
assert_eq!(one(0).unwrap(), vec!["a"]);
|
||||||
|
assert!(one(1).is_err());
|
||||||
|
assert!(one(2).is_err());
|
||||||
|
assert_eq!(one(3).unwrap(), vec!["ok"]);
|
||||||
|
assert!(d.read_string().is_err());
|
||||||
|
let seq = file.dataset("seq").unwrap();
|
||||||
|
let one = |i: u64| seq.read_vlen_selection::<i32>(&Selection::slice(&[i..i + 1]));
|
||||||
|
assert!(one(0).is_err());
|
||||||
|
assert_eq!(one(1).unwrap(), vec![vec![4]]);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn a_vl_element_at_the_undefined_heap_address_fails_like_h5py() {
|
||||||
|
// libhdf5 writes a null element with heap address 0 (h5py reads it as
|
||||||
|
// b''), and an empty string as a real zero-size heap object; neither
|
||||||
|
// uses the undefined address. An element of length 0 at the undefined
|
||||||
|
// address fails in libhdf5 ("addr undefined"); we returned "".
|
||||||
|
skip_if_no_python!();
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = dir.path().join("undef.h5");
|
||||||
|
let script = format!(
|
||||||
|
r#"
|
||||||
|
import struct, h5py, numpy as np
|
||||||
|
path = {path:?}
|
||||||
|
with h5py.File(path, 'w') as f:
|
||||||
|
f.create_dataset('d', data=np.array(['x', '', 'yz', ''], dtype=object), dtype=h5py.string_dtype())
|
||||||
|
off = f['d'].id.get_offset()
|
||||||
|
b = bytearray(open(path, 'rb').read())
|
||||||
|
# h5py's '' (element 3): length 0 at a real heap address, not 0 or all 0xff.
|
||||||
|
length, addr, _ = struct.unpack_from('<IQI', b, off + 48)
|
||||||
|
print('empty\t%d %d' % (length, addr not in (0, 2**64 - 1)))
|
||||||
|
struct.pack_into('<IQI', b, off + 16, 0, 2**64 - 1, 1)
|
||||||
|
open(path, 'wb').write(bytes(b))
|
||||||
|
with h5py.File(path, 'r') as f:
|
||||||
|
for i in range(4):
|
||||||
|
try:
|
||||||
|
print('d%d\t%s' % (i, f['d'][i].hex()))
|
||||||
|
except OSError as e:
|
||||||
|
print('d%d\terror %s' % (i, 'addr undefined' in str(e)))
|
||||||
|
"#,
|
||||||
|
path = path.display().to_string()
|
||||||
|
);
|
||||||
|
let expected = run_python(&script);
|
||||||
|
assert_eq!(expected["empty"], "0 1", "h5py writes '' at a real address");
|
||||||
|
assert_eq!(expected["d0"], "78");
|
||||||
|
assert_eq!(expected["d1"], "error True");
|
||||||
|
assert_eq!(expected["d2"], "797a");
|
||||||
|
assert_eq!(expected["d3"], "");
|
||||||
|
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
let d = file.dataset("d").unwrap();
|
||||||
|
let one = |i: u64| d.read_string_selection(&Selection::slice(&[i..i + 1]));
|
||||||
|
assert_eq!(one(0).unwrap(), vec!["x"]);
|
||||||
|
let e = one(1).unwrap_err().to_string();
|
||||||
|
assert!(e.contains("undefined global heap address"), "{e}");
|
||||||
|
assert_eq!(one(2).unwrap(), vec!["yz"]);
|
||||||
|
assert_eq!(one(3).unwrap(), vec![""]);
|
||||||
|
assert!(d.read_string().is_err());
|
||||||
|
assert!(d.read_string_bytes().is_err());
|
||||||
|
}
|
||||||
@@ -0,0 +1,126 @@
|
|||||||
|
//! Variable-length values in files with 4-byte offsets and lengths
|
||||||
|
//! (`sizeof_addr = 4`), checked against h5py/libhdf5 through the
|
||||||
|
//! `clawhdf5_format` decoders.
|
||||||
|
//!
|
||||||
|
//! These failed with `GlobalHeapObjectNotFound` or came back as
|
||||||
|
//! `AttrValue::Raw`: the VL datatype claimed 16-byte elements whatever the
|
||||||
|
//! file's offset size, and the global heap was read without the padding
|
||||||
|
//! libhdf5 puts after its collection and object headers. Skipped when
|
||||||
|
//! python3 with h5py is unavailable, unless `CLAWHDF5_REQUIRE_INTEROP=1`.
|
||||||
|
|
||||||
|
use std::process::Command;
|
||||||
|
|
||||||
|
use clawhdf5::{AttrValue, File, Selection};
|
||||||
|
use clawhdf5_format::data_read::{read_as_i64, read_compound_fields};
|
||||||
|
use clawhdf5_format::vl_data::{read_vl_bytes, read_vl_strings};
|
||||||
|
|
||||||
|
fn python() -> String {
|
||||||
|
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn interop_required() -> bool {
|
||||||
|
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
|
||||||
|
}
|
||||||
|
|
||||||
|
fn python_available() -> bool {
|
||||||
|
Command::new(python())
|
||||||
|
.args(["-c", "import h5py, numpy"])
|
||||||
|
.output()
|
||||||
|
.map(|o| o.status.success())
|
||||||
|
.unwrap_or(false)
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn vl_values_in_a_file_with_4_byte_offsets_read_like_h5py() {
|
||||||
|
if !python_available() {
|
||||||
|
assert!(
|
||||||
|
!interop_required(),
|
||||||
|
"CLAWHDF5_REQUIRE_INTEROP=1 but python3 with h5py is not available"
|
||||||
|
);
|
||||||
|
eprintln!("SKIP: python3 with h5py not available");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = tempfile::tempdir().unwrap();
|
||||||
|
let path = dir.path().join("offset4.h5");
|
||||||
|
let script = format!(
|
||||||
|
r#"
|
||||||
|
import h5py, numpy as np
|
||||||
|
S = h5py.string_dtype()
|
||||||
|
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE); fcpl.set_sizes(4, 4)
|
||||||
|
with h5py.File(h5py.h5f.create({path:?}.encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl)) as f:
|
||||||
|
f.attrs['vlattr'] = 'attr-value'
|
||||||
|
f.attrs.create('vlattr_arr', np.array(['p', 'qq', ''], dtype=object), dtype=S)
|
||||||
|
ct = np.dtype([('id', '<i4'), ('name', S), ('v', '<f8')])
|
||||||
|
arr = np.zeros(3, dtype=ct); arr['id'] = [1, 2, 3]; arr['name'] = ['one', '', 'three']
|
||||||
|
f.create_dataset('compound', data=arr)
|
||||||
|
f.attrs.create('vlen_attr', np.array([np.array([1, 2], dtype='<i8'), np.array([3], dtype='<i8')],
|
||||||
|
dtype=object), dtype=h5py.vlen_dtype(np.dtype('<i8')))
|
||||||
|
with h5py.File({path:?}, 'r') as f:
|
||||||
|
assert f.id.get_create_plist().get_sizes() == (4, 4)
|
||||||
|
print(f.attrs['vlattr'])
|
||||||
|
print(','.join(f.attrs['vlattr_arr']))
|
||||||
|
print(','.join(s.decode() for s in f['compound']['name']))
|
||||||
|
print(';'.join(' '.join(str(x) for x in s) for s in f.attrs['vlen_attr']))
|
||||||
|
"#,
|
||||||
|
path = path.display().to_string()
|
||||||
|
);
|
||||||
|
let output = Command::new(python())
|
||||||
|
.args(["-c", &script])
|
||||||
|
.output()
|
||||||
|
.expect("failed to run python");
|
||||||
|
assert!(
|
||||||
|
output.status.success(),
|
||||||
|
"python failed:\n{}",
|
||||||
|
String::from_utf8_lossy(&output.stderr)
|
||||||
|
);
|
||||||
|
let stdout = String::from_utf8(output.stdout).unwrap();
|
||||||
|
let lines: Vec<&str> = stdout.lines().collect();
|
||||||
|
let (vlattr, vlattr_arr, names, seqs) = (lines[0], lines[1], lines[2], lines[3]);
|
||||||
|
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
let sb = file.superblock();
|
||||||
|
assert_eq!((sb.offset_size, sb.length_size), (4, 4));
|
||||||
|
|
||||||
|
let attrs = file.root().attrs().unwrap();
|
||||||
|
match &attrs["vlattr"] {
|
||||||
|
AttrValue::String(s) => assert_eq!(s, vlattr),
|
||||||
|
other => panic!("vlattr: {other:?}"),
|
||||||
|
}
|
||||||
|
match &attrs["vlattr_arr"] {
|
||||||
|
AttrValue::StringArray(s) => assert_eq!(s.join(","), vlattr_arr),
|
||||||
|
other => panic!("vlattr_arr: {other:?}"),
|
||||||
|
}
|
||||||
|
|
||||||
|
// The compound's VL string member.
|
||||||
|
let ds = file.dataset("compound").unwrap();
|
||||||
|
let dt = ds.raw_datatype().unwrap();
|
||||||
|
assert_eq!(dt.type_size(), 24, "4 + 12-byte VL element + 8");
|
||||||
|
let raw = ds.read_selection(&Selection::All).unwrap();
|
||||||
|
let fields = read_compound_fields(&raw, &dt).unwrap();
|
||||||
|
let name = fields.iter().find(|f| f.name == "name").unwrap();
|
||||||
|
assert_eq!(name.datatype.type_size(), 12);
|
||||||
|
let got = read_vl_strings(file.as_bytes(), &name.raw_data, 3, 4, 4).unwrap();
|
||||||
|
assert_eq!(got.join(","), names);
|
||||||
|
let id = fields.iter().find(|f| f.name == "id").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
read_as_i64(&id.raw_data, &id.datatype).unwrap(),
|
||||||
|
vec![1, 2, 3]
|
||||||
|
);
|
||||||
|
|
||||||
|
// A VL sequence attribute.
|
||||||
|
let AttrValue::Raw { datatype, data, .. } = &attrs["vlen_attr"] else {
|
||||||
|
panic!("vlen_attr is Raw");
|
||||||
|
};
|
||||||
|
let clawhdf5_format::datatype::Datatype::VariableLength { base_type, .. } = datatype else {
|
||||||
|
panic!("vlen_attr is VL");
|
||||||
|
};
|
||||||
|
let got: Vec<String> = read_vl_bytes(file.as_bytes(), data, 2, 4, 4)
|
||||||
|
.unwrap()
|
||||||
|
.iter()
|
||||||
|
.map(|b| {
|
||||||
|
let v = read_as_i64(b, base_type).unwrap();
|
||||||
|
v.iter().map(i64::to_string).collect::<Vec<_>>().join(" ")
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
assert_eq!(got.join(";"), seqs);
|
||||||
|
}
|
||||||
+46
-6
@@ -164,13 +164,24 @@ fill-value item that did is fixed).
|
|||||||
is left out of `attrs()` (reported by `attrs_with_errors()`) instead of
|
is left out of `attrs()` (reported by `attrs_with_errors()`) instead of
|
||||||
failing the others.
|
failing the others.
|
||||||
- **Other readers:**
|
- **Other readers:**
|
||||||
- VL-string datasets are not readable through `File`.
|
- VL-string datasets are not readable through `File`. **Fixed
|
||||||
|
2026-09-26:** `read_string` reads them (also `read_string_bytes`,
|
||||||
|
`read_string_selection`, and on `MmapFile`/`LazyFile`), with h5py's
|
||||||
|
values: strings end at a NUL, null elements (heap address 0) are `""`,
|
||||||
|
and an element at the undefined heap address is an error as in libhdf5
|
||||||
|
(it read as `""` until 2026-09-26); VL sequences of
|
||||||
|
numbers read with `read_vlen::<T>()`, and VL values inside compounds or
|
||||||
|
`AttrValue::Raw` attributes decode with `File::decode_strings` /
|
||||||
|
`File::decode_vlen` (`crates/clawhdf5/tests/vl_data_interop.rs`).
|
||||||
- Variable-length values inside a compound (and VL-string attributes) in
|
- Variable-length values inside a compound (and VL-string attributes) in
|
||||||
a file with 4-byte offsets (`sizeof_addr = 4`) fail with
|
a file with 4-byte offsets (`sizeof_addr = 4`) fail with
|
||||||
`GlobalHeapObjectNotFound` or come back as `Raw`: these paths assume
|
`GlobalHeapObjectNotFound` or come back as `Raw`: these paths assume
|
||||||
the 16-byte element of an 8-byte-offset file. The datatype itself reads
|
the 16-byte element of an 8-byte-offset file. The datatype itself reads
|
||||||
(it was refused as "member overlaps with previous member" until
|
(it was refused as "member overlaps with previous member" until
|
||||||
2026-09-26).
|
2026-09-26). **Fixed 2026-09-26:** a VL type's element size is the one
|
||||||
|
its datatype message stores (12 with 4-byte offsets), and the global
|
||||||
|
heap is read with libhdf5's header padding
|
||||||
|
(`crates/clawhdf5/tests/vl_offset4_interop.rs`).
|
||||||
- Metadata cache images are not supported.
|
- Metadata cache images are not supported.
|
||||||
- x87 long double and binary128 are refused.
|
- x87 long double and binary128 are refused.
|
||||||
- N-Bit on 64-bit scale-offset data and some N-Bit parameter layouts fail.
|
- N-Bit on 64-bit scale-offset data and some N-Bit parameter layouts fail.
|
||||||
@@ -222,9 +233,10 @@ fill-value item that did is fixed).
|
|||||||
- (`cve-2024-32616` `/group1/dset3` and `cve-2025-2309`'s `Comp_OBJREF`
|
- (`cve-2024-32616` `/group1/dset3` and `cve-2025-2309`'s `Comp_OBJREF`
|
||||||
attribute are h5py/numpy type-mapping failures, not libhdf5 refusals.)
|
attribute are h5py/numpy type-mapping failures, not libhdf5 refusals.)
|
||||||
- `h5rs check` validates with the library's parsers, so it inherits what
|
- `h5rs check` validates with the library's parsers, so it inherits what
|
||||||
they accept: of the 150 CVE and fuzzer files, `check --data` passes 16,
|
they accept: of the 150 CVE and fuzzer files, `check --data` passes 15,
|
||||||
and h5dump 1.14.6 rejects 9 of those (tank, 2026-09-26; 28 and 21
|
and h5dump 1.14.6 rejects 8 of those (tank, 2026-09-26; 28 and 21
|
||||||
before these checks).
|
before these checks, 16 and 9 before a VL type's stored element size
|
||||||
|
was checked, which flags `cve-2024-32608`).
|
||||||
- **Writer:**
|
- **Writer:**
|
||||||
- Nested groups beyond one level: path-like names are now refused, not
|
- Nested groups beyond one level: path-like names are now refused, not
|
||||||
created.
|
created.
|
||||||
@@ -424,6 +436,32 @@ has produced more records than the file could physically hold.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Crafted global heaps exhaust the variable-length reader's memory
|
||||||
|
|
||||||
|
**Status:** fixed on `feat/p2-vl-strings` (2026-09-26). Not a regression of
|
||||||
|
that branch: every earlier release is affected through `read_vl_strings`.
|
||||||
|
|
||||||
|
Reading variable-length values kept an owned copy of every object of every
|
||||||
|
global heap collection visited, for the whole read. A file whose collections
|
||||||
|
nest inside one another's object data (32 bytes apart, each element pointing
|
||||||
|
at a different one) made retained memory O(elements × file size): a 744 KB
|
||||||
|
file reached 1.58 GB. Letting every collection's object chain jump to one
|
||||||
|
shared run of tiny objects made the parse time O(elements × objects) too.
|
||||||
|
libhdf5 refuses such files.
|
||||||
|
|
||||||
|
Now `VlResolver` caches where each object lies instead of a copy, drops its
|
||||||
|
cache past a 32 MiB budget, and refuses a collection that overlaps one it
|
||||||
|
has already read (libhdf5 gives each collection its own block, so only a
|
||||||
|
crafted file has them). `GlobalHeapCollection::parse` (and the new
|
||||||
|
`parse_index`) also refuse a collection that runs past the end of the file,
|
||||||
|
or an object that runs past the end of its collection. Guarded by
|
||||||
|
`crates/clawhdf5-format/tests/vl_heap_bounds.rs`, which measures peak heap
|
||||||
|
use with a counting allocator. Still open: a file may point many elements
|
||||||
|
at one large heap object, and a VL-*sequence* read then returns that
|
||||||
|
object once per element, as h5py would.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## Extensible Array chunk indexes read back wrong data past the inline elements
|
## Extensible Array chunk indexes read back wrong data past the inline elements
|
||||||
|
|
||||||
**Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to
|
**Status:** fixed on `main` (2026-09-20), after v2.6.0. **Every release up to
|
||||||
@@ -526,7 +564,9 @@ which is what libhdf5 itself writes.
|
|||||||
followed (no file system).
|
followed (no file system).
|
||||||
- Variable-length string datasets are read by decoding `read_selection`'s
|
- Variable-length string datasets are read by decoding `read_selection`'s
|
||||||
bytes with `clawhdf5_format::vl_data` in the wasm crate; `File` itself still
|
bytes with `clawhdf5_format::vl_data` in the wasm crate; `File` itself still
|
||||||
cannot (see the audit gaps above).
|
cannot (see the audit gaps above). (`File` can since 2026-09-26. Since
|
||||||
|
2026-09-26 the wasm crate resolves them with the same `VlResolver` as
|
||||||
|
`File` and `h5rs`, so all three return h5py's values.)
|
||||||
|
|
||||||
## The Node.js package (`packages/clawhdf5-node`) does not work
|
## The Node.js package (`packages/clawhdf5-node`) does not work
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user