Merge branch 'feat/p3-storage-trait' into feat/p3-range-zfp-edit

# Conflicts:
#	CHANGELOG.md
#	crates/clawhdf5-format/src/attribute.rs
#	crates/clawhdf5-format/src/btree_v1.rs
#	crates/clawhdf5-format/src/data_layout.rs
#	crates/clawhdf5-format/src/extensible_array.rs
#	crates/clawhdf5-format/src/fixed_array.rs
#	crates/clawhdf5-format/src/fractal_heap.rs
#	crates/clawhdf5-format/src/local_heap.rs
#	crates/clawhdf5-format/src/shared_message.rs
This commit is contained in:
osobh
2026-09-26 14:51:51 -05:00
26 changed files with 3480 additions and 497 deletions
+91
View File
@@ -58,6 +58,97 @@
truncating cast. The facade, `clawhdf5-io` and `clawhdf5-ann` are not truncating cast. The facade, `clawhdf5-io` and `clawhdf5-ann` are not
covered. covered.
### Range reads, milestone M1: the `Storage` trait (2026-09-26)
- **Breaking: `clawhdf5_format::error::FormatError` and `clawhdf5::Error`
are now `#[non_exhaustive]`.** An exhaustive `match` on either needs a
wildcard arm. `FormatError` has two new variants: `Storage(String)` (a
storage backend failed to serve a read) and
`ContiguousStorageRequired(&'static str)` (an operation not yet converted
to range reads was asked of a backend without the whole file in memory).
- New `clawhdf5_format::storage::Storage`, the synchronous, `no_std` read
interface of `docs/design/range-reads.md` option (a): `read_at(offset:
u64, len) -> Cow<[u8]>`, `read_ranges`, `len()` and an `as_contiguous()`
fast path; implemented for `[u8]`, `Vec<u8>`, and references, `Box`es
and (with `std`) `Arc`s of a `Storage`. Slices and `Vec`s serve borrowed
bytes, so parsing an in-memory file costs no copy.
- **The metadata parsers read through `Storage`.** Each converted parser has
an `*_in<S: Storage + ?Sized>(&S, ..)` core (a `&dyn Storage` works too),
and its `&[u8]` function is now a thin wrapper over it, so no caller
changes. The wrappers compile to a `[u8]` instance of the same code, so a
structure read in memory is a bounds check and a borrowed slice, with no
indirect call and no copy. Converted: the superblock
(`Superblock::parse_in`), its extension and cache image
(`read_superblock_extension_in`, `cache_image_state_in`), object headers
with their continuation chunks (`ObjectHeader::parse_in`), local and
global heaps, symbol-table nodes and the group B-tree (v1), fractal heaps,
fixed and extensible array chunk indexes, shared messages and the SOHM
table (`message_data_in`, `message_data_with_sohm_in`,
`load_sohm_table_in`, …), attributes (`extract_attributes_full_in`,
`extract_attributes_tolerant_in`, `AttributeMessage::parse_in_storage`),
fill values (`dataset_fill_value_from_storage`) and virtual-dataset
mappings (`DataLayout::resolve_vds_mappings_in`); also
`signature::find_signature_in`. Each structure is read with bounded reads
(a prefix, then the structure) instead of slicing the whole file; the
open-ended `&file_data[addr..]` slices in these modules are gone. Bounds
errors keep their values (absolute position, file length).
- Reads sized by untrusted fields are bounded by what the parser uses, so
a crafted size cannot turn one structure into a read of the rest of the
file on a range backend: local-heap names are read in growing pieces
(64 bytes first) rather than to the end of the data segment; a fractal
heap indirect block is read up to the entry covering the object (the
whole block only when that entry is unallocated); paged fixed and
extensible array data blocks over 1 MiB are read page by page, only the
pages in use; and a block under one checksum whose claimed size runs
past the end of the file fails its bounds check before any read (with
the `checksum` feature). An object header's prefix is one read (was
two).
- Structures still indexed by a v2 B-tree — dense attribute storage, a SOHM
B-tree index and huge fractal-heap objects found through their B-tree —
are not converted yet (the
v2 B-tree and dense groups come with milestone M3); over a backend without
the whole file in memory they are the clean `ContiguousStorageRequired`
error, never a partial result. Raw data, chunk B-tree (v1) indexes and VL
data are milestone M2.
- **No behaviour change**, checked three ways (2026-09-26, tank): every
existing test passes unchanged; the conformance sweep
(`conformance/run.sh --no-fetch`) gives a byte-identical `results.json`
at `f2ff2c4` and on this branch, and identical per-file probe output for
696 of the 697 files — the exception, `cve-2025-2310.h5`, reports one of
two errors depending on which parallel chunk decode fails first, at
`f2ff2c4` as on this branch; and a transcript of every converted `&[u8]`
function's result over the fixtures, the conformance corpus and the
h5py-written files below (748 files, 7 603 object headers) is
byte-identical between the two builds.
- **Speed on local files** (provisional: tank was shared with other jobs;
both builds linked into one binary and timed alternately, 200 rounds;
new Criterion bench `clawhdf5/benches/local_metadata_bench.rs` over a
400-group version-1 file, `clawhdf5-format/tests/fixtures/v1_groups_400.h5`):
against `f2ff2c4`, listing the file through the facade is 2.7% faster,
`ObjectHeader::parse` is within ±1%, symbol-table nodes and the group
B-tree walk are about 19% faster (their entry loops were tightened),
local-heap names and `resolve_group_children` 1.5–3% faster. The same
harness run on two copies of the old code differs by up to 2%. The
Criterion bench itself, old and new as separate binaries run alternately
(3 rounds), agrees except for `ObjectHeader::parse`, which it puts about
7% slower (25.9 vs 24.1 µs for 401 headers) while the listing that
parses those headers is 5–8% faster; that one remains unexplained and is
to be rechecked on an idle machine.
- New equivalence harness `clawhdf5-format/tests/storage_equivalence.rs`:
every converted parser runs over the file as a slice and over
`storage::CountingStorage` — a `Storage` that serves an in-memory buffer
through `read_at` only (`as_contiguous()` is `None`), copying what it
serves and counting reads — and must give identical results. It walks
the fixtures, files h5py writes for it (extensible arrays with super
blocks and paged data blocks, paged fixed arrays, large v1 and dense
groups, a user block, SOHM list and B-tree indexes, dense, shared and
committed-type attributes, fixed and extensible array data blocks over
1 MiB, whole and truncated), and with `CLAWHDF5_STORAGE_CORPUS=<dir>` a
corpus (all 653 HDF5 files of the conformance corpus pass). Only the
structures listed above as not converted may answer
`ContiguousStorageRequired`, and only in the checks that reach them; a
converted parser falling back to the whole file fails it. Milestones M2
and M3 extend it.
### Chunked full reads (2026-09-26) ### Chunked full reads (2026-09-26)
- **Chunks are decoded straight into the output, into reused buffers.** A - **Chunks are decoded straight into the output, into reused buffers.** A
full read of a chunked dataset faulted in about three times its size in full read of a chunked dataset faulted in about three times its size in
+158 -33
View File
@@ -17,6 +17,7 @@ use crate::fractal_heap::FractalHeapHeader;
use crate::message_type::MessageType; use crate::message_type::MessageType;
use crate::object_header::ObjectHeader; use crate::object_header::ObjectHeader;
use crate::shared_message; use crate::shared_message;
use crate::storage::{Storage, require_contiguous};
use crate::vl_data; use crate::vl_data;
/// A parsed HDF5 attribute message. /// A parsed HDF5 attribute message.
@@ -52,7 +53,7 @@ impl AttributeMessage {
/// ///
/// `length_size` is needed for dataspace dimension parsing. /// `length_size` is needed for dataspace dimension parsing.
pub fn parse(data: &[u8], length_size: u8) -> Result<AttributeMessage, FormatError> { pub fn parse(data: &[u8], length_size: u8) -> Result<AttributeMessage, FormatError> {
Self::parse_impl(data, length_size, None) Self::parse_impl(data, length_size, None::<(&[u8], u8)>)
} }
/// [`AttributeMessage::parse`] with access to the rest of the file, which /// [`AttributeMessage::parse`] with access to the rest of the file, which
@@ -67,13 +68,24 @@ impl AttributeMessage {
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<AttributeMessage, FormatError> { ) -> Result<AttributeMessage, FormatError> {
Self::parse_impl(data, length_size, Some((file_data, offset_size))) Self::parse_in_storage(data, file_data, offset_size, length_size)
} }
fn parse_impl( /// [`AttributeMessage::parse_in_file`] with the file behind any
/// [`Storage`].
pub fn parse_in_storage<S: Storage + ?Sized>(
data: &[u8],
file: &S,
offset_size: u8,
length_size: u8,
) -> Result<AttributeMessage, FormatError> {
Self::parse_impl(data, length_size, Some((file, offset_size)))
}
fn parse_impl<S: Storage + ?Sized>(
data: &[u8], data: &[u8],
length_size: u8, length_size: u8,
file: Option<(&[u8], u8)>, file: Option<(&S, u8)>,
) -> Result<AttributeMessage, FormatError> { ) -> Result<AttributeMessage, FormatError> {
ensure_len(data, 0, 2)?; ensure_len(data, 0, 2)?;
let version = data[0]; let version = data[0];
@@ -88,19 +100,19 @@ impl AttributeMessage {
/// The bytes of an embedded datatype/dataspace message, following the /// The bytes of an embedded datatype/dataspace message, following the
/// shared-message reference when `shared` is set. /// shared-message reference when `shared` is set.
fn embedded_message<'a>( fn embedded_message<'a, S: Storage + ?Sized>(
bytes: &'a [u8], bytes: &'a [u8],
shared: bool, shared: bool,
msg_type: MessageType, msg_type: MessageType,
length_size: u8, length_size: u8,
file: Option<(&[u8], u8)>, file: Option<(&S, u8)>,
) -> Result<Cow<'a, [u8]>, FormatError> { ) -> Result<Cow<'a, [u8]>, FormatError> {
if !shared { if !shared {
return Ok(Cow::Borrowed(bytes)); return Ok(Cow::Borrowed(bytes));
} }
let (file_data, offset_size) = file.ok_or(FormatError::UnresolvedSharedMessage)?; let (file_data, offset_size) = file.ok_or(FormatError::UnresolvedSharedMessage)?;
let shared_ref = shared_message::parse_shared_ref_sized(bytes, offset_size, length_size)?; let shared_ref = shared_message::parse_shared_ref_sized(bytes, offset_size, length_size)?;
shared_message::resolve_shared_message( shared_message::resolve_shared_message_in(
file_data, file_data,
&shared_ref, &shared_ref,
msg_type, msg_type,
@@ -145,10 +157,10 @@ impl AttributeMessage {
}) })
} }
fn parse_v2( fn parse_v2<S: Storage + ?Sized>(
data: &[u8], data: &[u8],
length_size: u8, length_size: u8,
file: Option<(&[u8], u8)>, file: Option<(&S, u8)>,
) -> Result<AttributeMessage, FormatError> { ) -> Result<AttributeMessage, FormatError> {
// Flags: bit 0 = datatype is shared, bit 1 = dataspace is shared. // Flags: bit 0 = datatype is shared, bit 1 = dataspace is shared.
let flags = data.get(1).copied().unwrap_or(0); let flags = data.get(1).copied().unwrap_or(0);
@@ -199,10 +211,10 @@ impl AttributeMessage {
}) })
} }
fn parse_v3( fn parse_v3<S: Storage + ?Sized>(
data: &[u8], data: &[u8],
length_size: u8, length_size: u8,
file: Option<(&[u8], u8)>, file: Option<(&S, u8)>,
) -> Result<AttributeMessage, FormatError> { ) -> Result<AttributeMessage, FormatError> {
// Flags: bit 0 = datatype is shared, bit 1 = dataspace is shared. // Flags: bit 0 = datatype is shared, bit 1 = dataspace is shared.
let flags = data.get(1).copied().unwrap_or(0); let flags = data.get(1).copied().unwrap_or(0);
@@ -418,7 +430,20 @@ pub fn extract_attributes_full(
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Vec<AttributeMessage>, FormatError> { ) -> Result<Vec<AttributeMessage>, FormatError> {
extract_attributes_with(file_data, header, offset_size, length_size, &mut Err) extract_attributes_full_in(file_data, header, offset_size, length_size)
}
/// [`extract_attributes_full`] over any [`Storage`]. Dense attribute
/// storage is indexed by a v2 B-tree, which is not read over [`Storage`]
/// yet: on a backend without the whole file in memory an object with dense
/// attributes is [`FormatError::ContiguousStorageRequired`].
pub fn extract_attributes_full_in<S: Storage + ?Sized>(
file: &S,
header: &ObjectHeader,
offset_size: u8,
length_size: u8,
) -> Result<Vec<AttributeMessage>, FormatError> {
extract_attributes_with(file, header, offset_size, length_size, &mut Err)
} }
/// Like [`extract_attributes_full`], but an attribute that cannot be read /// Like [`extract_attributes_full`], but an attribute that cannot be read
@@ -434,6 +459,17 @@ pub fn extract_attributes_tolerant(
header: &ObjectHeader, header: &ObjectHeader,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<(Vec<AttributeMessage>, Vec<FormatError>), FormatError> {
extract_attributes_tolerant_in(file_data, header, offset_size, length_size)
}
/// [`extract_attributes_tolerant`] over any [`Storage`] (see
/// [`extract_attributes_full_in`] for dense storage).
pub fn extract_attributes_tolerant_in<S: Storage + ?Sized>(
file_data: &S,
header: &ObjectHeader,
offset_size: u8,
length_size: u8,
) -> Result<(Vec<AttributeMessage>, Vec<FormatError>), FormatError> { ) -> Result<(Vec<AttributeMessage>, Vec<FormatError>), FormatError> {
let mut errors = Vec::new(); let mut errors = Vec::new();
let attrs = extract_attributes_with(file_data, header, offset_size, length_size, &mut |e| { let attrs = extract_attributes_with(file_data, header, offset_size, length_size, &mut |e| {
@@ -445,8 +481,8 @@ pub fn extract_attributes_tolerant(
/// Read every attribute; each one that fails goes to `on_error`, which /// Read every attribute; each one that fails goes to `on_error`, which
/// either stops the read (returns the error) or skips that attribute. /// either stops the read (returns the error) or skips that attribute.
fn extract_attributes_with( fn extract_attributes_with<S: Storage + ?Sized>(
file_data: &[u8], file_data: &S,
header: &ObjectHeader, header: &ObjectHeader,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
@@ -514,6 +550,19 @@ pub fn find_attribute_in_file(
name: &str, name: &str,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Option<AttributeMessage>, FormatError> {
find_attribute_in(file_data, header, name, offset_size, length_size)
}
/// [`find_attribute_in_file`] over any [`Storage`] (see
/// [`extract_attributes_full_in`] for dense storage, whose name index still
/// needs the whole file in memory).
pub fn find_attribute_in<S: Storage + ?Sized>(
file_data: &S,
header: &ObjectHeader,
name: &str,
offset_size: u8,
length_size: u8,
) -> Result<Option<AttributeMessage>, FormatError> { ) -> Result<Option<AttributeMessage>, FormatError> {
let attr_info = find_attribute_info(header, offset_size)?; let attr_info = find_attribute_info(header, offset_size)?;
let dense = attr_info let dense = attr_info
@@ -523,18 +572,19 @@ pub fn find_attribute_in_file(
// Compact only (or dense storage without a name index, which a // Compact only (or dense storage without a name index, which a
// listing reports): as a listing finds it. // listing reports): as a listing finds it.
return Ok( return Ok(
extract_attributes_tolerant(file_data, header, offset_size, length_size)? extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?
.0 .0
.into_iter() .into_iter()
.find(|a| a.name == name), .find(|a| a.name == name),
); );
}; };
let contiguous = require_contiguous(file_data, "dense attribute storage (a v2 B-tree)")?;
let btree_hdr = let btree_hdr =
BTreeV2Header::parse(file_data, to_usize(btree_addr)?, offset_size, length_size)?; BTreeV2Header::parse(contiguous, to_usize(btree_addr)?, offset_size, length_size)?;
let fh = FractalHeapHeader::parse(file_data, to_usize(fh_addr)?, offset_size, length_size)?; let fh = FractalHeapHeader::parse_in(file_data, fh_addr, offset_size, length_size)?;
if btree_hdr.tree_type != ATTRIBUTE_NAME_INDEX || btree_hdr.record_size < 4 { if btree_hdr.tree_type != ATTRIBUTE_NAME_INDEX || btree_hdr.record_size < 4 {
return Ok( return Ok(
extract_attributes_tolerant(file_data, header, offset_size, length_size)? extract_attributes_tolerant_in(file_data, header, offset_size, length_size)?
.0 .0
.into_iter() .into_iter()
.find(|a| a.name == name), .find(|a| a.name == name),
@@ -560,7 +610,7 @@ pub fn find_attribute_in_file(
// hash is the last field. // hash is the last field.
let hash = jenkins_lookup3(name.as_bytes()); let hash = jenkins_lookup3(name.as_bytes());
let hash_at = usize::from(btree_hdr.record_size) - 4; let hash_at = usize::from(btree_hdr.record_size) - 4;
let records = find_btree_v2_records(file_data, &btree_hdr, offset_size, &mut |r| match r let records = find_btree_v2_records(contiguous, &btree_hdr, offset_size, &mut |r| match r
.get(hash_at..hash_at + 4) .get(hash_at..hash_at + 4)
{ {
Some(h) => u32::from_le_bytes([h[0], h[1], h[2], h[3]]).cmp(&hash), Some(h) => u32::from_le_bytes([h[0], h[1], h[2], h[3]]).cmp(&hash),
@@ -572,8 +622,10 @@ pub fn find_attribute_in_file(
continue; continue;
}; };
let attr = fh let attr = fh
.read_managed_object(file_data, id_bytes, offset_size) .read_managed_object_in(file_data, id_bytes, offset_size)
.and_then(|d| AttributeMessage::parse_in_file(&d, file_data, offset_size, length_size)); .and_then(|d| {
AttributeMessage::parse_in_storage(&d, file_data, offset_size, length_size)
});
// One that cannot be read is left out, as from a listing. // One that cannot be read is left out, as from a listing.
if let Ok(attr) = attr if let Ok(attr) = attr
&& attr.name == name && attr.name == name
@@ -586,8 +638,8 @@ pub fn find_attribute_in_file(
/// The attributes stored in the object header itself (compact storage), and /// The attributes stored in the object header itself (compact storage), and
/// each one's creation order into `orders`. /// each one's creation order into `orders`.
fn extract_compact_attributes( fn extract_compact_attributes<S: Storage + ?Sized>(
file_data: &[u8], file_data: &S,
header: &ObjectHeader, header: &ObjectHeader,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
@@ -601,7 +653,7 @@ fn extract_compact_attributes(
// Shared attribute: resolve the reference to get actual attribute data // Shared attribute: resolve the reference to get actual attribute data
shared_message::parse_shared_ref_sized(&msg.data, offset_size, length_size) shared_message::parse_shared_ref_sized(&msg.data, offset_size, length_size)
.and_then(|shared_ref| { .and_then(|shared_ref| {
shared_message::resolve_shared_message( shared_message::resolve_shared_message_in(
file_data, file_data,
&shared_ref, &shared_ref,
MessageType::Attribute, MessageType::Attribute,
@@ -610,7 +662,7 @@ fn extract_compact_attributes(
) )
}) })
.and_then(|resolved| { .and_then(|resolved| {
AttributeMessage::parse_in_file( AttributeMessage::parse_in_storage(
&resolved, &resolved,
file_data, file_data,
offset_size, offset_size,
@@ -618,7 +670,7 @@ fn extract_compact_attributes(
) )
}) })
} else { } else {
AttributeMessage::parse_in_file(&msg.data, file_data, offset_size, length_size) AttributeMessage::parse_in_storage(&msg.data, file_data, offset_size, length_size)
}; };
let attr = attr.and_then(|a| check_in_header(a, header)); let attr = attr.and_then(|a| check_in_header(a, header));
match attr { match attr {
@@ -650,8 +702,8 @@ fn find_attribute_info(
/// Extract attributes from dense storage (fractal heap + B-tree v2), and /// Extract attributes from dense storage (fractal heap + B-tree v2), and
/// each one's creation order into `orders`. /// each one's creation order into `orders`.
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
fn extract_dense_attributes( fn extract_dense_attributes<S: Storage + ?Sized>(
file_data: &[u8], file_data: &S,
attr_info: &AttributeInfoMessage, attr_info: &AttributeInfoMessage,
fh_addr: u64, fh_addr: u64,
offset_size: u8, offset_size: u8,
@@ -661,7 +713,7 @@ fn extract_dense_attributes(
on_error: &mut dyn FnMut(FormatError) -> Result<(), FormatError>, on_error: &mut dyn FnMut(FormatError) -> Result<(), FormatError>,
) -> Result<(), FormatError> { ) -> Result<(), FormatError> {
// Parse fractal heap // Parse fractal heap
let fh = FractalHeapHeader::parse(file_data, to_usize(fh_addr)?, offset_size, length_size)?; let fh = FractalHeapHeader::parse_in(file_data, fh_addr, offset_size, length_size)?;
// Parse B-tree v2 for name index (type 8) // Parse B-tree v2 for name index (type 8)
let btree_addr = attr_info let btree_addr = attr_info
@@ -670,9 +722,10 @@ fn extract_dense_attributes(
expected: 1, expected: 1,
available: 0, available: 0,
})?; })?;
let contiguous = require_contiguous(file_data, "dense attribute storage (a v2 B-tree)")?;
let btree_hdr = let btree_hdr =
BTreeV2Header::parse(file_data, to_usize(btree_addr)?, offset_size, length_size)?; BTreeV2Header::parse(contiguous, to_usize(btree_addr)?, offset_size, length_size)?;
let records = collect_btree_v2_records(file_data, &btree_hdr, offset_size, length_size)?; let records = collect_btree_v2_records(contiguous, &btree_hdr, offset_size, length_size)?;
for record in &records { for record in &records {
// Per HDF5 spec, both type 8 and type 9 records start with heap_id: // Per HDF5 spec, both type 8 and type 9 records start with heap_id:
@@ -689,9 +742,9 @@ fn extract_dense_attributes(
// The data in the heap is a complete attribute message // The data in the heap is a complete attribute message
let attr = fh let attr = fh
.read_managed_object(file_data, id_bytes, offset_size) .read_managed_object_in(file_data, id_bytes, offset_size)
.and_then(|attr_data| { .and_then(|attr_data| {
AttributeMessage::parse_in_file(&attr_data, file_data, offset_size, length_size) AttributeMessage::parse_in_storage(&attr_data, file_data, offset_size, length_size)
}); });
match attr { match attr {
Ok(attr) => { Ok(attr) => {
@@ -1101,4 +1154,76 @@ mod tests {
let strs = attr.read_as_strings().unwrap(); let strs = attr.read_as_strings().unwrap();
assert_eq!(strs, vec!["abcd", "EFGH"]); assert_eq!(strs, vec!["abcd", "EFGH"]);
} }
/// Every object's attributes in h5py-written files read identically
/// through a read_at-only CountingStorage — compact ones, shared ones
/// and those behind an Attribute Info message — except dense storage,
/// whose v2 B-tree index is not read over Storage yet: that is the clean
/// ContiguousStorageRequired error, never a partial list. Through a
/// slice as Storage every object matches.
#[test]
fn storage_reads_match_slice_reads() {
use crate::storage::CountingStorage;
let files: [(&str, &[u8]); 5] = [
("attrs", include_bytes!("../tests/fixtures/attrs.h5")),
(
"mixed_attrs",
include_bytes!("../tests/fixtures/mixed_attrs.h5"),
),
(
"dense_attrs",
include_bytes!("../tests/fixtures/dense_attrs.h5"),
),
(
"dense_attrs_root",
include_bytes!("../tests/fixtures/dense_attrs_root.h5"),
),
(
"shared_fill_value",
include_bytes!("../tests/fixtures/shared_fill_value.h5"),
),
];
let (mut same, mut dense, mut attrs) = (0, 0, 0);
for (name, file) in files {
let sb = crate::superblock::Superblock::parse(file, 0).unwrap();
let (os, ls) = (sb.offset_size, sb.length_size);
let mut addrs = vec![sb.root_group_address];
addrs.extend(
crate::group_v2::resolve_group_children(file, &sb, sb.root_group_address)
.unwrap()
.iter()
.map(|e| e.object_header_address),
);
let storage = CountingStorage::new(file.to_vec());
for addr in addrs {
let header = ObjectHeader::parse(file, addr as usize, os, ls).unwrap();
let want = extract_attributes_full(file, &header, os, ls);
let slice_storage = extract_attributes_full_in(&file, &header, os, ls);
assert_eq!(format!("{slice_storage:?}"), format!("{want:?}"));
let got = extract_attributes_full_in(&storage, &header, os, ls);
let got_t = extract_attributes_tolerant_in(&storage, &header, os, ls);
let is_dense = find_attribute_info(&header, os)
.unwrap()
.is_some_and(|i| i.fractal_heap_address.is_some());
if is_dense {
let e = FormatError::ContiguousStorageRequired(
"dense attribute storage (a v2 B-tree)",
);
assert_eq!(got.unwrap_err(), e, "{name}");
assert_eq!(got_t.unwrap_err(), e, "{name}");
dense += 1;
} else {
attrs += want.as_ref().map_or(0, Vec::len);
assert_eq!(format!("{got:?}"), format!("{want:?}"), "{name}");
let want_t = extract_attributes_tolerant(file, &header, os, ls);
assert_eq!(format!("{got_t:?}"), format!("{want_t:?}"), "{name}");
same += 1;
}
}
}
assert!(
same >= 5 && dense >= 2 && attrs >= 5,
"{same} {dense} {attrs}"
);
}
} }
+89 -26
View File
@@ -3,8 +3,8 @@
#[cfg(not(feature = "std"))] #[cfg(not(feature = "std"))]
use alloc::vec::Vec; use alloc::vec::Vec;
use crate::addr::to_usize;
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::{Storage, read_exact_at};
/// A parsed B-tree v1 node. /// A parsed B-tree v1 node.
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
@@ -75,13 +75,28 @@ impl BTreeV1Node {
file_data: &[u8], file_data: &[u8],
offset: usize, offset: usize,
offset_size: u8, offset_size: u8,
length_size: u8,
) -> Result<BTreeV1Node, FormatError> {
Self::parse_in(file_data, offset as u64, offset_size, length_size)
}
/// [`Self::parse`] over any [`Storage`]: one read of the node's header,
/// one of its keys and children.
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
offset_size: u8,
_length_size: u8, _length_size: u8,
) -> Result<BTreeV1Node, FormatError> { ) -> Result<BTreeV1Node, FormatError> {
// signature(4) + node_type(1) + node_level(1) + entries_used(2) = 8 // signature(4) + node_type(1) + node_level(1) + entries_used(2) = 8
// + left_sibling(offset_size) + right_sibling(offset_size) // + left_sibling(offset_size) + right_sibling(offset_size)
let os = offset_size as usize; let os = offset_size as usize;
let header_size = 8 + os * 2; let header_size = 8 + os * 2;
ensure_len(file_data, offset, header_size)?; let header = read_exact_at(file, offset, header_size)?;
let file_data: &[u8] = &header;
// The header's read checked that `offset + header_size` fits.
let body_start = offset + header_size as u64;
let offset = 0usize;
if &file_data[offset..offset + 4] != b"TREE" { if &file_data[offset..offset + 4] != b"TREE" {
return Err(FormatError::InvalidBTreeSignature); return Err(FormatError::InvalidBTreeSignature);
@@ -103,31 +118,30 @@ impl BTreeV1Node {
} else { } else {
Some(read_offset(file_data, pos, offset_size)?) Some(read_offset(file_data, pos, offset_size)?)
}; };
pos += os;
// For type 0: keys are offset_size bytes, children are offset_size bytes // For type 0: keys are offset_size bytes, children are offset_size bytes
// Layout: key[0], child[0], key[1], child[1], ..., key[N-1], child[N-1], key[N] // Layout: key[0], child[0], key[1], child[1], ..., key[N-1], child[N-1], key[N]
let eu = entries_used as usize; let eu = entries_used as usize;
let key_size = os; // For type 0, key = offset_size let key_size = os; // For type 0, key = offset_size
let needed = eu * (key_size + os) + key_size; // eu children + (eu+1) keys let needed = eu * (key_size + os) + key_size; // eu children + (eu+1) keys
ensure_len(file_data, pos, needed)?; let body = read_exact_at(file, body_start, needed)?;
let file_data: &[u8] = &body;
let mut keys = Vec::with_capacity(eu + 1); let mut keys = Vec::with_capacity(eu + 1);
let mut children = Vec::with_capacity(eu); let mut children = Vec::with_capacity(eu);
for _i in 0..eu { if os == 0 {
// key[i] // What reading the first key reports (and keeps `chunks_exact`
let key = read_offset(file_data, pos, offset_size)?; // below from being given a zero size).
keys.push(key); return Err(FormatError::InvalidOffsetSize(offset_size));
pos += key_size;
// child[i]
let child = read_offset(file_data, pos, offset_size)?;
children.push(child);
pos += os;
} }
// final key // `needed` bytes: key[0], child[0], ..., child[eu - 1], key[eu].
let key = read_offset(file_data, pos, offset_size)?; let (pairs, last) = file_data.split_at(eu * (key_size + os));
keys.push(key); for pair in pairs.chunks_exact(key_size + os) {
keys.push(read_offset(pair, 0, offset_size)?);
children.push(read_offset(pair, key_size, offset_size)?);
}
keys.push(read_offset(last, 0, offset_size)?);
Ok(BTreeV1Node { Ok(BTreeV1Node {
node_type, node_type,
@@ -151,11 +165,21 @@ pub fn collect_symbol_table_nodes(
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Vec<u64>, FormatError> { ) -> Result<Vec<u64>, FormatError> {
collect_symbol_table_nodes_inner(file_data, btree_address, offset_size, length_size, 0) collect_symbol_table_nodes_in(file_data, btree_address, offset_size, length_size)
} }
fn collect_symbol_table_nodes_inner( /// [`collect_symbol_table_nodes`] over any [`Storage`]: two reads per node.
file_data: &[u8], pub fn collect_symbol_table_nodes_in<S: Storage + ?Sized>(
file: &S,
btree_address: u64,
offset_size: u8,
length_size: u8,
) -> Result<Vec<u64>, FormatError> {
collect_symbol_table_nodes_inner(file, btree_address, offset_size, length_size, 0)
}
fn collect_symbol_table_nodes_inner<S: Storage + ?Sized>(
file: &S,
btree_address: u64, btree_address: u64,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
@@ -165,12 +189,7 @@ fn collect_symbol_table_nodes_inner(
return Err(FormatError::NestingDepthExceeded); return Err(FormatError::NestingDepthExceeded);
} }
let node = BTreeV1Node::parse( let node = BTreeV1Node::parse_in(file, btree_address, offset_size, length_size)?;
file_data,
to_usize(btree_address)?,
offset_size,
length_size,
)?;
if node.node_type != 0 { if node.node_type != 0 {
return Err(FormatError::InvalidBTreeNodeType(node.node_type)); return Err(FormatError::InvalidBTreeNodeType(node.node_type));
@@ -184,7 +203,7 @@ fn collect_symbol_table_nodes_inner(
let mut result = Vec::new(); let mut result = Vec::new();
for &child_addr in &node.children { for &child_addr in &node.children {
let child_snods = collect_symbol_table_nodes_inner( let child_snods = collect_symbol_table_nodes_inner(
file_data, file,
child_addr, child_addr,
offset_size, offset_size,
length_size, length_size,
@@ -323,4 +342,48 @@ mod tests {
assert_eq!(node.entries_used, 1); assert_eq!(node.entries_used, 1);
assert_eq!(node.children, vec![0x50]); assert_eq!(node.children, vec![0x50]);
} }
/// Nodes and trees, cut at every length, parse identically through a
/// `read_at`-only storage.
#[test]
fn storage_parse_matches_slice_parse() {
use crate::storage::CountingStorage;
let nodes = [
build_btree_node(0, 0, &[0, 5, 10], &[0x100, 0x200], None, None, 8),
build_btree_node(0, 0, &[0, 5], &[0x100], Some(0x40), Some(0x80), 4),
build_btree_node(1, 2, &[0, 5], &[0x100], None, Some(0x80), 8),
];
for (n, node) in nodes.iter().enumerate() {
let os = if n == 1 { 4 } else { 8 };
for cut in 0..=node.len() {
let f = &node[..cut];
let storage = CountingStorage::new(f.to_vec());
let want = BTreeV1Node::parse(f, 0, os, 8);
let got = BTreeV1Node::parse_in(&storage, 0, os, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
}
}
let leaf1 = build_btree_node(0, 0, &[0, 5], &[0xA00], None, None, 8);
let leaf2 = build_btree_node(0, 0, &[5, 10], &[0xB00], None, None, 8);
let internal = build_btree_node(0, 1, &[0, 5, 10], &[0, 256], None, None, 8);
let mut file = vec![0u8; 512 + internal.len()];
file[..leaf1.len()].copy_from_slice(&leaf1);
file[256..256 + leaf2.len()].copy_from_slice(&leaf2);
file[512..].copy_from_slice(&internal);
for cut in [file.len(), 300, 260, 100, 10] {
let mut f = file.clone();
if cut < 512 {
// Truncate the leaves, keep the root.
f[cut..512].fill(0);
}
let storage = CountingStorage::new(f.clone());
assert_eq!(
collect_symbol_table_nodes_in(&storage, 512, 8, 8),
collect_symbol_table_nodes(&f, 512, 8, 8)
);
}
let storage = CountingStorage::new(file);
collect_symbol_table_nodes_in(&storage, 512, 8, 8).unwrap();
assert_eq!(storage.reads(), 6);
}
} }
+53 -5
View File
@@ -8,6 +8,7 @@ use std::string::String;
use crate::addr::to_usize; use crate::addr::to_usize;
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::Storage;
/// A single VDS (Virtual Dataset) source mapping. /// A single VDS (Virtual Dataset) source mapping.
/// ///
@@ -310,6 +311,16 @@ impl DataLayout {
&mut self, &mut self,
file_data: &[u8], file_data: &[u8],
length_size: u8, length_size: u8,
) -> Result<(), FormatError> {
self.resolve_vds_mappings_in(file_data, length_size)
}
/// [`Self::resolve_vds_mappings`] over any [`Storage`]: one read of the
/// global heap collection holding the mappings.
pub fn resolve_vds_mappings_in<S: Storage + ?Sized>(
&mut self,
file_data: &S,
length_size: u8,
) -> Result<(), FormatError> { ) -> Result<(), FormatError> {
if let DataLayout::Virtual { if let DataLayout::Virtual {
global_heap_address, global_heap_address,
@@ -319,11 +330,8 @@ impl DataLayout {
} = self } = self
&& let Some(addr) = *global_heap_address && let Some(addr) = *global_heap_address
{ {
let coll = crate::global_heap::GlobalHeapCollection::parse( let coll =
file_data, crate::global_heap::GlobalHeapCollection::parse_in(file_data, addr, length_size)?;
to_usize(addr)?,
length_size,
)?;
let obj = coll.get_object(*global_heap_index as u16).ok_or( let obj = coll.get_object(*global_heap_index as u16).ok_or(
FormatError::GlobalHeapObjectNotFound { FormatError::GlobalHeapObjectNotFound {
collection_address: addr, collection_address: addr,
@@ -1306,4 +1314,44 @@ mod tests {
let blob = [0x01u8, 0, 0, 0, 0, 0, 0, 0, 0]; let blob = [0x01u8, 0, 0, 0, 0, 0, 0, 0, 0];
assert!(parse_vds_mappings(&blob, 8).unwrap().is_empty()); assert!(parse_vds_mappings(&blob, 8).unwrap().is_empty());
} }
/// A virtual dataset's mappings resolve identically through a
/// read_at-only CountingStorage, in two reads of the global heap.
#[test]
fn vds_mappings_through_storage_match_slice() {
use crate::message_type::MessageType;
use crate::object_header::ObjectHeader;
use crate::storage::CountingStorage;
let file: &[u8] = include_bytes!("../tests/fixtures/vds_same_file.h5");
let sb = crate::superblock::Superblock::parse(file, 0).unwrap();
let (os, ls) = (sb.offset_size, sb.length_size);
let storage = CountingStorage::new(file.to_vec());
let mut virtuals = 0;
for child in
crate::group_v2::resolve_group_children(file, &sb, sb.root_group_address).unwrap()
{
let h =
ObjectHeader::parse(file, child.object_header_address as usize, os, ls).unwrap();
let Some(msg) = h
.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
else {
continue;
};
let mut want = DataLayout::parse(&msg.data, os, ls).unwrap();
if !matches!(want, DataLayout::Virtual { .. }) {
continue;
}
let mut got = want.clone();
want.resolve_vds_mappings(file, ls).unwrap();
storage.reset();
got.resolve_vds_mappings_in(&storage, ls).unwrap();
assert_eq!(format!("{got:?}"), format!("{want:?}"));
assert!(matches!(&got, DataLayout::Virtual { mappings, .. } if !mappings.is_empty()));
assert_eq!(storage.reads(), 2);
virtuals += 1;
}
assert!(virtuals >= 1);
}
} }
+21
View File
@@ -12,7 +12,11 @@ use std::string::String;
use core::fmt; use core::fmt;
/// Errors that can occur when parsing HDF5 binary format structures. /// Errors that can occur when parsing HDF5 binary format structures.
///
/// Non-exhaustive: new failure modes (new storage backends, new file
/// features) add variants, so a `match` needs a wildcard arm.
#[derive(Debug, Clone, PartialEq, Eq)] #[derive(Debug, Clone, PartialEq, Eq)]
#[non_exhaustive]
pub enum FormatError { pub enum FormatError {
/// The HDF5 magic signature was not found at any valid offset. /// The HDF5 magic signature was not found at any valid offset.
SignatureNotFound, SignatureNotFound,
@@ -243,6 +247,13 @@ pub enum FormatError {
/// A metadata cache image block libhdf5 refuses to load (the reason is /// A metadata cache image block libhdf5 refuses to load (the reason is
/// libhdf5's own error text). /// libhdf5's own error text).
InvalidCacheImage(&'static str), InvalidCacheImage(&'static str),
/// The [`Storage`](crate::storage::Storage) backend failed to serve a
/// read (an I/O or network error, or a short read inside the file).
Storage(String),
/// The operation still needs the whole file as one slice and the
/// [`Storage`](crate::storage::Storage) backend has no contiguous view
/// (`as_contiguous()` is `None`); the text names the operation.
ContiguousStorageRequired(&'static str),
} }
impl fmt::Display for FormatError { impl fmt::Display for FormatError {
@@ -529,6 +540,16 @@ impl fmt::Display for FormatError {
FormatError::InvalidCacheImage(why) => { FormatError::InvalidCacheImage(why) => {
write!(f, "invalid metadata cache image: {why}") write!(f, "invalid metadata cache image: {why}")
} }
FormatError::Storage(why) => {
write!(f, "storage read failed: {why}")
}
FormatError::ContiguousStorageRequired(what) => {
write!(
f,
"{what} needs the whole file in memory, which this storage backend does \
not provide"
)
}
} }
} }
} }
+215 -99
View File
@@ -13,16 +13,19 @@ use crate::addr::to_usize;
use crate::chunk_grid::ChunkGrid; use crate::chunk_grid::ChunkGrid;
use crate::chunked_read::ChunkInfo; use crate::chunked_read::ChunkInfo;
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::{PAGED_BLOCK_ONE_READ_MAX, Storage, Window, read_exact_at};
/// Verify the Jenkins lookup3 checksum stored immediately after /// Verify the Jenkins lookup3 checksum stored immediately after
/// `data[start..end]`, as every Extensible Array structure carries one. /// `data[start..end]`, as every Extensible Array structure carries one. `w`
/// is a window of the file and `start`/`end` are relative to it.
/// ///
/// A corrupt chunk index yields addresses pointing at the wrong bytes, so a /// A corrupt chunk index yields addresses pointing at the wrong bytes, so a
/// mismatch is an error: otherwise the damage surfaces as plausible data read /// mismatch is an error: otherwise the damage surfaces as plausible data read
/// from the wrong chunk. /// from the wrong chunk.
#[cfg(feature = "checksum")] #[cfg(feature = "checksum")]
fn verify_checksum(data: &[u8], start: usize, end: usize) -> Result<(), FormatError> { fn verify_checksum(w: &Window<'_>, start: usize, end: usize) -> Result<(), FormatError> {
ensure_len(data, end, 4)?; w.ensure(end, 4)?;
let data: &[u8] = &w.bytes;
let stored = u32::from_le_bytes([data[end], data[end + 1], data[end + 2], data[end + 3]]); let stored = u32::from_le_bytes([data[end], data[end + 1], data[end + 2], data[end + 3]]);
let computed = crate::checksum::jenkins_lookup3(&data[start..end]); let computed = crate::checksum::jenkins_lookup3(&data[start..end]);
if computed != stored { if computed != stored {
@@ -35,7 +38,7 @@ fn verify_checksum(data: &[u8], start: usize, end: usize) -> Result<(), FormatEr
} }
#[cfg(not(feature = "checksum"))] #[cfg(not(feature = "checksum"))]
fn verify_checksum(_data: &[u8], _start: usize, _end: usize) -> Result<(), FormatError> { fn verify_checksum(_w: &Window<'_>, _start: usize, _end: usize) -> Result<(), FormatError> {
Ok(()) Ok(())
} }
@@ -81,19 +84,6 @@ fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
}) })
} }
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> {
if offset
.checked_add(needed)
.is_none_or(|end| end > data.len())
{
return Err(FormatError::UnexpectedEof {
expected: offset.saturating_add(needed),
available: data.len(),
});
}
Ok(())
}
fn is_undefined_addr(addr: u64, offset_size: u8) -> bool { fn is_undefined_addr(addr: u64, offset_size: u8) -> bool {
match offset_size { match offset_size {
2 => addr == 0xFFFF, 2 => addr == 0xFFFF,
@@ -131,6 +121,16 @@ impl ExtensibleArrayHeader {
offset: usize, offset: usize,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Self, FormatError> {
Self::parse_in(file_data, offset as u64, offset_size, length_size)
}
/// [`Self::parse`] over any [`Storage`]: one read of the header.
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
offset_size: u8,
length_size: u8,
) -> Result<Self, FormatError> { ) -> Result<Self, FormatError> {
// EAHD: signature(4) + version(1) + client_id(1) + element_size(1) + // EAHD: signature(4) + version(1) + client_id(1) + element_size(1) +
// max_nelmts_bits(1) + idx_blk_elmts(1) + min_dblk_nelmts(1) + // max_nelmts_bits(1) + idx_blk_elmts(1) + min_dblk_nelmts(1) +
@@ -138,9 +138,10 @@ impl ExtensibleArrayHeader {
// 6 stats fields (each length_size) + index_block_address(offset_size) + checksum(4) // 6 stats fields (each length_size) + index_block_address(offset_size) + checksum(4)
let min_size = let min_size =
4 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 6 * length_size as usize + offset_size as usize + 4; 4 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 6 * length_size as usize + offset_size as usize + 4;
ensure_len(file_data, offset, min_size)?; let w = Window::read(file, offset, min_size)?;
w.ensure(0, min_size)?;
let d = &file_data[offset..]; let d: &[u8] = &w.bytes;
if &d[0..4] != b"EAHD" { if &d[0..4] != b"EAHD" {
return Err(FormatError::ChunkedReadError( return Err(FormatError::ChunkedReadError(
"invalid Extensible Array header signature".into(), "invalid Extensible Array header signature".into(),
@@ -173,7 +174,7 @@ impl ExtensibleArrayHeader {
pos += ls; // skip max_idx_set (6th stats field) pos += ls; // skip max_idx_set (6th stats field)
let index_block_address = read_offset(d, pos, offset_size)?; let index_block_address = read_offset(d, pos, offset_size)?;
pos += offset_size as usize; pos += offset_size as usize;
verify_checksum(file_data, offset, offset + pos)?; verify_checksum(&w, 0, pos)?;
Ok(ExtensibleArrayHeader { Ok(ExtensibleArrayHeader {
client_id, client_id,
@@ -194,11 +195,11 @@ impl ExtensibleArrayHeader {
} }
} }
/// Read a single element from the extensible array element data. /// Read a single element at offset `pos` of the window `w`.
/// Returns (chunk_info, bytes_consumed) or None if unallocated. /// Returns (chunk_info, bytes_consumed) or None if unallocated.
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
fn read_element( fn read_element(
data: &[u8], w: &Window<'_>,
pos: usize, pos: usize,
client_id: u8, client_id: u8,
element_size: u8, element_size: u8,
@@ -208,15 +209,11 @@ fn read_element(
grid: &ChunkGrid, grid: &ChunkGrid,
) -> Result<(Option<ChunkInfo>, usize), FormatError> { ) -> Result<(Option<ChunkInfo>, usize), FormatError> {
let os = offset_size as usize; let os = offset_size as usize;
let data: &[u8] = &w.bytes;
if client_id == 0 { if client_id == 0 {
// Non-filtered: just address // Non-filtered: just address
if pos + os > data.len() { w.ensure(pos, os)?;
return Err(FormatError::UnexpectedEof {
expected: pos + os,
available: data.len(),
});
}
if is_undefined(data, pos, offset_size) { if is_undefined(data, pos, offset_size) {
return Ok((None, os)); return Ok((None, os));
} }
@@ -244,15 +241,7 @@ fn read_element(
} }
let chunk_size_bytes = es - os - 4; let chunk_size_bytes = es - os - 4;
let elem_total = os + chunk_size_bytes + 4; let elem_total = os + chunk_size_bytes + 4;
if pos w.ensure(pos, elem_total)?;
.checked_add(elem_total)
.is_none_or(|end| end > data.len())
{
return Err(FormatError::UnexpectedEof {
expected: pos.saturating_add(elem_total),
available: data.len(),
});
}
if is_undefined(data, pos, offset_size) { if is_undefined(data, pos, offset_size) {
return Ok((None, elem_total)); return Ok((None, elem_total));
} }
@@ -316,9 +305,9 @@ fn page_nelmts(header: &ExtensibleArrayHeader) -> Option<usize> {
/// paged. The bitmap lives in the super block, not here — a paged data block /// paged. The bitmap lives in the super block, not here — a paged data block
/// stores only its prefix, then one slot per page. /// stores only its prefix, then one slot per page.
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
fn read_data_block_elements( fn read_data_block_elements<S: Storage + ?Sized>(
file_data: &[u8], file: &S,
db_offset: usize, db_offset: u64,
nelmts: usize, nelmts: usize,
header: &ExtensibleArrayHeader, header: &ExtensibleArrayHeader,
offset_size: u8, offset_size: u8,
@@ -331,21 +320,28 @@ fn read_data_block_elements(
// EADB: signature(4) + version(1) + client_id(1) + header_address(offset_size) // EADB: signature(4) + version(1) + client_id(1) + header_address(offset_size)
// + block offset(arr_off_size) // + block offset(arr_off_size)
let db_header_size = 4 + 1 + 1 + offset_size as usize + arr_off_size(header); let db_header_size = 4 + 1 + 1 + offset_size as usize + arr_off_size(header);
ensure_len(file_data, db_offset, db_header_size)?; let prefix = read_exact_at(file, db_offset, db_header_size)?;
if &file_data[db_offset..db_offset + 4] != b"EADB" { if &prefix[0..4] != b"EADB" {
return Err(FormatError::ChunkedReadError( return Err(FormatError::ChunkedReadError(
"invalid Extensible Array data block signature".into(), "invalid Extensible Array data block signature".into(),
)); ));
} }
let mut pos = db_offset + db_header_size; // Positions below are relative to the data block.
let mut pos = db_header_size;
let page = page_nelmts(header).ok_or_else(|| { let page = page_nelmts(header).ok_or_else(|| {
FormatError::Overflow("Extensible Array page element count overflows usize".into()) FormatError::Overflow("Extensible Array page element count overflows usize".into())
})?; })?;
let elem_bytes = if header.client_id == 0 {
offset_size as usize
} else {
header.element_size as usize
};
let mut chunks = Vec::new(); let mut chunks = Vec::new();
let read_run = |from: usize, let read_run = |w: &Window<'_>,
from: usize,
count: usize, count: usize,
first_index: usize, first_index: usize,
chunks: &mut Vec<ChunkInfo>| chunks: &mut Vec<ChunkInfo>|
@@ -353,7 +349,7 @@ fn read_data_block_elements(
let mut p = from; let mut p = from;
for i in 0..count { for i in 0..count {
let (info, consumed) = read_element( let (info, consumed) = read_element(
file_data, w,
p, p,
header.client_id, header.client_id,
header.element_size, header.element_size,
@@ -371,18 +367,19 @@ fn read_data_block_elements(
}; };
if nelmts <= page { if nelmts <= page {
// Prefix and elements are covered by one checksum. // Prefix and elements are covered by one checksum. One window holds
let elem_bytes = if header.client_id == 0 { // all of it (or ends at the end of the file), so its bounds checks
offset_size as usize // are the whole-file ones.
} else {
header.element_size as usize
};
let end = nelmts let end = nelmts
.checked_mul(elem_bytes) .checked_mul(elem_bytes)
.and_then(|b| pos.checked_add(b)) .and_then(|b| pos.checked_add(b))
.ok_or_else(|| FormatError::Overflow("Extensible Array data block span".into()))?; .ok_or_else(|| FormatError::Overflow("Extensible Array data block span".into()))?;
verify_checksum(file_data, db_offset, end)?; // The checksum's bounds check comes first: make it before reading.
read_run(pos, nelmts, start_index, &mut chunks)?; #[cfg(feature = "checksum")]
Window::check_extent(file, db_offset, end, 4)?;
let w = Window::read(file, db_offset, end.saturating_add(4))?;
verify_checksum(&w, 0, end)?;
read_run(&w, pos, nelmts, start_index, &mut chunks)?;
return Ok(chunks); return Ok(chunks);
} }
@@ -390,18 +387,32 @@ fn read_data_block_elements(
// each holding `page` elements followed by a checksum. Pages whose bit is // each holding `page` elements followed by a checksum. Pages whose bit is
// clear were never written; their slot still occupies the file, so stride // clear were never written; their slot still occupies the file, so stride
// over it rather than reading zeros as addresses. // over it rather than reading zeros as addresses.
verify_checksum(file_data, db_offset, pos)?; let npages = nelmts.div_ceil(page);
pos += 4; // The whole data block in one window when it is small: every position
let elem_bytes = if header.client_id == 0 { // checked below lies inside it (or past the end of the file). A larger
offset_size as usize // block is read as its prefix, then each page in use on its own.
let block_len = pos
.saturating_add(4)
.saturating_add(npages.saturating_mul(page.saturating_mul(elem_bytes).saturating_add(4)));
let whole = if block_len <= PAGED_BLOCK_ONE_READ_MAX {
Some(Window::read(file, db_offset, block_len)?)
} else { } else {
header.element_size as usize None
}; };
let head_w;
let head = match &whole {
Some(w) => w,
None => {
head_w = Window::read(file, db_offset, pos + 4)?;
&head_w
}
};
verify_checksum(head, 0, pos)?;
pos += 4;
let page_stride = page let page_stride = page
.checked_mul(elem_bytes) .checked_mul(elem_bytes)
.and_then(|b| b.checked_add(4)) .and_then(|b| b.checked_add(4))
.ok_or_else(|| FormatError::Overflow("Extensible Array page stride".into()))?; .ok_or_else(|| FormatError::Overflow("Extensible Array page stride".into()))?;
let npages = nelmts.div_ceil(page);
for p in 0..npages { for p in 0..npages {
// One bit per page across the whole super block, packed contiguously // One bit per page across the whole super block, packed contiguously
// and MSB-first within each byte, as H5VM_bit_get reads it. // and MSB-first within each byte, as H5VM_bit_get reads it.
@@ -411,10 +422,20 @@ fn read_data_block_elements(
.is_some_and(|byte| byte & (0x80 >> (bit % 8)) != 0); .is_some_and(|byte| byte & (0x80 >> (bit % 8)) != 0);
if initialised { if initialised {
let count = core::cmp::min(page, nelmts - p * page); let count = core::cmp::min(page, nelmts - p * page);
// `w` holds the page from `base` on (positions below are
// relative to it, and `pos` to the data block).
let page_w;
let (w, base) = match &whole {
Some(w) => (w, 0),
None => {
page_w = Window::read(file, db_offset.saturating_add(pos as u64), page_stride)?;
(&page_w, pos)
}
};
// Each page carries its own checksum, over a full page's worth of // Each page carries its own checksum, over a full page's worth of
// slots even when the last one holds fewer live elements. // slots even when the last one holds fewer live elements.
verify_checksum(file_data, pos, pos + page * elem_bytes)?; verify_checksum(w, pos - base, pos - base + page * elem_bytes)?;
read_run(pos, count, start_index + p * page, &mut chunks)?; read_run(w, pos - base, count, start_index + p * page, &mut chunks)?;
} }
pos = pos pos = pos
.checked_add(page_stride) .checked_add(page_stride)
@@ -436,6 +457,32 @@ pub fn read_extensible_array_chunks(
chunk_dimensions: &[u32], chunk_dimensions: &[u32],
element_size: u32, element_size: u32,
offset_size: u8, offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
read_extensible_array_chunks_in(
&file_data,
header,
dataset_dims,
max_dims,
chunk_dimensions,
element_size,
offset_size,
length_size,
)
}
/// [`read_extensible_array_chunks`] over any [`Storage`]: one read of the
/// index block's prefix, one of the whole index block, and the same for
/// every super block and data block it references.
#[allow(clippy::too_many_arguments)]
pub fn read_extensible_array_chunks_in<S: Storage + ?Sized>(
file: &S,
header: &ExtensibleArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
element_size: u32,
offset_size: u8,
_length_size: u8, _length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> { ) -> Result<Vec<ChunkInfo>, FormatError> {
let os = offset_size as usize; let os = offset_size as usize;
@@ -452,16 +499,17 @@ pub fn read_extensible_array_chunks(
// Parse index block (EAIB): signature(4) + version(1) + client_id(1) // Parse index block (EAIB): signature(4) + version(1) + client_id(1)
// + header address(offset_size), then the inline elements, then the // + header address(offset_size), then the inline elements, then the
// direct data block addresses, then the super block addresses. // direct data block addresses, then the super block addresses.
let ib_offset = to_usize(header.index_block_address)?; // Positions below are relative to the index block.
let ib_offset = header.index_block_address;
let ib_header_size = 4 + 1 + 1 + os; let ib_header_size = 4 + 1 + 1 + os;
ensure_len(file_data, ib_offset, ib_header_size)?; let prefix = read_exact_at(file, ib_offset, ib_header_size)?;
if &file_data[ib_offset..ib_offset + 4] != b"EAIB" { if &prefix[0..4] != b"EAIB" {
return Err(FormatError::ChunkedReadError( return Err(FormatError::ChunkedReadError(
"invalid Extensible Array index block signature".into(), "invalid Extensible Array index block signature".into(),
)); ));
} }
let mut pos = ib_offset + ib_header_size; let mut pos = ib_header_size;
let mut chunks = Vec::new(); let mut chunks = Vec::new();
let total_elements = to_usize(header.num_elements)?; let total_elements = to_usize(header.num_elements)?;
@@ -521,13 +569,19 @@ pub fn read_extensible_array_chunks(
.and_then(|n| n.checked_mul(os).and_then(|b| p.checked_add(b))) .and_then(|n| n.checked_mul(os).and_then(|b| p.checked_add(b)))
}) })
.ok_or_else(|| FormatError::Overflow("Extensible Array index block span".into()))?; .ok_or_else(|| FormatError::Overflow("Extensible Array index block span".into()))?;
verify_checksum(file_data, ib_offset, ib_end)?; // The whole index block in one window: every position read below is
// before `ib_end`.
// The checksum's bounds check comes first: make it before reading.
#[cfg(feature = "checksum")]
Window::check_extent(file, ib_offset, ib_end, 4)?;
let w = Window::read(file, ib_offset, ib_end.saturating_add(4))?;
verify_checksum(&w, 0, ib_end)?;
// 1. Elements stored inline in the index block. // 1. Elements stored inline in the index block.
let n_inline = (header.idx_blk_elmts as usize).min(total_elements); let n_inline = (header.idx_blk_elmts as usize).min(total_elements);
for i in 0..n_inline { for i in 0..n_inline {
let (info, consumed) = read_element( let (info, consumed) = read_element(
file_data, &w,
pos, pos,
header.client_id, header.client_id,
header.element_size, header.element_size,
@@ -551,8 +605,8 @@ pub fn read_extensible_array_chunks(
if global_index >= total_elements { if global_index >= total_elements {
return Ok(chunks); return Ok(chunks);
} }
ensure_len(file_data, pos, os)?; w.ensure(pos, os)?;
let addr = read_offset(file_data, pos, offset_size)?; let addr = read_offset(&w.bytes, pos, offset_size)?;
pos += os; pos += os;
if !is_undefined_addr(addr, offset_size) { if !is_undefined_addr(addr, offset_size) {
if dblk_nelmts > page_nelmts(header).unwrap_or(usize::MAX) { if dblk_nelmts > page_nelmts(header).unwrap_or(usize::MAX) {
@@ -563,8 +617,8 @@ pub fn read_extensible_array_chunks(
)); ));
} }
chunks.extend(read_data_block_elements( chunks.extend(read_data_block_elements(
file_data, file,
to_usize(addr)?, addr,
dblk_nelmts, dblk_nelmts,
header, header,
offset_size, offset_size,
@@ -584,16 +638,16 @@ pub fn read_extensible_array_chunks(
if global_index >= total_elements { if global_index >= total_elements {
break; break;
} }
ensure_len(file_data, pos, os)?; w.ensure(pos, os)?;
let sb_addr = read_offset(file_data, pos, offset_size)?; let sb_addr = read_offset(&w.bytes, pos, offset_size)?;
pos += os; pos += os;
let (ndblks, dblk_nelmts) = sblk_info(u, dmin).ok_or_else(|| { let (ndblks, dblk_nelmts) = sblk_info(u, dmin).ok_or_else(|| {
FormatError::Overflow("Extensible Array super block layout overflows usize".into()) FormatError::Overflow("Extensible Array super block layout overflows usize".into())
})?; })?;
if !is_undefined_addr(sb_addr, offset_size) { if !is_undefined_addr(sb_addr, offset_size) {
chunks.extend(read_super_block( chunks.extend(read_super_block(
file_data, file,
to_usize(sb_addr)?, sb_addr,
ndblks, ndblks,
dblk_nelmts, dblk_nelmts,
header, header,
@@ -618,9 +672,9 @@ pub fn read_extensible_array_chunks(
/// + block offset + the page-init bitmap for every data block it owns /// + block offset + the page-init bitmap for every data block it owns
/// + one address per data block + checksum. /// + one address per data block + checksum.
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
fn read_super_block( fn read_super_block<S: Storage + ?Sized>(
file_data: &[u8], file: &S,
sb_offset: usize, sb_offset: u64,
ndblks: usize, ndblks: usize,
dblk_nelmts: usize, dblk_nelmts: usize,
header: &ExtensibleArrayHeader, header: &ExtensibleArrayHeader,
@@ -631,9 +685,9 @@ fn read_super_block(
) -> Result<Vec<ChunkInfo>, FormatError> { ) -> Result<Vec<ChunkInfo>, FormatError> {
let os = offset_size as usize; let os = offset_size as usize;
let sb_header_size = 4 + 1 + 1 + os + arr_off_size(header); let sb_header_size = 4 + 1 + 1 + os + arr_off_size(header);
ensure_len(file_data, sb_offset, sb_header_size)?; let prefix = read_exact_at(file, sb_offset, sb_header_size)?;
if &file_data[sb_offset..sb_offset + 4] != b"EASB" { if &prefix[0..4] != b"EASB" {
return Err(FormatError::ChunkedReadError( return Err(FormatError::ChunkedReadError(
"invalid Extensible Array super block signature".into(), "invalid Extensible Array super block signature".into(),
)); ));
@@ -655,29 +709,38 @@ fn read_super_block(
let bitmap_bytes = per_dblk_bitmap let bitmap_bytes = per_dblk_bitmap
.checked_mul(ndblks) .checked_mul(ndblks)
.ok_or_else(|| FormatError::Overflow("Extensible Array page bitmap size".into()))?; .ok_or_else(|| FormatError::Overflow("Extensible Array page bitmap size".into()))?;
let bitmap_start = sb_offset + sb_header_size; // Positions below are relative to the super block, whose bytes (up to
ensure_len(file_data, bitmap_start, bitmap_bytes)?; // its checksum) are all in one window.
let bitmap = &file_data[bitmap_start..bitmap_start + bitmap_bytes]; let bitmap_start = sb_header_size;
// The bitmap's bounds check, then (with checksums) the checksum's, come
// before anything else is read from the block: make them before reading
// it, so size fields stretching it past the end of the file cost no read.
Window::check_extent(file, sb_offset, bitmap_start, bitmap_bytes)?;
let mut pos = bitmap_start + bitmap_bytes; let mut pos = bitmap_start + bitmap_bytes;
let mut chunks = Vec::new();
let mut global_idx = start_index;
// One checksum covers the prefix, the bitmap and every data block address. // One checksum covers the prefix, the bitmap and every data block address.
let sb_end = ndblks let sb_end = ndblks
.checked_mul(os) .checked_mul(os)
.and_then(|b| pos.checked_add(b)) .and_then(|b| pos.checked_add(b))
.ok_or_else(|| FormatError::Overflow("Extensible Array super block span".into()))?; .ok_or_else(|| FormatError::Overflow("Extensible Array super block span".into()))?;
verify_checksum(file_data, sb_offset, sb_end)?; #[cfg(feature = "checksum")]
Window::check_extent(file, sb_offset, sb_end, 4)?;
let w = Window::read(file, sb_offset, sb_end.saturating_add(4))?;
w.ensure(bitmap_start, bitmap_bytes)?;
let bitmap = &w.bytes[bitmap_start..bitmap_start + bitmap_bytes];
let mut chunks = Vec::new();
let mut global_idx = start_index;
verify_checksum(&w, 0, sb_end)?;
for i in 0..ndblks { for i in 0..ndblks {
ensure_len(file_data, pos, os)?; w.ensure(pos, os)?;
let addr = read_offset(file_data, pos, offset_size)?; let addr = read_offset(&w.bytes, pos, offset_size)?;
pos += os; pos += os;
if !is_undefined_addr(addr, offset_size) { if !is_undefined_addr(addr, offset_size) {
chunks.extend(read_data_block_elements( chunks.extend(read_data_block_elements(
file_data, file,
to_usize(addr)?, addr,
dblk_nelmts, dblk_nelmts,
header, header,
offset_size, offset_size,
@@ -888,11 +951,11 @@ mod tests {
assert_eq!(chunks[1].offsets, vec![20]); assert_eq!(chunks[1].offsets, vec![20]);
} }
/// Build a synthetic EA with inline elements + one direct data block. /// A synthetic EA with inline elements + one direct data block: the
#[test] /// file, with the header at 0x100 (8-byte offsets and lengths, 4 chunks
fn read_inline_plus_data_blocks() { /// of 10 elements from 0x1000 on).
fn build_inline_plus_data_blocks() -> Vec<u8> {
let os: u8 = 8; let os: u8 = 8;
let ls: u8 = 8;
let osv = os as usize; let osv = os as usize;
let chunk_byte_size = 10u64 * 8; // 10 elements × 8 bytes let chunk_byte_size = 10u64 * 8; // 10 elements × 8 bytes
let idx_blk_elmts = 2u8; let idx_blk_elmts = 2u8;
@@ -982,8 +1045,17 @@ mod tests {
dbpos += osv; dbpos += osv;
} }
stamp_checksum(&mut file_data, aedb_offset, dbpos); stamp_checksum(&mut file_data, aedb_offset, dbpos);
file_data
}
let header = ExtensibleArrayHeader::parse(&file_data, aehd_offset, os, ls).unwrap(); /// Build a synthetic EA with inline elements + one direct data block.
#[test]
fn read_inline_plus_data_blocks() {
let (os, ls) = (8u8, 8u8);
let chunk_byte_size = 10u64 * 8;
let base_addr = 0x1000u64;
let file_data = build_inline_plus_data_blocks();
let header = ExtensibleArrayHeader::parse(&file_data, 0x100, os, ls).unwrap();
let ds_dims = vec![40u64]; let ds_dims = vec![40u64];
let chunk_dims = vec![10u32]; let chunk_dims = vec![10u32];
let chunks = read_extensible_array_chunks( let chunks = read_extensible_array_chunks(
@@ -1019,7 +1091,8 @@ mod tests {
fn read_element_unallocated() { fn read_element_unallocated() {
let data = vec![0xFFu8; 16]; let data = vec![0xFFu8; 16];
let grid = ChunkGrid::fixed_array(&[50], None, &[10]).unwrap(); let grid = ChunkGrid::fixed_array(&[50], None, &[10]).unwrap();
let (info, consumed) = read_element(&data, 0, 0, 8, 8, 80, 0, &grid).unwrap(); let (info, consumed) =
read_element(&Window::whole(&data), 0, 0, 8, 8, 80, 0, &grid).unwrap();
assert!(info.is_none()); assert!(info.is_none());
assert_eq!(consumed, 8); assert_eq!(consumed, 8);
} }
@@ -1039,8 +1112,17 @@ mod tests {
data[12..16].copy_from_slice(&0u32.to_le_bytes()); data[12..16].copy_from_slice(&0u32.to_le_bytes());
let grid = ChunkGrid::fixed_array(&[50], None, &[10]).unwrap(); let grid = ChunkGrid::fixed_array(&[50], None, &[10]).unwrap();
let (info, consumed) = let (info, consumed) = read_element(
read_element(&data, 0, 1, elem_size as u8, os, 80, 2, &grid).unwrap(); &Window::whole(&data),
0,
1,
elem_size as u8,
os,
80,
2,
&grid,
)
.unwrap();
let ci = info.unwrap(); let ci = info.unwrap();
assert_eq!(ci.address, 0x2000); assert_eq!(ci.address, 0x2000);
assert_eq!(ci.chunk_size, 120); assert_eq!(ci.chunk_size, 120);
@@ -1048,4 +1130,38 @@ mod tests {
assert_eq!(ci.offsets, vec![20]); assert_eq!(ci.offsets, vec![20]);
assert_eq!(consumed, elem_size); assert_eq!(consumed, elem_size);
} }
/// The Storage path reads exactly what the slice path reads: the array
/// whole, cut at every length through its structures, and with a byte
/// damaged in each of them, through a read_at-only CountingStorage.
#[test]
fn storage_reads_match_slice_reads() {
use crate::storage::CountingStorage;
let full = build_inline_plus_data_blocks();
let mut files = Vec::new();
for cut in 0x100..0x340 {
files.push(full[..cut].to_vec());
}
for at in [0x104, 0x150, 0x204, 0x216, 0x230, 0x304, 0x318] {
let mut damaged = full.clone();
damaged[at] ^= 1;
files.push(damaged);
}
files.push(full);
let mut compared = 0;
for f in files {
let storage = CountingStorage::new(f.clone());
let want = ExtensibleArrayHeader::parse(&f, 0x100, 8, 8);
let got = ExtensibleArrayHeader::parse_in(&storage, 0x100, 8, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
let Ok(h) = want else { continue };
for dims in [&[40u64][..], &[25]] {
let want = read_extensible_array_chunks(&f, &h, dims, None, &[10], 8, 8, 8);
let got = read_extensible_array_chunks_in(&storage, &h, dims, None, &[10], 8, 8, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"), "{} bytes", f.len());
compared += 1;
}
}
assert!(compared > 100);
}
} }
+46 -1
View File
@@ -117,9 +117,21 @@ pub fn dataset_fill_value_in(
messages: &[HeaderMessage], messages: &[HeaderMessage],
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Option<Vec<u8>>, FormatError> {
dataset_fill_value_from_storage(&file_data, messages, offset_size, length_size)
}
/// [`dataset_fill_value_in`] with the file behind any
/// [`Storage`](crate::storage::Storage). (The trait is not imported here:
/// its `len` would shadow the slice method in this module.)
pub fn dataset_fill_value_from_storage(
file: &dyn crate::storage::Storage,
messages: &[HeaderMessage],
offset_size: u8,
length_size: u8,
) -> Result<Option<Vec<u8>>, FormatError> { ) -> Result<Option<Vec<u8>>, FormatError> {
fill_value_from(messages, |msg| { fill_value_from(messages, |msg| {
crate::shared_message::message_data_with_sohm(file_data, msg, offset_size, length_size) crate::shared_message::message_data_with_sohm_in(file, msg, offset_size, length_size)
.map(|data| data.into_owned()) .map(|data| data.into_owned())
}) })
} }
@@ -444,4 +456,37 @@ mod tests {
.collect(); .collect();
assert_eq!(filled, [2, 3, 7, 8]); assert_eq!(filled, [2, 3, 7, 8]);
} }
/// Fill values, shared ones in the SOHM heap included, resolve
/// identically through a read_at-only CountingStorage.
#[test]
fn storage_reads_match_slice_reads() {
use crate::object_header::ObjectHeader;
use crate::storage::CountingStorage;
let file: &[u8] = include_bytes!("../tests/fixtures/shared_fill_value.h5");
let sb = crate::superblock::Superblock::parse(file, 0).unwrap();
let (os, ls) = (sb.offset_size, sb.length_size);
let storage = CountingStorage::new(file.to_vec());
let mut shared = 0;
let children =
crate::group_v2::resolve_group_children(file, &sb, sb.root_group_address).unwrap();
assert!(children.len() >= 3);
for child in children {
let h =
ObjectHeader::parse(file, child.object_header_address as usize, os, ls).unwrap();
shared += h
.messages
.iter()
.filter(|m| {
m.msg_type == MessageType::FillValue
&& crate::shared_message::is_shared(m.flags)
})
.count();
let want = dataset_fill_value_in(file, &h.messages, os, ls);
assert_eq!(want, Ok(Some((-7i32).to_le_bytes().to_vec())));
let got = dataset_fill_value_from_storage(&storage, &h.messages, os, ls);
assert_eq!(got, want, "{}", child.name);
}
assert!(shared >= 2);
}
} }
+252 -72
View File
@@ -10,16 +10,19 @@ use crate::addr::to_usize;
use crate::chunk_grid::ChunkGrid; use crate::chunk_grid::ChunkGrid;
use crate::chunked_read::ChunkInfo; use crate::chunked_read::ChunkInfo;
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::{PAGED_BLOCK_ONE_READ_MAX, Storage, Window, len_usize, read_exact_at};
/// Verify the Jenkins lookup3 checksum stored immediately after /// Verify the Jenkins lookup3 checksum stored immediately after
/// `data[start..end]`, as every Fixed Array structure carries one. /// `data[start..end]`, as every Fixed Array structure carries one. `w` is
/// a window of the file and `start`/`end` are relative to it.
/// ///
/// A corrupt chunk index silently yields addresses pointing at the wrong /// A corrupt chunk index silently yields addresses pointing at the wrong
/// bytes, so a mismatch has to be an error rather than a shrug: without this /// bytes, so a mismatch has to be an error rather than a shrug: without this
/// the damage surfaces as plausible-looking data from the wrong chunk. /// the damage surfaces as plausible-looking data from the wrong chunk.
#[cfg(feature = "checksum")] #[cfg(feature = "checksum")]
fn verify_checksum(data: &[u8], start: usize, end: usize) -> Result<(), FormatError> { fn verify_checksum(w: &Window<'_>, start: usize, end: usize) -> Result<(), FormatError> {
ensure_len(data, end, 4)?; w.ensure(end, 4)?;
let data = &w.bytes;
let stored = u32::from_le_bytes([data[end], data[end + 1], data[end + 2], data[end + 3]]); let stored = u32::from_le_bytes([data[end], data[end + 1], data[end + 2], data[end + 3]]);
let computed = crate::checksum::jenkins_lookup3(&data[start..end]); let computed = crate::checksum::jenkins_lookup3(&data[start..end]);
if computed != stored { if computed != stored {
@@ -32,7 +35,7 @@ fn verify_checksum(data: &[u8], start: usize, end: usize) -> Result<(), FormatEr
} }
#[cfg(not(feature = "checksum"))] #[cfg(not(feature = "checksum"))]
fn verify_checksum(_data: &[u8], _start: usize, _end: usize) -> Result<(), FormatError> { fn verify_checksum(_w: &Window<'_>, _start: usize, _end: usize) -> Result<(), FormatError> {
Ok(()) Ok(())
} }
@@ -74,19 +77,6 @@ fn read_length(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
read_offset(data, pos, size) read_offset(data, pos, size)
} }
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> {
if offset
.checked_add(needed)
.is_none_or(|end| end > data.len())
{
return Err(FormatError::UnexpectedEof {
expected: offset.saturating_add(needed),
available: data.len(),
});
}
Ok(())
}
fn is_undefined(data: &[u8], pos: usize, size: u8) -> bool { fn is_undefined(data: &[u8], pos: usize, size: u8) -> bool {
let s = size as usize; let s = size as usize;
if pos + s > data.len() { if pos + s > data.len() {
@@ -102,13 +92,24 @@ impl FixedArrayHeader {
offset: usize, offset: usize,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Self, FormatError> {
Self::parse_in(file_data, offset as u64, offset_size, length_size)
}
/// [`Self::parse`] over any [`Storage`]: one read of the header.
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
offset_size: u8,
length_size: u8,
) -> Result<Self, FormatError> { ) -> Result<Self, FormatError> {
// FAHD signature(4) + version(1) + client_id(1) + element_size(1) + // FAHD signature(4) + version(1) + client_id(1) + element_size(1) +
// max_nelmts_bits(1) + num_elements(length_size) + data_block_addr(offset_size) + checksum(4) // max_nelmts_bits(1) + num_elements(length_size) + data_block_addr(offset_size) + checksum(4)
let min_size = 4 + 1 + 1 + 1 + 1 + length_size as usize + offset_size as usize + 4; let min_size = 4 + 1 + 1 + 1 + 1 + length_size as usize + offset_size as usize + 4;
ensure_len(file_data, offset, min_size)?; let w = Window::read(file, offset, min_size)?;
w.ensure(0, min_size)?;
let d = &file_data[offset..]; let d: &[u8] = &w.bytes;
if &d[0..4] != b"FAHD" { if &d[0..4] != b"FAHD" {
return Err(FormatError::ChunkedReadError( return Err(FormatError::ChunkedReadError(
"invalid Fixed Array header signature".into(), "invalid Fixed Array header signature".into(),
@@ -131,7 +132,7 @@ impl FixedArrayHeader {
pos += length_size as usize; pos += length_size as usize;
let data_block_address = read_offset(d, pos, offset_size)?; let data_block_address = read_offset(d, pos, offset_size)?;
pos += offset_size as usize; pos += offset_size as usize;
verify_checksum(file_data, offset, offset + pos)?; verify_checksum(&w, 0, pos)?;
Ok(FixedArrayHeader { Ok(FixedArrayHeader {
client_id, client_id,
@@ -157,15 +158,39 @@ pub fn read_fixed_array_chunks(
chunk_dimensions: &[u32], chunk_dimensions: &[u32],
element_size: u32, element_size: u32,
offset_size: u8, offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
read_fixed_array_chunks_in(
&file_data,
header,
dataset_dims,
max_dims,
chunk_dimensions,
element_size,
offset_size,
length_size,
)
}
/// [`read_fixed_array_chunks`] over any [`Storage`]: one read of the data
/// block's prefix, one of the whole data block (pages included).
#[allow(clippy::too_many_arguments)]
pub fn read_fixed_array_chunks_in<S: Storage + ?Sized>(
file: &S,
header: &FixedArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
element_size: u32,
offset_size: u8,
_length_size: u8, _length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> { ) -> Result<Vec<ChunkInfo>, FormatError> {
let file_len = len_usize(file);
let db_offset = to_usize(header.data_block_address)?; let db_offset = to_usize(header.data_block_address)?;
// Parse data block header: FADB(4) + version(1) + client_id(1) + header_address(offset_size) // Parse data block header: FADB(4) + version(1) + client_id(1) + header_address(offset_size)
let db_header_size = 4 + 1 + 1 + offset_size as usize; let db_header_size = 4 + 1 + 1 + offset_size as usize;
ensure_len(file_data, db_offset, db_header_size)?; let d = read_exact_at(file, db_offset as u64, db_header_size)?;
let d = &file_data[db_offset..];
if &d[0..4] != b"FADB" { if &d[0..4] != b"FADB" {
return Err(FormatError::ChunkedReadError( return Err(FormatError::ChunkedReadError(
"invalid Fixed Array data block signature".into(), "invalid Fixed Array data block signature".into(),
@@ -179,7 +204,7 @@ pub fn read_fixed_array_chunks(
// A chunk index cannot describe more elements than the file has bytes (each // A chunk index cannot describe more elements than the file has bytes (each
// element occupies at least `offset_size` bytes). Reject a corrupt count // element occupies at least `offset_size` bytes). Reject a corrupt count
// before it can drive a huge loop or overflow an offset computation. // before it can drive a huge loop or overflow an offset computation.
if num_elements > file_data.len() { if num_elements > file_len {
return Err(FormatError::ChunkedReadError( return Err(FormatError::ChunkedReadError(
"Fixed Array element count exceeds file size".into(), "Fixed Array element count exceeds file size".into(),
)); ));
@@ -209,30 +234,34 @@ pub fn read_fixed_array_chunks(
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64; chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
let mut chunks = Vec::new(); let mut chunks = Vec::new();
let push_element = // `rel` is relative to the data block, whose bytes are in `w`.
|i: usize, abs: usize, chunks: &mut Vec<ChunkInfo>| -> Result<(), FormatError> { let push_element = |w: &Window<'_>,
if let Some((address, chunk_size, filter_mask)) = parse_fa_element( i: usize,
file_data, rel: usize,
abs, chunks: &mut Vec<ChunkInfo>|
header.client_id, -> Result<(), FormatError> {
offset_size, if let Some((address, chunk_size, filter_mask)) = parse_fa_element(
header.element_size, w,
chunk_byte_size, rel,
)? { header.client_id,
// A slot beyond the current extent is ignored, as the offset_size,
// library does. header.element_size,
let Some(offsets) = grid.offsets(i as u64) else { chunk_byte_size,
return Ok(()); )? {
}; // A slot beyond the current extent is ignored, as the
chunks.push(ChunkInfo { // library does.
chunk_size, let Some(offsets) = grid.offsets(i as u64) else {
filter_mask, return Ok(());
offsets, };
address, chunks.push(ChunkInfo {
}); chunk_size,
} filter_mask,
Ok(()) offsets,
}; address,
});
}
Ok(())
};
// A data block is paged when it holds more elements than fit in one page. // A data block is paged when it holds more elements than fit in one page.
// `max_nelmts_bits` is an untrusted u8; a shift >= the pointer width would // `max_nelmts_bits` is an untrusted u8; a shift >= the pointer width would
@@ -247,10 +276,16 @@ pub fn read_fixed_array_chunks(
if !is_paged { if !is_paged {
// Non-paged: prefix, then `num_elements` elements packed directly, // Non-paged: prefix, then `num_elements` elements packed directly,
// then a checksum over both. // then a checksum over both. One window holds all of it (or ends at
verify_checksum(file_data, db_offset, elem_at(elements_start, num_elements)?)?; // the end of the file), so its bounds checks are the whole-file ones.
let end = elem_at(elements_start, num_elements)?;
// The checksum's bounds check comes first: make it before reading.
#[cfg(feature = "checksum")]
Window::check_extent(file, db_offset as u64, end - db_offset, 4)?;
let w = Window::read(file, db_offset as u64, end.saturating_add(4) - db_offset)?;
verify_checksum(&w, 0, end - db_offset)?;
for i in 0..num_elements { for i in 0..num_elements {
push_element(i, elem_at(elements_start, i)?, &mut chunks)?; push_element(&w, i, elem_at(elements_start, i)? - db_offset, &mut chunks)?;
} }
return Ok(chunks); return Ok(chunks);
} }
@@ -273,22 +308,40 @@ pub fn read_fixed_array_chunks(
.and_then(|x| x.checked_add(4)) .and_then(|x| x.checked_add(4))
.ok_or_else(stride_overflow)?; .ok_or_else(stride_overflow)?;
if bitmap_start + bitmap_size > file_data.len() { if bitmap_start + bitmap_size > file_len {
return Err(FormatError::UnexpectedEof { return Err(FormatError::UnexpectedEof {
expected: bitmap_start + bitmap_size, expected: bitmap_start + bitmap_size,
available: file_data.len(), available: file_len,
}); });
} }
// The whole data block in one window when it is small: every page slot
// is at most `page_stride` bytes, so every position checked below lies
// inside it (or past the end of the file). A larger block is read as its
// prefix and bitmap, then each page in use on its own.
let block_len = (pages_start - db_offset).saturating_add(npages.saturating_mul(page_stride));
let whole = if block_len <= PAGED_BLOCK_ONE_READ_MAX {
Some(Window::read(file, db_offset as u64, block_len)?)
} else {
None
};
let head_w;
let head = match &whole {
Some(w) => w,
None => {
head_w = Window::read(file, db_offset as u64, pages_start - db_offset)?;
&head_w
}
};
// The prefix and page bitmap are covered by their own checksum, and each // The prefix and page bitmap are covered by their own checksum, and each
// initialised page by one of its own. // initialised page by one of its own.
verify_checksum(file_data, db_offset, bitmap_start + bitmap_size)?; verify_checksum(head, 0, bitmap_start + bitmap_size - db_offset)?;
for p in 0..npages { for p in 0..npages {
let page_first = p * page_nelmts; // < num_elements, cannot overflow let page_first = p * page_nelmts; // < num_elements, cannot overflow
let page_count = core::cmp::min(page_nelmts, num_elements - page_first); let page_count = core::cmp::min(page_nelmts, num_elements - page_first);
// Check the page-init bit (MSB-first within each byte). // Check the page-init bit (MSB-first within each byte).
let bit_byte = file_data[bitmap_start + p / 8]; let bit_byte = head.bytes[bitmap_start + p / 8 - db_offset];
let bit_mask = 1u8 << (7 - (p % 8)); let bit_mask = 1u8 << (7 - (p % 8));
if bit_byte & bit_mask == 0 { if bit_byte & bit_mask == 0 {
continue; // entire page unallocated continue; // entire page unallocated
@@ -298,21 +351,33 @@ pub fn read_fixed_array_chunks(
.checked_mul(page_stride) .checked_mul(page_stride)
.and_then(|o| pages_start.checked_add(o)) .and_then(|o| pages_start.checked_add(o))
.ok_or_else(stride_overflow)?; .ok_or_else(stride_overflow)?;
verify_checksum(file_data, page_off, elem_at(page_off, page_count)?)?; let page_end = elem_at(page_off, page_count)?;
// `w` holds the page from `base` on (positions below are relative
// to it).
let page_w;
let (w, base) = match &whole {
Some(w) => (w, db_offset),
None => {
page_w =
Window::read(file, page_off as u64, page_end.saturating_add(4) - page_off)?;
(&page_w, page_off)
}
};
verify_checksum(w, page_off - base, page_end - base)?;
for e in 0..page_count { for e in 0..page_count {
push_element(page_first + e, elem_at(page_off, e)?, &mut chunks)?; push_element(w, page_first + e, elem_at(page_off, e)? - base, &mut chunks)?;
} }
} }
Ok(chunks) Ok(chunks)
} }
/// Parse a single Fixed Array element at absolute file offset `abs`. /// Parse a single Fixed Array element at offset `abs` of the window `w`.
/// ///
/// Returns `Some((address, chunk_size, filter_mask))` for an allocated chunk, or /// Returns `Some((address, chunk_size, filter_mask))` for an allocated chunk, or
/// `None` if the element is undefined (an unallocated chunk, address all-`0xFF`). /// `None` if the element is undefined (an unallocated chunk, address all-`0xFF`).
fn parse_fa_element( fn parse_fa_element(
file_data: &[u8], w: &Window<'_>,
abs: usize, abs: usize,
client_id: u8, client_id: u8,
offset_size: u8, offset_size: u8,
@@ -322,12 +387,8 @@ fn parse_fa_element(
let os = offset_size as usize; let os = offset_size as usize;
if client_id == 0 { if client_id == 0 {
// Non-filtered: element is just the chunk address. // Non-filtered: element is just the chunk address.
if abs + os > file_data.len() { w.ensure(abs, os)?;
return Err(FormatError::UnexpectedEof { let file_data: &[u8] = &w.bytes;
expected: abs + os,
available: file_data.len(),
});
}
if is_undefined(file_data, abs, offset_size) { if is_undefined(file_data, abs, offset_size) {
return Ok(None); return Ok(None);
} }
@@ -342,17 +403,14 @@ fn parse_fa_element(
)); ));
} }
let chunk_size_bytes = es - os - 4; let chunk_size_bytes = es - os - 4;
if abs + es > file_data.len() { w.ensure(abs, es)?;
return Err(FormatError::UnexpectedEof { let file_data: &[u8] = &w.bytes;
expected: abs + es,
available: file_data.len(),
});
}
if is_undefined(file_data, abs, offset_size) { if is_undefined(file_data, abs, offset_size) {
return Ok(None); return Ok(None);
} }
let address = read_offset(file_data, abs, offset_size)?; let address = read_offset(file_data, abs, offset_size)?;
let chunk_size = read_variable_length(&file_data[abs + os..], chunk_size_bytes)?; let chunk_size =
read_variable_length(&file_data[abs + os..abs + es - 4], chunk_size_bytes)?;
let fm_off = abs + os + chunk_size_bytes; let fm_off = abs + os + chunk_size_bytes;
let filter_mask = u32::from_le_bytes([ let filter_mask = u32::from_le_bytes([
file_data[fm_off], file_data[fm_off],
@@ -814,4 +872,126 @@ mod tests {
.collect(); .collect();
assert_eq!(got, expect); assert_eq!(got, expect);
} }
/// A fixed array (header at 0x100, data block at 0x200) of `n` chunks,
/// filtered or not, paged when `n` exceeds `1 << page_bits`; every
/// page initialised except page 1.
fn build_fixed_array(n: usize, filtered: bool, page_bits: u8) -> Vec<u8> {
let os = 8usize;
let es = if filtered { os + 4 + 4 } else { os };
let (fahd, db) = (0x100usize, 0x200usize);
let mut f = vec![0u8; 0x2000];
f[fahd..fahd + 4].copy_from_slice(b"FAHD");
f[fahd + 5] = u8::from(filtered);
f[fahd + 6] = es as u8;
f[fahd + 7] = page_bits;
f[fahd + 8..fahd + 16].copy_from_slice(&(n as u64).to_le_bytes());
f[fahd + 16..fahd + 24].copy_from_slice(&(db as u64).to_le_bytes());
stamp_checksum(&mut f, fahd, fahd + 24);
f[db..db + 4].copy_from_slice(b"FADB");
f[db + 5] = u8::from(filtered);
f[db + 6..db + 14].copy_from_slice(&(fahd as u64).to_le_bytes());
let elems = db + 6 + os;
let write = |f: &mut Vec<u8>, at: usize, i: usize| {
let addr = if i == 2 {
u64::MAX
} else {
0x1000 + i as u64 * 0x100
};
f[at..at + os].copy_from_slice(&addr.to_le_bytes());
if filtered {
f[at + os..at + os + 4].copy_from_slice(&(100 + i as u32).to_le_bytes());
f[at + os + 4..at + os + 8].copy_from_slice(&(i as u32 & 1).to_le_bytes());
}
};
let page = 1usize << page_bits;
if n <= page {
for i in 0..n {
write(&mut f, elems + i * es, i);
}
stamp_checksum(&mut f, db, elems + n * es);
} else {
let npages = n.div_ceil(page);
let bitmap = npages.div_ceil(8);
for p in 0..npages {
if p != 1 {
f[elems + p / 8] |= 0x80 >> (p % 8);
}
}
stamp_checksum(&mut f, db, elems + bitmap);
let pages_start = elems + bitmap + 4;
for p in (0..npages).filter(|&p| p != 1) {
let at = pages_start + p * (page * es + 4);
let count = page.min(n - p * page);
for e in 0..count {
write(&mut f, at + e * es, p * page + e);
}
stamp_checksum(&mut f, at, at + count * es);
}
}
f
}
/// Non-paged and paged, filtered and unfiltered arrays, cut at every
/// length through the data block and with a damaged byte, read
/// identically through a `read_at`-only storage.
#[test]
fn storage_reads_match_slice_reads() {
use crate::storage::CountingStorage;
for (n, filtered, bits) in [(3, false, 10), (3, true, 10), (11, false, 2), (11, true, 2)] {
let full = build_fixed_array(n, filtered, bits);
let es = if filtered { 16 } else { 8 };
let dims = [n as u64 * 20];
let h = FixedArrayHeader::parse(&full, 0x100, 8, 8).unwrap();
let chunks = read_fixed_array_chunks(&full, &h, &dims, None, &[20], 8, 8, 8).unwrap();
// Chunk 2 is unallocated, and so is page 1 of a paged array.
let expect = if n > 4 { n - 1 - 4 } else { n - 1 };
assert_eq!(chunks.len(), expect);
let mut files = Vec::new();
for cut in (0x100..0x200 + 40 + n * (es + 4) + 16).step_by(3) {
files.push(full[..cut].to_vec());
}
for at in [0x104, 0x210, 0x21a, 0x230] {
let mut damaged = full.clone();
damaged[at] ^= 1;
files.push(damaged);
}
files.push(full);
for f in files {
let storage = CountingStorage::new(f.clone());
let want = FixedArrayHeader::parse(&f, 0x100, 8, 8);
let got = FixedArrayHeader::parse_in(&storage, 0x100, 8, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
let Ok(h) = want else { continue };
let want = read_fixed_array_chunks(&f, &h, &dims, None, &[20], 8, 8, 8);
let got = read_fixed_array_chunks_in(&storage, &h, &dims, None, &[20], 8, 8, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"), "{} bytes", f.len());
}
}
}
/// A header whose element count stretches its data block (one checksum
/// over the whole block) far past the end of a 16 MiB file: the
/// checksum's bounds check fails before the block is read, with the
/// slice read's error.
#[cfg(feature = "checksum")]
#[test]
fn oversized_block_fails_before_reading() {
use crate::storage::CountingStorage;
let mut f = build_fixed_array(3, false, 10);
f.resize(16 << 20, 0);
let mut h = FixedArrayHeader::parse(&f, 0x100, 8, 8).unwrap();
h.max_nelmts_bits = 30;
h.num_elements = 4 << 20;
let dims = [h.num_elements * 20];
let want = read_fixed_array_chunks(&f, &h, &dims, None, &[20], 8, 8, 8);
assert!(
matches!(want, Err(FormatError::UnexpectedEof { .. })),
"{want:?}"
);
let storage = CountingStorage::new(f);
let got = read_fixed_array_chunks_in(&storage, &h, &dims, None, &[20], 8, 8, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
assert!(storage.bytes_read() < 64, "{} bytes", storage.bytes_read());
}
} }
+394 -60
View File
@@ -10,6 +10,7 @@ use crate::addr::to_usize;
use crate::btree_v2::{BTreeV2Header, find_btree_v2_records}; use crate::btree_v2::{BTreeV2Header, find_btree_v2_records};
use crate::error::FormatError; use crate::error::FormatError;
use crate::filter_pipeline::FilterPipeline; use crate::filter_pipeline::FilterPipeline;
use crate::storage::{Storage, Window, len_usize, read_exact_at, require_contiguous};
/// Parsed fractal heap header (signature "FRHP"). /// Parsed fractal heap header (signature "FRHP").
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
@@ -139,6 +140,36 @@ impl FractalHeapHeader {
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<FractalHeapHeader, FormatError> { ) -> Result<FractalHeapHeader, FormatError> {
Self::parse_in(file_data, offset as u64, offset_size, length_size)
}
/// [`Self::parse`] over any [`Storage`]: one read of the header (two
/// when it holds an I/O filter pipeline).
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
offset_size: u8,
length_size: u8,
) -> Result<FractalHeapHeader, FormatError> {
// Every field up to the checksum, without and with the filter
// information; the window holds all of it (or ends at the end of
// the file), so its bounds checks are the whole-file ones.
let (os, ls) = (usize::from(offset_size), usize::from(length_size));
let unfiltered_len = 26 + 12 * ls + 3 * os;
let mut w = Window::read(file, offset, unfiltered_len)?;
if w.bytes.len() == unfiltered_len {
let filter_len = usize::from(u16::from_le_bytes([w.bytes[7], w.bytes[8]]));
if filter_len > 0 {
w = Window::read(file, offset, unfiltered_len + ls + 4 + filter_len)?;
}
}
let ensure_len = |_: &[u8], pos: usize, needed: usize| w.ensure(pos, needed);
let read_offset = |_: &[u8], pos: usize, size: u8| {
w.ensure(pos, usize::from(size))?;
read_offset(&w.bytes, pos, size)
};
let file_data: &[u8] = &w.bytes;
let offset = 0usize;
ensure_len(file_data, offset, 5)?; ensure_len(file_data, offset, 5)?;
if &file_data[offset..offset + 4] != b"FRHP" { if &file_data[offset..offset + 4] != b"FRHP" {
return Err(FormatError::InvalidFractalHeapSignature); return Err(FormatError::InvalidFractalHeapSignature);
@@ -149,9 +180,6 @@ impl FractalHeapHeader {
return Err(FormatError::InvalidFractalHeapVersion(version)); return Err(FormatError::InvalidFractalHeapVersion(version));
} }
let os = offset_size as usize;
let ls = length_size as usize;
let mut pos = offset + 5; let mut pos = offset + 5;
ensure_len(file_data, pos, 2)?; ensure_len(file_data, pos, 2)?;
let heap_id_length = u16::from_le_bytes([file_data[pos], file_data[pos + 1]]); let heap_id_length = u16::from_le_bytes([file_data[pos], file_data[pos + 1]]);
@@ -354,6 +382,18 @@ impl FractalHeapHeader {
file_data: &[u8], file_data: &[u8],
id_bytes: &[u8], id_bytes: &[u8],
offset_size: u8, offset_size: u8,
) -> Result<Vec<u8>, FormatError> {
self.read_managed_object_in(file_data, id_bytes, offset_size)
}
/// [`Self::read_managed_object`] over any [`Storage`]. A huge object
/// found through the huge-object v2 B-tree still needs the whole file
/// in memory ([`FormatError::ContiguousStorageRequired`] otherwise).
pub fn read_managed_object_in<S: Storage + ?Sized>(
&self,
file_data: &S,
id_bytes: &[u8],
offset_size: u8,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
crate::lookup_stats::heap_object_read(); crate::lookup_stats::heap_object_read();
let Some(&first) = id_bytes.first() else { let Some(&first) = id_bytes.first() else {
@@ -385,7 +425,11 @@ impl FractalHeapHeader {
} }
/// Read a huge object (heap ID type 1). /// Read a huge object (heap ID type 1).
fn read_huge_object(&self, file_data: &[u8], id: &[u8]) -> Result<Vec<u8>, FormatError> { fn read_huge_object<S: Storage + ?Sized>(
&self,
file: &S,
id: &[u8],
) -> Result<Vec<u8>, FormatError> {
let os = usize::from(self.offset_size); let os = usize::from(self.offset_size);
let ls = usize::from(self.length_size); let ls = usize::from(self.length_size);
// (address, stored length, filter mask, decoded length); the last two // (address, stored length, filter mask, decoded length); the last two
@@ -416,18 +460,17 @@ impl FractalHeapHeader {
let key_len = (usize::from(self.heap_id_length).saturating_sub(1)).min(8); let key_len = (usize::from(self.heap_id_length).saturating_sub(1)).min(8);
ensure_len(id, 1, key_len)?; ensure_len(id, 1, key_len)?;
let key = le_uint(&id[1..1 + key_len]); let key = le_uint(&id[1..1 + key_len]);
self.find_huge_record(file_data, key)? self.find_huge_record(file, key)?
}; };
let start = usize::try_from(addr).map_err(|_| heap_error("huge object address"))?; let start = usize::try_from(addr).map_err(|_| heap_error("huge object address"))?;
let len = usize::try_from(stored_len).map_err(|_| heap_error("huge object length"))?; let len = usize::try_from(stored_len).map_err(|_| heap_error("huge object length"))?;
ensure_len(file_data, start, len)?; let stored = read_exact_at(file, start as u64, len)?;
let stored = &file_data[start..start + len];
match &self.filter_pipeline { match &self.filter_pipeline {
None => Ok(stored.to_vec()), None => Ok(stored.into_owned()),
Some(pipeline) => { Some(pipeline) => {
let mem = usize::try_from(mem_len).map_err(|_| heap_error("huge object size"))?; let mem = usize::try_from(mem_len).map_err(|_| heap_error("huge object size"))?;
let out = crate::filters::decompress_chunk_masked(stored, pipeline, mem, 1, mask)?; let out = crate::filters::decompress_chunk_masked(&stored, pipeline, mem, 1, mask)?;
if out.len() != mem { if out.len() != mem {
return Err(heap_error("filtered huge object decoded to the wrong size")); return Err(heap_error("filtered huge object decoded to the wrong size"));
} }
@@ -438,9 +481,9 @@ impl FractalHeapHeader {
/// Look up huge object `key` in the huge-object v2 B-tree, returning /// Look up huge object `key` in the huge-object v2 B-tree, returning
/// (address, stored length, filter mask, decoded length). /// (address, stored length, filter mask, decoded length).
fn find_huge_record( fn find_huge_record<S: Storage + ?Sized>(
&self, &self,
file_data: &[u8], file: &S,
key: u64, key: u64,
) -> Result<(u64, u64, u32, u64), FormatError> { ) -> Result<(u64, u64, u32, u64), FormatError> {
if is_undefined(self.huge_btree_address, self.offset_size) { if is_undefined(self.huge_btree_address, self.offset_size) {
@@ -448,6 +491,8 @@ impl FractalHeapHeader {
"huge object ID but the heap has no huge-object index", "huge object ID but the heap has no huge-object index",
)); ));
} }
// The v2 B-tree is read from a slice until it is converted.
let file_data = require_contiguous(file, "a huge fractal-heap object's B-tree")?;
let hdr = BTreeV2Header::parse( let hdr = BTreeV2Header::parse(
file_data, file_data,
to_usize(self.huge_btree_address)?, to_usize(self.huge_btree_address)?,
@@ -519,9 +564,9 @@ impl FractalHeapHeader {
} }
/// Read a managed object (heap ID type 0). /// Read a managed object (heap ID type 0).
fn read_heap_managed( fn read_heap_managed<S: Storage + ?Sized>(
&self, &self,
file_data: &[u8], file_data: &S,
id_bytes: &[u8], id_bytes: &[u8],
offset_size: u8, offset_size: u8,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
@@ -569,9 +614,9 @@ impl FractalHeapHeader {
/// header), so we just add it to the block address minus the block's heap /// header), so we just add it to the block address minus the block's heap
/// offset. A filtered heap stores each direct block (header included) /// offset. A filtered heap stores each direct block (header included)
/// through its filter pipeline, so the block is decoded first. /// through its filter pipeline, so the block is decoded first.
fn read_from_direct_block( fn read_from_direct_block<S: Storage + ?Sized>(
&self, &self,
file_data: &[u8], file: &S,
block: DirectBlock, block: DirectBlock,
target_offset: u64, target_offset: u64,
length: usize, length: usize,
@@ -587,9 +632,9 @@ impl FractalHeapHeader {
let stored_len = usize::try_from(block.filtered_size) let stored_len = usize::try_from(block.filtered_size)
.map_err(|_| heap_error("direct block size"))?; .map_err(|_| heap_error("direct block size"))?;
let size = usize::try_from(block.size).map_err(|_| heap_error("direct block size"))?; let size = usize::try_from(block.size).map_err(|_| heap_error("direct block size"))?;
ensure_len(file_data, block.addr, stored_len)?; let stored = read_exact_at(file, block.addr as u64, stored_len)?;
let decoded = crate::filters::decompress_chunk_masked( let decoded = crate::filters::decompress_chunk_masked(
&file_data[block.addr..block.addr + stored_len], &stored,
pipeline, pipeline,
size, size,
1, 1,
@@ -603,17 +648,16 @@ impl FractalHeapHeader {
.checked_add(local_offset) .checked_add(local_offset)
.ok_or(FormatError::UnexpectedEof { .ok_or(FormatError::UnexpectedEof {
expected: usize::MAX, expected: usize::MAX,
available: file_data.len(), available: len_usize(file),
})?; })?;
ensure_len(file_data, pos, length)?; Ok(read_exact_at(file, pos as u64, length)?.into_owned())
Ok(file_data[pos..pos + length].to_vec())
} }
/// Read an object by traversing an indirect block to find the right direct block. /// Read an object by traversing an indirect block to find the right direct block.
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
fn read_from_indirect_block( fn read_from_indirect_block<S: Storage + ?Sized>(
&self, &self,
file_data: &[u8], file: &S,
iblock_addr: usize, iblock_addr: usize,
nrows: u16, nrows: u16,
iblock_heap_offset: u64, iblock_heap_offset: u64,
@@ -627,29 +671,169 @@ impl FractalHeapHeader {
"fractal heap: maximum recursion depth exceeded".into(), "fractal heap: maximum recursion depth exceeded".into(),
)); ));
} }
// Parse indirect block header
ensure_len(file_data, iblock_addr, 4)?;
if &file_data[iblock_addr..iblock_addr + 4] != b"FHIB" {
return Err(FormatError::InvalidFractalHeapSignature);
}
let block_offset_bytes = (self.max_heap_size as usize).div_ceil(8); let block_offset_bytes = (self.max_heap_size as usize).div_ceil(8);
let iblock_header = 5 + offset_size as usize + block_offset_bytes; let iblock_header = 5 + offset_size as usize + block_offset_bytes;
let mut pos = iblock_addr + iblock_header;
let tw = self.table_width as u64;
let nrows_usize = nrows as usize; let nrows_usize = nrows as usize;
let mut current_heap_offset = iblock_heap_offset;
// Rows below max_direct_rows hold direct blocks; rows at/above hold // Rows below max_direct_rows hold direct blocks; rows at/above hold
// child indirect blocks. (NOT the FRHP "starting rows" field.) // child indirect blocks. (NOT the FRHP "starting rows" field.)
let start_indirect = self.max_direct_rows(); let start_indirect = self.max_direct_rows();
let max_direct_rows = nrows_usize.min(start_indirect); let max_direct_rows = nrows_usize.min(start_indirect);
// The block up to its last child entry. The walk below reads
// entries in order and stops at the one covering the target, which
// the geometry alone locates, so the first window ends there: a
// header claiming a huge table costs a read of the entries in front
// of the target, not of the rest of the file. Only when that entry
// is unallocated (or none covers the target) does the walk go on,
// over the whole block. Either window holds what it was asked for or
// ends at the end of the file, so its bounds checks are the
// whole-file ones.
let direct_entry = usize::from(offset_size)
+ if self.filter_pipeline.is_some() {
usize::from(self.length_size) + 4
} else {
0
};
let direct_entries = max_direct_rows.saturating_mul(usize::from(self.table_width));
let entries_len = |n: usize| {
n.min(direct_entries)
.saturating_mul(direct_entry)
.saturating_add(
n.saturating_sub(direct_entries)
.saturating_mul(usize::from(offset_size)),
)
};
let all_entries = direct_entries.saturating_add(
nrows_usize
.saturating_sub(start_indirect)
.saturating_mul(usize::from(self.table_width)),
);
let block_len = iblock_header.saturating_add(entries_len(all_entries));
let target_entry = self.indirect_entry_for(nrows_usize, iblock_heap_offset, target_offset);
let first_len = target_entry.map_or(block_len, |i| {
iblock_header
.saturating_add(entries_len(i.saturating_add(1)))
.min(block_len)
});
let mut next = self.walk_indirect_block(
&Window::read(file, iblock_addr as u64, first_len)?,
nrows_usize,
iblock_heap_offset,
target_offset,
offset_size,
target_entry.map_or(usize::MAX, |i| i.saturating_add(1)),
)?;
if next.is_none() && first_len < block_len {
next = self.walk_indirect_block(
&Window::read(file, iblock_addr as u64, block_len)?,
nrows_usize,
iblock_heap_offset,
target_offset,
offset_size,
usize::MAX,
)?;
}
match next {
Some(IndirectChild::Direct(block)) => {
self.read_from_direct_block(file, block, target_offset, length)
}
Some(IndirectChild::Indirect {
addr,
nrows,
heap_offset,
}) => self.read_from_indirect_block(
file,
addr,
nrows,
heap_offset,
target_offset,
length,
offset_size,
depth_remaining - 1,
),
None => Err(FormatError::UnexpectedEof {
expected: to_usize(target_offset)?.saturating_add(length),
available: len_usize(file),
}),
}
}
/// Which child entry of an indirect block (numbered in walk order:
/// direct rows, then indirect rows) covers `target_offset`, from the
/// doubling-table geometry alone — the entry
/// [`Self::walk_indirect_block`] stops at if it is allocated. `None`
/// when no entry does.
fn indirect_entry_for(&self, nrows: usize, heap_offset: u64, target: u64) -> Option<usize> {
// The walk adds block sizes with saturation; in u128 the same test
// is `cur <= target < cur + size` without it (a target of u64::MAX
// is never inside a saturated range).
if target == u64::MAX {
return None;
}
let (tw, target) = (u128::from(self.table_width), u128::from(target));
let mut cur = u128::from(heap_offset);
let mut before = 0usize;
for row in 0..nrows {
if target < cur {
return None;
}
// Direct and indirect rows alike span this row's block size per
// entry.
let size = u128::from(self.block_size_for_row(row));
let span = size * tw;
if size > 0 && target < cur + span {
let col = usize::try_from((target - cur) / size).ok()?;
return before.checked_add(col);
}
cur += span;
before = before.saturating_add(self.table_width as usize);
}
None
}
/// Walk an indirect block's child entries in order, in the window `w`
/// (the block from its signature on), and return the allocated child
/// covering `target_offset`, or `None` when no entry among the first
/// `limit` does.
fn walk_indirect_block(
&self,
w: &Window<'_>,
nrows: usize,
iblock_heap_offset: u64,
target_offset: u64,
offset_size: u8,
limit: usize,
) -> Result<Option<IndirectChild>, FormatError> {
let ensure_len = |_: &[u8], pos: usize, needed: usize| w.ensure(pos, needed);
let read_offset = |_: &[u8], pos: usize, size: u8| {
w.ensure(pos, usize::from(size))?;
read_offset(&w.bytes, pos, size)
};
let file_data: &[u8] = &w.bytes;
let block_offset_bytes = (self.max_heap_size as usize).div_ceil(8);
let iblock_header = 5 + offset_size as usize + block_offset_bytes;
let tw = self.table_width as u64;
let mut current_heap_offset = iblock_heap_offset;
let start_indirect = self.max_direct_rows();
let max_direct_rows = nrows.min(start_indirect);
let mut walked = 0usize;
// Parse indirect block header
ensure_len(file_data, 0, 4)?;
if &file_data[..4] != b"FHIB" {
return Err(FormatError::InvalidFractalHeapSignature);
}
let mut pos = iblock_header;
for row in 0..max_direct_rows { for row in 0..max_direct_rows {
let block_size = self.block_size_for_row(row); let block_size = self.block_size_for_row(row);
for _col in 0..tw { for _col in 0..tw {
if walked == limit {
return Ok(None);
}
walked += 1;
let child_addr = read_offset(file_data, pos, offset_size)?; let child_addr = read_offset(file_data, pos, offset_size)?;
pos += offset_size as usize; pos += offset_size as usize;
@@ -676,18 +860,13 @@ impl FractalHeapHeader {
&& target_offset >= current_heap_offset && target_offset >= current_heap_offset
&& target_offset < block_end && target_offset < block_end
{ {
return self.read_from_direct_block( return Ok(Some(IndirectChild::Direct(DirectBlock {
file_data, addr: to_usize(child_addr)?,
DirectBlock { size: block_size,
addr: to_usize(child_addr)?, heap_offset: current_heap_offset,
size: block_size, filtered_size,
heap_offset: current_heap_offset, filter_mask,
filtered_size, })));
filter_mask,
},
target_offset,
length,
);
} }
current_heap_offset = block_end; current_heap_offset = block_end;
} }
@@ -696,11 +875,15 @@ impl FractalHeapHeader {
// Rows at and above `start_indirect` hold child indirect blocks. A // Rows at and above `start_indirect` hold child indirect blocks. A
// child in row r spans exactly that row's block size of heap space, // child in row r spans exactly that row's block size of heap space,
// so it has as many rows as a table of that total size needs. // so it has as many rows as a table of that total size needs.
for row in start_indirect..nrows_usize { for row in start_indirect..nrows {
let child_space = self.block_size_for_row(row); let child_space = self.block_size_for_row(row);
let child_nrows = self.rows_for_size(child_space); let child_nrows = self.rows_for_size(child_space);
for _col in 0..tw { for _col in 0..tw {
if walked == limit {
return Ok(None);
}
walked += 1;
let child_addr = read_offset(file_data, pos, offset_size)?; let child_addr = read_offset(file_data, pos, offset_size)?;
pos += offset_size as usize; pos += offset_size as usize;
@@ -709,25 +892,16 @@ impl FractalHeapHeader {
&& target_offset >= current_heap_offset && target_offset >= current_heap_offset
&& target_offset < block_end && target_offset < block_end
{ {
return self.read_from_indirect_block( return Ok(Some(IndirectChild::Indirect {
file_data, addr: to_usize(child_addr)?,
to_usize(child_addr)?, nrows: child_nrows,
child_nrows, heap_offset: current_heap_offset,
current_heap_offset, }));
target_offset,
length,
offset_size,
depth_remaining - 1,
);
} }
current_heap_offset = block_end; current_heap_offset = block_end;
} }
} }
Ok(None)
Err(FormatError::UnexpectedEof {
expected: to_usize(target_offset)?.saturating_add(length),
available: file_data.len(),
})
} }
/// Number of rows in the doubling table whose block size is at most the /// Number of rows in the doubling table whose block size is at most the
@@ -771,6 +945,16 @@ impl FractalHeapHeader {
/// A managed direct block's location, extent and (for a filtered heap) its /// A managed direct block's location, extent and (for a filtered heap) its
/// stored size and filter mask. /// stored size and filter mask.
/// The child of an indirect block that covers a heap offset.
enum IndirectChild {
Direct(DirectBlock),
Indirect {
addr: usize,
nrows: u16,
heap_offset: u64,
},
}
struct DirectBlock { struct DirectBlock {
addr: usize, addr: usize,
size: u64, size: u64,
@@ -1030,4 +1214,154 @@ mod tests {
let id = [0x40u8, 0, 0, 0, 0, 0, 0]; let id = [0x40u8, 0, 0, 0, 0, 0, 0];
assert!(hdr.read_managed_object(&file_data, &id, 8).is_err()); assert!(hdr.read_managed_object(&file_data, &id, 8).is_err());
} }
/// Headers, and managed (in a direct root and through an indirect
/// root), huge and tiny objects read identically through a
/// `read_at`-only storage, for every truncation of the file.
#[test]
fn storage_reads_match_slice_reads() {
use crate::storage::CountingStorage;
let (mut file, header_end) = build_simple_heap(8, 8);
// An indirect root block at 600: row 0 holds the direct block at
// 256, then three undefined blocks.
file[600..604].copy_from_slice(b"FHIB");
let mut at = 600 + 5 + 8 + 2;
for addr in [256u64, u64::MAX, u64::MAX, u64::MAX] {
file[at..at + 8].copy_from_slice(&addr.to_le_bytes());
at += 8;
}
file[900..905].copy_from_slice(b"huge!");
let managed_id = |offset: u64, len: u64| {
let payload = offset | (len << 16);
let mut id = vec![0u8];
id.extend_from_slice(&payload.to_le_bytes()[..6]);
id
};
let mut huge = vec![0x10u8];
huge.extend_from_slice(&900u64.to_le_bytes());
huge.extend_from_slice(&5u64.to_le_bytes());
let ids = [
managed_id(15, 13),
managed_id(15, 200),
managed_id(130, 4),
huge,
vec![0x22, b'a', b'b', b'c', 0, 0, 0],
];
let mut cuts: Vec<usize> = (0..=header_end + 1).collect();
cuts.extend([256, 260, 271, 280, 600, 610, 620, 640, 900, 903, file.len()]);
for cut in cuts {
let f = &file[..cut];
let storage = CountingStorage::new(f.to_vec());
let want = FractalHeapHeader::parse(f, 0, 8, 8);
let got = FractalHeapHeader::parse_in(&storage, 0, 8, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"), "cut {cut}");
let Ok(direct) = want else { continue };
let mut indirect = direct.clone();
indirect.root_block_address = 600;
indirect.current_rows_in_root_indirect_block = 1;
let mut huge_ids = direct.clone();
huge_ids.heap_id_length = 17;
for hdr in [&direct, &indirect, &huge_ids] {
for id in &ids {
assert_eq!(
hdr.read_managed_object_in(&storage, id, 8),
hdr.read_managed_object(f, id, 8),
"cut {cut}"
);
}
}
}
}
/// A header claiming a huge doubling table (width 0xFFFF, 0xFFFF rows in
/// the root indirect block) in a 16 MiB file: reading an object from the
/// table's first block reads the entries up to it, not the rest of the
/// file, and gives what the slice read gives. When the covering entry is
/// unallocated the walk goes on over the whole block, still identically.
#[test]
fn huge_table_claims_read_only_what_the_walk_needs() {
use crate::storage::CountingStorage;
let (mut file, _) = build_simple_heap(8, 8);
file.resize(16 << 20, 0);
file[600..604].copy_from_slice(b"FHIB");
let first_entry = 600 + 5 + 8 + 2;
file[first_entry..first_entry + 8].copy_from_slice(&256u64.to_le_bytes());
let mut hdr = FractalHeapHeader::parse(&file, 0, 8, 8).unwrap();
hdr.table_width = 0xFFFF;
hdr.root_block_address = 600;
hdr.current_rows_in_root_indirect_block = 0xFFFF;
let managed_id = |offset: u64, len: u64| {
let payload = offset | (len << 16);
let mut id = vec![0u8];
id.extend_from_slice(&payload.to_le_bytes()[..6]);
id
};
let storage = CountingStorage::new(file.clone());
let id = managed_id(15, 13);
let want = hdr.read_managed_object(&file, &id, 8);
assert!(want.is_ok(), "{want:?}");
storage.reset();
assert_eq!(hdr.read_managed_object_in(&storage, &id, 8), want);
assert!(
storage.bytes_read() < 1024,
"{} bytes in {} reads",
storage.bytes_read(),
storage.reads()
);
// The second entry (heap offsets 128..256) is unallocated (zero is
// not the undefined address, so make it all ones).
file[first_entry + 8..first_entry + 16].fill(0xFF);
let storage = CountingStorage::new(file.clone());
let id = managed_id(130, 4);
let want = hdr.read_managed_object(&file, &id, 8);
assert_eq!(hdr.read_managed_object_in(&storage, &id, 8), want);
}
/// A huge object found through the huge-object B-tree needs the whole
/// file in memory until the B-tree reader is converted: a clean error
/// on other storage.
#[test]
fn huge_object_btree_needs_contiguous_storage() {
use crate::storage::CountingStorage;
let (file, _) = build_simple_heap(8, 8);
let mut hdr = FractalHeapHeader::parse(&file, 0, 8, 8).unwrap();
hdr.huge_btree_address = 700;
let storage = CountingStorage::new(file);
assert_eq!(
hdr.read_managed_object_in(&storage, &[0x10, 1, 0, 0, 0, 0, 0], 8),
Err(FormatError::ContiguousStorageRequired(
"a huge fractal-heap object's B-tree"
))
);
}
/// A header with an I/O filter pipeline (read in a second, longer
/// window) parses identically through a `read_at`-only storage, for
/// every truncation.
#[test]
fn filtered_header_parses_identically_through_storage() {
use crate::storage::CountingStorage;
let (simple, header_end) = build_simple_heap(8, 8);
let pipeline = [2u8, 1, 1, 0, 0, 0, 1, 0, 6, 0, 0, 0]; // deflate, level 6
let mut header = simple[..header_end - 4].to_vec();
header[7..9].copy_from_slice(&(pipeline.len() as u16).to_le_bytes());
header.extend_from_slice(&100u64.to_le_bytes()); // root block's stored size
header.extend_from_slice(&0u32.to_le_bytes()); // its filter mask
header.extend_from_slice(&pipeline);
let sum = crate::checksum::jenkins_lookup3(&header);
header.extend_from_slice(&sum.to_le_bytes());
let mut file = header.clone();
file.resize(256, 0);
let hdr = FractalHeapHeader::parse(&file, 0, 8, 8).unwrap();
assert!(hdr.filter_pipeline.is_some());
for cut in 0..=file.len() {
let f = &file[..cut];
let storage = CountingStorage::new(f.to_vec());
assert_eq!(
format!("{:?}", FractalHeapHeader::parse_in(&storage, 0, 8, 8)),
format!("{:?}", FractalHeapHeader::parse(f, 0, 8, 8)),
"cut {cut}"
);
}
}
} }
+113 -26
View File
@@ -1,9 +1,12 @@
//! HDF5 Global Heap collection parsing. //! HDF5 Global Heap collection parsing.
#[cfg(not(feature = "std"))] #[cfg(not(feature = "std"))]
use alloc::{format, string::String, vec::Vec}; use alloc::{borrow::Cow, format, string::String, vec::Vec};
#[cfg(feature = "std")]
use std::borrow::Cow;
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::{Storage, len_usize, read_exact_at};
/// Magic signature for global heap collections. /// Magic signature for global heap collections.
const GCOL_SIGNATURE: [u8; 4] = *b"GCOL"; const GCOL_SIGNATURE: [u8; 4] = *b"GCOL";
@@ -28,19 +31,20 @@ pub struct GlobalHeapObject {
pub data: Vec<u8>, pub data: Vec<u8>,
} }
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> { /// Checks that `[offset, offset + needed)` ends by `data_len`.
fn ensure_len(data_len: usize, offset: usize, needed: usize) -> Result<(), FormatError> {
match offset.checked_add(needed) { match offset.checked_add(needed) {
Some(end) if end <= data.len() => Ok(()), Some(end) if end <= data_len => Ok(()),
_ => Err(FormatError::UnexpectedEof { _ => Err(FormatError::UnexpectedEof {
expected: offset.saturating_add(needed), expected: offset.saturating_add(needed),
available: data.len(), available: data_len,
}), }),
} }
} }
fn read_length(data: &[u8], offset: usize, length_size: u8) -> Result<u64, FormatError> { fn read_length(data: &[u8], offset: usize, length_size: u8) -> Result<u64, FormatError> {
let s = length_size as usize; let s = length_size as usize;
ensure_len(data, offset, s)?; ensure_len(data.len(), offset, s)?;
let slice = &data[offset..offset + s]; let slice = &data[offset..offset + s];
Ok(match length_size { Ok(match length_size {
2 => u16::from_le_bytes([slice[0], slice[1]]) as u64, 2 => u16::from_le_bytes([slice[0], slice[1]]) as u64,
@@ -95,7 +99,17 @@ impl GlobalHeapCollection {
offset: usize, offset: usize,
length_size: u8, length_size: u8,
) -> Result<GlobalHeapCollection, FormatError> { ) -> Result<GlobalHeapCollection, FormatError> {
let index = Self::parse_index(file_data, offset, length_size)?; Self::parse_in(file_data, offset as u64, length_size)
}
/// [`Self::parse`] over any [`Storage`]: one read of the header, one of
/// the collection.
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
length_size: u8,
) -> Result<GlobalHeapCollection, FormatError> {
let (bytes, base, index) = Self::read_collection(file, offset, length_size)?;
Ok(GlobalHeapCollection { Ok(GlobalHeapCollection {
collection_size: index.collection_size, collection_size: index.collection_size,
objects: index objects: index
@@ -104,7 +118,7 @@ impl GlobalHeapCollection {
.map(|o| GlobalHeapObject { .map(|o| GlobalHeapObject {
index: o.index, index: o.index,
reference_count: o.reference_count, reference_count: o.reference_count,
data: file_data[o.offset..o.offset + o.size].to_vec(), data: bytes[o.offset - base..o.offset - base + o.size].to_vec(),
}) })
.collect(), .collect(),
}) })
@@ -122,43 +136,72 @@ impl GlobalHeapCollection {
offset: usize, offset: usize,
length_size: u8, length_size: u8,
) -> Result<GlobalHeapIndex, FormatError> { ) -> Result<GlobalHeapIndex, FormatError> {
Self::parse_index_in(file_data, offset as u64, length_size)
}
/// [`Self::parse_index`] over any [`Storage`]: one read of the header,
/// one of the collection. The object offsets are file offsets.
pub fn parse_index_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
length_size: u8,
) -> Result<GlobalHeapIndex, FormatError> {
Ok(Self::read_collection(file, offset, length_size)?.2)
}
/// Read the collection at `offset` and index its objects: the
/// collection's bytes, its offset as a `usize`, and the index (with
/// file offsets).
fn read_collection<S: Storage + ?Sized>(
file: &S,
offset: u64,
length_size: u8,
) -> Result<(Cow<'_, [u8]>, usize, GlobalHeapIndex), FormatError> {
let file_len = len_usize(file);
// signature(4) + version(1) + reserved(3) + collection_size(length_size), // signature(4) + version(1) + reserved(3) + collection_size(length_size),
// padded to a multiple of 8 as libhdf5 lays it out (`H5HG_SIZEOF_HDR`). // padded to a multiple of 8 as libhdf5 lays it out (`H5HG_SIZEOF_HDR`).
// With 8-byte lengths the padding is 0; with 4-byte lengths it is 4, // With 8-byte lengths the padding is 0; with 4-byte lengths it is 4,
// and reading without it put every object 4 bytes early. // and reading without it put every object 4 bytes early.
let header_size = pad8(8 + length_size as usize); let header_size = pad8(8 + length_size as usize);
ensure_len(file_data, offset, header_size)?; let header = read_exact_at(file, offset, header_size)?;
let offset = usize::try_from(offset).map_err(|_| FormatError::UnexpectedEof {
expected: usize::MAX,
available: file_len,
})?;
if file_data[offset..offset + 4] != GCOL_SIGNATURE { if header[..4] != GCOL_SIGNATURE {
return Err(FormatError::InvalidGlobalHeapSignature); return Err(FormatError::InvalidGlobalHeapSignature);
} }
let version = file_data[offset + 4]; let version = header[4];
if version != 1 { if version != 1 {
return Err(FormatError::InvalidGlobalHeapVersion(version)); return Err(FormatError::InvalidGlobalHeapVersion(version));
} }
let collection_size = read_length(file_data, offset + 8, length_size)?; let collection_size = read_length(&header, 8, length_size)?;
let collection_end = usize::try_from(collection_size) let collection_end = usize::try_from(collection_size)
.ok() .ok()
.and_then(|size| offset.checked_add(size)) .and_then(|size| offset.checked_add(size))
.ok_or(FormatError::UnexpectedEof { .ok_or(FormatError::UnexpectedEof {
expected: usize::MAX, expected: usize::MAX,
available: file_data.len(), available: file_len,
})?; })?;
if collection_end > file_data.len() { if collection_end > file_len {
return Err(FormatError::UnexpectedEof { return Err(FormatError::UnexpectedEof {
expected: collection_end, expected: collection_end,
available: file_data.len(), available: file_len,
}); });
} }
let collection = read_exact_at(file, offset as u64, collection_end - offset)?;
// Positions below are file offsets; `file_data(p)` is the byte at `p`.
let file_data = |p: usize| collection[p - offset];
let mut pos = offset + header_size; let mut pos = offset + header_size;
let mut objects = Vec::new(); let mut objects = Vec::new();
// Parse objects until we hit index 0 (free space) or run out of space // Parse objects until we hit index 0 (free space) or run out of space
while pos + 2 <= collection_end { while pos + 2 <= collection_end {
let object_index = u16::from_le_bytes([file_data[pos], file_data[pos + 1]]); let object_index = u16::from_le_bytes([file_data(pos), file_data(pos + 1)]);
if object_index == 0 { if object_index == 0 {
// Free space marker — done // Free space marker — done
@@ -168,11 +211,12 @@ impl GlobalHeapCollection {
// object_index(2) + reference_count(2) + reserved(4) + // object_index(2) + reference_count(2) + reserved(4) +
// object_size(length_size), padded to 8 (`H5HG_SIZEOF_OBJHDR`). // object_size(length_size), padded to 8 (`H5HG_SIZEOF_OBJHDR`).
let obj_header_size = pad8(8 + length_size as usize); let obj_header_size = pad8(8 + length_size as usize);
ensure_len(&file_data[..collection_end], pos, obj_header_size)?; ensure_len(collection_end, pos, obj_header_size)?;
let reference_count = u16::from_le_bytes([file_data[pos + 2], file_data[pos + 3]]); let reference_count = u16::from_le_bytes([file_data(pos + 2), file_data(pos + 3)]);
let object_size = usize::try_from(read_length(file_data, pos + 8, length_size)?) let object_size =
.map_err(|_| FormatError::Overflow("global heap object size".into()))?; usize::try_from(read_length(&collection[pos - offset..], 8, length_size)?)
.map_err(|_| FormatError::Overflow("global heap object size".into()))?;
pos += obj_header_size; pos += obj_header_size;
if pos if pos
@@ -197,10 +241,11 @@ impl GlobalHeapCollection {
pos = pos.saturating_add(pad8(object_size)); pos = pos.saturating_add(pad8(object_size));
} }
Ok(GlobalHeapIndex { let index = GlobalHeapIndex {
collection_size, collection_size,
objects, objects,
}) };
Ok((collection, offset, index))
} }
/// Get an object by its index. /// Get an object by its index.
@@ -226,7 +271,7 @@ mod tests {
let mut obj_size_total = 0usize; let mut obj_size_total = 0usize;
for (_, _, data) in objects { for (_, _, data) in objects {
let obj_header = pad8(8 + ls); let obj_header = pad8(8 + ls);
obj_size_total += obj_header + pad8(data.len()); obj_size_total += obj_header + pad8(<[u8]>::len(data));
} }
// Free space marker (2 bytes for index 0) // Free space marker (2 bytes for index 0)
obj_size_total += 2; obj_size_total += 2;
@@ -251,15 +296,17 @@ mod tests {
buf.extend_from_slice(&ref_count.to_le_bytes()); buf.extend_from_slice(&ref_count.to_le_bytes());
buf.extend_from_slice(&[0u8; 4]); // reserved buf.extend_from_slice(&[0u8; 4]); // reserved
match length_size { match length_size {
4 => buf.extend_from_slice(&(data.len() as u32).to_le_bytes()), // `<[u8]>::len`: with `Storage` in scope `data.len()` on a
8 => buf.extend_from_slice(&(data.len() as u64).to_le_bytes()), // `&&[u8]` resolves to `Storage::len` (a `u64`).
4 => buf.extend_from_slice(&(<[u8]>::len(data) as u32).to_le_bytes()),
8 => buf.extend_from_slice(&(<[u8]>::len(data) as u64).to_le_bytes()),
_ => panic!("unsupported"), _ => panic!("unsupported"),
} }
buf.resize(buf.len() + (pad8(8 + ls) - (8 + ls)), 0); buf.resize(buf.len() + (pad8(8 + ls) - (8 + ls)), 0);
buf.extend_from_slice(data); buf.extend_from_slice(data);
// Pad to 8 bytes // Pad to 8 bytes
let padded = pad8(data.len()); let padded = pad8(<[u8]>::len(data));
buf.resize(buf.len() + (padded - data.len()), 0); buf.resize(buf.len() + (padded - <[u8]>::len(data)), 0);
} }
// Free space marker // Free space marker
@@ -327,4 +374,44 @@ mod tests {
assert_eq!(coll.objects.len(), 1); assert_eq!(coll.objects.len(), 1);
assert_eq!(coll.objects[0].data, b"test"); assert_eq!(coll.objects[0].data, b"test");
} }
/// Collections, and every truncation of them, index and parse
/// identically through a `read_at`-only storage: two reads each.
#[test]
fn storage_parse_matches_slice_parse() {
use crate::storage::CountingStorage;
let objs: &[(u16, u16, &[u8])] = &[(1, 1, b"hello"), (2, 3, b"a longer object")];
for ls in [4u8, 8] {
let coll = build_collection(objs, ls);
let mut corrupt = coll.clone();
corrupt[8] = 200; // collection size past the end of the file
let mut overrun = coll.clone();
let size_at = pad8(8 + ls as usize) + 8;
overrun[size_at] = 250; // first object runs past the collection
for full in [coll, corrupt, overrun] {
for at in [0usize, 5] {
for cut in 0..=full.len() {
let mut f = vec![0u8; at];
f.extend_from_slice(&full[..cut]);
let storage = CountingStorage::new(f.clone());
let want = GlobalHeapCollection::parse(&f, at, ls);
let got = GlobalHeapCollection::parse_in(&storage, at as u64, ls);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
let want = GlobalHeapCollection::parse_index(&f, at, ls);
let got = GlobalHeapCollection::parse_index_in(&storage, at as u64, ls);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
}
}
}
}
let storage = CountingStorage::new(build_collection(objs, 8));
assert_eq!(
GlobalHeapCollection::parse_in(&storage, 0, 8)
.unwrap()
.objects
.len(),
2
);
assert_eq!(storage.reads(), 2);
}
} }
+1
View File
@@ -122,6 +122,7 @@ pub mod property_list;
pub mod selection; pub mod selection;
pub mod shared_message; pub mod shared_message;
pub mod signature; pub mod signature;
pub mod storage;
pub mod superblock; pub mod superblock;
pub mod superblock_ext; pub mod superblock_ext;
pub mod symbol_table; pub mod symbol_table;
+144 -36
View File
@@ -5,6 +5,7 @@ use alloc::string::String;
use crate::addr::to_usize; use crate::addr::to_usize;
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::{Storage, len_usize, read_exact_at};
/// Parsed HDF5 Local Heap header. /// Parsed HDF5 Local Heap header.
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
@@ -17,21 +18,6 @@ pub struct LocalHeap {
pub data_segment_address: u64, pub data_segment_address: u64,
} }
/// Checks that `[offset, offset + needed)` fits within `data`, guarding the
/// addition against `usize` overflow from a crafted near-`usize::MAX` offset.
fn ensure_len(data: &[u8], offset: usize, needed: usize) -> Result<(), FormatError> {
if offset
.checked_add(needed)
.is_none_or(|end| end > data.len())
{
return Err(FormatError::UnexpectedEof {
expected: offset.saturating_add(needed),
available: data.len(),
});
}
Ok(())
}
fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> { fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
let s = size as usize; let s = size as usize;
if pos.checked_add(s).is_none_or(|end| end > data.len()) { if pos.checked_add(s).is_none_or(|end| end > data.len()) {
@@ -51,6 +37,10 @@ fn read_offset(data: &[u8], pos: usize, size: u8) -> Result<u64, FormatError> {
}) })
} }
/// First read of a name on a backend without the file in memory: most link
/// names are shorter than this.
const NAME_READ_START: usize = 64;
impl LocalHeap { impl LocalHeap {
/// Parse a local heap header at the given offset in the file data. /// Parse a local heap header at the given offset in the file data.
pub fn parse( pub fn parse(
@@ -58,12 +48,24 @@ impl LocalHeap {
offset: usize, offset: usize,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<LocalHeap, FormatError> {
Self::parse_in(file_data, offset as u64, offset_size, length_size)
}
/// [`Self::parse`] over any [`Storage`]: one read of the header.
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
offset_size: u8,
length_size: u8,
) -> Result<LocalHeap, FormatError> { ) -> Result<LocalHeap, FormatError> {
// signature(4) + version(1) + reserved(3) = 8, then length_size*2 + offset_size // signature(4) + version(1) + reserved(3) = 8, then length_size*2 + offset_size
let ls = length_size as usize; let ls = length_size as usize;
let os = offset_size as usize; let os = offset_size as usize;
let total = 8 + ls * 2 + os; let total = 8 + ls * 2 + os;
ensure_len(file_data, offset, total)?; let header = read_exact_at(file, offset, total)?;
let file_data: &[u8] = &header;
let offset = 0usize;
if &file_data[offset..offset + 4] != b"HEAP" { if &file_data[offset..offset + 4] != b"HEAP" {
return Err(FormatError::InvalidLocalHeapSignature); return Err(FormatError::InvalidLocalHeapSignature);
@@ -100,6 +102,16 @@ impl LocalHeap {
/// The end of the list is `H5HL_FREE_NULL` (1); an all-ones value (the /// The end of the list is `H5HL_FREE_NULL` (1); an all-ones value (the
/// undefined address) is accepted as "no free list" too. /// undefined address) is accepted as "no free list" too.
pub fn validate_free_list(&self, file_data: &[u8], length_size: u8) -> Result<(), FormatError> { pub fn validate_free_list(&self, file_data: &[u8], length_size: u8) -> Result<(), FormatError> {
self.validate_free_list_in(file_data, length_size)
}
/// [`Self::validate_free_list`] over any [`Storage`]: two small reads
/// per free block.
pub fn validate_free_list_in<S: Storage + ?Sized>(
&self,
file: &S,
length_size: u8,
) -> Result<(), FormatError> {
const FREE_NULL: u64 = 1; const FREE_NULL: u64 = 1;
let ls = length_size as usize; let ls = length_size as usize;
let undefined = if ls >= 8 { let undefined = if ls >= 8 {
@@ -124,11 +136,12 @@ impl LocalHeap {
.and_then(|a| usize::try_from(a).ok()) .and_then(|a| usize::try_from(a).ok())
.ok_or(FormatError::InvalidLocalHeapFreeList)?; .ok_or(FormatError::InvalidLocalHeapFreeList)?;
let block_offset = next; let block_offset = next;
next = read_offset(file_data, at, length_size)?; next = read_offset(&read_exact_at(file, at as u64, ls)?, 0, length_size)?;
if next == 0 { if next == 0 {
return Err(FormatError::InvalidLocalHeapFreeList); return Err(FormatError::InvalidLocalHeapFreeList);
} }
let block_size = read_offset(file_data, at + ls, length_size)?; let block_size =
read_offset(&read_exact_at(file, (at + ls) as u64, ls)?, 0, length_size)?;
if block_offset if block_offset
.checked_add(block_size) .checked_add(block_size)
.is_none_or(|end| end > size) .is_none_or(|end| end > size)
@@ -141,6 +154,18 @@ impl LocalHeap {
/// Read a null-terminated string from the heap's data segment at the given byte offset. /// Read a null-terminated string from the heap's data segment at the given byte offset.
pub fn read_string(&self, file_data: &[u8], string_offset: u64) -> Result<String, FormatError> { pub fn read_string(&self, file_data: &[u8], string_offset: u64) -> Result<String, FormatError> {
self.read_string_in(file_data, string_offset)
}
/// [`Self::read_string`] over any [`Storage`]: one read of up to 64
/// bytes for a short name, more (each four times the last) up to the end
/// of the data segment for a longer one.
pub fn read_string_in<S: Storage + ?Sized>(
&self,
file: &S,
string_offset: u64,
) -> Result<String, FormatError> {
let file_len = len_usize(file);
let seg_addr = to_usize(self.data_segment_address)?; let seg_addr = to_usize(self.data_segment_address)?;
let str_start = let str_start =
seg_addr seg_addr
@@ -154,30 +179,40 @@ impl LocalHeap {
"local heap seg_addr + data_segment_size overflow".into(), "local heap seg_addr + data_segment_size overflow".into(),
))?; ))?;
if str_start >= file_data.len() || str_start >= seg_end { if str_start >= file_len || str_start >= seg_end {
return Err(FormatError::UnexpectedEof { return Err(FormatError::UnexpectedEof {
expected: str_start + 1, expected: str_start + 1,
available: file_data.len(), available: file_len,
}); });
} }
// Find null terminator // Find the null terminator, which lies before the end of the data
let search_end = seg_end.min(file_data.len()); // segment (or of the file). In memory that is one borrowed slice;
let mut end = str_start; // otherwise the bytes are read in growing pieces, so a name costs a
while end < search_end && file_data[end] != 0 { // read of about its own length, not of the rest of the segment
end += 1; // (whose size is an untrusted header field).
let search_end = seg_end.min(file_len);
let total = search_end - str_start;
let mut want = if file.as_contiguous().is_some() {
total
} else {
total.min(NAME_READ_START)
};
loop {
let rest = read_exact_at(file, str_start as u64, want)?;
if let Some(len) = rest.iter().position(|&b| b == 0) {
let s = core::str::from_utf8(&rest[..len])
.map_err(|_| FormatError::InvalidLocalHeapSignature)?;
return Ok(String::from(s));
}
if want == total {
return Err(FormatError::UnexpectedEof {
expected: search_end + 1,
available: search_end,
});
}
want = want.saturating_mul(4).min(total);
} }
if end >= search_end {
return Err(FormatError::UnexpectedEof {
expected: end + 1,
available: search_end,
});
}
let s = core::str::from_utf8(&file_data[str_start..end])
.map_err(|_| FormatError::InvalidLocalHeapSignature)?;
Ok(String::from(s))
} }
} }
@@ -346,4 +381,77 @@ mod tests {
let err = LocalHeap::parse(&file, 0, 8, 8).unwrap_err(); let err = LocalHeap::parse(&file, 0, 8, 8).unwrap_err();
assert_eq!(err, FormatError::InvalidLocalHeapVersion(1)); assert_eq!(err, FormatError::InvalidLocalHeapVersion(1));
} }
/// Header, free list and strings read identically through a
/// `read_at`-only storage, for every truncation of the file.
#[test]
fn storage_reads_match_slice_reads() {
use crate::storage::CountingStorage;
let plain = build_heap_file(0, 64, &["", "alpha", "beta"], 8, 8);
// A free block of 16 bytes at segment offset 12, ending the list.
let mut free = build_heap_file(0, 64, &["", "alpha", "beta", &"x".repeat(20)], 8, 8);
free[16..24].copy_from_slice(&12u64.to_le_bytes());
free[64 + 12..64 + 20].copy_from_slice(&1u64.to_le_bytes());
free[64 + 20..64 + 28].copy_from_slice(&16u64.to_le_bytes());
let mut bad_free = free.clone();
bad_free[64 + 20..64 + 28].copy_from_slice(&99u64.to_le_bytes());
for full in [plain, free, bad_free] {
for cut in 0..=full.len() {
let f = &full[..cut];
let storage = CountingStorage::new(f.to_vec());
let want = LocalHeap::parse(f, 0, 8, 8);
let got = LocalHeap::parse_in(&storage, 0, 8, 8);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
let Ok(heap) = want else { continue };
assert_eq!(
heap.validate_free_list_in(&storage, 8),
heap.validate_free_list(f, 8)
);
for off in [0u64, 1, 2, 6, 7, 11, 100] {
assert_eq!(heap.read_string_in(&storage, off), heap.read_string(f, off));
}
}
}
}
/// Names of every length around the first read's size, and one with no
/// terminator, read identically through a `read_at`-only storage; a
/// short name in a heap whose header claims a huge data segment costs
/// one small read, not a read of the rest of the file.
#[test]
fn long_names_and_hostile_segment_sizes() {
use crate::storage::CountingStorage;
let names: Vec<String> = [0usize, 1, 63, 64, 65, 255, 256, 257, 1000, 5000]
.iter()
.map(|&n| "n".repeat(n))
.collect();
let refs: Vec<&str> = names.iter().map(String::as_str).collect();
let mut file = build_heap_file(0, 64, &refs, 8, 8);
let heap = LocalHeap::parse(&file, 0, 8, 8).unwrap();
let storage = CountingStorage::new(file.clone());
let mut off = 0u64;
for name in &names {
let got = heap.read_string_in(&storage, off);
assert_eq!(got, heap.read_string(&file, off));
assert_eq!(got.unwrap(), *name);
off += name.len() as u64 + 1;
}
// The last name loses its terminator: both report the same error.
let seg_end = 64 + heap.data_segment_size as usize;
file[seg_end - 1] = b'n';
let storage = CountingStorage::new(file.clone());
let last = off - names[names.len() - 1].len() as u64 - 1;
let want = heap.read_string(&file, last);
assert!(want.is_err());
assert_eq!(heap.read_string_in(&storage, last), want);
// A 64 MiB file whose heap claims a data segment reaching its end.
let mut big = build_heap_file(0, 64, &["short", "names"], 8, 8);
big.resize(64 << 20, 0);
big[8..16].copy_from_slice(&((64u64 << 20) - 64).to_le_bytes());
let heap = LocalHeap::parse(&big, 0, 8, 8).unwrap();
let storage = CountingStorage::new(big.clone());
assert_eq!(heap.read_string_in(&storage, 6).unwrap(), "names");
assert_eq!((storage.reads(), storage.bytes_read()), (1, 64));
}
} }
+138 -42
View File
@@ -8,6 +8,7 @@ use byteorder::{ByteOrder, LittleEndian};
use crate::addr::to_usize; use crate::addr::to_usize;
use crate::error::FormatError; use crate::error::FormatError;
use crate::message_type::MessageType; use crate::message_type::MessageType;
use crate::storage::{Storage, Window, len_usize, read_exact_at};
/// OHDR signature for v2 object headers. /// OHDR signature for v2 object headers.
const OHDR_SIGNATURE: [u8; 4] = *b"OHDR"; const OHDR_SIGNATURE: [u8; 4] = *b"OHDR";
@@ -119,32 +120,52 @@ impl ObjectHeader {
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<ObjectHeader, FormatError> { ) -> Result<ObjectHeader, FormatError> {
ensure_len(data, offset, 4)?; Self::parse_in(data, offset as u64, offset_size, length_size)
if data[offset..offset + 4] == OHDR_SIGNATURE { }
Self::parse_v2(data, offset, offset_size, length_size)
/// [`Self::parse`] over any [`Storage`].
///
/// Reads the prefix (at most [`V2_PREFIX_MAX`] bytes, signature
/// included), then each chunk as one bounded read, continuation chunks
/// included.
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
offset_size: u8,
length_size: u8,
) -> Result<ObjectHeader, FormatError> {
// The longest prefix of either version, in one read. It holds the
// whole prefix or ends at the end of the file, so its bounds checks
// are the whole-file ones.
let prefix = Window::read(file, offset, V2_PREFIX_MAX)?;
prefix.ensure(0, 4)?;
if prefix.bytes[..4] == OHDR_SIGNATURE {
Self::parse_v2(file, offset, &prefix, offset_size, length_size)
} else { } else {
Self::parse_v1(data, offset, offset_size, length_size) Self::parse_v1(file, offset, &prefix, offset_size, length_size)
} }
} }
fn parse_v1( fn parse_v1<S: Storage + ?Sized>(
data: &[u8], file: &S,
offset: usize, offset: u64,
prefix: &Window<'_>,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<ObjectHeader, FormatError> { ) -> Result<ObjectHeader, FormatError> {
// version(1) + reserved(1) + num_messages(2) + ref_count(4) + header_size(4) = 12 // version(1) + reserved(1) + num_messages(2) + ref_count(4) + header_size(4) = 12
// then pad to 8-byte alignment from start of header // then pad to 8-byte alignment from start of header
ensure_len(data, offset, 12)?; prefix.ensure(0, 12)?;
let prefix = &prefix.bytes[..12];
let version = data[offset]; let version = prefix[0];
if version != 1 { if version != 1 {
return Err(FormatError::InvalidObjectHeaderVersion(version)); return Err(FormatError::InvalidObjectHeaderVersion(version));
} }
let num_messages = LittleEndian::read_u16(&data[offset + 2..offset + 4]) as usize; let num_messages = LittleEndian::read_u16(&prefix[2..4]) as usize;
let reference_count = LittleEndian::read_u32(&data[offset + 4..offset + 8]); let reference_count = LittleEndian::read_u32(&prefix[4..8]);
let header_data_size = LittleEndian::read_u32(&data[offset + 8..offset + 12]) as usize; let header_data_size = LittleEndian::read_u32(&prefix[8..12]) as usize;
// libhdf5 (H5O__prefix_deserialize): a header with messages needs room // libhdf5 (H5O__prefix_deserialize): a header with messages needs room
// for at least one message header, and one without has an empty chunk. // for at least one message header, and one without has an empty chunk.
@@ -162,14 +183,15 @@ impl ObjectHeader {
.checked_add(12 + padding) .checked_add(12 + padding)
.ok_or(FormatError::UnexpectedEof { .ok_or(FormatError::UnexpectedEof {
expected: usize::MAX, expected: usize::MAX,
available: data.len(), available: len_usize(file),
})?; })?;
ensure_len(data, msg_start, header_data_size)?; // parse_v1_chunk reads the chunk, with the bounds check that was here.
// The prefix's count (NIL messages included, capped: it is untrusted)
let mut messages = Vec::new(); // sizes the list once instead of growing it message by message.
let mut messages = Vec::with_capacity(num_messages.min(64));
let chunk0_count = Self::parse_v1_chunk( let chunk0_count = Self::parse_v1_chunk(
data, file,
msg_start, msg_start,
header_data_size, header_data_size,
offset_size, offset_size,
@@ -209,9 +231,9 @@ impl ObjectHeader {
/// end of the chunk, or leftover bytes too few for a message header (a /// end of the chunk, or leftover bytes too few for a message header (a
/// "gap", which only version 2 allows). /// "gap", which only version 2 allows).
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
fn parse_v1_chunk( fn parse_v1_chunk<S: Storage + ?Sized>(
data: &[u8], file: &S,
offset: usize, offset: u64,
length: usize, length: usize,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
@@ -221,9 +243,10 @@ impl ObjectHeader {
if depth_remaining == 0 { if depth_remaining == 0 {
return Err(FormatError::NestingDepthExceeded); return Err(FormatError::NestingDepthExceeded);
} }
ensure_len(data, offset, length)?; let chunk = read_exact_at(file, offset, length)?;
let end = offset + length; let data: &[u8] = &chunk;
let mut pos = offset; let end = length;
let mut pos = 0usize;
let mut count = 0usize; let mut count = 0usize;
while pos < end { while pos < end {
@@ -268,8 +291,8 @@ impl ObjectHeader {
let cont_offset = to_usize(read_offset(body, 0, offset_size)?)?; let cont_offset = to_usize(read_offset(body, 0, offset_size)?)?;
let cont_length = to_usize(read_offset(body, offset_size as usize, length_size)?)?; let cont_length = to_usize(read_offset(body, offset_size as usize, length_size)?)?;
Self::parse_v1_chunk( Self::parse_v1_chunk(
data, file,
cont_offset, cont_offset as u64,
cont_length, cont_length,
offset_size, offset_size,
length_size, length_size,
@@ -282,12 +305,22 @@ impl ObjectHeader {
Ok(count) Ok(count)
} }
fn parse_v2( fn parse_v2<S: Storage + ?Sized>(
data: &[u8], file: &S,
offset: usize, offset: u64,
prefix: &Window<'_>,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<ObjectHeader, FormatError> { ) -> Result<ObjectHeader, FormatError> {
// `ensure_len` checks positions relative to the header against the
// prefix window and reports them as the whole-file check did, with
// absolute positions and the file's length.
let data: &[u8] = &prefix.bytes;
let file_len = len_usize(file);
let base = usize::try_from(offset).unwrap_or(usize::MAX);
let abs = |rel: usize| base.saturating_add(rel);
let ensure_len = |_: &[u8], rel: usize, needed: usize| prefix.ensure(rel, needed);
let offset = 0usize;
// signature(4) + version(1) + flags(1) = 6 // signature(4) + version(1) + flags(1) = 6
ensure_len(data, offset, 6)?; ensure_len(data, offset, 6)?;
@@ -352,15 +385,20 @@ impl ObjectHeader {
} }
let chunk0_msg_start = pos; let chunk0_msg_start = pos;
let chunk0_msg_end = pos let Some(chunk0_abs_end) = abs(pos).checked_add(chunk0_size) else {
.checked_add(chunk0_size) return Err(FormatError::UnexpectedEof {
.ok_or(FormatError::UnexpectedEof {
expected: usize::MAX, expected: usize::MAX,
available: data.len(), available: file_len,
})?; });
};
let chunk0_msg_end = chunk0_abs_end - base;
// The whole first chunk, prefix to checksum, in one read (its
// bounds check is the one on the checksum's 4 bytes).
let chunk0 = read_exact_at(file, base as u64, chunk0_msg_end.saturating_add(4))?;
let data: &[u8] = &chunk0;
// Validate checksum: from OHDR signature through all messages (before checksum) // Validate checksum: from OHDR signature through all messages (before checksum)
ensure_len(data, chunk0_msg_end, 4)?;
#[cfg(feature = "checksum")] #[cfg(feature = "checksum")]
{ {
let stored = LittleEndian::read_u32(&data[chunk0_msg_end..chunk0_msg_end + 4]); let stored = LittleEndian::read_u32(&data[chunk0_msg_end..chunk0_msg_end + 4]);
@@ -395,8 +433,8 @@ impl ObjectHeader {
} }
cont_remaining -= 1; cont_remaining -= 1;
Self::parse_v2_continuation( Self::parse_v2_continuation(
data, file,
cont_offset, cont_offset as u64,
cont_length, cont_length,
has_creation_order, has_creation_order,
offset_size, offset_size,
@@ -495,9 +533,9 @@ impl ObjectHeader {
} }
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
fn parse_v2_continuation( fn parse_v2_continuation<S: Storage + ?Sized>(
data: &[u8], file: &S,
offset: usize, offset: u64,
length: usize, length: usize,
has_creation_order: bool, has_creation_order: bool,
offset_size: u8, offset_size: u8,
@@ -506,7 +544,9 @@ impl ObjectHeader {
continuations: &mut Vec<(usize, usize)>, continuations: &mut Vec<(usize, usize)>,
) -> Result<(), FormatError> { ) -> Result<(), FormatError> {
// OCHK signature(4) + messages + checksum(4) // OCHK signature(4) + messages + checksum(4)
ensure_len(data, offset, length)?; let chunk = read_exact_at(file, offset, length)?;
let data: &[u8] = &chunk;
let offset = 0usize;
if length < 8 { if length < 8 {
return Err(FormatError::UnexpectedEof { return Err(FormatError::UnexpectedEof {
expected: 8, expected: 8,
@@ -547,6 +587,10 @@ impl ObjectHeader {
} }
} }
/// Longest version-2 object header prefix: signature(4) + version(1) +
/// flags(1) + times(16) + attribute phase change(4) + chunk-0 size(8).
const V2_PREFIX_MAX: usize = 34;
/// Size of a version-1 message header: type(2) + size(2) + flags(1) + reserved(3). /// Size of a version-1 message header: type(2) + size(2) + flags(1) + reserved(3).
const V1_MSG_HEADER_SIZE: usize = 8; const V1_MSG_HEADER_SIZE: usize = 8;
@@ -755,13 +799,13 @@ mod tests {
let mut msg_bytes = Vec::new(); let mut msg_bytes = Vec::new();
for (mtype, mdata, mflags) in messages { for (mtype, mdata, mflags) in messages {
// v1 message sizes are multiples of 8 (the data is zero-padded). // v1 message sizes are multiples of 8 (the data is zero-padded).
let padded = mdata.len().div_ceil(8) * 8; let padded = <[u8]>::len(mdata).div_ceil(8) * 8;
msg_bytes.extend_from_slice(&mtype.to_le_bytes()); // type(2) msg_bytes.extend_from_slice(&mtype.to_le_bytes()); // type(2)
msg_bytes.extend_from_slice(&(padded as u16).to_le_bytes()); // size(2) msg_bytes.extend_from_slice(&(padded as u16).to_le_bytes()); // size(2)
msg_bytes.push(*mflags); // flags(1) msg_bytes.push(*mflags); // flags(1)
msg_bytes.extend_from_slice(&[0u8; 3]); // reserved(3) msg_bytes.extend_from_slice(&[0u8; 3]); // reserved(3)
msg_bytes.extend_from_slice(mdata); // data msg_bytes.extend_from_slice(mdata); // data
msg_bytes.resize(msg_bytes.len() + padded - mdata.len(), 0); msg_bytes.resize(msg_bytes.len() + padded - <[u8]>::len(mdata), 0);
} }
let mut buf = Vec::new(); let mut buf = Vec::new();
@@ -1281,4 +1325,56 @@ mod tests {
let err = ObjectHeader::parse(&data, 0, 8, 8).unwrap_err(); let err = ObjectHeader::parse(&data, 0, 8, 8).unwrap_err();
assert!(matches!(err, FormatError::UnexpectedEof { .. })); assert!(matches!(err, FormatError::UnexpectedEof { .. }));
} }
/// Every header, and every truncation of it, parses to the same result
/// (or the same error) through a `read_at`-only storage as from a slice;
/// a header in one chunk takes two reads (prefix, chunk).
#[test]
fn parse_in_matches_slice_parse() {
use crate::storage::CountingStorage;
let mut headers = vec![
build_v1_header(&[], 8, 8),
build_v1_header(&[(0x0001, &[1, 2, 3], 0), (0x0003, &[9; 8], 0)], 8, 8),
build_v2_header(0x00, &[(0x01, &[42], 0)], None),
build_v2_header(0x03, &[(0x01, &[1, 2], 0), (0x03, &[3], 0)], None),
build_v2_header(0x24, &[(0x01, &[1], 0)], Some((1, 2, 3, 4))),
build_v2_header(0x35, &[(0x01, &[1], 0)], Some((5, 6, 7, 8))),
];
// A v2 header with a continuation chunk at 256.
let mut ochk = OCHK_SIGNATURE.to_vec();
ochk.extend_from_slice(&[0x03, 2, 0, 0, 0xDE, 0xAD]);
let sum = crate::checksum::jenkins_lookup3(&ochk);
ochk.extend_from_slice(&sum.to_le_bytes());
let mut cont = 256u64.to_le_bytes().to_vec();
cont.extend_from_slice(&(ochk.len() as u64).to_le_bytes());
let main = build_v2_header(0x00, &[(0x01, &[42], 0), (0x10, &cont, 0)], None);
let mut with_cont = vec![0u8; 256 + ochk.len()];
with_cont[..main.len()].copy_from_slice(&main);
with_cont[256..].copy_from_slice(&ochk);
headers.push(with_cont);
for h in headers {
for at in [0usize, 3] {
for cut in 0..=h.len() {
let mut f = vec![0u8; at];
f.extend_from_slice(&h[..cut]);
if at == 0 && cut == h.len() {
f.resize(f.len() + 64, 0);
}
let want = ObjectHeader::parse(&f, at, 8, 8);
let storage = CountingStorage::new(f.clone());
let got = ObjectHeader::parse_in(&storage, at as u64, 8, 8);
assert_eq!(
format!("{got:?}"),
format!("{want:?}"),
"at {at}, cut {cut}"
);
}
}
}
let one_chunk = build_v2_header(0x00, &[(0x01, &[42], 0)], None);
let storage = CountingStorage::new(one_chunk);
ObjectHeader::parse_in(&storage, 0, 8, 8).unwrap();
assert_eq!(storage.reads(), 2);
}
} }
+227 -30
View File
@@ -23,12 +23,12 @@ use alloc::vec::Vec;
#[cfg(feature = "std")] #[cfg(feature = "std")]
use std::borrow::Cow; use std::borrow::Cow;
use crate::addr::to_usize;
use crate::btree_v2::{BTreeV2Header, collect_btree_v2_records}; use crate::btree_v2::{BTreeV2Header, collect_btree_v2_records};
use crate::error::FormatError; use crate::error::FormatError;
use crate::fractal_heap::FractalHeapHeader; use crate::fractal_heap::FractalHeapHeader;
use crate::message_type::MessageType; use crate::message_type::MessageType;
use crate::object_header::ObjectHeader; use crate::object_header::ObjectHeader;
use crate::storage::{Storage, Window, read_exact_at, require_contiguous};
/// Fractal heap ID length for SOHM entries (fixed at 8 bytes). /// Fractal heap ID length for SOHM entries (fixed at 8 bytes).
const FHEAP_ID_LEN: usize = 8; const FHEAP_ID_LEN: usize = 8;
@@ -254,17 +254,31 @@ pub fn parse_sohm_table(
nindexes: u8, nindexes: u8,
offset_size: u8, offset_size: u8,
) -> Result<SohmTable, FormatError> { ) -> Result<SohmTable, FormatError> {
ensure_len(file_data, table_addr, 4)?; parse_sohm_table_in(file_data, table_addr as u64, nindexes, offset_size)
if &file_data[table_addr..table_addr + 4] != b"SMTB" { }
/// [`parse_sohm_table`] over any [`Storage`]: one read of the signature,
/// one of every index entry.
pub fn parse_sohm_table_in<S: Storage + ?Sized>(
file: &S,
table_addr: u64,
nindexes: u8,
offset_size: u8,
) -> Result<SohmTable, FormatError> {
let sig = read_exact_at(file, table_addr, 4)?;
if *sig != *b"SMTB" {
return Err(FormatError::InvalidSohmTableSignature); return Err(FormatError::InvalidSohmTableSignature);
} }
let mut pos = table_addr + 4;
let os = offset_size as usize; let os = offset_size as usize;
let entry_size = 1 + 1 + 2 + 4 + 2 + 2 + 2 + os + os; // 14 + 2*offset_size let entry_size = 1 + 1 + 2 + 4 + 2 + 2 + 2 + os + os; // 14 + 2*offset_size
// Positions below are relative to the table.
let w = Window::read(file, table_addr, 4 + nindexes as usize * entry_size)?;
let file_data: &[u8] = &w.bytes;
let mut pos = 4;
let mut indexes = Vec::with_capacity(nindexes as usize); let mut indexes = Vec::with_capacity(nindexes as usize);
for _ in 0..nindexes { for _ in 0..nindexes {
ensure_len(file_data, pos, entry_size)?; w.ensure(pos, entry_size)?;
let version = file_data[pos]; let version = file_data[pos];
if version != 0 { if version != 0 {
return Err(FormatError::InvalidSohmTableVersion(version)); return Err(FormatError::InvalidSohmTableVersion(version));
@@ -370,16 +384,29 @@ pub fn parse_sohm_list(
num_messages: u16, num_messages: u16,
offset_size: u8, offset_size: u8,
) -> Result<Vec<SohmEntry>, FormatError> { ) -> Result<Vec<SohmEntry>, FormatError> {
ensure_len(file_data, list_addr, 4)?; parse_sohm_list_in(file_data, list_addr as u64, num_messages, offset_size)
if &file_data[list_addr..list_addr + 4] != b"SMLI" { }
/// [`parse_sohm_list`] over any [`Storage`]: one read of the signature, one
/// of every entry.
pub fn parse_sohm_list_in<S: Storage + ?Sized>(
file: &S,
list_addr: u64,
num_messages: u16,
offset_size: u8,
) -> Result<Vec<SohmEntry>, FormatError> {
let sig = read_exact_at(file, list_addr, 4)?;
if *sig != *b"SMLI" {
return Err(FormatError::InvalidSohmListSignature); return Err(FormatError::InvalidSohmListSignature);
} }
let entry_sz = sohm_entry_size(offset_size); let entry_sz = sohm_entry_size(offset_size);
let mut pos = list_addr + 4; // Positions below are relative to the list.
let w = Window::read(file, list_addr, 4 + num_messages as usize * entry_sz)?;
let mut pos = 4;
let mut entries = Vec::with_capacity(num_messages as usize); let mut entries = Vec::with_capacity(num_messages as usize);
for _ in 0..num_messages { for _ in 0..num_messages {
ensure_len(file_data, pos, entry_sz)?; w.ensure(pos, entry_sz)?;
let entry = parse_sohm_entry(&file_data[pos..], offset_size)?; let entry = parse_sohm_entry(&w.bytes[pos..], offset_size)?;
entries.push(entry); entries.push(entry);
pos += entry_sz; pos += entry_sz;
} }
@@ -393,6 +420,20 @@ pub fn parse_sohm_btree_entries(
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Vec<SohmEntry>, FormatError> { ) -> Result<Vec<SohmEntry>, FormatError> {
parse_sohm_btree_entries_in(file_data, btree_addr as u64, offset_size, length_size)
}
/// [`parse_sohm_btree_entries`] over any [`Storage`]. The v2 B-tree is not
/// read over [`Storage`] yet, so this needs the whole file in memory
/// ([`FormatError::ContiguousStorageRequired`] otherwise).
pub fn parse_sohm_btree_entries_in<S: Storage + ?Sized>(
file: &S,
btree_addr: u64,
offset_size: u8,
length_size: u8,
) -> Result<Vec<SohmEntry>, FormatError> {
let file_data = require_contiguous(file, "a shared-message B-tree index")?;
let btree_addr = usize::try_from(btree_addr).unwrap_or(usize::MAX);
let header = BTreeV2Header::parse(file_data, btree_addr, offset_size, length_size)?; let header = BTreeV2Header::parse(file_data, btree_addr, offset_size, length_size)?;
let records = collect_btree_v2_records(file_data, &header, offset_size, length_size)?; let records = collect_btree_v2_records(file_data, &header, offset_size, length_size)?;
let mut entries = Vec::with_capacity(records.len()); let mut entries = Vec::with_capacity(records.len());
@@ -414,15 +455,24 @@ pub fn load_sohm_table(
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Option<SohmTable>, FormatError> { ) -> Result<Option<SohmTable>, FormatError> {
let sig = crate::signature::find_signature(file_data)?; load_sohm_table_in(file_data, offset_size, length_size)
let sb = crate::superblock::Superblock::parse(file_data, sig)?; }
/// [`load_sohm_table`] over any [`Storage`].
pub fn load_sohm_table_in<S: Storage + ?Sized>(
file_data: &S,
offset_size: u8,
length_size: u8,
) -> Result<Option<SohmTable>, FormatError> {
let sig = crate::signature::find_signature_in(file_data)?;
let sb = crate::superblock::Superblock::parse_in(file_data, sig)?;
let Some(ext_addr) = sb let Some(ext_addr) = sb
.superblock_extension_address .superblock_extension_address
.filter(|&a| !is_undefined(a, offset_size)) .filter(|&a| !is_undefined(a, offset_size))
else { else {
return Ok(None); return Ok(None);
}; };
let ext = ObjectHeader::parse(file_data, to_usize(ext_addr)?, offset_size, length_size)?; let ext = ObjectHeader::parse_in(file_data, ext_addr, offset_size, length_size)?;
let Some(msg) = ext let Some(msg) = ext
.messages .messages
.iter() .iter()
@@ -431,9 +481,9 @@ pub fn load_sohm_table(
return Ok(None); return Ok(None);
}; };
let table_msg = parse_sohm_table_message(&msg.data, offset_size)?; let table_msg = parse_sohm_table_message(&msg.data, offset_size)?;
parse_sohm_table( parse_sohm_table_in(
file_data, file_data,
to_usize(table_msg.table_address)?, table_msg.table_address,
table_msg.nindexes, table_msg.nindexes,
offset_size, offset_size,
) )
@@ -447,17 +497,27 @@ pub fn message_data_with_sohm<'a>(
msg: &'a crate::object_header::HeaderMessage, msg: &'a crate::object_header::HeaderMessage,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Cow<'a, [u8]>, FormatError> {
message_data_with_sohm_in(file_data, msg, offset_size, length_size)
}
/// [`message_data_with_sohm`] over any [`Storage`].
pub fn message_data_with_sohm_in<'a, S: Storage + ?Sized>(
file_data: &S,
msg: &'a crate::object_header::HeaderMessage,
offset_size: u8,
length_size: u8,
) -> Result<Cow<'a, [u8]>, FormatError> { ) -> Result<Cow<'a, [u8]>, FormatError> {
if !is_shared(msg.flags) { if !is_shared(msg.flags) {
return Ok(Cow::Borrowed(&msg.data)); return Ok(Cow::Borrowed(&msg.data));
} }
let shared_ref = parse_shared_ref_sized(&msg.data, offset_size, length_size)?; let shared_ref = parse_shared_ref_sized(&msg.data, offset_size, length_size)?;
let table = if shared_ref.heap_id.is_some() { let table = if shared_ref.heap_id.is_some() {
load_sohm_table(file_data, offset_size, length_size)? load_sohm_table_in(file_data, offset_size, length_size)?
} else { } else {
None None
}; };
resolve_shared_message_with_sohm( resolve_shared_message_with_sohm_in(
file_data, file_data,
&shared_ref, &shared_ref,
msg.msg_type, msg.msg_type,
@@ -496,6 +556,25 @@ pub fn resolve_sohm_message(
target_msg_type: MessageType, target_msg_type: MessageType,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Vec<u8>, FormatError> {
resolve_sohm_message_in(
&file_data,
heap_id,
sohm_table,
target_msg_type,
offset_size,
length_size,
)
}
/// [`resolve_sohm_message`] over any [`Storage`].
pub fn resolve_sohm_message_in<S: Storage + ?Sized>(
file_data: &S,
heap_id: &[u8; FHEAP_ID_LEN],
sohm_table: &SohmTable,
target_msg_type: MessageType,
offset_size: u8,
length_size: u8,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
let index = find_index_for_msg_type(sohm_table, target_msg_type) let index = find_index_for_msg_type(sohm_table, target_msg_type)
.ok_or(FormatError::InvalidSharedMessageVersion(2))?; .ok_or(FormatError::InvalidSharedMessageVersion(2))?;
@@ -504,13 +583,9 @@ pub fn resolve_sohm_message(
return Err(FormatError::InvalidSharedMessageVersion(2)); return Err(FormatError::InvalidSharedMessageVersion(2));
} }
let fh_header = FractalHeapHeader::parse( let fh_header =
file_data, FractalHeapHeader::parse_in(file_data, index.heap_addr, offset_size, length_size)?;
to_usize(index.heap_addr)?, fh_header.read_managed_object_in(file_data, heap_id, offset_size)
offset_size,
length_size,
)?;
fh_header.read_managed_object(file_data, heap_id, offset_size)
} }
/// The payload of an object-header message, following the indirection if the /// The payload of an object-header message, following the indirection if the
@@ -527,12 +602,22 @@ pub fn message_data<'a>(
msg: &'a crate::object_header::HeaderMessage, msg: &'a crate::object_header::HeaderMessage,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Cow<'a, [u8]>, FormatError> {
message_data_in(file_data, msg, offset_size, length_size)
}
/// [`message_data`] over any [`Storage`].
pub fn message_data_in<'a, S: Storage + ?Sized>(
file_data: &S,
msg: &'a crate::object_header::HeaderMessage,
offset_size: u8,
length_size: u8,
) -> Result<Cow<'a, [u8]>, FormatError> { ) -> Result<Cow<'a, [u8]>, FormatError> {
if !is_shared(msg.flags) { if !is_shared(msg.flags) {
return Ok(Cow::Borrowed(&msg.data)); return Ok(Cow::Borrowed(&msg.data));
} }
let shared_ref = parse_shared_ref_sized(&msg.data, offset_size, length_size)?; let shared_ref = parse_shared_ref_sized(&msg.data, offset_size, length_size)?;
resolve_shared_message( resolve_shared_message_in(
file_data, file_data,
&shared_ref, &shared_ref,
msg.msg_type, msg.msg_type,
@@ -554,13 +639,30 @@ pub fn resolve_shared_message(
target_msg_type: MessageType, target_msg_type: MessageType,
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
) -> Result<Vec<u8>, FormatError> {
resolve_shared_message_in(
&file_data,
shared_ref,
target_msg_type,
offset_size,
length_size,
)
}
/// [`resolve_shared_message`] over any [`Storage`].
pub fn resolve_shared_message_in<S: Storage + ?Sized>(
file_data: &S,
shared_ref: &SharedMessageRef,
target_msg_type: MessageType,
offset_size: u8,
length_size: u8,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
let table = if shared_ref.heap_id.is_some() { let table = if shared_ref.heap_id.is_some() {
load_sohm_table(file_data, offset_size, length_size)? load_sohm_table_in(file_data, offset_size, length_size)?
} else { } else {
None None
}; };
resolve_shared_message_with_sohm( resolve_shared_message_with_sohm_in(
file_data, file_data,
shared_ref, shared_ref,
target_msg_type, target_msg_type,
@@ -578,6 +680,25 @@ pub fn resolve_shared_message_with_sohm(
offset_size: u8, offset_size: u8,
length_size: u8, length_size: u8,
sohm_table: Option<&SohmTable>, sohm_table: Option<&SohmTable>,
) -> Result<Vec<u8>, FormatError> {
resolve_shared_message_with_sohm_in(
&file_data,
shared_ref,
target_msg_type,
offset_size,
length_size,
sohm_table,
)
}
/// [`resolve_shared_message_with_sohm`] over any [`Storage`].
pub fn resolve_shared_message_with_sohm_in<S: Storage + ?Sized>(
file_data: &S,
shared_ref: &SharedMessageRef,
target_msg_type: MessageType,
offset_size: u8,
length_size: u8,
sohm_table: Option<&SohmTable>,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
// Dispatch on what the reference carries rather than on `ref_type`: v1/v2 // Dispatch on what the reference carries rather than on `ref_type`: v1/v2
// references are always an object-header address whatever their type // references are always an object-header address whatever their type
@@ -587,8 +708,7 @@ pub fn resolve_shared_message_with_sohm(
shared_ref.heap_id.as_ref(), shared_ref.heap_id.as_ref(),
) { ) {
(Some(addr), _) => { (Some(addr), _) => {
let target_header = let target_header = ObjectHeader::parse_in(file_data, addr, offset_size, length_size)?;
ObjectHeader::parse(file_data, to_usize(addr)?, offset_size, length_size)?;
for msg in &target_header.messages { for msg in &target_header.messages {
if msg.msg_type == target_msg_type && !is_shared(msg.flags) { if msg.msg_type == target_msg_type && !is_shared(msg.flags) {
return Ok(msg.data.clone()); return Ok(msg.data.clone());
@@ -615,7 +735,7 @@ pub fn resolve_shared_message_with_sohm(
} }
(None, Some(heap_id)) => { (None, Some(heap_id)) => {
let table = sohm_table.ok_or(FormatError::InvalidSharedMessageVersion(2))?; let table = sohm_table.ok_or(FormatError::InvalidSharedMessageVersion(2))?;
resolve_sohm_message( resolve_sohm_message_in(
file_data, file_data,
heap_id, heap_id,
table, table,
@@ -1053,4 +1173,81 @@ mod tests {
// With 2-byte offsets: OH=2+2=4, heap=12, entry=1+4+12=17 // With 2-byte offsets: OH=2+2=4, heap=12, entry=1+4+12=17
assert_eq!(sohm_entry_size(2), 17); assert_eq!(sohm_entry_size(2), 17);
} }
/// SOHM tables and lists parse identically through a read_at-only
/// CountingStorage: at two offsets, with 4- and 8-byte offsets, cut at
/// every length and with a bad signature.
#[test]
fn storage_reads_match_slice_reads() {
use crate::storage::CountingStorage;
let idx = |t: u8, n: u16| SohmIndex {
index_type: t,
mesg_types: 0x0008,
min_mesg_size: 50,
list_max: 50,
btree_min: 40,
num_messages: n,
index_addr: 0x3000,
heap_addr: 0x4000,
};
let heap_entry = |h: u32| SohmEntry {
location: 0,
hash: h,
heap_id: Some([1, 2, 3, 4, 5, 6, 7, h as u8]),
ref_count: Some(h),
mesg_index: None,
oh_addr: None,
};
let oh_entry = SohmEntry {
location: 1,
hash: 9,
heap_id: None,
ref_count: None,
mesg_index: Some(3),
oh_addr: Some(0x7000),
};
let mut compared = 0;
for os in [4u8, 8] {
let smtb = build_smtb(&[idx(0, 2), idx(1, 7)], os);
let smli = build_smli(&[heap_entry(1), oh_entry.clone(), heap_entry(2)], os);
for (body, n) in [(smtb, 2u16), (smli, 3)] {
let is_table = &body[..4] == b"SMTB";
for at in [0usize, 0x40] {
let mut full = vec![0u8; at];
full.extend_from_slice(&body);
let mut files = Vec::new();
for cut in at..=full.len() {
files.push(full[..cut].to_vec());
}
let mut bad = full.clone();
bad[at] = b'X';
files.push(bad);
for f in files {
let st = CountingStorage::new(f.clone());
let (want, got) = if is_table {
(
format!("{:?}", parse_sohm_table(&f, at, n as u8, os)),
format!("{:?}", parse_sohm_table_in(&st, at as u64, n as u8, os)),
)
} else {
(
format!("{:?}", parse_sohm_list(&f, at, n, os)),
format!("{:?}", parse_sohm_list_in(&st, at as u64, n, os)),
)
};
assert_eq!(got, want, "{} bytes", f.len());
assert!(st.reads() <= 2);
compared += 1;
}
}
}
}
assert!(compared > 200);
// The B-tree index is not read over Storage yet: a clean error.
let st = CountingStorage::new(vec![0u8; 64]);
assert_eq!(
parse_sohm_btree_entries_in(&st, 0, 8, 8).unwrap_err(),
FormatError::ContiguousStorageRequired("a shared-message B-tree index")
);
}
} }
+37
View File
@@ -1,6 +1,7 @@
//! HDF5 file signature (magic bytes) detection. //! HDF5 file signature (magic bytes) detection.
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::{Storage, read_exact_at};
/// The 8-byte HDF5 magic signature. /// The 8-byte HDF5 magic signature.
pub const HDF5_SIGNATURE: [u8; 8] = [0x89, b'H', b'D', b'F', b'\r', b'\n', 0x1A, b'\n']; pub const HDF5_SIGNATURE: [u8; 8] = [0x89, b'H', b'D', b'F', b'\r', b'\n', 0x1A, b'\n'];
@@ -39,6 +40,20 @@ pub fn find_signature(data: &[u8]) -> Result<usize, FormatError> {
Err(FormatError::SignatureNotFound) Err(FormatError::SignatureNotFound)
} }
/// [`find_signature`] over any [`Storage`]: one 8-byte read per candidate
/// offset.
pub fn find_signature_in<S: Storage + ?Sized>(file: &S) -> Result<u64, FormatError> {
let len = file.len();
let mut offset = 0u64;
while offset.checked_add(8).is_some_and(|end| end <= len) {
if *read_exact_at(file, offset, 8)? == HDF5_SIGNATURE {
return Ok(offset);
}
offset = if offset == 0 { 512 } else { offset * 2 };
}
Err(FormatError::SignatureNotFound)
}
/// Split a file into its user block and its HDF5 bytes. /// Split a file into its user block and its HDF5 bytes.
/// ///
/// Returns `(user_block, hdf5)`: `user_block` is everything before the /// Returns `(user_block, hdf5)`: `user_block` is everything before the
@@ -132,4 +147,26 @@ mod tests {
data[512..520].copy_from_slice(&HDF5_SIGNATURE); data[512..520].copy_from_slice(&HDF5_SIGNATURE);
assert_eq!(find_signature(&data), Ok(0)); assert_eq!(find_signature(&data), Ok(0));
} }
#[test]
fn find_signature_in_matches_slice_search() {
use crate::storage::CountingStorage;
for (len, at) in [
(0, None),
(7, None),
(8, Some(0)),
(600, Some(512)),
(5000, Some(4096)),
(3000, Some(2048)),
(3000, None),
] {
let mut data = vec![0u8; len];
if let Some(at) = at {
data[at..at + 8].copy_from_slice(&HDF5_SIGNATURE);
}
let want = find_signature(&data).map(|o| o as u64);
let got = find_signature_in(&CountingStorage::new(data));
assert_eq!(got, want, "{len} {at:?}");
}
}
} }
+466
View File
@@ -0,0 +1,466 @@
//! Where the parsers read the file from: the [`Storage`] trait.
//!
//! Every parser used to take the whole file as one `&[u8]`. [`Storage`] is
//! the abstraction that replaces it (see `docs/design/range-reads.md`,
//! option (a)): a parser asks for the bytes it needs, `[offset, offset +
//! len)`, with 64-bit offsets, and gets them back as a [`Cow`] — borrowed
//! when the backend holds the file in memory (a `Vec`, an mmap), owned when
//! it had to fetch them (a range request, a block cache).
//!
//! `impl Storage for [u8]` serves the in-memory case with no copy, and
//! [`Storage::as_contiguous`] lets a hot loop borrow the whole file at once
//! when the backend has it. Modules are converted one at a time: a converted
//! parser has an `*_in<S: Storage + ?Sized>(file: &S, ..)` core and keeps
//! its old `&[u8]` signature as a thin wrapper, so callers do not change.
//!
//! The cores are generic rather than taking `&dyn Storage` so that the
//! wrappers monomorphise for `[u8]`: the bounds check of each structure read
//! inlines to what the slice code did, with no indirect call and no copy,
//! which keeps local files as fast as before the migration. A `&dyn Storage`
//! still works (`S = dyn Storage`), and a remote backend pays one indirect
//! call per structure read.
//!
//! The trait is synchronous and `no_std`: parsing is CPU work, and a remote
//! backend bridges to its own I/O.
#[cfg(not(feature = "std"))]
use alloc::{borrow::Cow, boxed::Box, vec::Vec};
#[cfg(feature = "std")]
use std::{borrow::Cow, boxed::Box, vec::Vec};
use core::ops::Range;
use crate::error::FormatError;
/// A random-access source of file bytes.
///
/// Offsets are relative to the start of the HDF5 data (the superblock), like
/// every address in the file.
pub trait Storage {
/// Bytes `[offset, offset + len)`.
///
/// The result is shorter than `len` only when the range runs past the
/// end of the storage (and empty when `offset` is at or past the end);
/// a backend that cannot serve a range returns an error instead of a
/// short read.
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError>;
/// Current length of the storage in bytes.
fn len(&self) -> u64;
/// Whether the storage holds no bytes.
fn is_empty(&self) -> bool {
self.len() == 0
}
/// Several reads at once, in the order given. Backends that talk to a
/// remote store coalesce and parallelise these; the default reads them
/// one by one with [`Storage::read_at`].
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
ranges
.iter()
.map(|r| {
let len = usize::try_from(r.end.saturating_sub(r.start)).map_err(|_| {
FormatError::Overflow("read range longer than the address space".into())
})?;
self.read_at(r.start, len)
})
.collect()
}
/// The whole storage as one slice, when the backend has it in memory
/// (a `Vec`, an mmap). Hot loops use this to keep their zero-copy path;
/// `None` means every byte has to go through [`Storage::read_at`].
fn as_contiguous(&self) -> Option<&[u8]> {
None
}
}
impl Storage for [u8] {
#[inline]
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
let n = self.len();
let start = usize::try_from(offset).map_or(n, |o| o.min(n));
let end = start.saturating_add(len).min(n);
Ok(Cow::Borrowed(&self[start..end]))
}
#[inline]
fn len(&self) -> u64 {
<[u8]>::len(self) as u64
}
#[inline]
fn as_contiguous(&self) -> Option<&[u8]> {
Some(self)
}
}
impl Storage for Vec<u8> {
#[inline]
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
self.as_slice().read_at(offset, len)
}
#[inline]
fn len(&self) -> u64 {
Vec::len(self) as u64
}
#[inline]
fn as_contiguous(&self) -> Option<&[u8]> {
Some(self.as_slice())
}
}
impl<T: Storage + ?Sized> Storage for &T {
#[inline]
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
(**self).read_at(offset, len)
}
#[inline]
fn len(&self) -> u64 {
(**self).len()
}
#[inline]
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
(**self).read_ranges(ranges)
}
#[inline]
fn as_contiguous(&self) -> Option<&[u8]> {
(**self).as_contiguous()
}
}
impl<T: Storage + ?Sized> Storage for Box<T> {
#[inline]
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
(**self).read_at(offset, len)
}
#[inline]
fn len(&self) -> u64 {
(**self).len()
}
#[inline]
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
(**self).read_ranges(ranges)
}
#[inline]
fn as_contiguous(&self) -> Option<&[u8]> {
(**self).as_contiguous()
}
}
#[cfg(feature = "std")]
impl<T: Storage + ?Sized> Storage for std::sync::Arc<T> {
#[inline]
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
(**self).read_at(offset, len)
}
#[inline]
fn len(&self) -> u64 {
(**self).len()
}
#[inline]
fn read_ranges(&self, ranges: &[Range<u64>]) -> Result<Vec<Cow<'_, [u8]>>, FormatError> {
(**self).read_ranges(ranges)
}
#[inline]
fn as_contiguous(&self) -> Option<&[u8]> {
(**self).as_contiguous()
}
}
/// `storage.len()` as the `usize` the parsers' end-of-file errors report
/// (saturating on targets where the file is larger than the address space).
#[inline]
pub(crate) fn len_usize<S: Storage + ?Sized>(file: &S) -> usize {
usize::try_from(file.len()).unwrap_or(usize::MAX)
}
/// Bytes `[offset, offset + len)`, all of them.
///
/// A range that runs past the end of the storage is
/// [`FormatError::UnexpectedEof`] with `expected = offset + len` and
/// `available = storage length` — the error the `&[u8]` parsers give for
/// the same bounds check (`offset + len > file_data.len()`).
#[inline]
pub fn read_exact_at<S: Storage + ?Sized>(
file: &S,
offset: u64,
len: usize,
) -> Result<Cow<'_, [u8]>, FormatError> {
let eof = || FormatError::UnexpectedEof {
expected: usize::try_from(offset)
.unwrap_or(usize::MAX)
.saturating_add(len),
available: len_usize(file),
};
// In-memory fast path: plain slicing (for `S = [u8]` this inlines to
// the slice code's bounds check).
if let Some(all) = file.as_contiguous() {
return usize::try_from(offset)
.ok()
.and_then(|start| all.get(start..start.checked_add(len)?))
.map(Cow::Borrowed)
.ok_or_else(eof);
}
match offset.checked_add(len as u64) {
Some(end) if end <= file.len() => {}
_ => return Err(eof()),
}
let bytes = file.read_at(offset, len)?;
if bytes.len() < len {
// The storage shrank or the backend served a short read inside the
// file: never parse a partial structure.
return Err(short_read());
}
Ok(bytes)
}
#[cold]
#[inline(never)]
fn short_read() -> FormatError {
FormatError::Storage(
"short read inside the file (the storage shrank or the backend failed)".into(),
)
}
/// Largest paged data block (fixed or extensible array) read in one piece.
/// A bigger one is read as its prefix and then page by page, only the pages
/// in use, so a block whose size fields claim more than the file holds
/// costs no more than the pages it really has.
pub(crate) const PAGED_BLOCK_ONE_READ_MAX: usize = 1 << 20;
/// A window of the file: up to `max` bytes read at `base`, fewer only at
/// the end of the file. Its [`Window::ensure`] reports a bounds failure
/// exactly as the whole-file check `ensure_len(file_data, base + rel, n)`
/// did — with the absolute position and the file's length — as long as
/// every position checked lies within the `max` bytes the window was asked
/// for: then a position past the window is past the end of the file.
pub(crate) struct Window<'a> {
/// The bytes, from `base` on.
pub bytes: Cow<'a, [u8]>,
base: usize,
file_len: usize,
}
impl<'a> Window<'a> {
/// Read up to `max` bytes at `base`.
pub fn read<S: Storage + ?Sized>(
file: &'a S,
base: u64,
max: usize,
) -> Result<Self, FormatError> {
Ok(Window {
bytes: read_upto(file, base, max)?,
base: usize::try_from(base).unwrap_or(usize::MAX),
file_len: len_usize(file),
})
}
/// A whole in-memory file as one window (base 0).
#[cfg(test)]
pub fn whole(bytes: &'a [u8]) -> Self {
Window {
bytes: Cow::Borrowed(bytes),
base: 0,
file_len: bytes.len(),
}
}
/// [`Window::ensure`] for a window at `base` that has not been read:
/// whether `[rel, rel + needed)` lies in the file, with the same error.
/// Lets a parser whose first step is to check a structure's whole extent
/// (a checksum at its end) fail before reading a structure that a
/// hostile size field has stretched past the end of the file.
pub fn check_extent<S: Storage + ?Sized>(
file: &S,
base: u64,
rel: usize,
needed: usize,
) -> Result<(), FormatError> {
let base = usize::try_from(base).unwrap_or(usize::MAX);
let file_len = len_usize(file);
match base.checked_add(rel).and_then(|p| p.checked_add(needed)) {
Some(end) if end <= file_len => Ok(()),
_ => Err(FormatError::UnexpectedEof {
expected: base.saturating_add(rel).saturating_add(needed),
available: file_len,
}),
}
}
/// Check that `[rel, rel + needed)` (relative to `base`) is in the file.
#[inline]
pub fn ensure(&self, rel: usize, needed: usize) -> Result<(), FormatError> {
match rel.checked_add(needed) {
Some(end) if end <= self.bytes.len() => Ok(()),
_ => Err(FormatError::UnexpectedEof {
expected: self.base.saturating_add(rel).saturating_add(needed),
available: self.file_len,
}),
}
}
}
/// Up to `max` bytes from `offset` on: fewer only at the end of the
/// storage. For structures whose size is only known once their prefix has
/// been parsed and whose parsers bound-check what they are given.
#[inline]
pub fn read_upto<S: Storage + ?Sized>(
file: &S,
offset: u64,
max: usize,
) -> Result<Cow<'_, [u8]>, FormatError> {
if let Some(all) = file.as_contiguous() {
let start = usize::try_from(offset).map_or(all.len(), |o| o.min(all.len()));
let end = start.saturating_add(max).min(all.len());
return Ok(Cow::Borrowed(&all[start..end]));
}
let avail = file.len().saturating_sub(offset);
let len = usize::try_from(avail).map_or(max, |a| a.min(max));
let bytes = file.read_at(offset, len)?;
if bytes.len() < len {
return Err(short_read());
}
Ok(bytes)
}
/// Borrow the whole file for a code path that has not been converted to
/// [`Storage`] yet. On a backend without a contiguous view this is the
/// clean [`FormatError::ContiguousStorageRequired`] error, never a guess.
#[inline]
pub fn require_contiguous<'a, S: Storage + ?Sized>(
file: &'a S,
what: &'static str,
) -> Result<&'a [u8], FormatError> {
file.as_contiguous()
.ok_or(FormatError::ContiguousStorageRequired(what))
}
/// A [`Storage`] over an in-memory buffer that serves every byte through
/// [`Storage::read_at`] (its [`Storage::as_contiguous`] is `None`, so no
/// parser can take the whole-slice shortcut), copies what it serves (as a
/// remote backend would), and counts the reads and bytes.
///
/// It is the equivalence harness of the range-read migration: parsing a
/// file through it must give exactly what parsing the `&[u8]` gives, and
/// the counters are the request counts a cacheless range reader would make.
#[derive(Debug)]
pub struct CountingStorage {
data: Vec<u8>,
reads: portable_atomic::AtomicU64,
bytes: portable_atomic::AtomicU64,
}
impl CountingStorage {
/// Serve `data` (the file from the superblock on).
pub fn new(data: Vec<u8>) -> Self {
CountingStorage {
data,
reads: portable_atomic::AtomicU64::new(0),
bytes: portable_atomic::AtomicU64::new(0),
}
}
/// Number of `read_at` calls served so far.
pub fn reads(&self) -> u64 {
self.reads.load(portable_atomic::Ordering::Relaxed)
}
/// Number of bytes served so far.
pub fn bytes_read(&self) -> u64 {
self.bytes.load(portable_atomic::Ordering::Relaxed)
}
/// Reset both counters.
pub fn reset(&self) {
self.reads.store(0, portable_atomic::Ordering::Relaxed);
self.bytes.store(0, portable_atomic::Ordering::Relaxed);
}
}
impl Storage for CountingStorage {
fn read_at(&self, offset: u64, len: usize) -> Result<Cow<'_, [u8]>, FormatError> {
let got = self.data.as_slice().read_at(offset, len)?;
self.reads.fetch_add(1, portable_atomic::Ordering::Relaxed);
self.bytes
.fetch_add(got.len() as u64, portable_atomic::Ordering::Relaxed);
Ok(Cow::Owned(got.into_owned()))
}
fn len(&self) -> u64 {
self.data.len() as u64
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn slice_reads_are_borrowed_and_clamped() {
let data: Vec<u8> = (0u8..10).collect();
let s: &[u8] = &data;
let dynamic: &dyn Storage = &s;
assert_eq!(dynamic.len(), 10);
let r = dynamic.read_at(2, 3).unwrap();
assert!(matches!(r, Cow::Borrowed(_)));
assert_eq!(&*r, &[2, 3, 4]);
assert_eq!(&*dynamic.read_at(8, 5).unwrap(), &[8, 9]);
assert!(dynamic.read_at(10, 5).unwrap().is_empty());
assert!(dynamic.read_at(u64::MAX, 5).unwrap().is_empty());
assert_eq!(dynamic.as_contiguous(), Some(&data[..]));
let v: &dyn Storage = &data;
assert_eq!(v.as_contiguous(), Some(&data[..]));
}
#[test]
fn read_exact_matches_slice_bounds_errors() {
let data = [0u8; 10];
let s: &[u8] = &data;
assert_eq!(&*read_exact_at(&s, 4, 6).unwrap(), &[0; 6]);
assert_eq!(
read_exact_at(&s, 4, 7).unwrap_err(),
FormatError::UnexpectedEof {
expected: 11,
available: 10
}
);
assert!(read_exact_at(&s, u64::MAX, 1).is_err());
assert_eq!(read_upto(&s, 7, 100).unwrap().len(), 3);
assert_eq!(read_upto(&s, 70, 100).unwrap().len(), 0);
}
#[test]
fn counting_storage_counts_and_hides_the_slice() {
let c = CountingStorage::new((0u8..10).collect());
assert!(c.as_contiguous().is_none());
let r = c.read_at(3, 4).unwrap();
assert!(matches!(r, Cow::Owned(_)));
assert_eq!(&*r, &[3, 4, 5, 6]);
c.read_at(8, 4).unwrap();
assert_eq!((c.reads(), c.bytes_read()), (2, 6));
c.reset();
assert_eq!((c.reads(), c.bytes_read()), (0, 0));
}
#[test]
fn read_ranges_default_loops() {
let data: Vec<u8> = (0u8..10).collect();
let s: &[u8] = &data;
let got = s.read_ranges(&[1..3, 5..9]).unwrap();
assert_eq!(&*got[0], &[1, 2]);
assert_eq!(&*got[1], &[5, 6, 7, 8]);
}
}
+60 -8
View File
@@ -7,6 +7,11 @@ use byteorder::{ByteOrder, LittleEndian};
use crate::error::FormatError; use crate::error::FormatError;
use crate::signature::HDF5_SIGNATURE; use crate::signature::HDF5_SIGNATURE;
use crate::storage::{Storage, read_upto};
/// Bytes read to parse a superblock: more than the largest one (version 1
/// with 8-byte offsets and lengths, 100 bytes).
const SUPERBLOCK_READ_LEN: usize = 128;
/// Parsed HDF5 superblock (all versions). /// Parsed HDF5 superblock (all versions).
#[derive(Debug, Clone, PartialEq, Eq)] #[derive(Debug, Clone, PartialEq, Eq)]
@@ -161,7 +166,16 @@ impl Superblock {
file_data: &[u8], file_data: &[u8],
signature_offset: usize, signature_offset: usize,
) -> Result<u64, FormatError> { ) -> Result<u64, FormatError> {
let refreshed = Superblock::parse(file_data, signature_offset)?; self.refresh_eof_in(file_data, signature_offset as u64)
}
/// [`Self::refresh_eof`] over any [`Storage`].
pub fn refresh_eof_in<S: Storage + ?Sized>(
&mut self,
file: &S,
signature_offset: u64,
) -> Result<u64, FormatError> {
let refreshed = Superblock::parse_in(file, signature_offset)?;
self.eof_address = refreshed.eof_address; self.eof_address = refreshed.eof_address;
self.consistency_flags = refreshed.consistency_flags; self.consistency_flags = refreshed.consistency_flags;
Ok(self.eof_address) Ok(self.eof_address)
@@ -219,15 +233,23 @@ impl Superblock {
/// [`FormatError::UserBlockNotStripped`] because the addresses in the /// [`FormatError::UserBlockNotStripped`] because the addresses in the
/// returned superblock would otherwise be applied to the wrong bytes. /// returned superblock would otherwise be applied to the wrong bytes.
pub fn parse(data: &[u8], signature_offset: usize) -> Result<Superblock, FormatError> { pub fn parse(data: &[u8], signature_offset: usize) -> Result<Superblock, FormatError> {
Self::parse_in(data, signature_offset as u64)
}
/// [`Self::parse`] over any [`Storage`]: one read of the first
/// [`SUPERBLOCK_READ_LEN`] bytes (fewer when the file is shorter, which
/// is then refused with the same end-of-file errors as a short slice).
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
signature_offset: u64,
) -> Result<Superblock, FormatError> {
if signature_offset != 0 { if signature_offset != 0 {
return Err(FormatError::UserBlockNotStripped(signature_offset as u64)); return Err(FormatError::UserBlockNotStripped(signature_offset));
} }
let d = data // Every bounds check below needs at most 100 bytes, so on a longer
.get(signature_offset..) // file none of them can fail and the window's length does not show.
.ok_or(FormatError::UnexpectedEof { let window = read_upto(file, 0, SUPERBLOCK_READ_LEN)?;
expected: signature_offset + 1, let d: &[u8] = &window;
available: data.len(),
})?;
ensure_len(d, 9)?; // signature(8) + version(1) ensure_len(d, 9)?; // signature(8) + version(1)
// Verify signature // Verify signature
@@ -894,4 +916,34 @@ mod tests {
assert_eq!(parsed.version, 3); assert_eq!(parsed.version, 3);
assert_eq!(parsed.page_size, None); assert_eq!(parsed.page_size, None);
} }
/// Through a storage that serves only `read_at`, every version parses
/// to the same superblock, and every truncation to the same error, as
/// from a slice — in one read.
#[test]
fn parse_in_matches_slice_parse() {
use crate::storage::CountingStorage;
let mut files = vec![
build_v0_bytes(8),
build_v0_bytes(4),
build_v1_bytes(8),
build_v1_bytes(4),
build_v2_bytes(8, 2),
build_v2_bytes(4, 3),
];
for f in files.clone() {
let mut long = f.clone();
long.resize(4096, 0xAB);
files.push(long);
for cut in [0, 5, 9, 13, 20, 30, f.len() - 1] {
files.push(f[..cut.min(f.len())].to_vec());
}
}
for f in files {
let want = Superblock::parse(&f, 0);
let storage = CountingStorage::new(f.clone());
assert_eq!(Superblock::parse_in(&storage, 0), want, "{} bytes", f.len());
assert_eq!(storage.reads(), 1);
}
}
} }
+105 -12
View File
@@ -21,13 +21,14 @@
//! file is never copied whole. //! file is never copied whole.
#[cfg(not(feature = "std"))] #[cfg(not(feature = "std"))]
use alloc::{collections::BTreeSet, vec::Vec}; use alloc::{borrow::Cow, collections::BTreeSet, vec::Vec};
#[cfg(feature = "std")] #[cfg(feature = "std")]
use std::collections::BTreeSet; use std::{borrow::Cow, collections::BTreeSet};
use crate::error::FormatError; use crate::error::FormatError;
use crate::message_type::MessageType; use crate::message_type::MessageType;
use crate::object_header::ObjectHeader; use crate::object_header::ObjectHeader;
use crate::storage::{Storage, read_exact_at};
use crate::superblock::Superblock; use crate::superblock::Superblock;
/// Message type of the File Space Info message. /// Message type of the File Space Info message.
@@ -169,6 +170,15 @@ impl<'a> Cursor<'a> {
pub fn read_superblock_extension( pub fn read_superblock_extension(
data: &[u8], data: &[u8],
sb: &Superblock, sb: &Superblock,
) -> Result<Option<SuperblockExtension>, FormatError> {
read_superblock_extension_in(data, sb)
}
/// [`read_superblock_extension`] over any [`Storage`]; its length is the
/// end of file.
pub fn read_superblock_extension_in<S: Storage + ?Sized>(
file: &S,
sb: &Superblock,
) -> Result<Option<SuperblockExtension>, FormatError> { ) -> Result<Option<SuperblockExtension>, FormatError> {
let os = sb.offset_size; let os = sb.offset_size;
let ls = sb.length_size; let ls = sb.length_size;
@@ -181,8 +191,8 @@ pub fn read_superblock_extension(
return Ok(None); return Ok(None);
}; };
let addr = usize::try_from(addr).map_err(|_| ext_err("address out of range"))?; let addr = usize::try_from(addr).map_err(|_| ext_err("address out of range"))?;
let header = ObjectHeader::parse(data, addr, os, ls)?; let header = ObjectHeader::parse_in(file, addr as u64, os, ls)?;
let eoa = data.len() as u64; let eoa = file.len();
let mut ext = SuperblockExtension::default(); let mut ext = SuperblockExtension::default();
for msg in &header.messages { for msg in &header.messages {
@@ -340,12 +350,21 @@ impl CacheImage {
data: &[u8], data: &[u8],
location: CacheImageLocation, location: CacheImageLocation,
sb: &Superblock, sb: &Superblock,
) -> Result<Self, FormatError> {
Self::decode_in(data, location, sb)
}
/// [`Self::decode`] over any [`Storage`]: one read of the image block.
pub fn decode_in<S: Storage + ?Sized>(
file: &S,
location: CacheImageLocation,
sb: &Superblock,
) -> Result<Self, FormatError> { ) -> Result<Self, FormatError> {
let (offset_size, length_size) = (sb.offset_size, sb.length_size); let (offset_size, length_size) = (sb.offset_size, sb.length_size);
let bad = FormatError::InvalidCacheImage; let bad = FormatError::InvalidCacheImage;
let block = image_block(data, location)?; let block = image_block_in(file, location)?;
let eoa = data.len() as u64; let eoa = file.len();
let mut c = Cursor::new(block, bad(RAN_OFF)); let mut c = Cursor::new(&block, bad(RAN_OFF));
// Header: signature, version, flags, image data length, entry count. // Header: signature, version, flags, image data length, entry count.
if c.take(4)? != MDCI_SIGNATURE { if c.take(4)? != MDCI_SIGNATURE {
@@ -463,6 +482,14 @@ impl CacheImage {
image_block(data, self.location) image_block(data, self.location)
} }
/// [`Self::block`] over any [`Storage`].
pub fn block_in<'a, S: Storage + ?Sized>(
&self,
file: &'a S,
) -> Result<Cow<'a, [u8]>, FormatError> {
image_block_in(file, self.location)
}
/// Write every entry over `dst`, the file's bytes from the superblock /// Write every entry over `dst`, the file's bytes from the superblock
/// on (as long as the `data` the image was decoded from), taking the /// on (as long as the `data` the image was decoded from), taking the
/// entries from `block` (the image block, see [`Self::block`]). `block` /// entries from `block` (the image block, see [`Self::block`]). `block`
@@ -483,13 +510,32 @@ impl CacheImage {
} }
fn image_block(data: &[u8], location: CacheImageLocation) -> Result<&[u8], FormatError> { fn image_block(data: &[u8], location: CacheImageLocation) -> Result<&[u8], FormatError> {
let (start, len) = image_block_range(data.len() as u64, location)?;
let start = crate::addr::to_usize(start)?;
Ok(&data[start..start + len])
}
fn image_block_in<S: Storage + ?Sized>(
file: &S,
location: CacheImageLocation,
) -> Result<Cow<'_, [u8]>, FormatError> {
let (start, len) = image_block_range(file.len(), location)?;
read_exact_at(file, start, len)
}
/// Where the image block is, checked against a file of `file_len` bytes.
fn image_block_range(
file_len: u64,
location: CacheImageLocation,
) -> Result<(u64, usize), FormatError> {
let bad = FormatError::InvalidCacheImage; let bad = FormatError::InvalidCacheImage;
let start = usize::try_from(location.address).map_err(|_| bad("address out of range"))?; let start = usize::try_from(location.address).map_err(|_| bad("address out of range"))?;
let len = usize::try_from(location.length).map_err(|_| bad("length out of range"))?; let len = usize::try_from(location.length).map_err(|_| bad("length out of range"))?;
start start
.checked_add(len) .checked_add(len)
.and_then(|end| data.get(start..end)) .filter(|&end| end as u64 <= file_len)
.ok_or(bad("image block extends past the end of the file")) .ok_or(bad("image block extends past the end of the file"))?;
Ok((start as u64, len))
} }
/// What an opener must do before reading a file's metadata: check the /// What an opener must do before reading a file's metadata: check the
@@ -498,11 +544,19 @@ fn image_block(data: &[u8], location: CacheImageLocation) -> Result<&[u8], Forma
/// ([`CacheImage::decode`]). `data` is the file from the superblock on, up /// ([`CacheImage::decode`]). `data` is the file from the superblock on, up
/// to its recorded end of file. /// to its recorded end of file.
pub fn cache_image_state(data: &[u8], sb: &Superblock) -> Result<CacheImageState, FormatError> { pub fn cache_image_state(data: &[u8], sb: &Superblock) -> Result<CacheImageState, FormatError> {
match read_superblock_extension(data, sb)? { cache_image_state_in(data, sb)
}
/// [`cache_image_state`] over any [`Storage`].
pub fn cache_image_state_in<S: Storage + ?Sized>(
file: &S,
sb: &Superblock,
) -> Result<CacheImageState, FormatError> {
match read_superblock_extension_in(file, sb)? {
Some(SuperblockExtension { Some(SuperblockExtension {
cache_image: Some(location), cache_image: Some(location),
.. ..
}) => Ok(match CacheImage::decode(data, location, sb) { }) => Ok(match CacheImage::decode_in(file, location, sb) {
Ok(image) => CacheImageState::Loaded(image), Ok(image) => CacheImageState::Loaded(image),
Err(e) => CacheImageState::Unloadable(e), Err(e) => CacheImageState::Unloadable(e),
}), }),
@@ -566,7 +620,7 @@ mod tests {
/// holds the given messages, padded to `len` bytes. /// holds the given messages, padded to `len` bytes.
fn file_with_ext(messages: &[(u16, &[u8])], len: usize) -> Vec<u8> { fn file_with_ext(messages: &[(u16, &[u8])], len: usize) -> Vec<u8> {
let mut body = Vec::new(); let mut body = Vec::new();
for (t, d) in messages { for &(t, d) in messages {
let padded = d.len().div_ceil(8) * 8; let padded = d.len().div_ceil(8) * 8;
body.extend_from_slice(&t.to_le_bytes()); body.extend_from_slice(&t.to_le_bytes());
body.extend_from_slice(&(padded as u16).to_le_bytes()); body.extend_from_slice(&(padded as u16).to_le_bytes());
@@ -809,4 +863,43 @@ mod tests {
// An entry cannot be its own parent. // An entry cannot be its own parent.
assert!(load(image_with_deps(&[(40, b"C", 1, Some(40))])).is_err()); assert!(load(image_with_deps(&[(40, b"C", 1, Some(40))])).is_err());
} }
/// The extension and cache image decode identically through a
/// `read_at`-only storage, errors included.
#[test]
fn storage_parse_matches_slice_parse() {
use crate::storage::CountingStorage;
let img = image(&[(16, b"HEADER"), (40, b"NODE")]);
let mut with_image = file_with_ext(&[(MSG_MDCI, &mdci(256, img.len() as u64))], 256);
with_image.extend_from_slice(&img);
let mut bad_image = with_image.clone();
bad_image[256] = b'X';
let files = [
file_with_ext(&[(MSG_FSINFO, &fsinfo_v1(4096, false, 0))], 256),
file_with_ext(&[(MSG_FSINFO, &fsinfo_v1(256, false, 0))], 256),
file_with_ext(&[(MSG_MDCI, &mdci(128, 64))], 192),
file_with_ext(&[(MSG_MDCI, &mdci(0x10100, 0x1000_0000))], 2565),
file_with_ext(&[(MSG_FSINFO, &fsinfo_v1(4096, true, 12))], 60),
with_image,
bad_image,
];
for f in files {
let storage = CountingStorage::new(f.clone());
let sb = sb_v2(48);
assert_eq!(
read_superblock_extension_in(&storage, &sb),
read_superblock_extension(&f, &sb)
);
assert_eq!(
cache_image_state_in(&storage, &sb),
cache_image_state(&f, &sb)
);
if let Ok(CacheImageState::Loaded(image)) = cache_image_state(&f, &sb) {
assert_eq!(
&*image.block_in(&storage).unwrap(),
image.block(&f).unwrap()
);
}
}
}
} }
+52 -43
View File
@@ -4,6 +4,7 @@
use alloc::vec::Vec; use alloc::vec::Vec;
use crate::error::FormatError; use crate::error::FormatError;
use crate::storage::{Storage, read_exact_at};
/// Symbol Table message (type 0x0011) found in v1 group object headers. /// Symbol Table message (type 0x0011) found in v1 group object headers.
#[derive(Debug, Clone, PartialEq)] #[derive(Debug, Clone, PartialEq)]
@@ -79,65 +80,49 @@ impl SymbolTableNode {
offset: usize, offset: usize,
offset_size: u8, offset_size: u8,
) -> Result<SymbolTableNode, FormatError> { ) -> Result<SymbolTableNode, FormatError> {
// signature(4) + version(1) + reserved(1) + number_of_symbols(2) = 8 Self::parse_in(file_data, offset as u64, offset_size)
if offset }
.checked_add(8)
.is_none_or(|end| end > file_data.len())
{
return Err(FormatError::UnexpectedEof {
expected: offset.saturating_add(8),
available: file_data.len(),
});
}
if &file_data[offset..offset + 4] != b"SNOD" { /// [`Self::parse`] over any [`Storage`]: one read of the node's header,
/// one of its entries.
pub fn parse_in<S: Storage + ?Sized>(
file: &S,
offset: u64,
offset_size: u8,
) -> Result<SymbolTableNode, FormatError> {
// signature(4) + version(1) + reserved(1) + number_of_symbols(2) = 8
let header = read_exact_at(file, offset, 8)?;
if &header[..4] != b"SNOD" {
return Err(FormatError::InvalidSymbolTableNodeSignature); return Err(FormatError::InvalidSymbolTableNodeSignature);
} }
let version = file_data[offset + 4]; let version = header[4];
if version != 1 { if version != 1 {
return Err(FormatError::InvalidSymbolTableNodeVersion(version)); return Err(FormatError::InvalidSymbolTableNodeVersion(version));
} }
let num_symbols = let num_symbols = u16::from_le_bytes([header[6], header[7]]) as usize;
u16::from_le_bytes([file_data[offset + 6], file_data[offset + 7]]) as usize;
let os = offset_size as usize; let os = offset_size as usize;
// Each entry: link_name_offset(os) + obj_hdr_addr(os) + cache_type(4) + reserved(4) + scratch(16) // Each entry: link_name_offset(os) + obj_hdr_addr(os) + cache_type(4) + reserved(4) + scratch(16)
let entry_size = os + os + 4 + 4 + 16; let entry_size = os + os + 4 + 4 + 16;
let entries_start = offset + 8; // `offset + 8` fits: the header's read checked it. The entries'
let needed = entries_start.checked_add(num_symbols * entry_size).ok_or( // read is the bounds check (`offset + 8 + entries > file length`,
FormatError::UnexpectedEof { // which cannot overflow: at most 65535 entries of 40 bytes).
expected: usize::MAX, let body = read_exact_at(file, offset + 8, num_symbols * entry_size)?;
available: file_data.len(), let file_data: &[u8] = &body;
},
)?;
if needed > file_data.len() {
return Err(FormatError::UnexpectedEof {
expected: needed,
available: file_data.len(),
});
}
let mut entries = Vec::with_capacity(num_symbols); let mut entries = Vec::with_capacity(num_symbols);
let mut pos = entries_start; for entry in file_data.chunks_exact(entry_size) {
for _ in 0..num_symbols { let link_name_offset = read_offset(entry, 0, offset_size)?;
let link_name_offset = read_offset(file_data, pos, offset_size)?; let object_header_address = read_offset(entry, os, offset_size)?;
pos += os; let pos = 2 * os;
let object_header_address = read_offset(file_data, pos, offset_size)?; let cache_type =
pos += os; u32::from_le_bytes([entry[pos], entry[pos + 1], entry[pos + 2], entry[pos + 3]]);
let cache_type = u32::from_le_bytes([
file_data[pos],
file_data[pos + 1],
file_data[pos + 2],
file_data[pos + 3],
]);
pos += 4;
// reserved 4 bytes // reserved 4 bytes
pos += 4;
let mut scratch_pad = [0u8; 16]; let mut scratch_pad = [0u8; 16];
scratch_pad.copy_from_slice(&file_data[pos..pos + 16]); scratch_pad.copy_from_slice(&entry[pos + 8..pos + 24]);
pos += 16;
entries.push(SymbolTableEntry { entries.push(SymbolTableEntry {
link_name_offset, link_name_offset,
@@ -256,4 +241,28 @@ mod tests {
let result = SymbolTableNode::parse(&data, usize::MAX / 2, 8); let result = SymbolTableNode::parse(&data, usize::MAX / 2, 8);
assert!(result.is_err()); assert!(result.is_err());
} }
/// Nodes, cut at every length and at an offset, parse identically
/// through a `read_at`-only storage.
#[test]
fn storage_parse_matches_slice_parse() {
use crate::storage::CountingStorage;
for os in [4u8, 8] {
let node = build_snod(&[(0, 0x100, 0), (8, 0x200, 1), (16, 0x300, 2)], os);
let mut bad = node.clone();
bad[4] = 2;
for full in [node, bad] {
for at in [0usize, 7] {
for cut in 0..=full.len() {
let mut f = vec![0u8; at];
f.extend_from_slice(&full[..cut]);
let storage = CountingStorage::new(f.clone());
let want = SymbolTableNode::parse(&f, at, os);
let got = SymbolTableNode::parse_in(&storage, at as u64, os);
assert_eq!(format!("{got:?}"), format!("{want:?}"));
}
}
}
}
}
} }
Binary file not shown.
@@ -0,0 +1,701 @@
//! Equivalence harness for the range-read migration
//! (`docs/design/range-reads.md`, milestone M1).
//!
//! Every metadata parser converted to [`Storage`] must give exactly what its
//! `&[u8]` form gives. This walks real files — the fixtures, files h5py
//! writes to exercise the less common structures, and optionally the
//! conformance corpus — and, for every object, runs each converted parser
//! twice: over the file as a slice, and over a [`CountingStorage`] that
//! serves the same bytes through `read_at` only (`as_contiguous()` is
//! `None`, so no parser can fall back to the whole slice). The results must
//! be identical, value for value and error for error.
//!
//! The one allowed difference is [`FormatError::ContiguousStorageRequired`]
//! from the storage path, and only from the structures still indexed by a v2
//! B-tree (dense attributes, a SOHM B-tree index, huge fractal-heap objects;
//! see `CONTIGUOUS_REQUIRED`), which fail cleanly instead of reading the
//! whole file. Those are counted; the error from any other site or check
//! fails the harness.
//!
//! Milestones M2/M3 extend `check_object` with the raw-data and group
//! parsers as they are converted.
//!
//! - `CLAWHDF5_STORAGE_CORPUS=dir[:dir...]` adds every `.h5`/`.hdf5`/`.he5`/
//! `.nc`/`.h5ad` file under those directories (the conformance corpus is
//! `conformance/.cache/corpus`); `CLAWHDF5_STORAGE_REPORT=1` prints the
//! per-file read counts.
//! - The h5py-written files honour `CLAWHDF5_PYTHON` and
//! `CLAWHDF5_REQUIRE_INTEROP` like the facade's interop tests.
use std::collections::{HashSet, VecDeque};
use std::fmt::Debug;
use std::path::{Path, PathBuf};
use std::process::Command;
use clawhdf5_format::attribute::{
extract_attributes_full, extract_attributes_full_in, extract_attributes_tolerant,
extract_attributes_tolerant_in,
};
use clawhdf5_format::attribute_info::AttributeInfoMessage;
use clawhdf5_format::btree_v1::{collect_symbol_table_nodes, collect_symbol_table_nodes_in};
use clawhdf5_format::btree_v2::{BTreeV2Header, collect_btree_v2_records};
use clawhdf5_format::data_layout::DataLayout;
use clawhdf5_format::dataspace::Dataspace;
use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::error::FormatError;
use clawhdf5_format::extensible_array::{
ExtensibleArrayHeader, read_extensible_array_chunks, read_extensible_array_chunks_in,
};
use clawhdf5_format::fill_value::{dataset_fill_value_from_storage, dataset_fill_value_in};
use clawhdf5_format::fixed_array::{
FixedArrayHeader, read_fixed_array_chunks, read_fixed_array_chunks_in,
};
use clawhdf5_format::fractal_heap::FractalHeapHeader;
use clawhdf5_format::link_info::LinkInfoMessage;
use clawhdf5_format::local_heap::LocalHeap;
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::shared_message::{
self, load_sohm_table, load_sohm_table_in, message_data_with_sohm, message_data_with_sohm_in,
parse_sohm_btree_entries, parse_sohm_btree_entries_in, parse_sohm_list, parse_sohm_list_in,
};
use clawhdf5_format::signature::split_user_block;
use clawhdf5_format::storage::{CountingStorage, Storage};
use clawhdf5_format::superblock::Superblock;
use clawhdf5_format::superblock_ext::{
cache_image_state, cache_image_state_in, read_superblock_extension,
read_superblock_extension_in,
};
use clawhdf5_format::symbol_table::{SymbolTableMessage, SymbolTableNode};
/// Objects visited per file, heap objects read per heap: enough to cover
/// every structure kind while keeping a 35 000-group file fast.
const MAX_OBJECTS: usize = 1500;
const MAX_HEAP_IDS: usize = 200;
/// The structures that still need the whole file in memory, because they
/// are found through a version-2 B-tree (not converted yet), and the checks
/// that can reach each of them. Anything else answering
/// [`FormatError::ContiguousStorageRequired`] is a converted parser falling
/// back to the whole file, and fails the harness.
const CONTIGUOUS_REQUIRED: &[(&str, &[&str])] = &[
(
"dense attribute storage (a v2 B-tree)",
&["attributes", "attributes (tolerant)"],
),
(
"a shared-message B-tree index",
&[
"SOHM B-tree",
"shared message",
"fill value",
"attributes",
"attributes (tolerant)",
],
),
(
"a huge fractal-heap object's B-tree",
&["heap object", "attributes", "attributes (tolerant)"],
),
];
fn may_require_contiguous(check: &str, site: &str) -> bool {
CONTIGUOUS_REQUIRED
.iter()
.any(|(s, checks)| *s == site && checks.contains(&check))
}
#[derive(Default, Debug)]
struct Tally {
files: usize,
objects: usize,
checks: usize,
contiguous_required: usize,
reads: u64,
bytes: u64,
/// Chunk indexes (fixed and extensible arrays) read, and the most bytes
/// one of them took through the storage.
chunk_indexes: usize,
max_chunk_index_bytes: u64,
}
struct Walk<'a> {
slice: &'a [u8],
storage: &'a CountingStorage,
name: String,
tally: &'a mut Tally,
}
impl Walk<'_> {
/// The storage result must equal the slice result, or be the clean
/// "needs the whole file" error.
fn same<T: Debug>(
&mut self,
what: &str,
want: &Result<T, FormatError>,
got: &Result<T, FormatError>,
) {
self.tally.checks += 1;
if let Err(FormatError::ContiguousStorageRequired(site)) = got {
assert!(
may_require_contiguous(what, site),
"{}: {what} fell back to the whole file ({site}), which only the \
v2-B-tree-indexed structures may do",
self.name
);
self.tally.contiguous_required += 1;
return;
}
let (w, g) = (format!("{want:?}"), format!("{got:?}"));
assert!(
w == g,
"{}: {what} differs\n slice: {}\n storage: {}",
self.name,
&w[..w.len().min(600)],
&g[..g.len().min(600)]
);
}
fn index_read(&mut self, bytes_before: u64) {
self.tally.chunk_indexes += 1;
let bytes = self.storage.bytes_read() - bytes_before;
self.tally.max_chunk_index_bytes = self.tally.max_chunk_index_bytes.max(bytes);
}
fn st(&self) -> &dyn Storage {
self.storage
}
fn run(&mut self) {
let slice = self.slice;
let sb = Superblock::parse(slice, 0);
self.same("superblock", &sb, &Superblock::parse_in(self.st(), 0));
let Ok(sb) = sb else { return };
let (os, ls) = (sb.offset_size, sb.length_size);
let want = read_superblock_extension(slice, &sb);
self.same(
"superblock extension",
&want,
&read_superblock_extension_in(self.st(), &sb),
);
let want = cache_image_state(slice, &sb);
self.same("cache image", &want, &cache_image_state_in(self.st(), &sb));
let table = load_sohm_table(slice, os, ls);
self.same("SOHM table", &table, &load_sohm_table_in(self.st(), os, ls));
if let Ok(Some(table)) = &table {
for idx in &table.indexes {
if idx.index_type == 0 {
let want =
parse_sohm_list(slice, idx.index_addr as usize, idx.num_messages, os);
let got = parse_sohm_list_in(self.st(), idx.index_addr, idx.num_messages, os);
self.same("SOHM list", &want, &got);
} else {
let want = parse_sohm_btree_entries(slice, idx.index_addr as usize, os, ls);
let got = parse_sohm_btree_entries_in(self.st(), idx.index_addr, os, ls);
self.same("SOHM B-tree", &want, &got);
}
}
}
let mut seen = HashSet::new();
let mut queue = VecDeque::from([sb.root_group_address]);
while let Some(addr) = queue.pop_front() {
if seen.len() >= MAX_OBJECTS || !seen.insert(addr) {
continue;
}
self.tally.objects += 1;
self.check_object(&sb, addr);
// Traversal only (group lookups are milestone M0/M3 work).
if let Ok(children) =
clawhdf5_format::group_v2::resolve_group_children(slice, &sb, addr)
{
queue.extend(children.iter().map(|c| c.object_header_address));
}
}
}
fn check_object(&mut self, sb: &Superblock, addr: u64) {
let slice = self.slice;
let (os, ls) = (sb.offset_size, sb.length_size);
let header = ObjectHeader::parse(slice, addr as usize, os, ls);
self.same(
"object header",
&header,
&ObjectHeader::parse_in(self.st(), addr, os, ls),
);
let Ok(header) = header else { return };
let want = extract_attributes_full(slice, &header, os, ls);
self.same(
"attributes",
&want,
&extract_attributes_full_in(self.st(), &header, os, ls),
);
let want = extract_attributes_tolerant(slice, &header, os, ls);
let got = extract_attributes_tolerant_in(self.st(), &header, os, ls);
self.same("attributes (tolerant)", &want, &got);
let want = dataset_fill_value_in(slice, &header.messages, os, ls);
self.same(
"fill value",
&want,
&dataset_fill_value_from_storage(self.st(), &header.messages, os, ls),
);
for msg in &header.messages {
if shared_message::is_shared(msg.flags) {
let want = message_data_with_sohm(slice, msg, os, ls);
let got = message_data_with_sohm_in(self.st(), msg, os, ls);
self.same("shared message", &want, &got);
}
match msg.msg_type {
MessageType::SymbolTable => {
if let Ok(stm) = SymbolTableMessage::parse(&msg.data, os) {
self.check_v1_group(&stm, os, ls);
}
}
MessageType::LinkInfo => {
if let Ok(li) = LinkInfoMessage::parse(&msg.data, os) {
self.check_heap(
li.fractal_heap_address,
li.btree_name_index_address,
4,
os,
ls,
);
}
}
MessageType::AttributeInfo => {
if let Ok(ai) = AttributeInfoMessage::parse(&msg.data, os) {
self.check_heap(
ai.fractal_heap_address,
ai.btree_name_index_address,
0,
os,
ls,
);
}
}
_ => {}
}
}
self.check_layout(&header, os, ls);
}
/// A symbol-table group: its local heap, B-tree, nodes and names.
fn check_v1_group(&mut self, stm: &SymbolTableMessage, os: u8, ls: u8) {
let slice = self.slice;
let heap = LocalHeap::parse(slice, stm.local_heap_address as usize, os, ls);
self.same(
"local heap",
&heap,
&LocalHeap::parse_in(self.st(), stm.local_heap_address, os, ls),
);
let nodes = collect_symbol_table_nodes(slice, stm.btree_address, os, ls);
let got = collect_symbol_table_nodes_in(self.st(), stm.btree_address, os, ls);
self.same("group B-tree", &nodes, &got);
let Ok(heap) = heap else { return };
let want = heap.validate_free_list(slice, ls);
self.same(
"local heap free list",
&want,
&heap.validate_free_list_in(self.st(), ls),
);
let Ok(nodes) = nodes else { return };
for &node in nodes.iter().take(MAX_HEAP_IDS) {
let snod = SymbolTableNode::parse(slice, node as usize, os);
self.same(
"symbol table node",
&snod,
&SymbolTableNode::parse_in(self.st(), node, os),
);
let Ok(snod) = snod else { continue };
for e in &snod.entries {
let want = heap.read_string(slice, e.link_name_offset);
self.same(
"link name",
&want,
&heap.read_string_in(self.st(), e.link_name_offset),
);
}
}
}
/// A dense group's or dense attributes' fractal heap: the header, and
/// the objects its name index points at. `id_at` is where the heap ID
/// starts in a name-index record (after the hash for links).
fn check_heap(&mut self, heap: Option<u64>, index: Option<u64>, id_at: usize, os: u8, ls: u8) {
let slice = self.slice;
let Some(heap_addr) = heap else { return };
let fh = FractalHeapHeader::parse(slice, heap_addr as usize, os, ls);
self.same(
"fractal heap",
&fh,
&FractalHeapHeader::parse_in(self.st(), heap_addr, os, ls),
);
let (Ok(fh), Some(index)) = (fh, index) else {
return;
};
let Ok(bt) = BTreeV2Header::parse(slice, index as usize, os, ls) else {
return;
};
let Ok(records) = collect_btree_v2_records(slice, &bt, os, ls) else {
return;
};
let id_len = fh.heap_id_length as usize;
for rec in records.iter().take(MAX_HEAP_IDS) {
let Some(id) = rec.data.get(id_at..id_at + id_len) else {
continue;
};
let want = fh.read_managed_object(slice, id, os);
self.same(
"heap object",
&want,
&fh.read_managed_object_in(self.st(), id, os),
);
}
}
/// A dataset's layout: VDS mappings, and fixed/extensible array chunk
/// indexes.
fn check_layout(&mut self, header: &ObjectHeader, os: u8, ls: u8) {
let slice = self.slice;
let find = |t: MessageType| {
header
.messages
.iter()
.find(|m| m.msg_type == t)
.and_then(|m| shared_message::message_data_with_sohm(slice, m, os, ls).ok())
};
let Some(layout) = find(MessageType::DataLayout) else {
return;
};
let Ok(layout) = DataLayout::parse(&layout, os, ls) else {
return;
};
match &layout {
DataLayout::Virtual { .. } => {
let (mut want, mut got) = (layout.clone(), layout.clone());
let w = want.resolve_vds_mappings(slice, ls).map(|()| want);
let g = got.resolve_vds_mappings_in(self.st(), ls).map(|()| got);
self.same("VDS mappings", &w, &g);
}
DataLayout::Chunked {
chunk_dimensions,
btree_address: Some(addr),
version: 4,
chunk_index_type: Some(kind @ (3 | 4)),
..
} => {
let (Some(ds), Some(dt)) =
(find(MessageType::Dataspace), find(MessageType::Datatype))
else {
return;
};
let (Ok(ds), Ok((dt, _))) = (Dataspace::parse(&ds, ls), Datatype::parse(&dt))
else {
return;
};
let rank = ds.dimensions.len();
if chunk_dimensions.len() < rank {
return;
}
let dims = &chunk_dimensions[..rank];
let max = ds.max_dimensions.as_deref();
let es = dt.type_size();
if *kind == 3 {
let h = FixedArrayHeader::parse(slice, *addr as usize, os, ls);
self.same(
"fixed array header",
&h,
&FixedArrayHeader::parse_in(self.st(), *addr, os, ls),
);
let Ok(h) = h else { return };
let want =
read_fixed_array_chunks(slice, &h, &ds.dimensions, max, dims, es, os, ls);
let before = self.storage.bytes_read();
let got = read_fixed_array_chunks_in(
self.st(),
&h,
&ds.dimensions,
max,
dims,
es,
os,
ls,
);
self.index_read(before);
self.same("fixed array chunks", &want, &got);
} else {
let h = ExtensibleArrayHeader::parse(slice, *addr as usize, os, ls);
let got = ExtensibleArrayHeader::parse_in(self.st(), *addr, os, ls);
self.same("extensible array header", &h, &got);
let Ok(h) = h else { return };
let want = read_extensible_array_chunks(
slice,
&h,
&ds.dimensions,
max,
dims,
es,
os,
ls,
);
let before = self.storage.bytes_read();
let got = read_extensible_array_chunks_in(
self.st(),
&h,
&ds.dimensions,
max,
dims,
es,
os,
ls,
);
self.index_read(before);
self.same("extensible array chunks", &want, &got);
}
}
_ => {}
}
}
}
fn check_file(path: &Path, tally: &mut Tally) {
let Ok(bytes) = std::fs::read(path) else {
return;
};
check_bytes(&path.display().to_string(), &bytes, tally);
}
fn check_bytes(name: &str, bytes: &[u8], tally: &mut Tally) {
let Ok((_, hdf5)) = split_user_block(bytes) else {
return;
};
let storage = CountingStorage::new(hdf5.to_vec());
let before = (tally.checks, tally.objects);
let mut walk = Walk {
slice: hdf5,
storage: &storage,
name: name.to_string(),
tally,
};
walk.run();
tally.files += 1;
tally.reads += storage.reads();
tally.bytes += storage.bytes_read();
if std::env::var("CLAWHDF5_STORAGE_REPORT").is_ok_and(|v| v == "1") {
eprintln!(
"{:>6} objects {:>7} checks {:>8} reads {:>12} bytes {}",
tally.objects - before.1,
tally.checks - before.0,
storage.reads(),
storage.bytes_read(),
name
);
}
}
fn hdf5_files(dir: &Path, out: &mut Vec<PathBuf>) {
let Ok(entries) = std::fs::read_dir(dir) else {
return;
};
for e in entries.flatten() {
let p = e.path();
if p.is_dir() {
hdf5_files(&p, out);
} else if p
.extension()
.and_then(|x| x.to_str())
.is_some_and(|x| matches!(x, "h5" | "hdf5" | "he5" | "nc" | "h5ad" | "hdf"))
{
out.push(p);
}
}
}
#[test]
fn fixtures_parse_identically_through_storage() {
let mut files = Vec::new();
hdf5_files(
&Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures"),
&mut files,
);
files.sort();
assert!(files.len() >= 40, "{} fixtures", files.len());
let mut tally = Tally::default();
for f in &files {
check_file(f, &mut tally);
}
eprintln!("fixtures: {tally:?}");
assert!(tally.objects >= 150 && tally.checks >= 1000, "{tally:?}");
// Reads really went through read_at.
assert!(tally.reads > tally.objects as u64);
}
#[test]
fn corpus_parses_identically_through_storage() {
let Ok(dirs) = std::env::var("CLAWHDF5_STORAGE_CORPUS") else {
eprintln!("CLAWHDF5_STORAGE_CORPUS not set; skipping the corpus");
return;
};
let mut files = Vec::new();
for d in std::env::split_paths(&dirs) {
hdf5_files(&d, &mut files);
}
files.sort();
let mut tally = Tally::default();
for f in &files {
check_file(f, &mut tally);
}
eprintln!("corpus: {tally:?}");
assert!(tally.files > 0);
}
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn interop_required() -> bool {
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
}
/// Files h5py writes to cover what the fixtures do not: extensible arrays
/// deep enough for super blocks and paged data blocks, paged fixed arrays,
/// big symbol-table and dense groups, dense and shared attributes, a SOHM
/// list and a SOHM B-tree, committed datatypes, and a user block.
const GENERATE: &str = r#"
import ctypes, glob, os, sys, h5py, numpy as np
out = sys.argv[1]
def p(n): return os.path.join(out, n)
with h5py.File(p('ea.h5'), 'w', libver='latest') as f:
# One unlimited dimension: extensible array. 3000 chunks reach super
# blocks; 2-element chunks of int8 keep data small.
d = f.create_dataset('ea', shape=(6000,), maxshape=(None,), chunks=(2,), dtype='i1')
d[:] = np.arange(6000) % 100
d2 = f.create_dataset('ea_deflate', shape=(4000, 3), maxshape=(None, 3), chunks=(2, 3),
dtype='f4', compression='gzip')
d2[:] = np.random.default_rng(1).random((4000, 3))
d3 = f.create_dataset('ea_sparse', shape=(100000,), maxshape=(None,), chunks=(4,), dtype='i2')
d3[0:8] = 1; d3[50000:50004] = 2; d3[99996:] = 3
fa = f.create_dataset('fa_paged', shape=(5000,), chunks=(1,), dtype='u1')
fa[::3] = 7
fa2 = f.create_dataset('fa', shape=(40, 40), chunks=(8, 8), dtype='f8', compression='gzip')
fa2[:] = 1.5
for i in range(30):
f.attrs[f'a{i}'] = np.arange(i + 1)
g = f.create_group('dense')
for i in range(300):
g.create_group(f'child{i:04d}').attrs['i'] = i
f['committed'] = np.dtype([('x', 'i4'), ('y', 'f8')])
f.create_dataset('uses_committed', shape=(3,), dtype=f['committed'])
f['uses_committed'].attrs.create('ta', data=np.zeros(2, dtype=f['committed'].dtype), dtype=f['committed'])
f.attrs['vl'] = ['alpha', 'beta', 'gamma']
with h5py.File(p('v1_groups.h5'), 'w', libver='earliest', userblock_size=512) as f:
for i in range(400):
g = f.create_group(f'g{i:04d}')
g.attrs['n'] = i
f.create_dataset('x', data=np.arange(10))
# Paged chunk indexes whose data blocks are bigger than the storage reads
# in one piece (1 MiB), with two chunks written: a fixed array of 300 000
# chunks (a 2.4 MB data block) and an extensible array grown to 1.2e9 (its
# last data block holds 131 072 chunks: over 1 MiB).
with h5py.File(p('big_paged.h5'), 'w', libver='latest') as f:
d = f.create_dataset('fa', shape=(300000,), chunks=(1,), dtype='u1')
d[5] = 1; d[250000] = 2
e = f.create_dataset('ea', shape=(1,), maxshape=(None,), chunks=(1,), dtype='u1')
e.resize((1200000000,)); e[10] = 1; e[1100000000] = 3
libs = glob.glob(os.path.join(os.path.dirname(h5py.__file__), '..', 'h5py.libs', 'libhdf5-*.so*'))
if libs:
lib = ctypes.CDLL(libs[0])
lib.H5Pset_shared_mesg_nindexes.argtypes = [ctypes.c_int64, ctypes.c_uint]
lib.H5Pset_shared_mesg_index.argtypes = [ctypes.c_int64, ctypes.c_uint, ctypes.c_uint, ctypes.c_uint]
lib.H5Pset_shared_mesg_phase_change.argtypes = [ctypes.c_int64, ctypes.c_uint, ctypes.c_uint]
for name, list_max in [('sohm_list.h5', 50), ('sohm_btree.h5', 0)]:
fcpl = h5py.h5p.create(h5py.h5p.FILE_CREATE)
assert lib.H5Pset_shared_mesg_nindexes(fcpl.id, 1) >= 0
# datatype, dataspace, fill value, filter pipeline, attribute
assert lib.H5Pset_shared_mesg_index(fcpl.id, 0, 0x02 | 0x04 | 0x08 | 0x10 | 0x20, 1) >= 0
assert lib.H5Pset_shared_mesg_phase_change(fcpl.id, list_max, 0) >= 0
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
fapl.set_libver_bounds(h5py.h5f.LIBVER_LATEST, h5py.h5f.LIBVER_LATEST)
fid = h5py.h5f.create(p(name).encode(), h5py.h5f.ACC_TRUNC, fcpl=fcpl, fapl=fapl)
with h5py.File(fid) as f:
for i in range(20):
d = f.create_dataset(f'd{i}', shape=(10, i + 1), dtype='f4', fillvalue=-1.0,
chunks=(5, 1), compression='gzip')
d.attrs['units'] = 'metres per second, a long enough string to share'
d.attrs['scale'] = np.arange(20, dtype='f8')
print('ok')
"#;
#[test]
fn h5py_files_parse_identically_through_storage() {
let dir = Path::new(env!("CARGO_TARGET_TMPDIR")).join("storage_equivalence");
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).unwrap();
let out = Command::new(python())
.args(["-c", GENERATE, dir.to_str().unwrap()])
.output();
match out {
Ok(o) if o.status.success() => {}
Ok(o) if interop_required() => panic!(
"h5py generation failed:\n{}\n{}",
String::from_utf8_lossy(&o.stdout),
String::from_utf8_lossy(&o.stderr)
),
Err(e) if interop_required() => panic!("python not available: {e}"),
_ => {
eprintln!("python3 with h5py unavailable; skipping");
return;
}
}
let mut files = Vec::new();
hdf5_files(&dir, &mut files);
files.sort();
let names: Vec<_> = files
.iter()
.map(|f| f.file_name().unwrap().to_string_lossy().into_owned())
.collect();
for want in ["ea.h5", "v1_groups.h5", "big_paged.h5"] {
assert!(names.iter().any(|n| n == want), "{names:?}");
}
if interop_required() {
assert!(names.iter().any(|n| n == "sohm_btree.h5"), "{names:?}");
}
let mut tally = Tally::default();
for f in &files {
if f.ends_with("big_paged.h5") {
// Only the pages in use are read, not the whole data blocks:
// read in one piece, the fixed array took 2.4 MB and the
// extensible array 1.2 MB (its super block's page bitmap and
// block addresses are most of what remains).
let mut big = Tally::default();
check_file(f, &mut big);
eprintln!("big paged blocks: {big:?}");
assert_eq!(big.chunk_indexes, 2, "{big:?}");
assert!(big.max_chunk_index_bytes < 256 << 10, "{big:?}");
// Truncated anywhere, the page-by-page reads still agree with
// the slice reads (errors included).
let bytes = std::fs::read(f).unwrap();
for cut in (0..bytes.len()).step_by(bytes.len() / 97) {
check_bytes(
&format!("big_paged.h5 cut at {cut}"),
&bytes[..cut],
&mut big,
);
}
eprintln!("big paged blocks, truncated: {big:?}");
assert!(big.chunk_indexes > 100, "{big:?}");
}
check_file(f, &mut tally);
}
eprintln!("h5py files: {tally:?}");
assert!(tally.objects >= 700, "{tally:?}");
// Dense attributes and the SOHM B-tree are the known clean errors.
assert!(tally.contiguous_required > 0, "{tally:?}");
}
+1
View File
@@ -82,6 +82,7 @@ pub(crate) fn to_py_err(e: clawhdf5_rs::Error) -> PyErr {
| Error::ZeroCopyUnaligned { .. } => { | Error::ZeroCopyUnaligned { .. } => {
PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string()) PyErr::new::<pyo3::exceptions::PyValueError, _>(e.to_string())
} }
_ => PyErr::new::<pyo3::exceptions::PyOSError, _>(e.to_string()),
} }
} }
+4
View File
@@ -31,6 +31,10 @@ harness = false
name = "parallel_bench" name = "parallel_bench"
harness = false harness = false
[[bench]]
name = "local_metadata_bench"
harness = false
[features] [features]
default = ["mmap", "provenance", "lzf"] default = ["mmap", "provenance", "lzf"]
mmap = ["clawhdf5-io/mmap"] mmap = ["clawhdf5-io/mmap"]
@@ -0,0 +1,91 @@
//! Metadata parsing over an in-memory file: the local fast path that the
//! range-read `Storage` migration must not slow down
//! (`docs/design/range-reads.md`, "Keeping the local fast path").
//!
//! Only the `&[u8]` APIs are used, so the same file builds against older
//! revisions for an A/B comparison. The input is a version-1 (symbol table)
//! file with 400 groups, written by h5py with `libver='earliest'` and a
//! 512-byte user block (`clawhdf5-format/tests/fixtures/v1_groups_400.h5`).
use clawhdf5::{File, Group};
use clawhdf5_format::btree_v1::collect_symbol_table_nodes;
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
use clawhdf5_format::superblock::Superblock;
use clawhdf5_format::symbol_table::{SymbolTableMessage, SymbolTableNode};
use criterion::{Criterion, criterion_group, criterion_main};
use std::hint::black_box;
const FIXTURE: &str = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5-format/tests/fixtures/v1_groups_400.h5"
);
fn walk(g: &Group<'_>, objs: &mut usize) {
for name in g.datasets().unwrap_or_default() {
*objs += 1;
if let Ok(ds) = g.dataset(&name) {
let _ = black_box(ds.shape());
let _ = black_box(ds.dtype());
let _ = black_box(ds.attrs());
}
}
for name in g.groups().unwrap_or_default() {
*objs += 1;
if let Ok(sub) = g.group(&name) {
walk(&sub, objs);
}
}
}
fn bench_local_metadata(c: &mut Criterion) {
let bytes = std::fs::read(FIXTURE).unwrap();
let (_, f) = clawhdf5_format::signature::split_user_block(&bytes).unwrap();
let sb = Superblock::parse(f, 0).unwrap();
let (os, ls) = (sb.offset_size, sb.length_size);
let root = ObjectHeader::parse(f, sb.root_group_address as usize, os, ls).unwrap();
let stm = root
.messages
.iter()
.find(|m| m.msg_type == MessageType::SymbolTable)
.map(|m| SymbolTableMessage::parse(&m.data, os).unwrap())
.unwrap();
let nodes = collect_symbol_table_nodes(f, stm.btree_address, os, ls).unwrap();
let headers: Vec<u64> = nodes
.iter()
.flat_map(|&a| SymbolTableNode::parse(f, a as usize, os).unwrap().entries)
.map(|e| e.object_header_address)
.collect();
assert_eq!(headers.len(), 401);
let mut g = c.benchmark_group("local_metadata");
g.bench_function("object_header_parse_x401", |b| {
b.iter(|| {
for &a in &headers {
black_box(ObjectHeader::parse(f, a as usize, os, ls).unwrap());
}
})
});
g.bench_function("snod_parse_all", |b| {
b.iter(|| {
for &a in &nodes {
black_box(SymbolTableNode::parse(f, a as usize, os).unwrap());
}
})
});
g.bench_function("btree_v1_walk", |b| {
b.iter(|| black_box(collect_symbol_table_nodes(f, stm.btree_address, os, ls).unwrap()))
});
let file = File::open(FIXTURE).unwrap();
g.bench_function("facade_list_400_groups", |b| {
b.iter(|| {
let mut n = 0;
walk(&file.root(), &mut n);
assert_eq!(n, 401);
})
});
g.finish();
}
criterion_group!(benches, bench_local_metadata);
criterion_main!(benches);
+5
View File
@@ -6,7 +6,12 @@ use clawhdf5_format::error::FormatError;
use clawhdf5_format::message_type::MessageType; use clawhdf5_format::message_type::MessageType;
/// Errors that can occur when using the high-level API. /// Errors that can occur when using the high-level API.
///
/// Non-exhaustive: new failure modes add variants, so a `match` needs a
/// wildcard arm. Failures of a [`Storage`](clawhdf5_format::storage::Storage)
/// backend arrive as `Error::Format(FormatError::Storage(..))`.
#[derive(Debug)] #[derive(Debug)]
#[non_exhaustive]
pub enum Error { pub enum Error {
/// I/O error from the filesystem. /// I/O error from the filesystem.
Io(std::io::Error), Io(std::io::Error),
+16 -4
View File
@@ -1,7 +1,9 @@
# Design: range reads (reading HDF5 without holding the whole file) # Design: range reads (reading HDF5 without holding the whole file)
Status: proposal, 2026-09-26. No library code has changed; this document is Status: proposal, 2026-09-26; the plan for Phase 3's largest architectural
the plan for Phase 3's largest architectural change. Every count below was change. Progress: M1, first part (the `Storage` trait and the metadata
parsers listed in `CHANGELOG.md` under "Range reads, milestone M1") is done;
group B-tree v2 lookups, dense groups and the facade are not converted yet. Every count below was
taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given taken on `tank` on 2026-09-26 at commit `de2a53f`, with the commands given
next to it. No timing numbers appear here on purpose: the machine was shared next to it. No timing numbers appear here on purpose: the machine was shared
with other build jobs when this was written. with other build jobs when this was written.
@@ -259,6 +261,11 @@ every `file_data[a..b]` becomes `file.read_at(a, b - a)?`.
(binary size matters for wasm); `&dyn Storage` costs one indirect call per (binary size matters for wasm); `&dyn Storage` costs one indirect call per
structure read, negligible next to parsing. Hot raw-data loops keep their structure read, negligible next to parsing. Hot raw-data loops keep their
speed through `as_contiguous()`. speed through `as_contiguous()`.
*M1 outcome:* `&dyn` was not negligible for small structures — with it,
`ObjectHeader::parse` was ~25% and a 400-group facade listing ~14% slower
than the slice code (provisional, shared machine). The cores are now
generic (`S: Storage + ?Sized`), so the `&[u8]` wrappers get a `[u8]`
instance and `dyn Storage` is one more instance, not one per backend.
### (b) A page-cache "virtual slice" ### (b) A page-cache "virtual slice"
@@ -400,7 +407,7 @@ fast path within benchmark noise.
B-tree v1/v2, fractal heap, fixed/extensible array, symbol table, group B-tree v1/v2, fractal heap, fixed/extensible array, symbol table, group
v1/v2, shared messages, attributes, fill value, data layout — one commit v1/v2, shared messages, attributes, fill value, data layout — one commit
each. The old `&[u8]` signature stays as a thin wrapper over the new one each. The old `&[u8]` signature stays as a thin wrapper over the new one
(`fn parse(data: &[u8], ..) { parse_in(data as &dyn Storage, ..) }`), so (`fn parse(data: &[u8], ..) { parse_in(data, ..) }`, generic core), so
callers and the other crates don't move yet. callers and the other crates don't move yet.
- Replace the 5 open-ended slices and 38 `len()` checks with bounded reads. - Replace the 5 open-ended slices and 38 `len()` checks with bounded reads.
@@ -441,8 +448,13 @@ Total: roughly 6–10 engineer-weeks for M0–M4 (estimate, not measured).
- `impl Storage for [u8]` returns `Cow::Borrowed` — no copy, no allocation. - `impl Storage for [u8]` returns `Cow::Borrowed` — no copy, no allocation.
- Hot loops (raw-data copies, contiguous typed reads, `read_selection_native`) - Hot loops (raw-data copies, contiguous typed reads, `read_selection_native`)
branch once on `as_contiguous()` and then run today's code. branch once on `as_contiguous()` and then run today's code.
- `&dyn` dispatch is per structure, not per byte; parse code keeps working on - The parser cores are generic over `S: Storage + ?Sized`, so the in-memory
instance has no dispatch at all; a remote backend behind `&dyn Storage`
pays one indirect call per structure read. Parse code keeps working on
the returned slice. the returned slice.
- Reads sized by untrusted fields cover what the parser uses (see
`CHANGELOG.md`, M1), so a hostile size costs no read of the rest of the
file.
- Gate: `crates/clawhdf5/benches/mmap_bench.rs`, the concurrent-read benches in - Gate: `crates/clawhdf5/benches/mmap_bench.rs`, the concurrent-read benches in
`clawhdf5-bench`, and the conformance run time, before and after each M1/M2 `clawhdf5-bench`, and the conformance run time, before and after each M1/M2
commit, on an otherwise idle machine. Anything outside noise blocks the commit, on an otherwise idle machine. Anything outside noise blocks the