From e5d6f59e125f400ce953d9ad4896707cd685fc2f Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 20:38:51 -0500 Subject: [PATCH 1/7] clawhdf5-netcdf4: phony dimensions, skipped types and order as netCDF-C Read a file's metadata the way netCDF-C 4.9.3 does (libhdf5/hdf5open.c), for the whole file on first use (src/model.rs, replacing src/scope.rs): - links in creation order when the group tracks it, else name order; a group's datasets before its subgroups; dimension ids file-wide; - variables' dimensions from _Netcdf4Coordinates (file-wide ids), else the scales DIMENSION_LIST attaches when the first axis has one, else netCDF-C's phony dimensions phony_dim_ (create_phony_dims: shared by length and unlimitedness within a group, not between two axes of one variable, numbered subgroups first, a zero length unlimited); - datasets of types netCDF-C cannot represent are not variables (references, bit fields, time, arrays, compounds/enums/VLENs over them), replaying netCDF-C's file-wide type list, failed types included; - unlimited lengths as nc4_find_dim_len (its group and below). NcType gains Enum, Compound, VLen, Opaque and is #[non_exhaustive]; Variable::nc_type is netCDF-C's type (1-byte strings NC_CHAR). New clawhdf5_format::group_v2::links_in_creation_order_in. Tests compare with netCDF-C itself (tests/netcdf_c_view.py calls the libnetcdf netCDF4-python bundles through ctypes): new interop cases for h5py files without dimension scales, every type class, link order; and the gated corpus_vs_netcdf_c (CLAWHDF5_NETCDF_CORPUS): 420 of the 429 conformance-corpus files netCDF-C opens match (main: 68); the other 9 are explained in tests/corpus_known_differences.txt and known-issues. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 54 ++ README.md | 2 +- crates/clawhdf5-format/src/group_v2.rs | 43 + crates/clawhdf5-netcdf4/README.md | 65 +- crates/clawhdf5-netcdf4/src/dimension.rs | 211 +---- crates/clawhdf5-netcdf4/src/group.rs | 66 +- crates/clawhdf5-netcdf4/src/lib.rs | 71 +- crates/clawhdf5-netcdf4/src/model.rs | 853 ++++++++++++++++++ crates/clawhdf5-netcdf4/src/scope.rs | 231 ----- crates/clawhdf5-netcdf4/src/types.rs | 31 +- crates/clawhdf5-netcdf4/src/variable.rs | 27 +- crates/clawhdf5-netcdf4/tests/common/mod.rs | 281 ++++++ .../tests/corpus_known_differences.txt | 26 + .../tests/corpus_vs_netcdf_c.rs | 173 ++++ .../clawhdf5-netcdf4/tests/interop_tests.rs | 266 ++++++ .../clawhdf5-netcdf4/tests/netcdf4_tests.rs | 10 +- .../clawhdf5-netcdf4/tests/netcdf_c_view.py | 238 +++++ docs/known-issues.md | 139 ++- 18 files changed, 2263 insertions(+), 524 deletions(-) create mode 100644 crates/clawhdf5-netcdf4/src/model.rs delete mode 100644 crates/clawhdf5-netcdf4/src/scope.rs create mode 100644 crates/clawhdf5-netcdf4/tests/common/mod.rs create mode 100644 crates/clawhdf5-netcdf4/tests/corpus_known_differences.txt create mode 100644 crates/clawhdf5-netcdf4/tests/corpus_vs_netcdf_c.rs create mode 100644 crates/clawhdf5-netcdf4/tests/netcdf_c_view.py diff --git a/CHANGELOG.md b/CHANGELOG.md index 88e255b..843c58b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,60 @@ ## Unreleased +### NetCDF-4: phony dimensions, skipped types and order as in netCDF-C (2026-09-29) +- `clawhdf5-netcdf4` now reads a file's metadata as netCDF-C 4.9.3 does + (`libhdf5/hdf5open.c`), for the whole file on first use + (`src/model.rs`): links in creation order when the group tracks it, + else in name order; a group's datasets before its subgroups; dimension + ids file-wide (`_Netcdf4Dimid`, else the next free id); then variables' + dimensions, subgroups first: `_Netcdf4Coordinates` ids (looked up + file-wide), else the scales `DIMENSION_LIST` attaches (when the first + axis has one), else phony dimensions. +- **Files without dimension scales** get netCDF-C's phony dimensions, + `phony_dim_` (`create_phony_dims`): per axis, the first dimension of + the variable's group of the same length and unlimitedness not used by an + earlier axis of the variable (real dimensions included), else a new one; + a length of 0 is always unlimited. They used to be one dimension per 1-D + dataset, named after it, and `dim_` for other axes. +- **Datasets netCDF-C skips are not variables:** references, bit fields, + time and array types, and compounds, enums and VLENs whose members or + base type are not netCDF atomic types (4- and 8-byte floats only) or a + type read before — replaying netCDF-C's file-wide type list, which also + keeps a type it failed to read, so the second dataset of a compound with + a reference member is a variable, as in netCDF-C. +- `NcType` gains `Enum`, `Compound`, `VLen` and `Opaque` and is now + `#[non_exhaustive]` (breaking for exhaustive matches); + `Variable::nc_type` is the type netCDF-C gives the variable: a 1-byte + fixed-length string is `Char`, a longer one `String` (both were + `String`); user-defined types have their class (they were `Char`). + `dtype_to_nctype` maps compounds and enums to their classes. +- Groups (`group_names`), variables (`variables`, `variable_names`) and + dimensions (`dimensions`, by id) come in netCDF-C's order; a zero-length + dimension scale is unlimited; an unlimited dimension's length is the + longest extent along it of the variables in its group and below + (`nc4_find_dim_len`; was: of the variables its `REFERENCE_LIST` names). + `NetCDF4File::group` and `NetCDF4Group::group` take a path (`"a/b"`). +- Deliberate differences: floats of other than 4 or 8 bytes are + `Float`/`Double` (netCDF-C 4.9.3 on libhdf5 1.14.6 labels them + `NC_STRING`); an axis netCDF-C leaves without a dimension (where it + reads uninitialised memory) gets one by the phony rule. +- New `clawhdf5_format::group_v2::links_in_creation_order_in`: a group's + link names in creation order, or `None` when it does not track it. +- Tests compare with netCDF-C itself (`tests/netcdf_c_view.py` calls the + libnetcdf netCDF4-python bundles through ctypes, since netCDF4-python + hides variables of types it does not support): `interop_tests` cases of + h5py files without dimension scales (sharing, unlimited and zero-length + axes, subgroups, a real dimension taken by length), of every HDF5 type + class, of creation and name order in compact and dense groups, and a + netCDF4-python file; and the gated `tests/corpus_vs_netcdf_c.rs` + (`CLAWHDF5_NETCDF_CORPUS=`): over the conformance corpus 420 of the + 429 files netCDF-C 4.9.3 opens match in groups, dimensions, variables, + types, shapes and numeric values (up to 5000 elements); `main` at + `4260af4` matched 68. The other 9 (external links, values the HDF5 + reader refuses) are explained in `tests/corpus_known_differences.txt` + and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4). + Affected v2.1.0 to v2.7.0. + ### NetCDF-4: variables' dimensions come from the file (2026-09-28) - `clawhdf5-netcdf4` gave each variable the first unused dimension of equal size (else an anonymous `dim_`), so a variable on an unlimited diff --git a/README.md b/README.md index d7f61da..79182af 100644 --- a/README.md +++ b/README.md @@ -386,7 +386,7 @@ filter). | `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) | | `clawhdf5-io` | I/O helpers: mmap, async, an HSDS client, `mpi-io` (not collective I/O) | | `clawhdf5-remote` | HTTP(S) and object-store files through a block cache | -| `clawhdf5-netcdf4` | NetCDF-4 dimensions, variables, CF attributes | +| `clawhdf5-netcdf4` | NetCDF-4 dimensions, variables, CF attributes, as netCDF-C reports them (phony dimensions included) | | `clawhdf5-derive` | Derive macros for HDF5-serialisable structs | | `clawhdf5-tools` | `h5rs`: `ls`, `dump`, `stat`, `diff`, `check` | | **Bindings** | | diff --git a/crates/clawhdf5-format/src/group_v2.rs b/crates/clawhdf5-format/src/group_v2.rs index e902239..087f819 100644 --- a/crates/clawhdf5-format/src/group_v2.rs +++ b/crates/clawhdf5-format/src/group_v2.rs @@ -731,6 +731,49 @@ fn group_children( Ok(entries) } +/// The names of the links of the group at `group_address` in creation +/// order — the order libhdf5 iterates them in with `H5_INDEX_CRT_ORDER` +/// (`H5Literate`) — when the group tracks the creation order of its links; +/// `None` when it does not (a version-1 group never does), in which case +/// libhdf5 can only iterate by name (`H5_INDEX_NAME`: byte order of the +/// names). Hard, soft and external links are listed (user-defined ones, +/// which cannot be followed, are not); links without a creation order +/// value, which a tracking group should not have, come last in the order +/// they are stored. +pub fn links_in_creation_order_in( + file_data: &S, + superblock: &Superblock, + group_address: u64, +) -> Result>, FormatError> { + let os = superblock.offset_size; + let ls = superblock.length_size; + let header = ObjectHeader::parse_in(file_data, checked_addr(group_address)?, os, ls)?; + if !is_v2_group(&header) { + return Ok(None); + } + let link_info = find_link_info(&header, os)?; + if link_info.max_creation_order.is_none() { + return Ok(None); + } + let mut links: Vec<(u64, String)> = Vec::new(); + let mut visit = |link: LinkMessage| { + links.push((link.creation_order.unwrap_or(u64::MAX), link.name)); + }; + if let Some(fh_addr) = link_info.fractal_heap_address { + for_each_dense_link(file_data, &link_info, fh_addr, os, ls, false, visit)?; + } else { + for msg in &header.messages { + if msg.msg_type == MessageType::Link + && let Some(link) = parse_link(&msg.data, os)? + { + visit(link); + } + } + } + links.sort_by_key(|&(order, _)| order); + Ok(Some(links.into_iter().map(|(_, name)| name).collect())) +} + /// Soft links followed while resolving one path. Guards against link cycles /// (`a -> b -> a`), which are legal to create. const MAX_SOFT_LINK_DEPTH: u8 = 16; diff --git a/crates/clawhdf5-netcdf4/README.md b/crates/clawhdf5-netcdf4/README.md index 426e823..936f627 100644 --- a/crates/clawhdf5-netcdf4/README.md +++ b/crates/clawhdf5-netcdf4/README.md @@ -36,23 +36,46 @@ let values: Vec = temp.read_f64()?; | Item | What | |---|---| -| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | +| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable` (a name or a path into a subgroup), `global_attrs`, `group` (a name or a path), `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | | `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `variable_names`, `attrs`, nested `group`) | | `Variable` | `name`, `shape`, `stored_shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` | | `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) | | `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` | -| `NcType` | the NetCDF type of a variable | +| `NcType` | the NetCDF type of a variable: an atomic type (`Byte` ... `UInt64`, `Float`, `Double`, `Char`, `String`) or the class of a user-defined one (`Enum`, `Compound`, `VLen`, `Opaque`); `#[non_exhaustive]` | -Variables and dimensions follow netCDF-C: +Groups, dimensions and variables are what netCDF-C 4.9.3 reports +(`libhdf5/hdf5open.c`), names, order and types included. The first call +that needs them reads the metadata of the whole file, as `nc_open` does: -- A variable's dimensions are the ones the file names: the ids in its - `_Netcdf4Coordinates` attribute, else the dimension scales its - `DIMENSION_LIST` references, found in its group or a parent group. Only - an axis the file names no dimension for (an HDF5 file not written by a - netCDF library) gets the first dimension of the group of the same size, - else an anonymous `dim_`. -- Dimension scales that are only dimensions are not variables; a dataset +- Order: a group's links in creation order when it tracks it (netCDF-4 + files do), else in name order (h5py's default); groups and variables in + that order, dimensions by id. +- A dimension scale defines a dimension (its `_Netcdf4Dimid`, else the + next free id; unlimited when its first axis is or its length is 0); + scales that are only dimensions are not variables, and a dataset `_nc4_non_coord_` is the variable ``. +- A variable's dimensions are the ones the file names: the ids in its + `_Netcdf4Coordinates` attribute (any group), else — when its first axis + has one — the dimension scales its `DIMENSION_LIST` attaches, found in + its group or a parent group. +- Otherwise (an HDF5 file not written by a netCDF library) its axes get + netCDF-C's **phony dimensions** `phony_dim_`: per axis, the first + dimension of the variable's group with the same length and + unlimitedness that an earlier axis of the variable does not use, else a + new one. The numbers run file-wide, a group's subgroups before its own + variables. +- Datasets of types netCDF-C cannot represent are not variables: + references, bit fields, time and array types, and compounds, enums and + VLENs over members or base types that are not netCDF atomic types (4- + and 8-byte floats only) or types netCDF-C has read before in the file. + Enum, compound, VLEN and opaque datasets are variables of those classes + (netCDF4-python itself leaves out opaque ones). +- Two deliberate differences: floats of other than 4 or 8 bytes (half, + bfloat16, 4/6/8-bit floats, `long double`) are `Float`/`Double` and + read as numbers, where netCDF-C 4.9.3 on libhdf5 1.14.6 labels them + `NC_STRING`; and an axis netCDF-C leaves without a dimension (an id or + scale it cannot find, or no scale on an axis after a first one that has + one, where it reads uninitialised memory) gets one by the phony rule. - A variable along an unlimited dimension has the dimension's length: `shape` is that length, and the reads return that many values, the records the variable has not written as its fill value (`_FillValue`, @@ -60,11 +83,23 @@ Variables and dimensions follow netCDF-C: `stored_shape` is the HDF5 dataset's extent. No cargo features. Tests compare against files written by netCDF4-python, -h5py dimension scales, h5netcdf and xarray, variable by variable with what -netCDF4-python reads (`tests/interop_tests.rs`; the CI job requires them -with `CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is -not installed). What the HDF5 reader underneath cannot read is listed in -[`docs/known-issues.md`](../../docs/known-issues.md). +h5py (with and without dimension scales, every HDF5 type class), h5netcdf +and xarray, with what netCDF4-python reads and with what netCDF-C itself +reports (`tests/interop_tests.rs`; `tests/netcdf_c_view.py` calls the +libnetcdf netCDF4-python bundles; the CI job requires them with +`CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is not +installed). `tests/corpus_vs_netcdf_c.rs` compares every file of a corpus +netCDF-C opens, when `CLAWHDF5_NETCDF_CORPUS` names one: over the +conformance corpus, 420 of 429 files match (tank, 2026-09-29), and the +other 9 are explained in `tests/corpus_known_differences.txt`: + +```sh +CLAWHDF5_NETCDF_CORPUS=conformance/.cache/corpus CLAWHDF5_PYTHON=$PWD/.venv/bin/python \ + cargo test -p clawhdf5-netcdf4 --test corpus_vs_netcdf_c -- --nocapture +``` + +The differences from netCDF-C, and what the HDF5 reader underneath cannot +read, are listed in [`docs/known-issues.md`](../../docs/known-issues.md). ## License diff --git a/crates/clawhdf5-netcdf4/src/dimension.rs b/crates/clawhdf5-netcdf4/src/dimension.rs index 4f88819..9262d32 100644 --- a/crates/clawhdf5-netcdf4/src/dimension.rs +++ b/crates/clawhdf5-netcdf4/src/dimension.rs @@ -1,16 +1,15 @@ //! NetCDF-4 dimension representation. //! //! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE -//! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via -//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited); their length -//! is the largest extent of the variables attached to them. +//! attribute and a `_Netcdf4Dimid` attribute; variables name theirs in +//! `_Netcdf4Coordinates` and `DIMENSION_LIST`. A file without them gets +//! netCDF-C's phony dimensions. How they are put together is in +//! `crate::model`. use std::collections::HashMap; use clawhdf5::AttrValue; -use crate::error::Error; - /// A NetCDF-4 dimension. #[derive(Debug, Clone, PartialEq, Eq)] pub struct Dimension { @@ -22,202 +21,10 @@ pub struct Dimension { pub is_unlimited: bool, } -/// A dimension scale of one group: the dataset that defines a dimension. -#[derive(Debug, Clone)] -pub(crate) struct Scale { - /// Object header address of the scale's dataset (what a variable's - /// `DIMENSION_LIST` references). - pub address: u64, - /// Its `_Netcdf4Dimid` (what a variable's `_Netcdf4Coordinates` lists). - pub dimid: Option, - /// Index of its dimension in [`GroupDims::dims`]. - pub dim: usize, -} - -/// The dimensions a group defines, with the scales that define them. -#[derive(Debug, Clone, Default)] -pub(crate) struct GroupDims { - /// The group's dimensions, in `_Netcdf4Dimid` order (then discovery order). - pub dims: Vec, - /// The dimension scales behind `dims`; empty when the group has no - /// dimension scale and `dims` were inferred from 1-D datasets. - pub scales: Vec, -} - -impl GroupDims { - /// The dimension defined by the scale at `address`. - pub fn by_address(&self, address: u64) -> Option<&Dimension> { - self.scales - .iter() - .find(|s| s.address == address) - .map(|s| &self.dims[s.dim]) - } - - /// The dimension whose scale has `_Netcdf4Dimid` `id`. - pub fn by_dimid(&self, id: i64) -> Option<&Dimension> { - self.scales - .iter() - .find(|s| s.dimid == Some(id)) - .map(|s| &self.dims[s.dim]) - } -} - -/// The dimensions of an HDF5 group (root or subgroup). -/// -/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed -/// dimension's size is the dataset's first (and typically only) shape extent. -/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace; -/// their size is computed by `unlimited_len`. A group with no dimension -/// scale at all (not written by a netCDF library) gets one dimension per -/// 1-D dataset instead. -pub(crate) fn group_dims( - file: &clawhdf5::File, - group: &clawhdf5::Group<'_>, -) -> Result { - let addresses: HashMap = group.entries()?.into_iter().collect(); - let dataset_names = group.datasets()?; - // (dimid, dimension, scale address), in discovery order. - let mut found: Vec<(Option, Dimension, u64)> = Vec::new(); - - for ds_name in &dataset_names { - let ds = group.dataset(ds_name)?; - let attrs = ds.attrs()?; - if !is_dimension_scale(&attrs) { - continue; - } - let Some(&address) = addresses.get(ds_name) else { - continue; - }; - let shape = ds.shape()?; - let is_unlimited = is_unlimited(&ds); - let size = if is_unlimited { - unlimited_len(file, &attrs, &shape) - } else { - shape.first().copied().unwrap_or(0) - }; - let dim = Dimension { - name: ds_name.clone(), - size, - is_unlimited, - }; - found.push((get_dimid(&attrs), dim, address)); - } - - if found.is_empty() { - // Fallback: infer dimensions from dataset shapes and names. - // In NetCDF-4, coordinate variables are datasets whose name matches - // a dimension name. If there are no explicit DIMENSION_SCALE attributes, - // we look for 1-D datasets that might be coordinate variables. - let mut dims = Vec::new(); - for ds_name in &dataset_names { - let ds = group.dataset(ds_name)?; - let shape = ds.shape()?; - if shape.len() == 1 { - dims.push(Dimension { - name: ds_name.clone(), - size: shape[0], - is_unlimited: is_unlimited(&ds), - }); - } - } - return Ok(GroupDims { - dims, - scales: Vec::new(), - }); - } - - // By dimid; scales without one keep their discovery order after those - // with one (the sort is stable). - found.sort_by_key(|(id, ..)| (id.is_none(), id.unwrap_or(0))); - let mut out = GroupDims::default(); - for (i, (dimid, dim, address)) in found.into_iter().enumerate() { - out.dims.push(dim); - out.scales.push(Scale { - address, - dimid, - dim: i, - }); - } - Ok(out) -} - /// The start of the `NAME` attribute netCDF-C gives a dimension scale that /// is only a dimension, not also a (coordinate) variable. const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable"; -/// The current length of an unlimited dimension, as netCDF-C reports it -/// (`NC4_inq_dim` → `nc4_find_dim_len`): the largest current extent, along -/// the dimension, of the variables that use it, in any group; 0 when none -/// has been written. netCDF-C does not extend a dimension scale that is not -/// also a variable, so such a scale's own extent (0) is not counted; a -/// coordinate variable's is. The variables are the scale's attachments, -/// listed with the axis they use in its `REFERENCE_LIST` attribute (the -/// mirror of each variable's `DIMENSION_LIST`). Attachments that cannot be -/// read are skipped; without a readable `REFERENCE_LIST` the length is the -/// scale's own extent, as before. -fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap, shape: &[u64]) -> u64 { - let own = shape.first().copied().unwrap_or(0); - let is_variable = !is_pure_dimension(attrs); - let Some(refs) = reference_list(file, attrs) else { - return own; - }; - refs.into_iter() - .filter_map(|(address, axis)| { - let shape = file.dataset_at(address).ok()?.shape().ok()?; - shape.get(usize::try_from(axis).ok()?).copied() - }) - .chain(is_variable.then_some(own)) - .max() - .unwrap_or(0) -} - -/// The `(dataset address, axis)` pairs of a dimension scale's -/// `REFERENCE_LIST` attribute (HDF5 dimension scales: a compound of an -/// object reference `dataset` and an integer `dimension`), or `None` when it -/// is missing or not in that form. -fn reference_list( - file: &clawhdf5::File, - attrs: &HashMap, -) -> Option> { - use clawhdf5_format::data_read::{read_compound_field, read_object_references}; - use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder}; - let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("REFERENCE_LIST") else { - return None; - }; - let dataset = read_compound_field(data, datatype, "dataset").ok()?; - let addresses = read_object_references( - &dataset.raw_data, - &dataset.datatype, - file.superblock().offset_size, - ) - .ok()?; - let dimension = read_compound_field(data, datatype, "dimension").ok()?; - let Datatype::FixedPoint { - size, byte_order, .. - } = dimension.datatype - else { - return None; - }; - let size = usize::try_from(size).ok().filter(|s| (1..=8).contains(s))?; - let axes = dimension.raw_data.chunks_exact(size).map(|b| { - let mut v = [0u8; 8]; - match byte_order { - DatatypeByteOrder::BigEndian => { - v[8 - size..].copy_from_slice(b); - u64::from_be_bytes(v) - } - _ => { - v[..size].copy_from_slice(b); - u64::from_le_bytes(v) - } - } - }); - if axes.len() != addresses.len() { - return None; - } - Some(addresses.into_iter().map(|r| r.address).zip(axes).collect()) -} - /// The dimension scale attached to each axis of a variable, from its /// `DIMENSION_LIST` attribute (HDF5 dimension scales: one variable-length /// sequence of object references per axis) — the address of the scale @@ -281,13 +88,9 @@ pub(crate) fn is_dimension_scale(attrs: &HashMap) -> bool { pub(crate) fn get_dimid(attrs: &HashMap) -> Option { match attrs.get("_Netcdf4Dimid") { Some(AttrValue::I64(id)) => Some(*id), - Some(AttrValue::U64(id)) => Some(*id as i64), + Some(AttrValue::U64(id)) => i64::try_from(*id).ok(), + Some(AttrValue::I64Array(ids)) if ids.len() == 1 => Some(ids[0]), + Some(AttrValue::U64Array(ids)) if ids.len() == 1 => i64::try_from(ids[0]).ok(), _ => None, } } - -/// Whether a dataset's first axis is unlimited (`max_dimensions[0] == -/// u64::MAX` in its dataspace). -fn is_unlimited(ds: &clawhdf5::Dataset<'_>) -> bool { - matches!(ds.max_dimensions(), Ok(Some(max_dims)) if max_dims.first() == Some(&u64::MAX)) -} diff --git a/crates/clawhdf5-netcdf4/src/group.rs b/crates/clawhdf5-netcdf4/src/group.rs index 4178ee5..7bc9812 100644 --- a/crates/clawhdf5-netcdf4/src/group.rs +++ b/crates/clawhdf5-netcdf4/src/group.rs @@ -7,36 +7,36 @@ use std::collections::HashMap; use clawhdf5::AttrValue; -use crate::dimension::{self, Dimension}; +use crate::dimension::Dimension; use crate::error::Error; -use crate::scope::{self, Scope}; +use crate::model::Model; use crate::variable::Variable; /// A NetCDF-4 group corresponding to an HDF5 group. pub struct NetCDF4Group<'f> { /// Group name. name: String, - /// Path of the group from the root (`/`-separated). - path: String, /// Underlying HDF5 file. file: &'f clawhdf5::File, - /// Underlying HDF5 group. - hdf5_group: clawhdf5::Group<'f>, + /// The file's groups, dimensions and variables. + model: &'f Model, + /// This group's index in `model`. + index: usize, } impl<'f> NetCDF4Group<'f> { - /// Create a new NetCDF4Group from an HDF5 group. + /// The group `index` of `model`. pub(crate) fn new( name: String, - path: String, file: &'f clawhdf5::File, - hdf5_group: clawhdf5::Group<'f>, + model: &'f Model, + index: usize, ) -> Self { Self { name, - path, file, - hdf5_group, + model, + index, } } @@ -45,52 +45,56 @@ impl<'f> NetCDF4Group<'f> { &self.name } - /// List dimensions defined in this group (not those of its parent - /// groups, which its variables can also use). + /// List dimensions defined in this group, in dimension-id order (not + /// those of its parent groups, which its variables can also use). pub fn dimensions(&self) -> Result, Error> { - Ok(dimension::group_dims(self.file, &self.hdf5_group)?.dims) + Ok(self.model.dimensions(self.index)) } - /// List variables in this group: its datasets, except the dimension - /// scales that are only dimensions. Their dimensions can be defined in - /// this group or a parent group. + /// List variables in this group, in netCDF-C's order: its datasets, + /// except the dimension scales that are only dimensions and datasets + /// of types netCDF-C cannot represent. Their dimensions can be defined + /// in this group or a parent group. pub fn variables(&self) -> Result>, Error> { - Scope::new(self.file, &self.path)?.variables() + self.model.variables(self.file, self.index) } - /// Get a specific variable by name. + /// Get a specific variable by name (or by path, `"sub/var"`). pub fn variable(&self, name: &str) -> Result, Error> { - scope::variable_at(self.file, &self.path, name) + self.model.variable(self.file, self.index, name) } /// Read all attributes of this group. pub fn attrs(&self) -> Result, Error> { - Ok(self.hdf5_group.attrs()?) + Ok(self + .file + .group_at(self.model.group_address(self.index)) + .attrs()?) } - /// List subgroup names. + /// List subgroup names, in netCDF-C's order. pub fn group_names(&self) -> Result, Error> { - Ok(self.hdf5_group.groups()?) + Ok(self.model.group_names(self.index)) } - /// Get a subgroup by name. + /// Get a subgroup by name (or by path, `"a/b"`). pub fn group(&self, name: &str) -> Result, Error> { - let hdf5_group = self - .hdf5_group - .group(name) - .map_err(|_| Error::GroupNotFound(name.to_string()))?; + let index = self + .model + .find_group(self.index, name) + .ok_or_else(|| Error::GroupNotFound(name.to_string()))?; Ok(NetCDF4Group::new( name.to_string(), - format!("{}/{name}", self.path), self.file, - hdf5_group, + self.model, + index, )) } /// The names of this group's variables (see /// [`variables`](Self::variables)). pub fn variable_names(&self) -> Result, Error> { - Scope::new(self.file, &self.path)?.variable_names() + Ok(self.model.variable_names(self.index)) } } diff --git a/crates/clawhdf5-netcdf4/src/lib.rs b/crates/clawhdf5-netcdf4/src/lib.rs index 0e22773..3ebc2da 100644 --- a/crates/clawhdf5-netcdf4/src/lib.rs +++ b/crates/clawhdf5-netcdf4/src/lib.rs @@ -27,7 +27,7 @@ pub mod cf; pub mod dimension; pub mod error; pub mod group; -mod scope; +mod model; pub mod types; pub mod variable; @@ -40,26 +40,49 @@ pub use types::NcType; pub use variable::Variable; use std::collections::HashMap; +use std::sync::OnceLock; + +use model::Model; /// A NetCDF-4 file reader. /// /// Wraps a clawhdf5 File and provides NetCDF-4 semantics: dimensions, /// variables with CF attributes, groups, and type mapping. +/// +/// The groups, dimensions and variables are what netCDF-C reports for the +/// file (see the crate README): the first call that needs them reads the +/// metadata of the whole file, as `nc_open` does, and keeps it; the values +/// are read when asked for. pub struct NetCDF4File { hdf5: clawhdf5::File, + model: OnceLock, } impl NetCDF4File { /// Open a NetCDF-4 file from a filesystem path. pub fn open>(path: P) -> Result { - let hdf5 = clawhdf5::File::open(path)?; - Ok(Self { hdf5 }) + Ok(Self::from_hdf5(clawhdf5::File::open(path)?)) } /// Open a NetCDF-4 file from in-memory bytes. pub fn from_bytes(data: Vec) -> Result { - let hdf5 = clawhdf5::File::from_bytes(data)?; - Ok(Self { hdf5 }) + Ok(Self::from_hdf5(clawhdf5::File::from_bytes(data)?)) + } + + fn from_hdf5(hdf5: clawhdf5::File) -> Self { + Self { + hdf5, + model: OnceLock::new(), + } + } + + /// The file's groups, dimensions and variables, read on first use. + fn model(&self) -> Result<&Model, Error> { + if let Some(model) = self.model.get() { + return Ok(model); + } + let model = Model::build(&self.hdf5)?; + Ok(self.model.get_or_init(|| model)) } /// Get the _NCProperties root attribute, if present. @@ -74,27 +97,29 @@ impl NetCDF4File { } } - /// List dimensions defined in the root group. + /// List dimensions defined in the root group, in dimension-id order. pub fn dimensions(&self) -> Result, Error> { - Ok(dimension::group_dims(&self.hdf5, &self.hdf5.root())?.dims) + Ok(self.model()?.dimensions(0)) } - /// List all variables in the root group: its datasets, except the - /// dimension scales that are only dimensions (netCDF-C does not list - /// them either). + /// List the variables of the root group, in netCDF-C's order: its + /// datasets, except the dimension scales that are only dimensions and + /// datasets of types netCDF-C cannot represent (see + /// [`NcType`]). pub fn variables(&self) -> Result>, Error> { - scope::Scope::new(&self.hdf5, "/")?.variables() + self.model()?.variables(&self.hdf5, 0) } /// The names of the root group's variables (see /// [`variables`](Self::variables)). pub fn variable_names(&self) -> Result, Error> { - scope::Scope::new(&self.hdf5, "/")?.variable_names() + Ok(self.model()?.variable_names(0)) } - /// Get a specific variable by name from the root group. + /// Get a specific variable by name from the root group; the name may be + /// a path into a subgroup (`"sub/var"`). pub fn variable(&self, name: &str) -> Result, Error> { - scope::variable_at(&self.hdf5, "", name) + self.model()?.variable(&self.hdf5, 0, name) } /// Read all global (root group) attributes. @@ -102,22 +127,22 @@ impl NetCDF4File { Ok(self.hdf5.root().attrs()?) } - /// List subgroup names in the root group. + /// List subgroup names in the root group, in netCDF-C's order. pub fn group_names(&self) -> Result, Error> { - Ok(self.hdf5.root().groups()?) + Ok(self.model()?.group_names(0)) } - /// Get a subgroup by name. + /// Get a subgroup by name (or by path, `"a/b"`). pub fn group(&self, name: &str) -> Result, Error> { - let hdf5_group = self - .hdf5 - .group(name) - .map_err(|_| Error::GroupNotFound(name.to_string()))?; + let model = self.model()?; + let index = model + .find_group(0, name) + .ok_or_else(|| Error::GroupNotFound(name.to_string()))?; Ok(NetCDF4Group::new( - name.to_string(), name.to_string(), &self.hdf5, - hdf5_group, + model, + index, )) } diff --git a/crates/clawhdf5-netcdf4/src/model.rs b/crates/clawhdf5-netcdf4/src/model.rs new file mode 100644 index 0000000..5c0cd4b --- /dev/null +++ b/crates/clawhdf5-netcdf4/src/model.rs @@ -0,0 +1,853 @@ +//! A file's groups, dimensions and variables, as netCDF-C reads them. +//! +//! netCDF-C (`libhdf5/hdf5open.c`, 4.9.3) reads a file in two passes, and +//! the names, order and sharing of the dimensions depend on both, so this +//! module replays them for the whole file at once: +//! +//! 1. `rec_read_metadata`: each group's links in creation order when the +//! group tracks it, else in name order; a group's datasets and named +//! datatypes before its subgroups, which follow in the same order. A +//! dimension scale (`CLASS` `DIMENSION_SCALE`) defines a dimension with +//! the id in its `_Netcdf4Dimid`, else the next free id (ids are +//! file-wide), its first extent as length, unlimited when that axis is +//! (or the length is 0); it is a variable too unless its `NAME` says it +//! is only a dimension. Every other dataset is a variable, unless its +//! type is one netCDF-C cannot represent ([`VarTypes::nc_type`]); a +//! dataset `_nc4_non_coord_` is the variable ``. +//! 2. `rec_match_dimscales`, subgroups first, then the group's variables in +//! order: a variable gets the dimensions whose ids its +//! `_Netcdf4Coordinates` lists (looked up file-wide), else — when its +//! `DIMENSION_LIST` attaches a scale to its first axis — the scales that +//! list attaches, looked up in its group and then each parent, else +//! "phony" dimensions (`create_phony_dims`): per axis, the first +//! dimension of the variable's group of that length and unlimitedness +//! not already used by an earlier axis of the variable, else a new one +//! called `phony_dim_`. +//! +//! An unlimited dimension's length is the largest extent, along it, of the +//! variables that use it in its group and the groups below +//! (`nc4_find_dim_len`). +//! +//! Where netCDF-C 4.9.3 leaves an axis without a dimension (an id or a scale +//! it cannot find, or an axis without a scale of a variable whose first +//! axis has one — there it reads uninitialised memory), netCDF4-python +//! cannot open the file; this crate gives such an axis a dimension by the +//! phony rule instead. + +use std::collections::HashMap; + +use clawhdf5::AttrValue; +use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder}; +use clawhdf5_format::object_header::{ObjectClass, ObjectHeader}; + +use crate::dimension::{self, Dimension}; +use crate::error::Error; +use crate::types::NcType; +use crate::variable::Variable; + +/// The prefix netCDF-C gives the dataset of a variable that has a +/// dimension's name but is not that dimension's coordinate variable (the +/// dimension's scale holds the name). +const NON_COORD_PREFIX: &str = "_nc4_non_coord_"; + +/// Groups read at most, a guard against files whose groups are hard-linked +/// into each other many times over (netCDF-C reads each link as its own +/// group). +const MAX_GROUPS: usize = 100_000; + +/// A file as netCDF-C sees it. +#[derive(Debug)] +pub(crate) struct Model { + /// Every group; the root is the first. + groups: Vec, + /// Every dimension, in the order they were created. + dims: Vec, +} + +#[derive(Debug)] +struct Group { + /// Object header address of the HDF5 group. + address: u64, + /// Its subgroups, `(name, index in Model::groups)`, in netCDF-C's order. + children: Vec<(String, usize)>, + parent: Option, + /// The dimensions it defines (indexes in `Model::dims`), in creation + /// order. + dims: Vec, + /// Its variables, in netCDF-C's order. + vars: Vec, +} + +#[derive(Debug)] +struct Dim { + id: i64, + name: String, + /// The length it was created with (for an unlimited dimension, replaced + /// by its current length once every variable has its dimensions). + len: u64, + unlimited: bool, + /// Object header address of its dimension scale; `None` for a phony + /// dimension. + scale: Option, +} + +#[derive(Debug)] +struct Var { + /// The netCDF name. + name: String, + /// Whether its dataset is `_nc4_non_coord_`. + non_coord: bool, + address: u64, + nc_type: NcType, + extent: Vec, + /// Whether each axis is unlimited in the dataspace. + unlimited: Vec, + /// How netCDF-C finds its dimensions. + source: DimSource, + /// Its dimensions (indexes in `Model::dims`), one per axis, once found. + dims: Vec, +} + +#[derive(Debug)] +enum DimSource { + /// A one-dimensional coordinate variable: its own scale's dimension. + OwnScale(usize), + /// The ids in `_Netcdf4Coordinates`; for a multi-dimensional coordinate + /// variable, also its own dimension (for the first axis, should an id + /// not be found). + Coordinates(Vec, Option), + /// The scale `DIMENSION_LIST` attaches to each axis (the first has one). + Scales(Vec>, Option), + /// None: phony dimensions. + Phony(Option), +} + +impl Model { + /// Read the metadata of `file` as netCDF-C does. + pub fn build(file: &clawhdf5::File) -> Result { + let mut builder = Builder { + file, + model: Model { + groups: Vec::new(), + dims: Vec::new(), + }, + next_id: 0, + types: VarTypes::default(), + }; + let root = file.superblock().root_group_address; + builder.model.groups.push(Group { + address: root, + children: Vec::new(), + parent: None, + dims: Vec::new(), + vars: Vec::new(), + }); + builder.read_group(0, &mut vec![root])?; + builder.match_dims(0); + builder.unlimited_lengths(); + Ok(builder.model) + } + + /// The group at `path` (`/`-separated names), from the group `from`. + pub fn find_group(&self, from: usize, path: &str) -> Option { + path.split('/') + .filter(|p| !p.is_empty()) + .try_fold(from, |g, name| { + self.groups[g] + .children + .iter() + .find(|(n, _)| n == name) + .map(|&(_, i)| i) + }) + } + + /// The object header address of group `g`. + pub fn group_address(&self, g: usize) -> u64 { + self.groups[g].address + } + + /// The names of group `g`'s subgroups, in netCDF-C's order. + pub fn group_names(&self, g: usize) -> Vec { + self.groups[g] + .children + .iter() + .map(|(n, _)| n.clone()) + .collect() + } + + /// The dimensions group `g` defines, in id order (as `nc_inq_dimids`). + pub fn dimensions(&self, g: usize) -> Vec { + let mut dims: Vec<&Dim> = self.groups[g].dims.iter().map(|&d| &self.dims[d]).collect(); + dims.sort_by_key(|d| d.id); + dims.into_iter().map(Dim::dimension).collect() + } + + /// The names of group `g`'s variables, in netCDF-C's order. + pub fn variable_names(&self, g: usize) -> Vec { + self.groups[g].vars.iter().map(|v| v.name.clone()).collect() + } + + /// Group `g`'s variables, in netCDF-C's order. + pub fn variables<'f>( + &self, + file: &'f clawhdf5::File, + g: usize, + ) -> Result>, Error> { + self.groups[g] + .vars + .iter() + .map(|v| self.open(file, v)) + .collect() + } + + /// The variable `name` of group `g`; `name` may be a path (`"sub/var"`) + /// relative to it. Of two variables of one name (datasets `` and + /// `_nc4_non_coord_`), the second. + pub fn variable<'f>( + &self, + file: &'f clawhdf5::File, + g: usize, + name: &str, + ) -> Result, Error> { + let not_found = || Error::VariableNotFound(name.to_string()); + let trimmed = name.trim_start_matches('/'); + let (g, leaf) = match trimmed.rsplit_once('/') { + Some((dir, leaf)) => (self.find_group(g, dir).ok_or_else(not_found)?, leaf), + None => (g, trimmed), + }; + let vars = &self.groups[g].vars; + let var = vars + .iter() + .find(|v| v.name == leaf && v.non_coord) + .or_else(|| vars.iter().find(|v| v.name == leaf)) + .ok_or_else(not_found)?; + self.open(file, var) + } + + fn open<'f>(&self, file: &'f clawhdf5::File, var: &Var) -> Result, Error> { + let ds = file.dataset_at(var.address)?; + let attrs = ds.attrs().unwrap_or_default(); + let dims = var.dims.iter().map(|&d| self.dims[d].dimension()).collect(); + Ok(Variable::new( + var.name.clone(), + ds, + dims, + attrs, + var.nc_type, + )) + } +} + +impl Dim { + fn dimension(&self) -> Dimension { + Dimension { + name: self.name.clone(), + size: self.len, + is_unlimited: self.unlimited, + } + } +} + +struct Builder<'f> { + file: &'f clawhdf5::File, + model: Model, + /// netCDF-C's `next_dimid`. + next_id: i64, + types: VarTypes, +} + +impl Builder<'_> { + /// Pass 1 for group `g` and, after its own links, its subgroups. + /// `ancestors` holds the addresses of the groups from the root to `g`, + /// so that a group linked into itself is not read forever. + fn read_group(&mut self, g: usize, ancestors: &mut Vec) -> Result<(), Error> { + let address = self.model.groups[g].address; + let sb = self.file.superblock(); + let mut subgroups = Vec::new(); + for (name, child) in ordered_entries(self.file, address)? { + let Ok(header) = + ObjectHeader::parse_in(self.file.storage(), child, sb.offset_size, sb.length_size) + else { + continue; + }; + match header.object_class() { + Some(ObjectClass::Dataset) => self.read_dataset(g, name, child), + Some(ObjectClass::NamedDatatype) => { + if let Some(dt) = header_datatype(&header) { + self.types.named(&dt); + } + } + _ if is_group(&header) => subgroups.push((name, child)), + _ => {} + } + } + for (name, child) in subgroups { + if ancestors.contains(&child) || self.model.groups.len() >= MAX_GROUPS { + continue; + } + let index = self.model.groups.len(); + self.model.groups.push(Group { + address: child, + children: Vec::new(), + parent: Some(g), + dims: Vec::new(), + vars: Vec::new(), + }); + self.model.groups[g].children.push((name, index)); + ancestors.push(child); + let read = self.read_group(index, ancestors); + ancestors.pop(); + read?; + } + Ok(()) + } + + /// `read_dataset`: a dimension for a dimension scale, a variable for + /// the rest. A dataset that cannot be opened is skipped. + fn read_dataset(&mut self, g: usize, name: String, address: u64) { + let Ok(ds) = self.file.dataset_at(address) else { + return; + }; + let attrs = ds.attrs().unwrap_or_default(); + let Ok(extent) = ds.shape() else { + return; + }; + let max = ds.max_dimensions().ok().flatten(); + let unlimited: Vec = (0..extent.len()) + .map(|i| matches!(&max, Some(m) if m.get(i) == Some(&u64::MAX))) + .collect(); + + let mut own = None; + if dimension::is_dimension_scale(&attrs) && !extent.is_empty() { + // read_scale + let id = match dimension::get_dimid(&attrs) { + Some(id) => { + if id >= self.next_id { + self.next_id = id.saturating_add(1); + } + id + } + None => self.take_id(), + }; + let len = extent[0]; + own = Some(self.add_dim( + g, + Dim { + id, + name: name.clone(), + len, + unlimited: unlimited[0] || len == 0, + scale: Some(address), + }, + )); + if dimension::is_pure_dimension(&attrs) { + return; + } + } + + // read_var: a type netCDF-C cannot represent drops the variable + // (the dimension of a scale stays). + let Some(nc_type) = ds + .raw_datatype() + .ok() + .and_then(|dt| self.types.nc_type(&dt)) + else { + return; + }; + let rank = extent.len(); + let coordinates = coordinates(&attrs).filter(|ids| ids.len() == rank && rank > 0); + let source = match (own, coordinates) { + (Some(dim), _) if rank == 1 => DimSource::OwnScale(dim), + (_, Some(ids)) => DimSource::Coordinates(ids, own), + _ => match dimension::dimension_list(self.file, &attrs) { + Some(scales) + if scales.len() == rank && scales.first().is_some_and(Option::is_some) => + { + DimSource::Scales(scales, own) + } + _ => DimSource::Phony(own), + }, + }; + let (name, non_coord) = match name.strip_prefix(NON_COORD_PREFIX) { + Some(rest) if !rest.is_empty() => (rest.to_string(), true), + _ => (name, false), + }; + self.model.groups[g].vars.push(Var { + name, + non_coord, + address, + nc_type, + extent, + unlimited, + source, + dims: Vec::new(), + }); + } + + fn take_id(&mut self) -> i64 { + let id = self.next_id; + self.next_id = self.next_id.saturating_add(1); + id + } + + fn add_dim(&mut self, g: usize, dim: Dim) -> usize { + let index = self.model.dims.len(); + self.model.dims.push(dim); + self.model.groups[g].dims.push(index); + index + } + + /// Pass 2 (`rec_match_dimscales`): subgroups first, then the group's + /// variables in order. + fn match_dims(&mut self, g: usize) { + let children: Vec = self.model.groups[g] + .children + .iter() + .map(|&(_, c)| c) + .collect(); + for child in children { + self.match_dims(child); + } + for v in 0..self.model.groups[g].vars.len() { + let var = &self.model.groups[g].vars[v]; + let rank = var.extent.len(); + let mut found: Vec> = vec![None; rank]; + match &var.source { + DimSource::OwnScale(dim) => found[0] = Some(*dim), + DimSource::Coordinates(ids, own) => { + for (slot, id) in found.iter_mut().zip(ids) { + *slot = self.dim_by_id(*id); + } + if found[0].is_none() { + found[0] = *own; + } + } + DimSource::Scales(scales, own) => { + for (slot, scale) in found.iter_mut().zip(scales) { + *slot = scale.and_then(|s| self.dim_by_scale(g, s)); + } + if found[0].is_none() { + found[0] = *own; + } + } + DimSource::Phony(own) => { + if rank > 0 { + found[0] = *own; + } + } + } + let mut dims: Vec = Vec::with_capacity(rank); + for (axis, found) in found.into_iter().enumerate() { + let dim = match found { + Some(dim) => dim, + None => { + let var = &self.model.groups[g].vars[v]; + let (len, unlimited) = (var.extent[axis], var.unlimited[axis]); + self.phony_dim(g, len, unlimited, &dims) + } + }; + dims.push(dim); + } + self.model.groups[g].vars[v].dims = dims; + } + } + + /// The dimension with id `id`: the last one created with it + /// (`nc4_find_dim` looks ids up in a file-wide table, where a later + /// dimension of the same id replaces an earlier one). + fn dim_by_id(&self, id: i64) -> Option { + self.model.dims.iter().rposition(|d| d.id == id) + } + + /// The dimension of the scale at `address`, in group `g` or the nearest + /// parent that has it. + fn dim_by_scale(&self, g: usize, address: u64) -> Option { + let mut group = Some(g); + while let Some(i) = group { + let found = self.model.groups[i] + .dims + .iter() + .copied() + .find(|&d| self.model.dims[d].scale == Some(address)); + if found.is_some() { + return found; + } + group = self.model.groups[i].parent; + } + None + } + + /// `create_phony_dims` for one axis: the first dimension of group `g` + /// of length `len` and unlimitedness `unlimited` that no earlier axis + /// of the variable uses (`taken`), else a new `phony_dim_`. + fn phony_dim(&mut self, g: usize, len: u64, unlimited: bool, taken: &[usize]) -> usize { + let dims = &self.model.dims; + let existing = self.model.groups[g].dims.iter().copied().find(|&d| { + let dim = &dims[d]; + dim.len == len + && dim.unlimited == unlimited + && !taken.iter().any(|&t| dims[t].id == dim.id) + }); + if let Some(dim) = existing { + return dim; + } + let id = self.take_id(); + self.add_dim( + g, + Dim { + id, + name: format!("phony_dim_{id}"), + len, + // `nc4_dim_list_add`: a length of 0 is NC_UNLIMITED. + unlimited: unlimited || len == 0, + scale: None, + }, + ) + } + + /// Each unlimited dimension's length: the largest extent along it of + /// the variables using it in its group and the groups below. + fn unlimited_lengths(&mut self) { + let mut owner = vec![0usize; self.model.dims.len()]; + for (g, group) in self.model.groups.iter().enumerate() { + for &d in &group.dims { + owner[d] = g; + } + } + let mut lens = vec![0u64; self.model.dims.len()]; + for (g, group) in self.model.groups.iter().enumerate() { + for var in &group.vars { + for (&d, &e) in var.dims.iter().zip(&var.extent) { + if self.model.dims[d].unlimited && self.is_within(g, owner[d]) { + lens[d] = lens[d].max(e); + } + } + } + } + for (dim, len) in self.model.dims.iter_mut().zip(lens) { + if dim.unlimited { + dim.len = len; + } + } + } + + /// Whether group `g` is `ancestor` or below it. + fn is_within(&self, g: usize, ancestor: usize) -> bool { + let mut group = Some(g); + while let Some(i) = group { + if i == ancestor { + return true; + } + group = self.model.groups[i].parent; + } + false + } +} + +/// A group's entries in netCDF-C's order: creation order when the group +/// tracks it, else the byte order of the names. +fn ordered_entries(file: &clawhdf5::File, address: u64) -> Result, Error> { + let mut entries = file.group_at(address).entries()?; + let order = clawhdf5_format::group_v2::links_in_creation_order_in( + file.storage(), + file.superblock(), + address, + ) + .ok() + .flatten(); + match order { + Some(names) => { + let position: HashMap<&str, usize> = names + .iter() + .enumerate() + .map(|(i, n)| (n.as_str(), i)) + .collect(); + entries.sort_by_key(|(n, _)| position.get(n.as_str()).copied().unwrap_or(usize::MAX)); + } + None => entries.sort_by(|a, b| a.0.as_bytes().cmp(b.0.as_bytes())), + } + Ok(entries) +} + +fn is_group(header: &ObjectHeader) -> bool { + use clawhdf5_format::message_type::MessageType; + header.messages.iter().any(|m| { + matches!( + m.msg_type, + MessageType::LinkInfo | MessageType::Link | MessageType::SymbolTable + ) + }) +} + +/// The datatype a named datatype's header holds. +fn header_datatype(header: &ObjectHeader) -> Option { + use clawhdf5_format::message_type::MessageType; + let msg = header + .messages + .iter() + .find(|m| m.msg_type == MessageType::Datatype)?; + Datatype::parse_in_header(&msg.data, header.version) + .ok() + .map(|(dt, _)| dt) +} + +/// A variable's `_Netcdf4Coordinates`: the id of the dimension of each +/// axis. +fn coordinates(attrs: &HashMap) -> Option> { + match attrs.get("_Netcdf4Coordinates")? { + AttrValue::I64Array(ids) => Some(ids.clone()), + AttrValue::I64(id) => Some(vec![*id]), + AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(), + AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]), + _ => None, + } +} + +/// The user-defined types netCDF-C has read so far, and the netCDF type of +/// a dataset. +/// +/// netCDF-C keeps every type `read_type` meets in a file-wide list and +/// finds one again by `H5Tequal` on the native types — also a type it +/// failed to read: `read_type` adds the type (with its class, for a +/// compound, enum, variable-length or opaque type) before it looks at the +/// members or base type, and does not remove it when one of those fails. +/// So a dataset of a type netCDF-C skipped once becomes a variable the +/// second time (a compound with a reference member, say), and a compound +/// member of a type it skipped (a bit field) is accepted. This replays +/// that, except that a dataset whose skipped type has no class (a bit +/// field, time, array or complex type) stays hidden: netCDF-C lists the +/// second such dataset with an invalid type (class 0). +#[derive(Debug, Default)] +pub(crate) struct VarTypes { + /// Every type read, with its class when it has one netCDF knows. + known: Vec<(Datatype, Option)>, +} + +impl VarTypes { + /// The netCDF type netCDF-C gives a dataset of type `dt` + /// (`get_type_info2`), or `None` when it skips the dataset + /// (`NC_EBADTYPID`): a reference, bit field, time or array type, a + /// compound with a member, or an enum or variable-length type with a + /// base type, that is not a netCDF atomic type or a type read before + /// (see the type docs). + /// + /// Floats other than 4 and 8 bytes (half floats, bfloat16, the 4-, 6- + /// and 8-bit floats, `long double`) are `Float` (up to 4 bytes) and + /// `Double` here; netCDF-C 4.9.3 on libhdf5 1.14.6 labels them + /// `NC_STRING` (their native type matches none of its own). + pub fn nc_type(&mut self, dt: &Datatype) -> Option { + match dt { + Datatype::FixedPoint { size, signed, .. } => int_type(*size, *signed), + Datatype::FloatingPoint { size, .. } => Some(if *size <= 4 { + NcType::Float + } else { + NcType::Double + }), + Datatype::String { size, .. } => Some(if *size > 1 { + NcType::String + } else { + NcType::Char + }), + Datatype::VariableLength { + is_string: true, .. + } => Some(NcType::String), + _ => match self.find(dt) { + Some(class) => class, + None => self.read_type(dt), + }, + } + } + + /// A named datatype of the file (`read_type` on a committed type). + pub fn named(&mut self, dt: &Datatype) { + if self.find(dt).is_none() { + self.read_type(dt); + } + } + + /// The type read before that is `dt`, by its class. + fn find(&self, dt: &Datatype) -> Option> { + self.known + .iter() + .find(|(k, _)| native_eq(k, dt)) + .map(|&(_, class)| class) + } + + /// `read_type`: remember `dt` (not a reference type), then its class if + /// netCDF-C can represent its members or base type. + fn read_type(&mut self, dt: &Datatype) -> Option { + let class = match dt { + Datatype::Reference { .. } => return None, + Datatype::Compound { .. } => Some(NcType::Compound), + Datatype::VariableLength { .. } => Some(NcType::VLen), + Datatype::Opaque { .. } => Some(NcType::Opaque), + Datatype::Enumeration { .. } => Some(NcType::Enum), + _ => None, + }; + self.known.push((dt.clone(), class)); + let parts_ok = match dt { + Datatype::Compound { members, .. } => members.iter().all(|m| match &m.datatype { + Datatype::Array { base_type, .. } => self.is_member_type(base_type), + other => self.is_member_type(other), + }), + Datatype::VariableLength { base_type, .. } + | Datatype::Enumeration { base_type, .. } => self.is_member_type(base_type), + _ => true, + }; + class.filter(|_| parts_ok) + } + + /// `get_netcdf_type`: whether netCDF-C takes `dt` as the type of a + /// compound member or the base of an enum or variable-length type — an + /// atomic type (4- and 8-byte floats only) or a type read before. + fn is_member_type(&self, dt: &Datatype) -> bool { + match dt { + Datatype::FixedPoint { size, signed, .. } => int_type(*size, *signed).is_some(), + Datatype::FloatingPoint { size: 4 | 8, .. } + | Datatype::String { .. } + | Datatype::VariableLength { + is_string: true, .. + } => true, + _ => self.find(dt).is_some(), + } + } +} + +/// The netCDF integer type of an HDF5 integer: the native integer libhdf5 +/// converts it to (`H5Tget_native_type`: the smallest at least as wide). +fn int_type(size: u32, signed: bool) -> Option { + Some(match (size, signed) { + (1, true) => NcType::Byte, + (1, false) => NcType::UByte, + (2, true) => NcType::Short, + (2, false) => NcType::UShort, + (3..=4, true) => NcType::Int, + (3..=4, false) => NcType::UInt, + (5..=8, true) => NcType::Int64, + (5..=8, false) => NcType::UInt64, + _ => return None, + }) +} + +/// Whether two datatypes have the same native type (`H5Tequal` after +/// `H5Tget_native_type`): byte order, padding and compound member offsets +/// do not count. +fn native_eq(a: &Datatype, b: &Datatype) -> bool { + use Datatype as D; + match (a, b) { + ( + D::FixedPoint { + size: s1, + signed: g1, + .. + }, + D::FixedPoint { + size: s2, + signed: g2, + .. + }, + ) => g1 == g2 && int_type(*s1, *g1) == int_type(*s2, *g2), + (D::FloatingPoint { size: s1, .. }, D::FloatingPoint { size: s2, .. }) => s1 == s2, + ( + D::String { + size: s1, + padding: p1, + charset: c1, + }, + D::String { + size: s2, + padding: p2, + charset: c2, + }, + ) => s1 == s2 && p1 == p2 && c1 == c2, + ( + D::VariableLength { + is_string: i1, + base_type: b1, + charset: c1, + .. + }, + D::VariableLength { + is_string: i2, + base_type: b2, + charset: c2, + .. + }, + ) => i1 == i2 && if *i1 { c1 == c2 } else { native_eq(b1, b2) }, + (D::Compound { members: m1, .. }, D::Compound { members: m2, .. }) => { + m1.len() == m2.len() + && m1 + .iter() + .zip(m2) + .all(|(x, y)| x.name == y.name && native_eq(&x.datatype, &y.datatype)) + } + ( + D::Enumeration { + base_type: b1, + members: m1, + .. + }, + D::Enumeration { + base_type: b2, + members: m2, + .. + }, + ) => { + native_eq(b1, b2) + && m1.len() == m2.len() + && m1.iter().zip(m2).all(|(x, y)| { + x.name == y.name && enum_value(&x.value, b1) == enum_value(&y.value, b2) + }) + } + (D::Opaque { size: s1, tag: t1 }, D::Opaque { size: s2, tag: t2 }) => s1 == s2 && t1 == t2, + ( + D::Array { + base_type: b1, + dimensions: d1, + }, + D::Array { + base_type: b2, + dimensions: d2, + }, + ) => d1 == d2 && native_eq(b1, b2), + ( + D::Reference { + size: s1, + ref_type: r1, + }, + D::Reference { + size: s2, + ref_type: r2, + }, + ) => s1 == s2 && r1 == r2, + (D::BitField { size: s1, .. }, D::BitField { size: s2, .. }) + | (D::Time { size: s1, .. }, D::Time { size: s2, .. }) => s1 == s2, + ( + D::Complex { + size: s1, + base_type: b1, + }, + D::Complex { + size: s2, + base_type: b2, + }, + ) => s1 == s2 && native_eq(b1, b2), + _ => false, + } +} + +/// An enum member's value, in the order of its base type's bytes. +fn enum_value(value: &[u8], base: &Datatype) -> Vec { + let big = matches!( + base, + Datatype::FixedPoint { + byte_order: DatatypeByteOrder::BigEndian, + .. + } + ); + let mut v = value.to_vec(); + if big { + v.reverse(); + } + v +} diff --git a/crates/clawhdf5-netcdf4/src/scope.rs b/crates/clawhdf5-netcdf4/src/scope.rs deleted file mode 100644 index d74213b..0000000 --- a/crates/clawhdf5-netcdf4/src/scope.rs +++ /dev/null @@ -1,231 +0,0 @@ -//! A group's variables and the dimensions they are defined on. -//! -//! netCDF-C (`libhdf5/hdf5open.c`) gives a variable its dimensions from the -//! file, never by size: the dimension ids in its `_Netcdf4Coordinates` -//! attribute (each dimension scale's `_Netcdf4Dimid`), else the dimension -//! scales its `DIMENSION_LIST` attribute references, looked up in the -//! variable's group and then each parent group up to the root. Only an axis -//! with neither (a file not written by a netCDF library) gets a dimension -//! by size. Dimension scales that are only dimensions are not variables, and -//! a variable stored as `_nc4_non_coord_` (a variable sharing a -//! dimension's name without being its coordinate variable) is ``. - -use std::collections::{HashMap, HashSet}; - -use clawhdf5::AttrValue; - -use crate::dimension::{self, Dimension, GroupDims}; -use crate::error::Error; -use crate::variable::Variable; - -/// The prefix netCDF-C gives the dataset of a variable that has a -/// dimension's name but is not that dimension's coordinate variable (the -/// dimension's scale holds the name). -const NON_COORD_PREFIX: &str = "_nc4_non_coord_"; - -/// A group, with the dimensions visible from it. -pub(crate) struct Scope<'f> { - file: &'f clawhdf5::File, - group: clawhdf5::Group<'f>, - /// This group's dimensions, then its parent's, and so on to the root's. - levels: Vec, -} - -impl<'f> Scope<'f> { - /// The group at `path` (`/`-separated from the root; `""` or `"/"` is - /// the root). - pub fn new(file: &'f clawhdf5::File, path: &str) -> Result { - let parts: Vec<&str> = path.split('/').filter(|p| !p.is_empty()).collect(); - let mut levels = Vec::with_capacity(parts.len() + 1); - for n in (0..=parts.len()).rev() { - let group = file.group(&parts[..n].join("/"))?; - levels.push(dimension::group_dims(file, &group)?); - } - let group = file.group(&parts.join("/"))?; - Ok(Self { - file, - group, - levels, - }) - } - - /// The group's datasets, as `(dataset name, object header address)` in - /// listing order. - fn datasets(&self) -> Result, Error> { - let datasets: HashSet = self.group.datasets()?.into_iter().collect(); - Ok(self - .group - .entries()? - .into_iter() - .filter(|(name, _)| datasets.contains(name)) - .collect()) - } - - /// The group's variables: every dataset but the dimension scales that - /// are only dimensions. - pub fn variables(&self) -> Result>, Error> { - let mut variables = Vec::new(); - for (ds_name, address) in self.datasets()? { - let ds = self.file.dataset_at(address)?; - let attrs = ds.attrs()?; - if dimension::is_pure_dimension(&attrs) { - continue; - } - variables.push(self.variable_from(nc_name(&ds_name), address, ds, attrs)?); - } - Ok(variables) - } - - /// The names of the group's variables. - pub fn variable_names(&self) -> Result, Error> { - let mut names = Vec::new(); - for (ds_name, address) in self.datasets()? { - let attrs = self.file.dataset_at(address)?.attrs()?; - if !dimension::is_pure_dimension(&attrs) { - names.push(nc_name(&ds_name)); - } - } - Ok(names) - } - - /// The variable called `name`: the dataset `_nc4_non_coord_` if - /// there is one, else the dataset `` unless it is only a - /// dimension. - pub fn variable(&self, name: &str) -> Result, Error> { - let not_found = || Error::VariableNotFound(name.to_string()); - let datasets = self.datasets()?; - let prefixed = format!("{NON_COORD_PREFIX}{name}"); - let address = datasets - .iter() - .find(|(n, _)| *n == prefixed) - .or_else(|| datasets.iter().find(|(n, _)| n == name)) - .map(|&(_, address)| address) - .ok_or_else(not_found)?; - let ds = self.file.dataset_at(address)?; - let attrs = ds.attrs()?; - if dimension::is_pure_dimension(&attrs) { - return Err(not_found()); - } - self.variable_from(nc_name(name), address, ds, attrs) - } - - fn variable_from( - &self, - name: String, - address: u64, - ds: clawhdf5::Dataset<'f>, - attrs: HashMap, - ) -> Result, Error> { - let shape = ds.shape()?; - let dims = self.variable_dims(address, &attrs, &shape); - Ok(Variable::new(name, ds, dims, attrs)) - } - - /// The first dimension, searching this group and then its ancestors, - /// that `find` picks. - fn find<'a>( - &'a self, - find: impl Fn(&'a GroupDims) -> Option<&'a Dimension>, - ) -> Option { - self.levels.iter().find_map(find).cloned() - } - - /// The dimensions of the dataset at `address`, one per axis of `shape`, - /// as netCDF-C resolves them (see the module docs). - fn variable_dims( - &self, - address: u64, - attrs: &HashMap, - shape: &[u64], - ) -> Vec { - let rank = shape.len(); - let mut dims: Vec> = vec![None; rank]; - if rank == 0 { - return Vec::new(); - } - // A coordinate variable is the scale of its (first) dimension. - dims[0] = self.levels[0].by_address(address).cloned(); - - if let Some(ids) = coordinates(attrs).filter(|ids| ids.len() == rank) { - for (slot, id) in dims.iter_mut().zip(ids) { - if slot.is_none() { - *slot = self.find(|level| level.by_dimid(id)); - } - } - } - if dims.iter().any(Option::is_none) - && let Some(scales) = - dimension::dimension_list(self.file, attrs).filter(|s| s.len() == rank) - { - for (slot, scale) in dims.iter_mut().zip(scales) { - if slot.is_none() - && let Some(scale) = scale - { - *slot = self.find(|level| level.by_address(scale)); - } - } - } - - // Neither: the first dimension of this group of the same size not - // already taken by another such axis, else an anonymous one. - let own = &self.levels[0].dims; - let mut used = vec![false; own.len()]; - dims.into_iter() - .zip(shape) - .map(|(dim, &size)| { - dim.unwrap_or_else(|| { - match own - .iter() - .enumerate() - .find(|&(i, d)| !used[i] && d.size == size) - { - Some((i, d)) => { - used[i] = true; - d.clone() - } - None => Dimension { - name: format!("dim_{size}"), - size, - is_unlimited: false, - }, - } - }) - }) - .collect() - } -} - -/// The variable `name` of the group at `group_path`; `name` may itself be -/// a path (`"sub/var"`), relative to that group. -pub(crate) fn variable_at<'f>( - file: &'f clawhdf5::File, - group_path: &str, - name: &str, -) -> Result, Error> { - match name.trim_start_matches('/').rsplit_once('/') { - Some((dir, leaf)) => Scope::new(file, &format!("{group_path}/{dir}")) - .map_err(|_| Error::VariableNotFound(name.to_string()))? - .variable(leaf), - None => Scope::new(file, group_path)?.variable(name.trim_start_matches('/')), - } -} - -/// The netCDF name of the dataset `ds_name`. -fn nc_name(ds_name: &str) -> String { - ds_name - .strip_prefix(NON_COORD_PREFIX) - .unwrap_or(ds_name) - .to_string() -} - -/// A variable's `_Netcdf4Coordinates`: the `_Netcdf4Dimid` of the dimension -/// of each axis. -fn coordinates(attrs: &HashMap) -> Option> { - match attrs.get("_Netcdf4Coordinates")? { - AttrValue::I64Array(ids) => Some(ids.clone()), - AttrValue::I64(id) => Some(vec![*id]), - AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(), - AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]), - _ => None, - } -} diff --git a/crates/clawhdf5-netcdf4/src/types.rs b/crates/clawhdf5-netcdf4/src/types.rs index 70945c8..da2ad09 100644 --- a/crates/clawhdf5-netcdf4/src/types.rs +++ b/crates/clawhdf5-netcdf4/src/types.rs @@ -4,8 +4,16 @@ use clawhdf5::DType; -/// NetCDF-4 data types corresponding to the standard NetCDF type system. +/// NetCDF-4 data types corresponding to the standard NetCDF type system: +/// the atomic types, and the class of a user-defined type. +/// +/// A dataset whose type netCDF-C cannot represent — a reference, bit +/// field, time or array type; a compound with such a member, a +/// half-precision float member, or a member of a user-defined type not +/// read before it; an enum or variable-length type over such a base — is +/// not a variable, as in netCDF-C. #[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] +#[non_exhaustive] pub enum NcType { /// NC_BYTE: signed 8-bit integer Byte, @@ -29,8 +37,18 @@ pub enum NcType { Double, /// NC_STRING: variable-length string String, - /// NC_CHAR: fixed-length string / character data + /// NC_CHAR: a fixed-length string of one byte (longer ones are + /// `String`, as netCDF-C reads them) Char, + /// NC_ENUM: an enumeration (user-defined type) + Enum, + /// NC_COMPOUND: a compound (user-defined type) + Compound, + /// NC_VLEN: a variable-length sequence (user-defined type) + VLen, + /// NC_OPAQUE: an opaque type (user-defined type; netCDF4-python skips + /// such variables) + Opaque, } impl std::fmt::Display for NcType { @@ -48,11 +66,16 @@ impl std::fmt::Display for NcType { NcType::Double => write!(f, "NC_DOUBLE"), NcType::String => write!(f, "NC_STRING"), NcType::Char => write!(f, "NC_CHAR"), + NcType::Enum => write!(f, "NC_ENUM"), + NcType::Compound => write!(f, "NC_COMPOUND"), + NcType::VLen => write!(f, "NC_VLEN"), + NcType::Opaque => write!(f, "NC_OPAQUE"), } } } -/// Map a clawhdf5 DType to a NetCDF type. +/// Map a clawhdf5 DType to a NetCDF type. A variable's own type, as +/// netCDF-C reads it, is [`Variable::nc_type`](crate::Variable::nc_type). pub fn dtype_to_nctype(dtype: &DType) -> NcType { match dtype { DType::I8 => NcType::Byte, @@ -66,6 +89,8 @@ pub fn dtype_to_nctype(dtype: &DType) -> NcType { DType::F32 => NcType::Float, DType::F64 => NcType::Double, DType::String | DType::VariableLengthString => NcType::String, + DType::Compound(_) => NcType::Compound, + DType::Enum(_) => NcType::Enum, _ => NcType::Char, // fallback for other types } } diff --git a/crates/clawhdf5-netcdf4/src/variable.rs b/crates/clawhdf5-netcdf4/src/variable.rs index 51d529e..da2e286 100644 --- a/crates/clawhdf5-netcdf4/src/variable.rs +++ b/crates/clawhdf5-netcdf4/src/variable.rs @@ -16,7 +16,7 @@ use clawhdf5::AttrValue; use crate::cf::{self, CfAttributes, FillValue}; use crate::dimension::Dimension; use crate::error::Error; -use crate::types::{NcType, dtype_to_nctype}; +use crate::types::NcType; /// A NetCDF-4 variable backed by an HDF5 dataset. pub struct Variable<'f> { @@ -28,6 +28,8 @@ pub struct Variable<'f> { dims: Vec, /// The dataset's attributes. attrs: HashMap, + /// Its netCDF type. + nc_type: NcType, } impl<'f> Variable<'f> { @@ -37,12 +39,14 @@ impl<'f> Variable<'f> { dataset: clawhdf5::Dataset<'f>, dims: Vec, attrs: HashMap, + nc_type: NcType, ) -> Self { Self { name, dataset, dims, attrs, + nc_type, } } @@ -51,11 +55,12 @@ impl<'f> Variable<'f> { &self.name } - /// The dimensions of this variable, one per axis: the ones the file - /// gives it (`_Netcdf4Coordinates`, else `DIMENSION_LIST`), found in its - /// group or a parent group. An axis the file gives no dimension (a file - /// not written by a netCDF library) gets the first dimension of the - /// variable's group of the same size, else an anonymous `dim_`. + /// The dimensions of this variable, one per axis, as netCDF-C gives + /// them: the ones the file names (`_Netcdf4Coordinates`, else the + /// scales `DIMENSION_LIST` attaches), found in its group or a parent + /// group; for a dataset without them (a file not written by a netCDF + /// library), netCDF-C's phony dimensions `phony_dim_`, shared by + /// the variables of a group by length. pub fn dimensions(&self) -> &[Dimension] { &self.dims } @@ -75,10 +80,11 @@ impl<'f> Variable<'f> { Ok(self.dataset.shape()?) } - /// The NetCDF data type of this variable. + /// The NetCDF data type of this variable, as netCDF-C reads it (a + /// fixed-length string longer than one byte is `String`, one byte + /// long `Char`). pub fn nc_type(&self) -> Result { - let dtype = self.dataset.dtype()?; - Ok(dtype_to_nctype(&dtype)) + Ok(self.nc_type) } /// Read all attributes as a HashMap. @@ -297,6 +303,9 @@ fn default_fill(nc_type: NcType) -> FillValue { NcType::Double => FillValue::Float(9.969_209_968_386_869e36), NcType::String => FillValue::String(String::new()), NcType::Char => FillValue::Int(0), + // User-defined types have no default fill value in netCDF-C; their + // values are not read through these methods. + _ => FillValue::Int(0), } } diff --git a/crates/clawhdf5-netcdf4/tests/common/mod.rs b/crates/clawhdf5-netcdf4/tests/common/mod.rs new file mode 100644 index 0000000..491f3eb --- /dev/null +++ b/crates/clawhdf5-netcdf4/tests/common/mod.rs @@ -0,0 +1,281 @@ +//! clawhdf5-netcdf4 compared with netCDF-C itself: `tests/netcdf_c_view.py` +//! prints what netCDF-C (the libnetcdf netCDF4-python bundles) reports for +//! a file, and [`differences`] lists how clawhdf5-netcdf4's reading differs. +#![allow(dead_code)] + +use std::collections::{BTreeMap, BTreeSet}; +use std::path::{Path, PathBuf}; +use std::process::Command; +use std::sync::atomic::{AtomicUsize, Ordering}; + +use clawhdf5_netcdf4::{NcType, NetCDF4File, NetCDF4Group, Variable}; + +/// The Python with netCDF4-python (`CLAWHDF5_PYTHON`, else `python3`). +pub fn python() -> String { + std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string()) +} + +/// netCDF-C's view of one file. +#[derive(Default)] +pub struct View { + /// `G`, `D` and `V` lines, in order. + pub meta: Vec, + /// `(group, variable)` → values, from the `X` lines. + pub values: BTreeMap<(String, String), Vec>, +} + +/// netCDF-C's view of each file (`None`: netCDF-C cannot open it), from +/// `tests/netcdf_c_view.py`. +pub fn netcdf_c_views(files: &[PathBuf]) -> BTreeMap> { + let dir = tempfile::tempdir().unwrap(); + let list = dir.path().join("files.txt"); + let mut paths = String::new(); + for f in files { + paths.push_str(&f.to_string_lossy()); + paths.push('\n'); + } + std::fs::write(&list, paths).unwrap(); + let script = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/netcdf_c_view.py"); + let out = Command::new(python()) + .arg(&script) + .arg("--list") + .arg(&list) + .output() + .expect("failed to run python"); + assert!( + out.status.success(), + "netcdf_c_view.py failed: {}", + String::from_utf8_lossy(&out.stderr) + ); + parse_views(&String::from_utf8_lossy(&out.stdout)) +} + +/// The views in the output of `tests/netcdf_c_view.py`. +pub fn parse_views(text: &str) -> BTreeMap> { + let mut views = BTreeMap::new(); + let mut current: Option<(PathBuf, Option)> = None; + for line in text.lines() { + let fields: Vec<&str> = line.split('\t').collect(); + match fields[0] { + "FILE" => current = Some((PathBuf::from(fields[1]), Some(View::default()))), + "END" => { + let (path, view) = current.take().expect("END without FILE"); + views.insert(path, view); + } + "ERROR" => current.as_mut().expect("ERROR without FILE").1 = None, + "X" => { + let view = current.as_mut().and_then(|c| c.1.as_mut()).unwrap(); + let vals = if fields[3] == "-" { + Vec::new() + } else { + fields[3] + .split(' ') + .map(|v| v.parse().expect("value")) + .collect() + }; + view.values + .insert((fields[1].to_string(), fields[2].to_string()), vals); + } + _ => { + let view = current.as_mut().and_then(|c| c.1.as_mut()).unwrap(); + view.meta.push(line.to_string()); + } + } + } + views +} + +/// Variables whose type is labelled as netCDF-C labels it rather than as +/// clawhdf5-netcdf4 does (see [`nc_type_label`]). +pub static RELABELLED: AtomicUsize = AtomicUsize::new(0); +/// Values not compared because this build of the HDF5 reader lacks a +/// filter (a cargo feature). +pub static FILTER_SKIPPED: AtomicUsize = AtomicUsize::new(0); + +/// The same `G`/`D`/`V` lines from clawhdf5-netcdf4. +pub fn our_meta(file: &NetCDF4File) -> Result, String> { + let e = |e: clawhdf5_netcdf4::Error| e.to_string(); + let mut out = Vec::new(); + out.push("G\t/".to_string()); + push_group( + &mut out, + file.hdf5_file(), + "/", + file.dimensions().map_err(e)?, + file.variables().map_err(e)?, + )?; + for name in file.group_names().map_err(e)? { + walk( + &mut out, + file.hdf5_file(), + &format!("/{name}"), + &file.group(&name).map_err(e)?, + )?; + } + Ok(out) +} + +fn walk( + out: &mut Vec, + hdf5: &clawhdf5::File, + path: &str, + group: &NetCDF4Group<'_>, +) -> Result<(), String> { + let e = |e: clawhdf5_netcdf4::Error| e.to_string(); + out.push(format!("G\t{path}")); + push_group( + out, + hdf5, + path, + group.dimensions().map_err(e)?, + group.variables().map_err(e)?, + )?; + for name in group.group_names().map_err(e)? { + walk( + out, + hdf5, + &format!("{path}/{name}"), + &group.group(&name).map_err(e)?, + )?; + } + Ok(()) +} + +/// The type of variable `name` of the group at `path` as the `V` line +/// shows it: its `NcType`, except for the one deliberate difference from +/// netCDF-C — a float of other than 4 or 8 bytes (half, bfloat16, the 4-, +/// 6- and 8-bit floats, `long double`), `NC_FLOAT`/`NC_DOUBLE` here, is +/// labelled `NC_STRING` by netCDF-C 4.9.3 on libhdf5 1.14.6; such a +/// variable is shown as netCDF-C shows it, and counted. +fn nc_type_label(hdf5: &clawhdf5::File, path: &str, var: &Variable<'_>) -> Result { + let nc_type = var.nc_type().map_err(|e| e.to_string())?; + if matches!(nc_type, NcType::Float | NcType::Double) { + let dataset = format!("{}/{}", path.trim_end_matches('/'), var.name()); + if let Ok(ds) = hdf5.dataset(&dataset) + && let Ok(clawhdf5_format::datatype::Datatype::FloatingPoint { size, .. }) = + ds.raw_datatype() + && size != 4 + && size != 8 + { + RELABELLED.fetch_add(1, Ordering::Relaxed); + return Ok("NC_STRING".to_string()); + } + } + Ok(nc_type.to_string()) +} + +fn push_group( + out: &mut Vec, + hdf5: &clawhdf5::File, + path: &str, + dims: Vec, + vars: Vec>, +) -> Result<(), String> { + for d in dims { + out.push(format!( + "D\t{path}\t{}\t{}\t{}", + d.name, + d.size, + u8::from(d.is_unlimited) + )); + } + for v in vars { + let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect(); + let shape: Vec = v + .shape() + .map_err(|e| e.to_string())? + .iter() + .map(u64::to_string) + .collect(); + let or_dash = |s: String| if s.is_empty() { "-".to_string() } else { s }; + out.push(format!( + "V\t{path}\t{}\t{}\t{}\t{}", + v.name(), + nc_type_label(hdf5, path, &v)?, + or_dash(dims.join(",")), + or_dash(shape.join(",")) + )); + } + Ok(()) +} + +/// The values of variable `name` of group `group`. +fn our_values(file: &NetCDF4File, group: &str, name: &str) -> Result, String> { + let var = if group == "/" { + file.variable(name) + } else { + file.variable(&format!("{}/{name}", group.trim_start_matches('/'))) + } + .map_err(|e| e.to_string())?; + match var.nc_type().map_err(|e| e.to_string())? { + NcType::String | NcType::Char => Err("not numeric".into()), + _ => var.read_raw_f64().map_err(|e| e.to_string()), + } +} + +/// How clawhdf5-netcdf4's reading of `path` differs from `want`; empty +/// when it does not. +pub fn differences(path: &Path, want: &View) -> Vec { + let file = match NetCDF4File::open(path) { + Ok(f) => f, + Err(e) => return vec![format!("netCDF-C opens it, clawhdf5-netcdf4 does not: {e}")], + }; + let meta = match std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| our_meta(&file))) { + Ok(Ok(meta)) => meta, + Ok(Err(e)) => return vec![format!("error: {e}")], + Err(_) => return vec!["panic".to_string()], + }; + let mut diffs = Vec::new(); + if meta != want.meta { + let ours: BTreeSet<&String> = meta.iter().collect(); + let theirs: BTreeSet<&String> = want.meta.iter().collect(); + for line in want.meta.iter().filter(|l| !ours.contains(l)) { + diffs.push(format!("netCDF-C: {line}")); + } + for line in meta.iter().filter(|l| !theirs.contains(l)) { + diffs.push(format!("ours: {line}")); + } + if diffs.is_empty() { + diffs.push("same lines, different order".to_string()); + } + } + for ((group, name), want) in &want.values { + let got = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { + our_values(&file, group, name) + })) + .unwrap_or_else(|_| Err("panic".to_string())); + match got { + Ok(got) => { + let same = got.len() == want.len() + && got + .iter() + .zip(want) + .all(|(a, b)| a.to_bits() == b.to_bits() || (a.is_nan() && b.is_nan())); + if !same { + diffs.push(format!("values of {group} {name} differ")); + } + } + Err(e) if e.contains("this build lacks the") => { + FILTER_SKIPPED.fetch_add(1, Ordering::Relaxed); + } + Err(e) => diffs.push(format!("values of {group} {name}: {e}")), + } + } + diffs +} + +/// clawhdf5-netcdf4 reads the file at `path` as netCDF-C does: the same +/// groups, dimensions, variables and values (see [`differences`]). +pub fn assert_matches_netcdf_c(path: &Path) { + let views = netcdf_c_views(&[path.to_path_buf()]); + let Some(Some(want)) = views.get(path) else { + panic!("netCDF-C cannot open {}", path.display()); + }; + let diffs = differences(path, want); + assert!( + diffs.is_empty(), + "{} differs from netCDF-C:\n{}", + path.display(), + diffs.join("\n") + ); +} diff --git a/crates/clawhdf5-netcdf4/tests/corpus_known_differences.txt b/crates/clawhdf5-netcdf4/tests/corpus_known_differences.txt new file mode 100644 index 0000000..7a413bb --- /dev/null +++ b/crates/clawhdf5-netcdf4/tests/corpus_known_differences.txt @@ -0,0 +1,26 @@ +# Files of the conformance corpus (conformance/.cache/corpus) that netCDF-C +# opens and clawhdf5-netcdf4 reads differently, with the reason. +# +# Read by tests/corpus_vs_netcdf_c.rs; see docs/known-issues.md +# ("NetCDF-4: differences from netCDF-C"). +# +# External links: libhdf5 follows them into the other file (present in the +# corpus); clawhdf5 does not follow external links (known-issues: "External +# links and external raw data are not followed"), so the linked groups and +# datasets are missing and, in files without dimension scales, the phony +# dimension numbers after them shift. +hdf5/test/testfiles/be_extlink1.h5 external link not followed +hdf5/test/testfiles/le_extlink1.h5 external link not followed +hdf5/tools/test/testfiles/h5diff_ext2softlink_src.h5 external link not followed +hdf5/tools/test/testfiles/h5diff_grp_recurse_ext2-1.h5 external link not followed +hdf5/tools/test/testfiles/h5diff_grp_recurse_ext2-2.h5 external link not followed +# +# Values the HDF5 reader refuses and libhdf5 1.14.6 (the netCDF4-python +# wheel's) returns; metadata matches. The first three are the scale-offset +# and N-Bit ref-bug / our-error files of CONFORMANCE.md (libhdf5 reads past +# the stored data); bad_nbit_decompress.h5 is not in the conformance run and +# is not investigated yet. +cve_hdf5/cvefiles/cve-2025-2308.h5 values: scale-offset chunk shorter than its values (libhdf5 over-read) +cve_hdf5/cvefiles/cve-2025-44904.h5 values: unfiltered chunks shorter than a chunk (libhdf5 over-read) +hdf5/test/testfiles/bad_nbit_parms_walk.h5 values: N-Bit parameters too short (conformance our-error/ref-bug) +hdf5/test/testfiles/bad_nbit_decompress.h5 values: N-Bit chunk refused ("element count exceeds chunk size"); not investigated diff --git a/crates/clawhdf5-netcdf4/tests/corpus_vs_netcdf_c.rs b/crates/clawhdf5-netcdf4/tests/corpus_vs_netcdf_c.rs new file mode 100644 index 0000000..d0db1a9 --- /dev/null +++ b/crates/clawhdf5-netcdf4/tests/corpus_vs_netcdf_c.rs @@ -0,0 +1,173 @@ +//! Every file of a corpus that netCDF-C opens, read by clawhdf5-netcdf4 and +//! compared with what netCDF-C reports: groups (order), dimensions (names, +//! lengths, unlimited, order), variables (names, order, types, dimensions, +//! shapes) and the values of numeric variables of at most 5000 elements. +//! +//! Gated: set `CLAWHDF5_NETCDF_CORPUS` to a directory (relative to the +//! workspace root, or absolute) — the conformance corpus +//! (`conformance/.cache/corpus`) or any tree of HDF5/netCDF-4 files — and +//! `CLAWHDF5_PYTHON` to a Python with netCDF4-python (whose bundled +//! libnetcdf `tests/netcdf_c_view.py` calls). Without the variable the test +//! does nothing. netCDF-C reads each file in its own process (4 at a time, +//! `CLAWHDF5_NETCDF_JOBS`), under a 60 s timeout and a 4 GiB address-space +//! limit: 40 s to 3 minutes for the conformance corpus on tank. +//! `CLAWHDF5_NETCDF_VIEW` may name the saved output of an earlier +//! `netcdf_c_view.py --list ` run over the same (absolute) +//! paths, to skip that. +//! +//! A file whose differences are explained is listed, with the reason, in +//! `tests/corpus_known_differences.txt` (and in `docs/known-issues.md`); a +//! listed file that matches fails the test too, so the list stays true. +//! +//! ```sh +//! CLAWHDF5_NETCDF_CORPUS=conformance/.cache/corpus CLAWHDF5_PYTHON=$PWD/.venv/bin/python \ +//! cargo test -p clawhdf5-netcdf4 --test corpus_vs_netcdf_c -- --nocapture +//! ``` + +mod common; + +use std::collections::BTreeMap; +use std::path::{Path, PathBuf}; +use std::sync::atomic::Ordering; + +use common::{FILTER_SKIPPED, RELABELLED, differences, netcdf_c_views, parse_views}; + +/// The file extensions `conformance/list_files.py` sweeps. +const EXTS: [&str; 7] = ["h5", "hdf5", "he5", "nc", "nc4", "hdf", "h5f"]; + +/// The files of the corpus: those with an HDF5/netCDF-4 extension, except +/// netCDF classic files (magic `CDF`), following directory symlinks, in +/// byte order of their paths. +fn corpus_files(root: &Path) -> Vec { + fn walk(dir: &Path, out: &mut Vec, depth: usize) { + let Ok(entries) = std::fs::read_dir(dir) else { + return; + }; + for entry in entries.flatten() { + let path = entry.path(); + if path.file_name().is_some_and(|n| n == ".git") { + continue; + } + let Ok(meta) = std::fs::metadata(&path) else { + continue; + }; + if meta.is_dir() && depth < 32 { + walk(&path, out, depth + 1); + } else if meta.is_file() + && path + .extension() + .and_then(|e| e.to_str()) + .is_some_and(|e| EXTS.contains(&e.to_ascii_lowercase().as_str())) + { + let mut magic = [0u8; 3]; + let classic = std::fs::File::open(&path) + .and_then(|mut f| std::io::Read::read_exact(&mut f, &mut magic)) + .is_ok() + && &magic == b"CDF"; + if !classic { + out.push(path); + } + } + } + } + let mut out = Vec::new(); + walk(root, &mut out, 0); + out.sort_by(|a, b| { + a.as_os_str() + .as_encoded_bytes() + .cmp(b.as_os_str().as_encoded_bytes()) + }); + out +} + +/// The files `tests/corpus_known_differences.txt` explains: path relative +/// to the corpus root → reason. +fn known_differences() -> BTreeMap { + let list = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/corpus_known_differences.txt"); + std::fs::read_to_string(list) + .expect("tests/corpus_known_differences.txt") + .lines() + .filter(|l| !l.trim().is_empty() && !l.starts_with('#')) + .map(|l| { + let (path, reason) = l.split_once('\t').unwrap_or((l, "")); + (path.trim().to_string(), reason.trim().to_string()) + }) + .collect() +} + +#[test] +fn corpus_matches_netcdf_c() { + let Ok(root) = std::env::var("CLAWHDF5_NETCDF_CORPUS") else { + eprintln!("SKIP: set CLAWHDF5_NETCDF_CORPUS to a corpus directory"); + return; + }; + // A relative path is from the workspace root (cargo runs the test in + // the crate's directory). + let root = Path::new(env!("CARGO_MANIFEST_DIR")) + .join("../..") + .join(root); + let root = std::fs::canonicalize(&root).unwrap_or(root); + let files = corpus_files(&root); + assert!(!files.is_empty(), "no files under {}", root.display()); + // `CLAWHDF5_NETCDF_VIEW`: the output of an earlier run of + // `tests/netcdf_c_view.py` over the same files, which takes a few + // minutes over the conformance corpus. + let views = match std::env::var("CLAWHDF5_NETCDF_VIEW") { + Ok(saved) => parse_views(&std::fs::read_to_string(saved).expect("CLAWHDF5_NETCDF_VIEW")), + Err(_) => netcdf_c_views(&files), + }; + let known = known_differences(); + + let (mut opened, mut matched) = (0, 0); + let mut explained = Vec::new(); + let mut unexplained = Vec::new(); + let mut stale = Vec::new(); + for path in &files { + let Some(Some(want)) = views.get(path) else { + continue; + }; + opened += 1; + let rel = path + .strip_prefix(&root) + .unwrap_or(path) + .to_string_lossy() + .into_owned(); + let diffs = differences(path, want); + match (diffs.is_empty(), known.get(&rel)) { + (true, None) => matched += 1, + (true, Some(_)) => stale.push(rel), + (false, Some(reason)) => explained.push((rel, reason.clone(), diffs)), + (false, None) => unexplained.push((rel, diffs)), + } + } + eprintln!( + "{} files; netCDF-C opens {opened}; {matched} match; {} differ as explained; \ + {} differ unexplained; {} listed but match; {} variables relabelled \ + (floats of other than 4 or 8 bytes); {} variables' values not compared \ + (filter not in this build)", + files.len(), + explained.len(), + unexplained.len(), + stale.len(), + RELABELLED.load(Ordering::Relaxed), + FILTER_SKIPPED.load(Ordering::Relaxed), + ); + for (rel, reason, diffs) in &explained { + eprintln!("explained: {rel} ({reason}): {} differences", diffs.len()); + } + for (rel, diffs) in &unexplained { + eprintln!("DIFFERS: {rel}"); + for d in diffs.iter().take(20) { + eprintln!(" {d}"); + } + } + assert!( + unexplained.is_empty(), + "{} files differ from netCDF-C unexplained", + unexplained.len() + ); + assert!( + stale.is_empty(), + "listed in corpus_known_differences.txt but match: {stale:?}" + ); +} diff --git a/crates/clawhdf5-netcdf4/tests/interop_tests.rs b/crates/clawhdf5-netcdf4/tests/interop_tests.rs index 2beffca..5656cb1 100644 --- a/crates/clawhdf5-netcdf4/tests/interop_tests.rs +++ b/crates/clawhdf5-netcdf4/tests/interop_tests.rs @@ -2,6 +2,8 @@ //! //! Tests are skipped if python3 or netCDF4/xarray Python packages are not available. +mod common; + use std::process::Command; use clawhdf5_netcdf4::{AttrValue, NcType, NetCDF4File}; @@ -842,3 +844,267 @@ ds.to_netcdf({path:?}, engine={engine:?}, unlimited_dims=["time"]) assert_same_view(&path); } } + +// =========================================================================== +// Compared with netCDF-C itself (tests/netcdf_c_view.py): groups, +// dimensions, variables, types and values, in netCDF-C's order +// =========================================================================== + +/// The dimensions of the variable `name` (a path) of `file`. +fn dim_names(file: &NetCDF4File, name: &str) -> Vec { + file.variable(name) + .unwrap() + .dimensions() + .iter() + .map(|d| d.name.clone()) + .collect() +} + +/// An h5py file without dimension scales: netCDF-C's phony dimensions, +/// numbered file-wide (subgroups before their parent's variables, each +/// group's datasets in name order, as h5py does not track creation order), +/// shared by length within a group — but not between two axes of one +/// variable, nor between a fixed and an unlimited axis — with a length of 0 +/// always unlimited and never shared with a fixed axis, and the real +/// dimension of a scale taken by length too. +#[test] +fn phony_dimensions_match_netcdf_c() { + skip_if_no_netcdf4!(); + let dir = tempfile::tempdir().unwrap(); + let path = dir.path().join("phony.h5"); + run_python(&format!( + r#" +import h5py +import numpy as np +with h5py.File({path:?}, "w") as f: + f["zz"] = np.arange(12.0).reshape(2, 2, 3) + f["aa"] = np.arange(6, dtype="i4").reshape(3, 2) + f.create_dataset("un", data=np.ones((2, 3), "f4"), maxshape=(None, 3)) + f.create_dataset("un2", data=np.ones(2, "f4"), maxshape=(None,)) + f["zero"] = np.zeros((0,)) + f.create_dataset("zero_un", (0,), "f4", maxshape=(None,)) + f["zero2"] = np.zeros((0, 2)) + f["s"] = 1.5 + g = f.create_group("g") + g["x"] = np.arange(5, dtype="i2") + g["y"] = np.arange(2, dtype="u1") + g.create_group("h")["q"] = np.arange(7.0) + f.create_group("b")["w"] = np.arange(18, dtype="i8").reshape(2, 9) + f["sc"] = np.arange(6.0) + f["sc"].make_scale("sc") + f["second_only"] = np.zeros((2, 6)) + f["second_only"].dims[1].attach_scale(f["sc"]) +"#, + path = path.display().to_string() + )); + common::assert_matches_netcdf_c(&path); + + let file = NetCDF4File::open(&path).unwrap(); + // `sc` is dimension 0; /b gets 1 and 2, /g/h 3, /g 4 and 5, / from 6. + assert_eq!(dim_names(&file, "b/w"), ["phony_dim_1", "phony_dim_2"]); + assert_eq!(dim_names(&file, "g/h/q"), ["phony_dim_3"]); + let h = file.group("g/h").unwrap(); + assert_eq!(h.variable_names().unwrap(), ["q"]); + assert_eq!(h.dimensions().unwrap()[0].name, "phony_dim_3"); + assert_eq!(dim_names(&file, "aa"), ["phony_dim_6", "phony_dim_7"]); + assert_eq!( + dim_names(&file, "zz"), + ["phony_dim_7", "phony_dim_11", "phony_dim_6"] + ); + assert_eq!(dim_names(&file, "second_only"), ["phony_dim_7", "sc"]); + assert_eq!(dim_names(&file, "un"), ["phony_dim_8", "phony_dim_6"]); + assert_eq!(dim_names(&file, "un2"), ["phony_dim_8"]); + assert_eq!(dim_names(&file, "zero"), ["phony_dim_9"]); + assert_eq!(dim_names(&file, "zero_un"), ["phony_dim_9"]); + assert_eq!(dim_names(&file, "zero2"), ["phony_dim_10", "phony_dim_7"]); + let root: Vec<(String, u64, bool)> = file + .dimensions() + .unwrap() + .into_iter() + .map(|d| (d.name, d.size, d.is_unlimited)) + .collect(); + assert_eq!(root[0], ("sc".to_string(), 6, false)); + assert!(root.contains(&("phony_dim_8".to_string(), 2, true))); + assert!(root.contains(&("phony_dim_9".to_string(), 0, true))); + assert!(root.contains(&("phony_dim_10".to_string(), 0, true))); + assert_eq!( + file.variable_names().unwrap(), + [ + "aa", + "s", + "sc", + "second_only", + "un", + "un2", + "zero", + "zero2", + "zero_un", + "zz" + ] + ); +} + +/// Datasets of types netCDF-C cannot represent are not variables: +/// references, bit fields, array types, a compound with a reference member +/// or a half-float member, a compound nesting a compound not seen before; +/// enum, compound, variable-length and opaque types are variables of those +/// classes. netCDF-C also remembers a type it failed to read, so the second +/// dataset of a compound with a reference member is a variable, and a +/// nested compound is accepted once a dataset of the inner type was read. +#[test] +fn types_netcdf_c_skips_are_not_variables() { + skip_if_no_netcdf4!(); + let dir = tempfile::tempdir().unwrap(); + let path = dir.path().join("types.h5"); + run_python(&format!( + r#" +import h5py +import numpy as np +inner = np.dtype([("x", "i2"), ("y", "f4")]) +with h5py.File({path:?}, "w") as f: + f["i4"] = np.arange(3, dtype="i4") + f["f2"] = np.arange(3, dtype="f2") + f["s1"] = np.array([b"a", b"b"], dtype="S1") + f["s5"] = np.array([b"abc", b"b"], dtype="S5") + f["vs"] = np.array(["x", "yy"], dtype=h5py.string_dtype()) + f["bool"] = np.array([True, False]) + f["cmp"] = np.array([(1, 2.0)], dtype=[("a", "i4"), ("b", "f8")]) + f["en"] = np.array([0, 1], dtype=h5py.enum_dtype({{"A": 0, "B": 1}}, basetype="i1")) + d = f.create_dataset("vl", (2,), dtype=h5py.vlen_dtype("i4")) + d[0] = [1, 2] + d[1] = [3] + f["op"] = np.array([b"ab", b"cd"], dtype="V2") + f.create_dataset("ref", (1,), dtype=h5py.ref_dtype)[0] = f["i4"].ref + f["cref1"] = np.array([(1, f["i4"].ref)], dtype=[("a", "i4"), ("r", h5py.ref_dtype)]) + f["cref2"] = np.array([(2, f["i4"].ref)], dtype=[("a", "i4"), ("r", h5py.ref_dtype)]) + f["cmp_f2"] = np.zeros(2, dtype=[("a", "f2")]) + f["in1"] = np.zeros(2, dtype=inner) + f["nested"] = np.zeros(2, dtype=[("a", "i4"), ("in", inner)]) + f["a_nested"] = np.zeros(2, dtype=[("b", "i4"), ("in", np.dtype([("p", "i1")]))]) + sid = h5py.h5s.create_simple((2,)) + h5py.h5d.create(f.id, b"bitf", h5py.h5t.STD_B8LE.copy(), sid) + h5py.h5d.create(f.id, b"arr", h5py.h5t.array_create(h5py.h5t.NATIVE_INT32, (3,)), sid) +"#, + path = path.display().to_string() + )); + common::assert_matches_netcdf_c(&path); + + let file = NetCDF4File::open(&path).unwrap(); + assert_eq!( + file.variable_names().unwrap(), + [ + "bool", "cmp", "cref2", "en", "f2", "i4", "in1", "nested", "op", "s1", "s5", "vl", "vs" + ] + ); + let nc_type = |name: &str| file.variable(name).unwrap().nc_type().unwrap(); + assert_eq!(nc_type("bool"), NcType::Enum); + assert_eq!(nc_type("cmp"), NcType::Compound); + assert_eq!(nc_type("vl"), NcType::VLen); + assert_eq!(nc_type("op"), NcType::Opaque); + assert_eq!(nc_type("s1"), NcType::Char); + assert_eq!(nc_type("s5"), NcType::String); + // netCDF-C 4.9.3 on libhdf5 1.14.6 says NC_STRING (deliberate + // difference, see the README). + assert_eq!(nc_type("f2"), NcType::Float); + assert_eq!( + file.variable("f2").unwrap().read_raw_f32().unwrap(), + [0.0, 1.0, 2.0] + ); + assert!(matches!( + file.variable("ref"), + Err(clawhdf5_netcdf4::Error::VariableNotFound(_)) + )); +} + +/// The order of groups and variables: creation order where the group +/// tracks it (every netCDF-4 file; also h5py with `track_order`), compact +/// or dense (more than 8 links), else name order (h5py by default). +#[test] +fn group_and_variable_order_match_netcdf_c() { + skip_if_no_netcdf4!(); + let dir = tempfile::tempdir().unwrap(); + let nc = dir.path().join("order.nc"); + let h5 = dir.path().join("order.h5"); + run_python(&format!( + r#" +import h5py +import netCDF4 as nc +import numpy as np +names = ["zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c"] +with nc.Dataset({nc:?}, "w") as f: + f.createDimension("x", 2) + for i, n in enumerate(names): + f.createVariable(n, "i4", ("x",))[:] = [i, i + 1] + for n in ["gz", "ga", "gm"]: + f.createGroup(n).createVariable("v", "f8", ("x",))[:] = [1, 2] + small = f.createGroup("small") + for n in ["q", "b", "p"]: + small.createVariable(n, "i2", ("x",))[:] = [3, 4] +with h5py.File({h5:?}, "w") as f: + for i, n in enumerate(names): + f[n] = np.arange(i + 1) + t = f.create_group("tracked", track_order=True) + for i, n in enumerate(names): + t[n] = np.arange(3, dtype="i1") + for n in ["gz", "ga"]: + f.create_group(n)["v"] = np.arange(4.0) +"#, + nc = nc.display().to_string(), + h5 = h5.display().to_string() + )); + common::assert_matches_netcdf_c(&nc); + common::assert_matches_netcdf_c(&h5); + + let file = NetCDF4File::open(&nc).unwrap(); + assert_eq!( + file.variable_names().unwrap(), + [ + "zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c" + ] + ); + assert_eq!(file.group_names().unwrap(), ["gz", "ga", "gm", "small"]); + let file = NetCDF4File::open(&h5).unwrap(); + assert_eq!(file.group_names().unwrap(), ["ga", "gz", "tracked"]); + assert_eq!( + file.variable_names().unwrap(), + [ + "a", "alpha", "beta", "c", "gamma", "k", "mid", "omega", "zeta", "zz" + ] + ); + assert_eq!( + file.group("tracked").unwrap().variable_names().unwrap(), + [ + "zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c" + ] + ); +} + +/// The files of the earlier tests, compared with netCDF-C itself too: +/// dimension scales of h5py, netCDF4-python files with groups, unlimited +/// dimensions and non-coordinate variables named like a dimension. +#[test] +fn netcdf4_python_files_match_netcdf_c() { + skip_if_no_netcdf4!(); + let dir = tempfile::tempdir().unwrap(); + let path = dir.path().join("mixed.nc"); + run_python(&format!( + r#" +import netCDF4 as nc +import numpy as np +with nc.Dataset({path:?}, "w") as f: + f.createDimension("time", None) + f.createDimension("p", 2) + f.createDimension("q", 2) + f.createVariable("a", "i4", ("time",))[0:2] = [1, 2] + f.createVariable("b", "f4", ("time", "q"))[0:5, :] = np.arange(10).reshape(5, 2) + f.createVariable("q", "f4", ("q",))[:] = [0, 1] + f.createVariable("p", "f4", ("q", "p"))[:] = np.array([[0, 1], [2, 3]]) + g = f.createGroup("g") + g.createDimension("r", 3) + g.createVariable("w", "i4", ("r", "time"))[:, 0:1] = np.ones((3, 1)) + g.createGroup("h").createVariable("z", "i8", ("p", "r"))[:] = np.arange(6).reshape(2, 3) +"#, + path = path.display().to_string() + )); + common::assert_matches_netcdf_c(&path); +} diff --git a/crates/clawhdf5-netcdf4/tests/netcdf4_tests.rs b/crates/clawhdf5-netcdf4/tests/netcdf4_tests.rs index a093759..882f42a 100644 --- a/crates/clawhdf5-netcdf4/tests/netcdf4_tests.rs +++ b/crates/clawhdf5-netcdf4/tests/netcdf4_tests.rs @@ -776,9 +776,12 @@ fn test_pure_dimensions_hidden_and_non_coord_names() { .with_shape(&[2]); let file = NetCDF4File::from_bytes(b.finish().unwrap()).unwrap(); + // `x`, and the phony dimension netCDF-C gives the 3 values of + // `_nc4_non_coord_x` (no dimension scale attached). let dims = file.dimensions().unwrap(); - assert_eq!(dims.len(), 1); - assert_eq!(dims[0].name, "x"); + let names: Vec<&str> = dims.iter().map(|d| d.name.as_str()).collect(); + assert_eq!(names, ["x", "phony_dim_1"]); + assert_eq!(dims[1].size, 3); let mut names = file.variable_names().unwrap(); names.sort(); assert_eq!(names, ["v", "x"]); @@ -786,7 +789,8 @@ fn test_pure_dimensions_hidden_and_non_coord_names() { assert_eq!(x.name(), "x"); assert_eq!(x.read_raw_f64().unwrap(), [1.0, 2.0, 3.0]); assert!(!x.is_coordinate()); - // No DIMENSION_LIST: `v` gets `x` by size, as before. + // No DIMENSION_LIST: `v` gets the group's first dimension of its length + // (netCDF-C's phony rule, which also takes real dimensions). assert_eq!(file.variable("v").unwrap().dimensions()[0].name, "x"); } diff --git a/crates/clawhdf5-netcdf4/tests/netcdf_c_view.py b/crates/clawhdf5-netcdf4/tests/netcdf_c_view.py new file mode 100644 index 0000000..d913b20 --- /dev/null +++ b/crates/clawhdf5-netcdf4/tests/netcdf_c_view.py @@ -0,0 +1,238 @@ +#!/usr/bin/env python3 +"""netcdf_c_view.py FILE...: print each file as netCDF-C sees it. + +The metadata comes from netCDF-C itself (the libnetcdf that netCDF4-python +bundles, called through ctypes), not from netCDF4-python's objects, which +leave out variables of types netCDF-C supports but netCDF4-python does not +(opaque, compounds of vlen strings, ...). Values come from netCDF4-python. +Each file is read in its own process under a timeout, so a file that crashes +or hangs libnetcdf only costs that file. + +Output, per file, fields separated by tabs: + + FILE + ERROR netCDF-C cannot open it; nothing else + G every group, pre-order, children in + netCDF-C's order ("/" is the root) + D <0|1> its dimensions in dimension-id order + (1: unlimited) + V + its variables in netCDF-C's order; + and comma-separated, + "-" when there are none; as + clawhdf5_netcdf4::NcType prints it + X the values of a numeric variable of + at most MAX_VALUES elements, as + space-separated Python float reprs + ("-" when there are none) + END + +Used by tests/corpus_vs_netcdf_c.rs (CLAWHDF5_NETCDF_CORPUS). +""" +import concurrent.futures +import ctypes +import glob +import os +import subprocess +import sys + +MAX_VALUES = 5000 +TIMEOUT = 60 +MEMORY = 4 << 30 # address space of each file's process + +ATOMIC = { + 1: "NC_BYTE", 2: "NC_CHAR", 3: "NC_SHORT", 4: "NC_INT", 5: "NC_FLOAT", + 6: "NC_DOUBLE", 7: "NC_UBYTE", 8: "NC_USHORT", 9: "NC_UINT", + 10: "NC_INT64", 11: "NC_UINT64", 12: "NC_STRING", +} +USER_CLASS = {13: "NC_VLEN", 14: "NC_OPAQUE", 15: "NC_ENUM", 16: "NC_COMPOUND"} +NUMERIC = {"NC_BYTE", "NC_SHORT", "NC_INT", "NC_FLOAT", "NC_DOUBLE", + "NC_UBYTE", "NC_USHORT", "NC_UINT", "NC_INT64", "NC_UINT64"} +NAME = 257 # NC_MAX_NAME + 1 +MAXDIMS = 1024 + + +def libnetcdf(): + """The libnetcdf netCDF4-python is linked with (same process, same copy).""" + import netCDF4 + site = os.path.dirname(os.path.dirname(netCDF4.__file__)) + found = glob.glob(os.path.join(site, "netcdf4.libs", "libnetcdf*.so*")) + \ + glob.glob(os.path.join(site, "netCDF4.libs", "libnetcdf*.so*")) + if found: + return ctypes.CDLL(found[0]) + return ctypes.CDLL("libnetcdf.so") + + +def view(path, out): + nc = libnetcdf() + ncid = ctypes.c_int() + rc = nc.nc_open(path.encode(), 0, ctypes.byref(ncid)) # NC_NOWRITE + if rc != 0: + nc.nc_strerror.restype = ctypes.c_char_p + out.append("ERROR\t" + nc.nc_strerror(rc).decode(errors="replace")) + return + numeric = [] # (group path, variable name) + try: + walk(nc, ncid.value, "/", out, numeric) + finally: + nc.nc_close(ncid) + values(path, numeric, out) + + +def check(rc, what): + if rc != 0: + raise RuntimeError(f"{what} failed: {rc}") + + +def name_of(fn, *args): + buf = ctypes.create_string_buffer(NAME) + check(fn(*args, buf), fn.__name__) + return buf.value.decode(errors="replace") + + +def walk(nc, gid, path, out, numeric): + out.append(f"G\t{path}") + n = ctypes.c_int() + ids = (ctypes.c_int * MAXDIMS)() + check(nc.nc_inq_dimids(gid, ctypes.byref(n), ids, 0), "nc_inq_dimids") + for dimid in ids[: n.value]: + length = ctypes.c_size_t() + dname = ctypes.create_string_buffer(NAME) + check(nc.nc_inq_dim(gid, dimid, dname, ctypes.byref(length)), "nc_inq_dim") + out.append(f"D\t{path}\t{dname.value.decode(errors='replace')}\t{length.value}\t" + f"{int(is_unlimited(nc, gid, dimid))}") + nvars = ctypes.c_int() + varids = (ctypes.c_int * 65536)() + check(nc.nc_inq_varids(gid, ctypes.byref(nvars), varids), "nc_inq_varids") + for varid in varids[: nvars.value]: + vname = ctypes.create_string_buffer(NAME) + xtype = ctypes.c_int() + ndims = ctypes.c_int() + dimids = (ctypes.c_int * MAXDIMS)() + natts = ctypes.c_int() + check(nc.nc_inq_var(gid, varid, vname, ctypes.byref(xtype), ctypes.byref(ndims), + dimids, ctypes.byref(natts)), "nc_inq_var") + names, shape = [], [] + for dimid in dimids[: ndims.value]: + dname = ctypes.create_string_buffer(NAME) + length = ctypes.c_size_t() + if nc.nc_inq_dim(gid, dimid, dname, ctypes.byref(length)) != 0: + names.append("?") + shape.append("?") + continue + names.append(dname.value.decode(errors="replace")) + shape.append(str(length.value)) + tname = type_name(nc, gid, xtype.value) + vn = vname.value.decode(errors="replace") + out.append(f"V\t{path}\t{vn}\t{tname}\t{','.join(names) or '-'}\t{','.join(shape) or '-'}") + if tname in NUMERIC and "?" not in shape: + count = 1 + for s in shape: + count *= int(s) + if count <= MAX_VALUES: + numeric.append((path, vn)) + ngrps = ctypes.c_int() + grps = (ctypes.c_int * 65536)() + check(nc.nc_inq_grps(gid, ctypes.byref(ngrps), grps), "nc_inq_grps") + for child in grps[: ngrps.value]: + cname = ctypes.create_string_buffer(NAME) + check(nc.nc_inq_grpname(child, cname), "nc_inq_grpname") + cpath = path.rstrip("/") + "/" + cname.value.decode(errors="replace") + walk(nc, child, cpath, out, numeric) + + +def is_unlimited(nc, gid, dimid): + n = ctypes.c_int() + ids = (ctypes.c_int * MAXDIMS)() + # Unlimited dimensions visible from this group include its parents'. + if nc.nc_inq_unlimdims(gid, ctypes.byref(n), ids) != 0: + return False + return dimid in ids[: n.value] + + +def type_name(nc, gid, xtype): + if xtype in ATOMIC: + return ATOMIC[xtype] + size = ctypes.c_size_t() + base = ctypes.c_int() + nfields = ctypes.c_size_t() + klass = ctypes.c_int() + tname = ctypes.create_string_buffer(NAME) + if nc.nc_inq_user_type(gid, xtype, tname, ctypes.byref(size), ctypes.byref(base), + ctypes.byref(nfields), ctypes.byref(klass)) != 0: + return f"type{xtype}" + return USER_CLASS.get(klass.value, f"class{klass.value}") + + +def values(path, numeric, out): + if not numeric: + return + import numpy as np + import netCDF4 + try: + ds = netCDF4.Dataset(path) + except Exception: # noqa: BLE001 - netCDF4-python refuses some files netCDF-C opens + return + with ds: + for gpath, name in numeric: + try: + group = ds if gpath == "/" else ds[gpath] + var = group.variables[name] + var.set_auto_maskandscale(False) + # A whole-variable read through netCDF-C 4.9.3 lays out a + # variable shorter than an unlimited dimension that is not its + # first wrongly (written values first); reads of one index of + # the leading axis are right. + if var.ndim >= 2: + data = np.stack([np.asarray(var[i]) for i in range(var.shape[0])]) \ + if var.shape[0] else np.zeros(var.shape) + else: + data = np.asarray(var[...]) + flat = np.asarray(data, dtype=np.float64).ravel() + except Exception: # noqa: BLE001 + continue + vals = " ".join(repr(float(v)) for v in flat) or "-" + out.append(f"X\t{gpath}\t{name}\t{vals}") + + +def limit_memory(): + import resource + resource.setrlimit(resource.RLIMIT_AS, (MEMORY, MEMORY)) + + +def one(path): + """Run view() for one file in a child process.""" + try: + proc = subprocess.run([sys.executable, __file__, "--one", path], + capture_output=True, timeout=TIMEOUT, text=True, + preexec_fn=limit_memory) + except subprocess.TimeoutExpired: + return [f"FILE\t{path}", "ERROR\ttimeout", "END"] + lines = proc.stdout.splitlines() + if proc.returncode != 0 or not lines or lines[-1] != "END": + err = (proc.stderr.strip().splitlines() or [f"exit {proc.returncode}"])[-1] + return [f"FILE\t{path}", f"ERROR\tchild failed: {err}", "END"] + return lines + + +def main(argv): + if argv and argv[0] == "--one": + out = [f"FILE\t{argv[1]}"] + try: + view(argv[1], out) + except Exception as e: # noqa: BLE001 + out = [f"FILE\t{argv[1]}", f"ERROR\t{e}"] + out.append("END") + sys.stdout.write("\n".join(out) + "\n") + return + if argv and argv[0] == "--list": + with open(argv[1]) as fh: + argv = [line.rstrip("\n") for line in fh if line.strip()] + jobs = int(os.environ.get("CLAWHDF5_NETCDF_JOBS", "4")) + with concurrent.futures.ThreadPoolExecutor(jobs) as pool: + for lines in pool.map(one, argv): + sys.stdout.write("\n".join(lines) + "\n") + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/docs/known-issues.md b/docs/known-issues.md index 09fd916..5410a74 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -15,10 +15,8 @@ Checked against `main` at `9b5803f` on 2026-09-28. | Issue | Kind | Since | |---|---|---| - -| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 | +| [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 | | [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 | -| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 | | [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 | | [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | | [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | @@ -33,6 +31,83 @@ Checked against `main` at `9b5803f` on 2026-09-28. --- +## NetCDF-4: differences from netCDF-C + +**Status:** open (documented 2026-09-29): deliberate differences, and the +corpus files that still read differently. `clawhdf5-netcdf4` reports a +file's groups, dimensions and variables — names, order, types, shapes — as +netCDF-C 4.9.3 does (`libhdf5/hdf5open.c`; see the crate README), and +`crates/clawhdf5-netcdf4/tests/corpus_vs_netcdf_c.rs` compares it with +netCDF-C itself (the libnetcdf netCDF4-python bundles, called through +ctypes by `tests/netcdf_c_view.py`) over a corpus. Over the conformance +corpus (tank, 2026-09-29, netCDF4-python 1.7.4: netCDF-C 4.9.3 on libhdf5 +1.14.6; `CLAWHDF5_NETCDF_CORPUS=
/conformance/.cache/corpus +CLAWHDF5_PYTHON=$PWD/.venv/bin/python cargo test -p clawhdf5-netcdf4 --test +corpus_vs_netcdf_c -- --nocapture`): of 692 files, netCDF-C opens 429; +420 match in groups, dimensions (names, lengths, unlimited, order), +variables (names, order, types, dimensions, shapes) and the values of +every numeric variable of up to 5000 elements; the other 9 are listed +below and in `tests/corpus_known_differences.txt`, which the test holds to. + +Deliberate: + +- **Floats of other than 4 or 8 bytes** (half, bfloat16, the 4-, 6- and + 8-bit floats, `long double`) are `NcType::Float` (up to 4 bytes) or + `Double`, and read as numbers. netCDF-C 4.9.3 on libhdf5 1.14.6 labels + them `NC_STRING`: their native type (libhdf5 1.14.6 has a native + `_Float16`) matches none of its own. 27 corpus variables; the comparison + shows them as netCDF-C does and counts them. +- **An axis netCDF-C leaves without a dimension** gets one by the phony + rule (the group's first dimension of its length, else a new + `phony_dim_`): a `_Netcdf4Coordinates` id or a `DIMENSION_LIST` scale + it cannot find, and an axis without a scale of a variable whose first + axis has one — there netCDF-C 4.9.3 reads uninitialised memory + (`get_attached_info` mallocs the object ids and `dimscale_visitor` never + fills that axis's): in one run it gave a 2-long axis the group's first + phony dimension, 3 long; in another netCDF4-python could not open the + file. No corpus file has such an axis. +- **Types netCDF-C remembers after failing to read them.** netCDF-C adds a + type to its list before reading its members, and keeps it when that + fails, so the second dataset of such a type is a variable. For a + compound, enum or variable-length type it has a class and is shown here + too (`interop_tests::types_netcdf_c_skips_are_not_variables`); a bit + field, time, array or complex type has none, and netCDF-C lists the + second dataset with an invalid type (class 0): it stays hidden here. +- **Files netCDF-C refuses** are read as well as they can be: a + multi-dimensional dimension scale without `_Netcdf4Coordinates`, a + `_Netcdf4Coordinates` of the wrong length, a named datatype netCDF-C + cannot represent, an unreadable dataset or attribute (netCDF-C fails + `nc_open`; this crate skips it). +- **Cost:** the first call that needs the metadata reads that of the whole + file (every group, and every dataset's attributes), as `nc_open` does; + it was one group's per call. Not measured; worth measuring on a file + with many groups and variables before a release. +- A whole-variable read of a variable shorter than an unlimited dimension + that is not its first: see [the entry of #27](#netcdf-4-variables-dimensions-are-guessed-from-sizes). + +Corpus files that differ (2026-09-29): + +- **External links** (`hdf5/test/testfiles/be_extlink1.h5`, + `le_extlink1.h5`, `hdf5/tools/test/testfiles/h5diff_ext2softlink_src.h5`, + `h5diff_grp_recurse_ext2-1.h5`, `h5diff_grp_recurse_ext2-2.h5`): libhdf5 + follows them into the other file; clawhdf5 does not + ([External links](#external-links-and-external-raw-data-are-not-followed)), + so the linked objects are missing and the phony dimension numbers after + them shift. +- **Values the HDF5 reader refuses** and libhdf5 1.14.6 returns (metadata + matches): `cve_hdf5/cvefiles/cve-2025-2308.h5`, `cve-2025-44904.h5` and + `hdf5/test/testfiles/bad_nbit_parms_walk.h5`, the scale-offset and N-Bit + files of CONFORMANCE.md's ref-bug/our-error list (libhdf5 reads past the + stored data); and `hdf5/test/testfiles/bad_nbit_decompress.h5`, whose + `/Nbit_float_data_le` clawhdf5 refuses ("nbit: element count exceeds + chunk size", also `h5rs dump`) while libhdf5 1.14.6 and h5py 3.16 + (libhdf5 2.0.0) return values. That file is not in the conformance + report and has not been investigated. +- Not compared: the values of 18 variables whose filter (SZIP, Blosc, + bzip2) this crate's default build of the HDF5 reader lacks. Blosc and + bzip2 are pure-Rust cargo features of `clawhdf5` (`blosc`, `bzip2`) that + a dependent can turn on; SZIP needs libaec. + ## In-place modification (`FileEditor`) limits **Status:** open (documented 2026-09-26, updated when version-2 B-tree @@ -539,6 +614,61 @@ Newest first. "Before any release" means no tagged release (v2.7.0 and earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date given. +## NetCDF-4: files without dimension scales, types and order differed from netCDF-C + +**Status:** fixed 2026-09-29 (branch `fix/netcdf-phony-dims`). Affected +every release (v2.1.0 to v2.7.0) and `main` after #27. Wrong metadata only +(names, order and types of dimensions and variables); values were read +right. Users: the dimensions of a file without dimension scales are now +netCDF-C's `phony_dim_`, so code that matched `dim_` or took a +1-D dataset's name for a dimension must use the new names; datasets of +types netCDF-C skips (references, bit fields, array types, compounds of +those, ...) are no longer variables; lists come in netCDF-C's order. +`NcType` has four new variants (`Enum`, `Compound`, `VLen`, `Opaque`) and +is `#[non_exhaustive]`: an exhaustive `match` needs a wildcard arm. + +Found 2026-09-28 by the one-off corpus comparison of #27 (52 of the 78 +files netCDF4-python opened matched; netCDF4-python itself hides variables +of opaque and other types it does not support, so the comparison now asks +netCDF-C through its C API). Under that comparison `main` at `4260af4` +matched 68 of the 429 corpus files netCDF-C opens (tank, 2026-09-29, the +command above run on a copy of `4260af4` with the new test files). What +differed: + +- **Files without dimension scales.** netCDF-C gives every axis a phony + dimension `phony_dim_` (`create_phony_dims`), ids file-wide after the + real dimensions, subgroups numbered before their parent's variables, + shared by length and unlimitedness within a group but not between two + axes of one variable, a length of 0 always unlimited. The crate reported + one dimension per 1-D dataset, named after it (an h5py file with `x(5)` + and `v(5)` had dimensions `x` and `v`, and `v` on `x`), and other axes as + `dim_`. +- **Datasets netCDF-C skips** (`NC_EBADTYPID`: references, bit fields, + time and array types, a compound with such a member or a half-float + member, an enum or VLEN over one) were variables, with `NcType::Char`; + so were enums, compounds, VLENs and opaque types, and 1-byte strings + were `String` where netCDF-C says `NC_CHAR` (longer ones `NC_STRING`). +- **Order.** Groups and variables came in the order of the HDF5 links as + stored (hash order in a dense group); netCDF-C uses creation order when + the group tracks it, else name order. Dimensions without + `_Netcdf4Dimid` came after the others instead of taking the next id in + reading order; a zero-length dimension scale was not unlimited. +- `_Netcdf4Coordinates` ids were looked up only in the variable's group + and its parents (netCDF-C: file-wide), and an unlimited dimension's + length came from its scale's `REFERENCE_LIST` in any group + (`nc4_find_dim_len`: the variables in its group and below). + +The crate now reads the metadata of the whole file in netCDF-C's two +passes (`crates/clawhdf5-netcdf4/src/model.rs`), with a new +`clawhdf5_format::group_v2::links_in_creation_order_in` for the order. +Tests: `interop_tests::phony_dimensions_match_netcdf_c`, +`types_netcdf_c_skips_are_not_variables`, +`group_and_variable_order_match_netcdf_c` and +`netcdf4_python_files_match_netcdf_c` compare h5py- and +netCDF4-python-written files with netCDF-C itself, and +`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in +[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)). + ## NetCDF-4: variables' dimensions are guessed from sizes @@ -590,7 +720,8 @@ dimension scales, and h5netcdf 1.8.1 and xarray files (tank, 2026-09-28, `CLAWHDF5_PYTHON= cargo test -p clawhdf5-netcdf4`). Files without dimension scales still get dimensions by size (netCDF-C gives them `phony_dim_`), as before; see `CHANGELOG.md` for a -comparison over the conformance corpus's netCDF-readable files. +comparison over the conformance corpus's netCDF-readable files. (Fixed +2026-09-29: see the entry above.) ## HDF5 1.8 could not read the files we wrote -- 2.54.0 From e4ba09946f76d027463c2d8937a5d0675b69fd4e Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 20:33:13 -0500 Subject: [PATCH 2/7] HDF5 2.0 native complex as a first-class type on read; Python libver= - Datatype::parse returns Datatype::Complex for class 11 (also inside compounds, arrays and VL types) instead of the {r, i} compound view. - Facade: DType::Complex(Box); read_complex_f32/f64 accept it. - h5rs dump/ls/diff print native complex as h5dump/h5ls/h5diff 2.2.0 do (checked against a fixture written by h5py 3.16 / libhdf5 2.0.0); dump --json keeps the {r, i} compound (hdf5-json has no complex class). - clawhdf5-wasm reads native complex datasets as [re, im] pairs. - Python: clawhdf5.File(path, 'w', libver=...) with h5py's values, mapped to FileBuilder::libver_bounds; 'v108' output opens in HDF5 1.8.23. - Docs: known-issues entry moved to Fixed (history), CHANGELOG, READMEs. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 63 ++++++ README.md | 4 +- crates/clawhdf5-format/src/datatype.rs | 70 +++--- crates/clawhdf5-py/README.md | 17 +- crates/clawhdf5-py/src/file.rs | 92 +++++++- crates/clawhdf5-py/tests/test_libver.py | 143 +++++++++++++ crates/clawhdf5-tools/src/diff.rs | 25 ++- crates/clawhdf5-tools/src/dtype.rs | 40 +++- crates/clawhdf5-tools/src/value.rs | 36 +++- .../tests/native_complex_dump.rs | 202 ++++++++++++++++++ crates/clawhdf5-wasm/src/core.rs | 92 +++++++- crates/clawhdf5/src/reader.rs | 5 +- crates/clawhdf5/src/types.rs | 11 + .../tests/fixtures/gen_native_complex.py | 75 +++++++ .../tests/fixtures/native_complex_hdf5_2.ddl | 80 +++++++ .../tests/fixtures/native_complex_hdf5_2.h5 | Bin 0 -> 4791 bytes crates/clawhdf5/tests/integration_tests.rs | 24 ++- docs/known-issues.md | 47 +++- examples/wasm-viewer/README.md | 8 +- examples/wasm-viewer/test/browser.sh | 6 + examples/wasm-viewer/test/make_fixture.py | 25 ++- examples/wasm-viewer/test/test.mjs | 4 +- 22 files changed, 988 insertions(+), 81 deletions(-) create mode 100644 crates/clawhdf5-py/tests/test_libver.py create mode 100644 crates/clawhdf5-tools/tests/native_complex_dump.rs create mode 100644 crates/clawhdf5/tests/fixtures/gen_native_complex.py create mode 100644 crates/clawhdf5/tests/fixtures/native_complex_hdf5_2.ddl create mode 100644 crates/clawhdf5/tests/fixtures/native_complex_hdf5_2.h5 diff --git a/CHANGELOG.md b/CHANGELOG.md index 843c58b..2d9e329 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -56,6 +56,69 @@ and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4). Affected v2.1.0 to v2.7.0. +### HDF5 2.0 native complex is its own type on read; Python `libver=` (2026-09-29) +- **`Datatype::parse` returns `Datatype::Complex { size, base_type }`** for + a class-11 message, also as a compound member, array base or + variable-length base. It used to return the equivalent `{r, i}` + compound (`docs/known-issues.md`, "HDF5 2.0 native complex numbers read + as a `{r, i}` compound"). **Breaking** for code that matched the + compound view of a native complex type: `Dataset::raw_datatype()` and + attribute datatypes are now `Datatype::Complex`; + `Datatype::complex_as_compound(size, base)` still gives the `{r, i}` view, + and `data_read::read_compound_fields` accepts a `Complex` directly. + Re-serializing a parsed type (e.g. copying it to another file) now writes + class 11 again instead of a compound. +- **Facade:** `DType::Complex(Box)` (new variant, **breaking** for + exhaustive `match`es over `DType`): `Complex(F32)` is numpy `complex64`, + `Complex(F64)` `complex128`, a binary16 part `Complex(Other("float16"))` + (the same `Other` name small floats already use); `Display` prints + `complex`. h5py's own complex encoding, the compound `{r, i}`, is + still `DType::Compound`. `read_complex_f64`/`read_complex_f32` read both. +- **`h5rs`:** `dump` prints native complex as h5dump 2.2.0 does: + `H5T_COMPLEX_IEEE_F{16,32,64}{LE,BE}` (else `H5T_COMPLEX { }`), in + arrays and compounds too, and values as `1.5-2i` (`%g%+gi`; for binary16 + parts, which h5dump has no C type for, `1+-2i` as it prints them). `ls` + shows `complex64`, `complex128-be`, `complex32` (and `complex` for + non-IEEE parts); `ls -v` shows h5ls 2.2.0's `complex number of` / + `IEEE 64-bit big-endian float`. `diff` compares complex values part by + part and prints the difference as h5diff 2.2.0 does (`1+0i`); a native + complex and an `{r, i}` compound are no longer comparable (different + classes, as in h5diff). `dump --json` keeps the `{r, i}` compound and + `[re, im]` values: hdf5-json (h5json 2.0.0) has no complex class. +- **Browser (`clawhdf5-wasm`):** native complex datasets are readable, as + `[re, im]` pairs: `info().elementShape` ends in `2`, `read()` returns the + parts interleaved in a `Float32Array`/`Float64Array` (binary16 parts + widened to f32), and `dtype` is `complex` (`complex`, `array[2]>`). The viewer shows each element + as `[re, im]` with no change to the page. h5py's `{r, i}` compound is + still refused, as every compound is. +- **Python:** `clawhdf5.File(path, 'w', libver=...)` takes h5py's values: + `'v108'`, `'v110'`, `'v112'`, `'v114'`, `'v200'`, `'latest'` (the low + bound; high `'latest'`, as h5py) or a `(low, high)` tuple, mapped to + `FileBuilder::libver_bounds`. `'earliest'` as the low bound writes the + 1.8 format with a `UserWarning` (clawhdf5 cannot write the pre-1.8 + format); as the high bound it is a `ValueError`, as are unknown names and + a low bound above the high one. It is ignored for `'r'` and refused + (`NotImplementedError`) for `'r+'`/`'a'`. Native complex already read as + numpy `complex64`/`complex128` and still does + (`test_read_native_complex_from_h5py`). +- Tests (tank, 2026-09-29): `crates/clawhdf5-tools/tests/native_complex_dump.rs` + compares `h5rs dump` with h5dump 2.2.0's output stored next to a fixture + written by h5py 3.16 / libhdf5 2.0.0 + (`crates/clawhdf5/tests/fixtures/gen_native_complex.py`: + `F32LE`/`F64LE`/`F64BE`/`F16LE`, scalar, compound member, array, root + attribute), line for line outside the values and value for value within + them, and `ls`/`diff` with h5ls/h5diff 2.2.0; + `integration_tests::complex_datasets_and_attributes_round_trip` + (`DType::Complex`, `raw_datatype()`); `clawhdf5-wasm` unit tests on the + fixture, `h5py_interop` and `examples/wasm-viewer/test` (Node: + `test.mjs`, 274 + 1324 checks; the page in headless Chromium for + `/native_c128`, `/native_c64_be`, `/pairs`) with two native complex + datasets added to `make_fixture.py` (when h5py's libhdf5 is 2.0+); + `crates/clawhdf5-py/tests/test_libver.py` (every bound; `'v108'` output + opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer + hard-codes the fixture's length. + ### NetCDF-4: variables' dimensions come from the file (2026-09-28) - `clawhdf5-netcdf4` gave each variable the first unused dimension of equal size (else an anonymous `dim_`), so a variable on an unlimited diff --git a/README.md b/README.md index 79182af..9acde11 100644 --- a/README.md +++ b/README.md @@ -115,12 +115,12 @@ Limits and open issues, with dates, are in |---|---|---|---| | **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) | | **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | -| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 | +| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 | | **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more | | **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | | **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) | | **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | -| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) | +| **Bindings** | Python (read, `'w'` for numeric and complex arrays with h5py's `libver=`, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`; HDF5 2.0 native complex as `[re, im]` pairs); no Zstd/SZIP/pcodec, no compound (h5py's `{r, i}` complex included), reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) | Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`, `blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure diff --git a/crates/clawhdf5-format/src/datatype.rs b/crates/clawhdf5-format/src/datatype.rs index 5e9a678..e3ac607 100644 --- a/crates/clawhdf5-format/src/datatype.rs +++ b/crates/clawhdf5-format/src/datatype.rs @@ -145,13 +145,16 @@ pub enum Datatype { /// imaginary, in rectangular form. `size` is twice the base size and the /// base is an IEEE float. /// - /// This variant exists for **writing** (see + /// [`Datatype::parse`] returns this variant for every class-11 message, + /// also inside compounds, arrays and variable-length types. Readers that + /// want the member view (`{r, i}`, as h5py writes complex numbers by + /// default) use [`Datatype::complex_as_compound`]; `read_compound_fields` + /// accepts this variant directly. + /// + /// Writing it is opt-in (see /// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and - /// newer can read class 11, so it is opt-in and h5py's compound `{r, i}` - /// stays the default complex encoding. [`Datatype::parse`] still - /// surfaces a class-11 message as that equivalent `{r, i}` compound, so - /// every compound reader handles both encodings; parsing what this - /// variant serializes therefore yields a `Compound`, not a `Complex`. + /// newer can read class 11, so h5py's compound `{r, i}` stays the + /// default complex encoding. Complex { size: u32, base_type: Box }, } @@ -857,9 +860,7 @@ impl Datatype { // Complex number (HDF5 2.0, datatype version 5). The properties // are a single base floating-point datatype message; an element // is two consecutive base-type values (real, imaginary). There - // is no member list. Surface it as the equivalent two-member - // compound `{r, i}` — the same shape h5py writes for numpy - // complex dtypes — so downstream compound readers work as-is. + // is no member list. if version != 5 { return Err(FormatError::InvalidDatatypeVersion { class: class_id, @@ -875,7 +876,13 @@ impl Datatype { actual: size as usize, }); } - Ok((Self::complex_as_compound(size, &base_type), pos)) + Ok(( + Datatype::Complex { + size, + base_type: Box::new(base_type), + }, + pos, + )) } _ => Err(FormatError::InvalidDatatypeClass(class_id)), } @@ -1174,9 +1181,8 @@ impl Datatype { /// The `{r, i}` compound equivalent to a native complex type of `size` /// bytes over `base_type`: `r` at offset 0, `i` right after it — the - /// shape h5py writes for numpy complex dtypes, and what [`Self::parse`] - /// returns for a class-11 message. Readers that meet a - /// [`Datatype::Complex`] handle it through this view. + /// shape h5py writes for numpy complex dtypes. Readers that want a + /// [`Datatype::Complex`] as members handle it through this view. pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype { let base_size = base_type.type_size(); Datatype::Compound { @@ -1849,21 +1855,24 @@ mod tests { fn test_complex_v5_from_hdf5_2_0() { let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap(); assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len()); - match dt { - Datatype::Compound { size, members } => { - assert_eq!(size, 16); - assert_eq!(members.len(), 2); - assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0)); - assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8)); - for m in &members { - assert!(matches!( - m.datatype, - Datatype::FloatingPoint { size: 8, .. } - )); - } + match &dt { + Datatype::Complex { size, base_type } => { + assert_eq!(*size, 16); + assert!(matches!( + base_type.as_ref(), + Datatype::FloatingPoint { size: 8, .. } + )); } - other => panic!("expected Compound, got {other:?}"), + other => panic!("expected Complex, got {other:?}"), } + // The member view is h5py's `{r, i}` compound. + let Datatype::Compound { members, .. } = + Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type()) + else { + unreachable!() + }; + assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0)); + assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8)); } #[test] @@ -1887,7 +1896,7 @@ mod tests { assert_eq!(members.len(), 2); assert!(matches!( &members[0].datatype, - Datatype::Compound { size: 16, members } if members.len() == 2 + Datatype::Complex { size: 16, .. } )); assert_eq!( (members[1].name.as_str(), members[1].byte_offset), @@ -1918,12 +1927,9 @@ mod tests { assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0); assert_eq!(dt.type_size(), 16); dt.check_encodable().unwrap(); - // Parsing surfaces class 11 as the equivalent `{r, i}` compound. + // Parsing returns the native complex type unchanged. let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap(); - assert_eq!( - parsed, - Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type()) - ); + assert_eq!(parsed, dt); let f32c = make_native_complex_f32_type().serialize(); assert_eq!( diff --git a/crates/clawhdf5-py/README.md b/crates/clawhdf5-py/README.md index c05f97e..58ac9db 100644 --- a/crates/clawhdf5-py/README.md +++ b/crates/clawhdf5-py/README.md @@ -104,7 +104,22 @@ writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and `complex128` arrays; the file is written on `close()`. Complex arrays are stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads (not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the -Rust API writes that on request). +Rust API writes that on request). Both forms read back as numpy +`complex64`/`complex128`. + +`libver=` sets the library version bounds as in h5py: `'v108'`, +`'v110'`, `'v112'`, `'v114'`, `'v200'` or `'latest'` (the low bound; the +high bound is then `'latest'`), or a `(low, high)` tuple. With `'v108'` +the file is written in the HDF5 1.8 format (version-2 superblock, +version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the +HDF5 1.10 format. `'earliest'` as the low bound writes the 1.8 format too, +with a `UserWarning`: clawhdf5 cannot write the pre-1.8 format. The +argument is ignored when reading and refused for `'r+'`/`'a'`. + +```python +with clawhdf5.File("old.h5", "w", libver="v108") as f: # HDF5 1.8 reads it + f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip") +``` ## Editing a file in place diff --git a/crates/clawhdf5-py/src/file.rs b/crates/clawhdf5-py/src/file.rs index bfb6d57..1b86090 100644 --- a/crates/clawhdf5-py/src/file.rs +++ b/crates/clawhdf5-py/src/file.rs @@ -13,10 +13,13 @@ use crate::attrs::PyAttrs; use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group}; use crate::handle::Handle; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err}; +use clawhdf5_rs::LibVer; /// Internal state for write mode. struct WriteState { path: PathBuf, + /// `libver=` as (low, high); `None` keeps the writer's default. + libver: Option<(LibVer, LibVer)>, root_datasets: Vec, root_attrs: Arc>>, groups: Vec>>, @@ -80,10 +83,32 @@ impl PyFile { /// `az://`; which schemes work depends on how the wheel was built) /// to read the file remotely with default options (see `open_url`) /// mode: 'r' for read (default), 'w' for write + /// libver: library version bounds for mode 'w', as h5py's: one of + /// 'earliest', 'v108', 'v110', 'v112', 'v114', 'v200', 'latest' (the + /// low bound; the high bound is then 'latest') or a (low, high) + /// tuple of them. The low bound is the oldest HDF5 release whose + /// format the file uses ('v108': HDF5 1.8 can read it); the high + /// bound the newest whose features it may use. clawhdf5 cannot write + /// the pre-1.8 format, so a low bound of 'earliest' writes the 1.8 + /// format (with a warning) and a high bound of 'earliest' is an + /// error. Default (None): the HDF5 1.10 format clawhdf5 has always + /// written. Ignored for reading; not supported with 'r+' / 'a'. #[new] - #[pyo3(signature = (path, mode="r"))] - fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult { + #[pyo3(signature = (path, mode="r", libver=None))] + fn new( + py: Python<'_>, + path: &str, + mode: &str, + libver: Option<&Bound<'_, PyAny>>, + ) -> PyResult { let filename = path.to_string(); + let libver = libver.map(|v| parse_libver(py, v)).transpose()?; + if libver.is_some() && matches!(mode, "r+" | "a") { + return Err(PyNotImplementedError::new_err(format!( + "libver with mode '{mode}': clawhdf5's in-place editor keeps the format \ + versions the file already uses" + ))); + } if is_url(path) { if mode != "r" { return Err(PyValueError::new_err(format!( @@ -113,6 +138,7 @@ impl PyFile { // Absolute now: the file is written at close, possibly // after the working directory changed. path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)), + libver, root_datasets: Vec::new(), root_attrs: Arc::new(Mutex::new(Vec::new())), groups: Vec::new(), @@ -460,10 +486,71 @@ fn parse_compression( } } +/// One of h5py's `libver` names as a bound; `high` says which end it is. +/// `Ok(None)` is 'earliest' as the low bound: the pre-1.8 format, which +/// clawhdf5 cannot write. +fn libver_name(name: &str, high: bool) -> PyResult> { + Ok(Some(match name { + "earliest" if high => { + return Err(PyValueError::new_err( + "libver high bound 'earliest' (the pre-1.8 format) cannot be written by \ + clawhdf5; the oldest format it writes is 'v108'", + )); + } + "earliest" => return Ok(None), + "v108" => LibVer::V18, + "v110" => LibVer::V110, + "v112" => LibVer::V112, + "v114" => LibVer::V114, + "v200" => LibVer::V200, + "latest" => LibVer::Latest, + other => { + return Err(PyValueError::new_err(format!( + "unknown libver '{other}'; expected 'earliest', 'v108', 'v110', 'v112', \ + 'v114', 'v200' or 'latest'" + ))); + } + })) +} + +/// h5py's `libver=`: a name (the low bound, high bound 'latest') or a +/// `(low, high)` pair. +fn parse_libver(py: Python<'_>, v: &Bound<'_, PyAny>) -> PyResult<(LibVer, LibVer)> { + let (low, high): (String, String) = match v.extract::() { + Ok(name) => (name, "latest".into()), + Err(_) => v.extract().map_err(|_| { + PyValueError::new_err("libver must be a string or a (low, high) tuple of strings") + })?, + }; + let Some(high) = libver_name(&high, true)? else { + unreachable!("a high bound is never None") + }; + let low = match libver_name(&low, false)? { + Some(low) => low, + None => { + PyModule::import(py, "warnings")?.getattr("warn")?.call1(( + "libver 'earliest': clawhdf5 cannot write the pre-1.8 format; the file \ + is written in the HDF5 1.8 format ('v108') instead", + py.get_type::(), + ))?; + LibVer::V18 + } + }; + if low > high { + return Err(PyValueError::new_err(format!( + "libver low bound {low} is newer than the high bound {high}" + ))); + } + Ok((low, high)) +} + /// Build and write the HDF5 file from accumulated write state. fn finalize_write(state: WriteState) -> PyResult<()> { crate::no_panic(|| { let mut builder = clawhdf5_rs::FileBuilder::new(); + if let Some((low, high)) = state.libver { + builder.libver_bounds(low, high); + } // Root attributes let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner()); @@ -520,6 +607,7 @@ mod tests { let state = WriteState { path: path.clone(), + libver: None, root_datasets: vec![DatasetSpec { name: "data".into(), data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]), diff --git a/crates/clawhdf5-py/tests/test_libver.py b/crates/clawhdf5-py/tests/test_libver.py new file mode 100644 index 0000000..5eb862a --- /dev/null +++ b/crates/clawhdf5-py/tests/test_libver.py @@ -0,0 +1,143 @@ +"""`clawhdf5.File(path, 'w', libver=...)`: h5py's library version bounds. + +A file written with libver='v108' must open in HDF5 1.8. Its h5dump is +found through CLAWHDF5_H5DUMP18 or at ~/.cache/hdf5-1.8.23/bin/h5dump +(scripts/build-hdf5-1.8.sh builds it); without it that check is skipped. +""" + +import os +import subprocess + +import numpy as np +import pytest + +import clawhdf5 + + +def superblock_version(path): + with open(path, "rb") as f: + head = f.read(9) + assert head[:8] == b"\x89HDF\r\n\x1a\n" + return head[8] + + +def h5dump18(): + path = os.environ.get("CLAWHDF5_H5DUMP18") or os.path.expanduser( + "~/.cache/hdf5-1.8.23/bin/h5dump" + ) + try: + out = subprocess.run([path, "--version"], capture_output=True, text=True) + except OSError: + return None + return path if "1.8." in out.stdout else None + + +def write_sample(path, libver): + with clawhdf5.File(path, "w", libver=libver) as f: + f.create_dataset("x", data=np.arange(10, dtype=np.float64)) + f.create_dataset( + "chunked", + data=np.arange(100, dtype=np.int32), + chunks=(30,), + compression="gzip", + ) + f.create_dataset("z", data=np.array([1 + 2j, -3j], dtype=np.complex128)) + g = f.create_group("g") + g.create_dataset("y", data=np.ones(3, dtype=np.float32)) + f.attrs["version"] = 1 + + +@pytest.mark.parametrize( + "libver, sb", + [ + (None, 3), + ("v108", 2), + (("v108", "v108"), 2), + (("v108", "latest"), 2), + ("v110", 3), + ("v112", 3), + ("v114", 3), + ("v200", 3), + ("latest", 3), + (("v110", "v200"), 3), + ], +) +def test_libver_bounds(h5py, tmp_path, libver, sb): + path = str(tmp_path / "libver.h5") + write_sample(path, libver) + # v108 low bound: HDF5 1.8's version-2 superblock; 1.10 and later: 3. + assert superblock_version(path) == sb + with h5py.File(path, "r") as f: + np.testing.assert_array_equal(f["x"][:], np.arange(10.0)) + np.testing.assert_array_equal(f["chunked"][:], np.arange(100)) + np.testing.assert_array_equal(f["z"][:], [1 + 2j, -3j]) + np.testing.assert_array_equal(f["g/y"][:], np.ones(3)) + assert f.attrs["version"] == 1 + with clawhdf5.File(path, "r") as f: + np.testing.assert_array_equal(f["chunked"][:], np.arange(100)) + + +def test_libver_v108_opens_in_hdf5_1_8(tmp_path): + h5dump = h5dump18() + if h5dump is None: + pytest.skip("no HDF5 1.8 h5dump (CLAWHDF5_H5DUMP18)") + path = str(tmp_path / "v108.h5") + write_sample(path, "v108") + out = subprocess.run([h5dump, path], capture_output=True, text=True) + assert out.returncode == 0, out.stderr + assert "h5dump error" not in out.stderr + assert 'DATASET "y"' in out.stdout + # The chunked, deflated dataset (a version-1 B-tree) reads in full. + out = subprocess.run( + [h5dump, "-w", "0", "-d", "/chunked", path], capture_output=True, text=True + ) + assert out.returncode == 0, out.stderr + assert "(0): " + ", ".join(str(k) for k in range(100)) + "\n" in out.stdout + # The default (1.10 format) does not open in 1.8: the bound matters. + path = str(tmp_path / "default.h5") + write_sample(path, None) + out = subprocess.run([h5dump, path], capture_output=True, text=True) + assert out.returncode != 0 + + +def test_libver_earliest_writes_v108_with_a_warning(h5py, tmp_path): + path = str(tmp_path / "earliest.h5") + with pytest.warns(UserWarning, match="pre-1.8"): + write_sample(path, "earliest") + assert superblock_version(path) == 2 + with h5py.File(path, "r") as f: + np.testing.assert_array_equal(f["x"][:], np.arange(10.0)) + + +@pytest.mark.parametrize( + "libver, match", + [ + (("v108", "earliest"), "high bound 'earliest'"), + ("v109", "unknown libver 'v109'"), + (("latest", "v108"), "newer than the high bound"), + (("v108",), "tuple"), + (3, "tuple"), + ], +) +def test_libver_errors(tmp_path, libver, match): + with pytest.raises(ValueError, match=match): + clawhdf5.File(str(tmp_path / "bad.h5"), "w", libver=libver) + + +def test_libver_v108_high_bound_writes_everything(tmp_path): + # Native complex numbers need HDF5 2.0; the Python writer stores complex + # as h5py's {r, i} compound, which 1.8 reads, so the 1.8 high bound is + # fine for everything clawhdf5.File writes. + path = str(tmp_path / "v18only.h5") + write_sample(path, ("v108", "v108")) + assert superblock_version(path) == 2 + + +def test_libver_not_for_editing(tmp_path): + path = str(tmp_path / "e.h5") + write_sample(path, None) + with pytest.raises(NotImplementedError, match="libver"): + clawhdf5.File(path, "r+", libver="latest") + # Reading ignores it, as the bounds only affect what is written. + with clawhdf5.File(path, "r", libver="v108") as f: + assert f["x"].shape == (10,) diff --git a/crates/clawhdf5-tools/src/diff.rs b/crates/clawhdf5-tools/src/diff.rs index d02bb8c..f3d1db5 100644 --- a/crates/clawhdf5-tools/src/diff.rs +++ b/crates/clawhdf5-tools/src/diff.rs @@ -603,7 +603,22 @@ impl Diff { let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) { (Some(x), Some(y), ..) => x.abs_diff(y).to_string(), (_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64), - _ => String::new(), + // Each part's absolute difference, as h5diff 2.x prints + // it (`1+0i`). + _ => match (&va, &vb) { + ( + Value::Complex { re: ar, im: ai, .. }, + Value::Complex { re: br, im: bi, .. }, + ) => match (number(ar), number(ai), number(br), number(bi)) { + (Some(ar), Some(ai), Some(br), Some(bi)) => format!( + "{}+{}i", + value::fmt_float((ar - br).abs(), 64), + value::fmt_float((ai - bi).abs(), 64) + ), + _ => String::new(), + }, + _ => String::new(), + }, }; rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}")); } @@ -713,6 +728,9 @@ impl Diff { (Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => { p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v)) } + (Value::Complex { re: pr, im: pi, .. }, Value::Complex { re: qr, im: qi, .. }) => { + self.equal(a, b, pr, qr) && self.equal(a, b, pi, qi) + } (Value::Ref(None), Value::Ref(None)) => true, (Value::Ref(Some(p)), Value::Ref(Some(q))) => { // Addresses mean nothing across files: compare the paths the @@ -846,7 +864,10 @@ fn types_comparable(a: &Datatype, b: &Datatype) -> bool { ( Datatype::Enumeration { base_type: x, .. }, Datatype::Enumeration { base_type: y, .. }, - ) => types_comparable(x, y), + ) + | (Datatype::Complex { base_type: x, .. }, Datatype::Complex { base_type: y, .. }) => { + types_comparable(x, y) + } _ => true, } } diff --git a/crates/clawhdf5-tools/src/dtype.rs b/crates/clawhdf5-tools/src/dtype.rs index f7cea2e..28491f2 100644 --- a/crates/clawhdf5-tools/src/dtype.rs +++ b/crates/clawhdf5-tools/src/dtype.rs @@ -135,9 +135,16 @@ pub fn short(dt: &Datatype) -> String { return f.short.into(); } match dt { - Datatype::Complex { size, base_type } => { - short(&Datatype::complex_as_compound(*size, base_type)) - } + // numpy's names (`complex64` is two `float32`) for IEEE parts, + // `complex` otherwise. + Datatype::Complex { size, base_type } => match base_type.as_ref() { + Datatype::FloatingPoint { byte_order, .. } if is_ieee(base_type) => format!( + "complex{}{}", + u64::from(*size) * 8, + if be(byte_order) { "-be" } else { "" } + ), + _ => format!("complex<{}>", short(base_type)), + }, Datatype::FixedPoint { size, signed, @@ -216,8 +223,9 @@ pub fn long(dt: &Datatype) -> String { return f.long.into(); } match dt { - Datatype::Complex { size, base_type } => { - long(&Datatype::complex_as_compound(*size, base_type)) + // As h5ls 2.x prints a complex type it has no native name for. + Datatype::Complex { base_type, .. } => { + format!("complex number of\n {}", long(base_type)) } Datatype::FixedPoint { size, @@ -361,6 +369,19 @@ fn atomic_ddl(dt: &Datatype) -> Option { )), // As h5dump 2.x names them (checked against h5dump 2.2.0). Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()), + // h5dump 2.x's predefined complex names: IEEE binary16/32/64 parts. + Datatype::Complex { base_type, .. } => match base_type.as_ref() { + Datatype::FloatingPoint { + size, byte_order, .. + } if is_ieee(base_type) && !matches!(byte_order, DatatypeByteOrder::Vax) => { + Some(format!( + "H5T_COMPLEX_IEEE_F{}{}", + u64::from(*size) * 8, + order_suffix(byte_order) + )) + } + _ => None, + }, Datatype::BitField { size, byte_order, .. } => Some(format!( @@ -418,6 +439,10 @@ pub fn ddl(dt: &Datatype, ind: usize) -> String { Datatype::VariableLength { base_type, .. } => { format!("H5T_VLEN {{ {} }}", ddl(base_type, ind)) } + // Parts h5dump has no predefined complex name for. + Datatype::Complex { base_type, .. } => { + format!("H5T_COMPLEX {{ {} }}", ddl(base_type, ind)) + } Datatype::Opaque { size, tag } => { let tag = String::from_utf8_lossy(tag); format!( @@ -498,6 +523,9 @@ fn string_ddl( /// hdf5-json type object. pub fn json(dt: &Datatype) -> J { match dt { + // hdf5-json (h5json 2.0.0) has no complex class: a native complex + // is written as the `{r, i}` compound h5py uses for complex numbers, + // its values as `[re, im]` pairs. Datatype::Complex { size, base_type } => { json(&Datatype::complex_as_compound(*size, base_type)) } @@ -598,7 +626,7 @@ fn pad_json(p: &StringPadding) -> &'static str { /// Class name used to decide whether two datatypes can be compared. pub fn class(dt: &Datatype) -> &'static str { match dt { - Datatype::Complex { .. } => "compound", + Datatype::Complex { .. } => "complex", Datatype::FixedPoint { .. } => "integer", Datatype::FloatingPoint { .. } => "float", Datatype::Time { .. } => "time", diff --git a/crates/clawhdf5-tools/src/value.rs b/crates/clawhdf5-tools/src/value.rs index 8412e8a..06e7b0d 100644 --- a/crates/clawhdf5-tools/src/value.rs +++ b/crates/clawhdf5-tools/src/value.rs @@ -27,6 +27,15 @@ pub enum Value { /// An enum member (name, when the value matches one) and its value. Enum(Option, i128), Compound(Vec<(String, Value)>), + /// A native complex number (HDF5 2.0 class 11): real and imaginary + /// part. `sign_free` is set when h5dump joins the parts with a bare `+` + /// (parts it has no native C type for: binary16, non-IEEE), giving + /// `1+-2i`; otherwise the imaginary part carries its sign (`1-2i`). + Complex { + re: Box, + im: Box, + sign_free: bool, + }, Array(Vec), /// A variable-length sequence. Seq(Vec), @@ -183,11 +192,18 @@ impl<'a> Decoder<'a> { return Value::Error("short element".into()); }; match dt { - Datatype::Complex { size, base_type } => self.decode( - &Datatype::complex_as_compound(*size, base_type), - b, - depth + 1, - ), + Datatype::Complex { base_type, .. } => { + let part = base_type.type_size() as usize; + let (re, im) = b.split_at(part.min(b.len())); + Value::Complex { + re: Box::new(self.decode(base_type, re, depth + 1)), + im: Box::new(self.decode(base_type, im, depth + 1)), + // h5dump prints `float`/`double` complex (whatever the + // byte order) with `%g%+gi`, anything else as + // `+i`. + sign_free: !(dtype::is_ieee(base_type) && matches!(part, 4 | 8)), + } + } Datatype::FixedPoint { .. } => match decode_int(dt, b) { Some(v) => Value::Int(v), None => Value::Bytes(b.to_vec()), @@ -354,6 +370,15 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option) -> String { .collect::>() .join(", ") ), + Value::Complex { re, im, sign_free } => { + let (re, im) = (text(re, h5paths), text(im, h5paths)); + let plus = if *sign_free || !im.starts_with('-') { + "+" + } else { + "" + }; + format!("{re}{plus}{im}i") + } Value::Array(vs) => format!( "[ {} ]", vs.iter() @@ -408,6 +433,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option) -> J { Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)), Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths), Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()), + Value::Complex { re, im, .. } => J::Array(vec![to_json(re, h5paths), to_json(im, h5paths)]), Value::Array(vs) | Value::Seq(vs) => { J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect()) } diff --git a/crates/clawhdf5-tools/tests/native_complex_dump.rs b/crates/clawhdf5-tools/tests/native_complex_dump.rs new file mode 100644 index 0000000..cd81dca --- /dev/null +++ b/crates/clawhdf5-tools/tests/native_complex_dump.rs @@ -0,0 +1,202 @@ +//! `h5rs dump` and `h5rs ls` of HDF5 2.0 native complex numbers (datatype +//! class 11), against the output of h5dump 2.2.0 and h5ls 2.2.0 (the Debian +//! h5dump CI installs, 1.14.x, predates the type). The h5dump output is +//! stored next to the fixture; see +//! `crates/clawhdf5/tests/fixtures/gen_native_complex.py`. + +use std::path::PathBuf; +use std::process::Command; + +fn fixture(name: &str) -> PathBuf { + PathBuf::from(env!("CARGO_MANIFEST_DIR")) + .join("../clawhdf5/tests/fixtures") + .join(name) +} + +fn h5rs(args: &[&str]) -> String { + let out = Command::new(env!("CARGO_BIN_EXE_h5rs")) + .args(args) + .output() + .unwrap(); + assert!(out.status.success(), "h5rs {args:?}: {out:?}"); + String::from_utf8(out.stdout).unwrap() +} + +/// Split a dump into its lines, with the data lines (`(i): ...`) replaced +/// by their index, and the data values in order. +fn split(ddl: &str) -> (Vec, Vec) { + let (mut frame, mut data) = (Vec::new(), Vec::new()); + for line in ddl.lines() { + match line.split_once("):") { + Some((idx, body)) if idx.trim_start().starts_with('(') => { + frame.push(format!("{idx}):")); + // Values are split at `, `; array elements come bracketed. + data.extend( + body.split(", ") + .map(|s| s.trim().trim_matches(|c| c == '[' || c == ']' || c == ',')) + .map(str::trim) + .filter(|s| !s.is_empty() && *s != "{") + .map(String::from), + ); + } + _ => { + let t = line.trim().trim_end_matches(','); + if t.ends_with('i') || t.parse::().is_ok() { + // A compound member's value on its own line. + data.push(t.to_string()); + } else { + frame.push(line.to_string()); + } + } + } + } + (frame, data) +} + +/// `a+bi`, `a-bi` or `a+-bi` as its two parts (a real number as one). +fn parts(v: &str) -> Vec { + let Some(z) = v.strip_suffix('i') else { + return vec![v.parse().unwrap()]; + }; + let b = z.as_bytes(); + let at = (1..b.len()) + .find(|&k| (b[k] == b'+' || b[k] == b'-') && !matches!(b[k - 1], b'e' | b'E' | b'+')) + .unwrap_or_else(|| panic!("not a complex value: {v}")); + let (re, im) = z.split_at(at); + let im = im.strip_prefix('+').unwrap_or(im); + vec![re.parse().unwrap(), im.parse().unwrap()] +} + +#[test] +fn dump_matches_h5dump_2_2() { + let path = fixture("native_complex_hdf5_2.h5"); + let ours = h5rs(&["dump", path.to_str().unwrap()]); + let reference = std::fs::read_to_string(fixture("native_complex_hdf5_2.ddl")).unwrap(); + let (mut our_frame, our_data) = split(&ours); + let (mut ref_frame, ref_data) = split(&reference); + // The first line names the file as it was given. + our_frame.remove(0); + ref_frame.remove(0); + // Everything but the values is h5dump's byte for byte: the types print + // as H5T_COMPLEX_IEEE_F64LE, H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }, ... + assert_eq!(our_frame, ref_frame); + assert_eq!(our_data.len(), ref_data.len(), "{our_data:?}\n{ref_data:?}"); + assert_eq!(our_data.len(), 36); + // Values: h5dump prints `%g` (6 digits), inf/nan, and the binary16 + // parts it has no C type for as `+i` (`1+-2i`); h5rs prints + // the shortest round-trip string and Inf/NaN, joined the same way. + for (i, (o, r)) in our_data.iter().zip(&ref_data).enumerate() { + assert_eq!(o.contains("+-"), r.contains("+-"), "[{i}] {o} vs {r}"); + let (op, rp) = (parts(o), parts(r)); + assert_eq!(op.len(), rp.len(), "[{i}] {o} vs {r}"); + for (x, y) in op.iter().zip(&rp) { + let same = if y.is_nan() || y.is_infinite() { + x.is_nan() == y.is_nan() && (y.is_nan() || x == y) + } else { + *x as f32 == *y as f32 || (x - y).abs() <= 5e-6 * y.abs() + }; + assert!(same, "[{i}] h5rs {o}, h5dump {r}"); + assert_eq!( + x.is_sign_negative(), + y.is_sign_negative(), + "[{i}] {o} vs {r}" + ); + } + } +} + +#[test] +fn ls_describes_complex_like_h5ls_2_2() { + let path = fixture("native_complex_hdf5_2.h5"); + // h5ls 2.2.0 -v, `Type:` of the types it has no native C name for (it + // prints `native double _Complex` for the others, as it prints `native + // double` where h5rs prints `IEEE 64-bit little-endian float`). + for (name, long) in [ + ( + "c128be", + "complex number of\n IEEE 64-bit big-endian float", + ), + ( + "c128", + "complex number of\n IEEE 64-bit little-endian float", + ), + ( + "c32", + "complex number of\n IEEE 16-bit little-endian float", + ), + ( + "array", + "[2] complex number of\n IEEE 32-bit little-endian float", + ), + ] { + let target = format!("{}/{name}", path.display()); + let verbose = h5rs(&["ls", "-v", &target]); + assert!( + verbose.contains(&format!(" Type: {long}\n")), + "{name}:\n{verbose}" + ); + } + let listing = h5rs(&["ls", path.to_str().unwrap()]); + for (name, short) in [ + ("c64", "complex64"), + ("c128", "complex128"), + ("c128be", "complex128-be"), + ("c32", "complex32"), + ("scalar", "complex128"), + ("array", "array[2]"), + ("compound", "compound{z: complex128, k: int64}"), + ] { + assert!( + listing + .lines() + .any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))), + "{name}:\n{listing}" + ); + } +} + +#[test] +fn json_dump_uses_the_r_i_compound() { + // hdf5-json (h5json 2.0.0) has no complex class: the type is written as + // the `{r, i}` compound h5py uses, the values as `[re, im]` pairs. + let path = fixture("native_complex_hdf5_2.h5"); + let j: serde_json::Value = + serde_json::from_str(&h5rs(&["dump", "--json", path.to_str().unwrap()])).unwrap(); + let ds = j["datasets"] + .as_object() + .unwrap() + .values() + .find(|d| d["alias"][0] == "/scalar") + .unwrap(); + assert_eq!(ds["type"]["class"], "H5T_COMPOUND"); + assert_eq!(ds["type"]["fields"][0]["name"], "r"); + assert_eq!(ds["type"]["fields"][1]["type"]["base"], "H5T_IEEE_F64LE"); + assert_eq!(ds["value"], serde_json::json!([2.5, -0.5])); +} + +#[test] +fn diff_compares_complex_values() { + let path = fixture("native_complex_hdf5_2.h5"); + let p = path.to_str().unwrap(); + // Same values, different byte order: no differences. + let out = Command::new(env!("CARGO_BIN_EXE_h5rs")) + .args(["diff", p, p, "/c128", "/c128be"]) + .output() + .unwrap(); + assert!(out.status.success(), "{out:?}"); + // One element differs in its imaginary part: one difference, exit + // status 1, as h5diff 2.2.0 reports it. + let out = Command::new(env!("CARGO_BIN_EXE_h5rs")) + .args(["diff", "-r", p, p, "/c128", "/c128x"]) + .output() + .unwrap(); + assert_eq!(out.status.code(), Some(1), "{out:?}"); + let text = String::from_utf8_lossy(&out.stdout); + assert!(text.contains("1 difference(s) found"), "{text}"); + // h5diff 2.2.0: `[ 1 1 ] 0+1i 1+1i 1+0i`. + let row = text.lines().find(|l| l.starts_with("[ 1 1 ]")).unwrap(); + assert_eq!( + row.split_whitespace().collect::>(), + ["[", "1", "1", "]", "0+1i", "1+1i", "1+0i"] + ); +} diff --git a/crates/clawhdf5-wasm/src/core.rs b/crates/clawhdf5-wasm/src/core.rs index ea92fb5..c86c4dc 100644 --- a/crates/clawhdf5-wasm/src/core.rs +++ b/crates/clawhdf5-wasm/src/core.rs @@ -312,11 +312,14 @@ impl Reader { let base = array_base(dt); let is_array = !std::ptr::eq(base, dt); Ok(match base { + // Read as the part type: an array or complex element is its + // parts stored one after another (a complex number's real, then + // imaginary part). Datatype::FloatingPoint { size, .. } if *size <= 4 => { - Data::F32(data_read::read_as_f32(raw, dt).map_err(err)?) + Data::F32(data_read::read_as_f32(raw, base).map_err(err)?) } Datatype::FloatingPoint { .. } => { - Data::F64(data_read::read_as_f64(raw, dt).map_err(err)?) + Data::F64(data_read::read_as_f64(raw, base).map_err(err)?) } Datatype::FixedPoint { size, signed, .. } => { let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err); @@ -393,10 +396,13 @@ fn narrow>(v: Vec) -> Result &Datatype { match dt { - Datatype::Array { base_type, .. } => array_base(base_type), + Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => { + array_base(base_type) + } _ => dt, } } @@ -411,6 +417,12 @@ fn element_shape(dt: &Datatype) -> Vec { dims.extend(element_shape(base_type)); dims } + // `[re, im]`: a complex element reads as its two parts. + Datatype::Complex { base_type, .. } => { + let mut dims = vec![2]; + dims.extend(element_shape(base_type)); + dims + } _ => Vec::new(), } } @@ -487,9 +499,7 @@ pub fn describe(dt: &Datatype) -> String { } } match dt { - Datatype::Complex { size, base_type } => { - describe(&Datatype::complex_as_compound(*size, base_type)) - } + Datatype::Complex { base_type, .. } => format!("complex<{}>", describe(base_type)), Datatype::FixedPoint { size, signed, @@ -681,6 +691,74 @@ mod tests { assert!(e.contains("not supported"), "{e}"); } + #[test] + fn native_complex_reads_as_re_im_pairs() { + let z = [[1.0f64, -2.0], [0.5, 0.0], [-0.0, 3.25], [f64::MAX, 1e-300]]; + let mut b = FileBuilder::new(); + b.create_dataset("z") + .with_native_complex_f64_data(&z) + .with_shape(&[2, 2]); + b.create_dataset("z32") + .with_native_complex_f32_data(&[[1.5f32, -2.5]]); + let r = Reader::open(b.finish().unwrap()).unwrap(); + let i = r.info("z").unwrap(); + assert_eq!( + (i.shape, i.dtype, i.element_shape), + (vec![2, 2], "complex".to_string(), vec![2]) + ); + let all = r.read("z", None).unwrap(); + assert_eq!(all.shape, vec![2, 2, 2]); + assert_eq!(all.data, Data::F64(z.iter().flatten().copied().collect())); + let slab = Hyperslab { + start: vec![1, 0], + count: vec![1, 2], + stride: None, + block: None, + }; + let part = r.read("z", Some(&slab)).unwrap(); + assert_eq!(part.shape, vec![1, 2, 2]); + assert_eq!(part.data, Data::F64(vec![-0.0, 3.25, f64::MAX, 1e-300])); + let z32 = r.read("z32", None).unwrap(); + assert_eq!( + (z32.shape, z32.data), + (vec![1, 2], Data::F32(vec![1.5, -2.5])) + ); + } + + #[test] + fn native_complex_written_by_libhdf5_2() { + // Written by h5py 3.16 / libhdf5 2.0.0 (gen_native_complex.py). + let path = concat!( + env!("CARGO_MANIFEST_DIR"), + "/../clawhdf5/tests/fixtures/native_complex_hdf5_2.h5" + ); + let r = Reader::open(std::fs::read(path).unwrap()).unwrap(); + let be = r.read("c128be", None).unwrap(); + assert_eq!(be.shape, vec![2, 3, 2]); + assert_eq!(be.data, r.read("c128", None).unwrap().data); + assert_eq!(r.info("c128be").unwrap().dtype, "complex"); + // binary16 parts widen to f32. + assert_eq!( + r.read("c32", None).unwrap().data, + Data::F32(vec![1.0, -2.0, 0.5, 65504.0, 0.0, -0.0]) + ); + // An array of complex: array dimensions, then the two parts. + let i = r.info("array").unwrap(); + assert_eq!( + (i.dtype.as_str(), i.element_shape), + ("array[2]>", vec![2, 2]) + ); + assert_eq!( + r.read("array", None).unwrap().data, + Data::F32(vec![1.0, 2.0, 3.0, 4.0, -1.0, -1.0, 0.0, 0.0]) + ); + let s = r.read("scalar", None).unwrap(); + assert_eq!((s.shape, s.data), (vec![2], Data::F64(vec![2.5, -0.5]))); + // A compound holding a complex member is still refused. + let e = r.read("compound", None).unwrap_err(); + assert!(e.contains("compound{z: complex, k: i64}"), "{e}"); + } + #[test] fn garbage_is_an_error() { assert!(Reader::open(vec![0u8; 64]).is_err()); diff --git a/crates/clawhdf5/src/reader.rs b/crates/clawhdf5/src/reader.rs index 4426d39..763ef71 100644 --- a/crates/clawhdf5/src/reader.rs +++ b/crates/clawhdf5/src/reader.rs @@ -1342,7 +1342,10 @@ impl<'f> Dataset<'f> { ) -> Result<(Vec, Vec), Error> { let dt = self.datatype()?; let is_complex = match &dt { - // Class 11 parses to this same `{r, i}` compound. + Datatype::Complex { size, base_type } => { + matches!(base_type.as_ref(), Datatype::FloatingPoint { .. }) + && *size == 2 * base_type.type_size() + } Datatype::Compound { size, members } => { matches!(members.as_slice(), [r, i] if r.name == "r" && i.name == "i" diff --git a/crates/clawhdf5/src/types.rs b/crates/clawhdf5/src/types.rs index 77bbe63..6e71a16 100644 --- a/crates/clawhdf5/src/types.rs +++ b/crates/clawhdf5/src/types.rs @@ -41,6 +41,13 @@ pub enum DType { Array(Box, Vec), /// Variable-length UTF-8 string stored via a global heap reference. VariableLengthString, + /// HDF5 2.0 native complex number (datatype class 11) over the given + /// floating-point part type: `Complex(F32)` is numpy `complex64`, + /// `Complex(F64)` `complex128`, and a half-precision part is + /// `Complex(Other("float16"))`. h5py's default complex encoding, the + /// compound `{r, i}`, stays a [`DType::Compound`]; `read_complex_f64` + /// and `read_complex_f32` read both. + Complex(Box), /// Catch-all for HDF5 datatypes that do not map to a specific variant. Other(String), } @@ -60,6 +67,7 @@ impl fmt::Display for DType { DType::U64 => write!(f, "u64"), DType::String => write!(f, "string"), DType::VariableLengthString => write!(f, "vlen_string"), + DType::Complex(base) => write!(f, "complex<{base}>"), DType::Compound(fields) => { write!(f, "compound{{")?; for (i, (name, dt)) in fields.iter().enumerate() { @@ -148,6 +156,9 @@ pub(crate) fn classify_datatype(dt: &clawhdf5_format::datatype::Datatype) -> DTy base_type, dimensions, } => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()), + Datatype::Complex { base_type, .. } => { + DType::Complex(Box::new(classify_datatype(base_type))) + } _ => DType::Other(format!("{dt:?}")), } } diff --git a/crates/clawhdf5/tests/fixtures/gen_native_complex.py b/crates/clawhdf5/tests/fixtures/gen_native_complex.py new file mode 100644 index 0000000..a43bf53 --- /dev/null +++ b/crates/clawhdf5/tests/fixtures/gen_native_complex.py @@ -0,0 +1,75 @@ +"""Generate native_complex_hdf5_2.h5: HDF5 2.0 native complex (class 11). + +Datasets (every one class 11, or a type holding it): + + c64 H5T_COMPLEX_IEEE_F32LE, 4 values incl. -0, inf, nan, subnormal + c128 H5T_COMPLEX_IEEE_F64LE, 2 x 3 + c128be H5T_COMPLEX_IEEE_F64BE, the same values + c128x H5T_COMPLEX_IEEE_F64LE, c128 with element (1,1) changed + c32 H5T_COMPLEX_IEEE_F16LE, 3 values + scalar H5T_COMPLEX_IEEE_F64LE, scalar dataspace + compound {z: H5T_COMPLEX_IEEE_F64LE @0, k: H5T_STD_I64LE @16} + array H5T_ARRAY [2] of H5T_COMPLEX_IEEE_F32LE, 2 elements + +plus the root attribute `zattr` (H5T_COMPLEX_IEEE_F64LE, 2 values). + +Written with h5py 3.16 (libhdf5 2.0.0) through its low-level API, the file +type as the memory type (no conversion). Re-run only if the fixture ever +needs regenerating (generated 2026-09-29): + + python gen_native_complex.py native_complex_hdf5_2.h5 + +The h5dump / h5ls 2.2.0 reference output the tests compare with was made +with the tools of the libhdf5 2.2.0 build described in gen_mx_floats.py: + + h5dump native_complex_hdf5_2.h5 > native_complex_hdf5_2.ddl + h5ls -r -v native_complex_hdf5_2.h5 > native_complex_hdf5_2.ls +""" +import sys + +import h5py +import numpy as np +from h5py import h5a, h5d, h5s, h5t + +out = sys.argv[1] +f = h5py.File(out, "w", libver="latest") + + +def ds(name, tid, arr, shape=None): + shape = arr.shape if shape is None else shape + sp = h5s.create_simple(shape) if shape != () else h5s.create(h5s.SCALAR) + d = h5d.create(f.id, name.encode(), tid, sp) + d.write(h5s.ALL, h5s.ALL, np.ascontiguousarray(arr), mtype=tid) + + +nan, inf = float("nan"), float("inf") +c64 = np.array( + [1.5 - 2j, complex(0.0, -0.0), complex(inf, nan), complex(-inf, 1e-40)], + dtype=np.complex64, +) +ds("c64", h5t.COMPLEX_IEEE_F32LE, c64) +c128 = np.array( + [[1 + 2j, -3.5 + 4e300j, 0.1 + 0.2j], [5e-324 - 0j, 1j, -1]], dtype=np.complex128 +) +ds("c128", h5t.COMPLEX_IEEE_F64LE, c128) +ds("c128be", h5t.COMPLEX_IEEE_F64BE, c128.astype(">c16")) +c128x = c128.copy() +c128x[1, 1] = 1 + 1j +ds("c128x", h5t.COMPLEX_IEEE_F64LE, c128x) +half = np.array([1.0, -2.0, 0.5, 65504.0, 0.0, -0.0], dtype=np.float16) +ds("c32", h5t.COMPLEX_IEEE_F16LE, half, shape=(3,)) +ds("scalar", h5t.COMPLEX_IEEE_F64LE, np.array(2.5 - 0.5j), shape=()) + +ct = h5t.create(h5t.COMPOUND, 24) +ct.insert(b"z", 0, h5t.COMPLEX_IEEE_F64LE) +ct.insert(b"k", 16, h5t.STD_I64LE) +cd = np.array([(1 + 1j, 7), (-2.5 + 0j, -8)], dtype=[("z", "T)LjoE0<(tf#_ujs5-6ZWd@$zos>D&y@_`0mZUF>6+1c@e8=4Jd zF_H(`W1rB0BoqR$O{MQw04CZr@RQqEox#R>Mg@~UQcK<~wB8z7nt zS-JVky2(`$Hy`e%6D7Ifcq4nMfE;fq&V-{4p!Qe@EI(Kn2v_YPXOr41LD@1@6mN`{ zw~u;`bRVZsexpuJV3_A2XTZVs-F`Mb){Xv>R3cYI(?-J238s!hl_#%7sXn|m5&K0{ znqfKtO*#Pw&wu0K%*NqlB9+Jt@tJ9ECC@-Iefdgybm$x~3-l~vd~DO%M$*(sN1mL~ zg3k|+JoBnu8^QHdfGl3v&c%|=zuV4-H7t`4iR1l3(48HH>G)#r)XaU7Ia}d^EiIHE z^;ay?&+@!j?PaAoY;bueO!Y3tr(tx~p7REnu@@Awo0eroK>C^I{=}|O#{J*>m-YoL z+&{#TyI6h_$H}U(3`b#(eF|-C-Xp3S43`yH;b($f5YG-?Sh7^G4=OdOF;HWm#z2jM z8Uy8JpkzkI$*qrv6Vy9FJB@#KpnU>>-m59~_|v=MC49-pR9|KL&Z@NcEaz@*GMQVV zR0@XPf8V)bZuWw91}8v-cyHSO`qkkk(-D=~KMgUhhDad6UFZ}iA*Lj9>L6& rB-oExJeUvxTtE|vk&(>0YlObn$kGy29pxw$LRndx{y29&9{Twka&9Rw literal 0 HcmV?d00001 diff --git a/crates/clawhdf5/tests/integration_tests.rs b/crates/clawhdf5/tests/integration_tests.rs index f978a3a..042124a 100644 --- a/crates/clawhdf5/tests/integration_tests.rs +++ b/crates/clawhdf5/tests/integration_tests.rs @@ -1064,12 +1064,26 @@ fn complex_datasets_and_attributes_round_trip() { let ds = file.dataset(name).unwrap(); assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}"); assert_eq!(ds.shape().unwrap(), vec![3]); - // Both encodings read as the same {r, i} compound. - assert_eq!( - ds.dtype().unwrap(), + // h5py's encoding is a compound, the native one a complex type. + let want = if name == "native64" { + DType::Complex(Box::new(DType::F32)) + } else { DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)]) - ); + }; + assert_eq!(ds.dtype().unwrap(), want, "{name}"); } + assert_eq!( + file.dataset("native128").unwrap().dtype().unwrap(), + DType::Complex(Box::new(DType::F64)) + ); + assert_eq!( + file.dataset("native128").unwrap().raw_datatype().unwrap(), + clawhdf5::make_native_complex_f64_type() + ); + assert_eq!( + DType::Complex(Box::new(DType::F64)).to_string(), + "complex" + ); for name in ["compound128", "native128"] { same( file.dataset(name).unwrap().read_complex_f64().unwrap(), @@ -1094,7 +1108,7 @@ fn complex_datasets_and_attributes_round_trip() { data, }) => { assert!(shape.is_empty()); - assert_eq!(datatype.type_size(), 16); + assert_eq!(datatype, clawhdf5::make_native_complex_f64_type()); let fields = clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap(); let part = |k: usize| { diff --git a/docs/known-issues.md b/docs/known-issues.md index 5410a74..9c17945 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -315,15 +315,12 @@ wrong data. decoded, and external references are an error (object references decode). A multi-dimensional numeric attribute is returned as a flat array (its shape is not reported; `AttrValue::Raw` carries the shape). - HDF5 2.0's native complex type (class 11) is read as the equivalent - `{r, i}` compound: values are right (`Dataset::read_complex_f64`, and - numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs - ls`/`dump` and the browser reader show a compound where h5dump 2.x - prints `H5T_COMPLEX_IEEE_F64LE`, so `h5rs dump` of such a file does not - match h5dump 2.x (h5dump 1.14 cannot read it at all). Writing class 11 - (added 2026-09-28) is opt-in (`with_native_complex_f64_data`, - `make_native_complex_f64_type`); the Python bindings write complex - arrays as h5py's compound only. + HDF5 2.0's native complex type (class 11) is read as its own type since + 2026-09-29 ([fixed](#hdf5-20-native-complex-numbers-read-as-a-r-i-compound)); + writing it is opt-in (`with_native_complex_f64_data`, + `make_native_complex_f64_type`), and the Python bindings write complex + arrays as h5py's compound only. `h5rs dump --json` writes it as the + `{r, i}` compound: hdf5-json (h5json 2.0.0) has no complex class. - **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in that: libhdf5 fails only the first metadata read of an image it cannot load and then reads the file's own (possibly stale) metadata, where we @@ -533,8 +530,10 @@ browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27. again retries it; what was fetched stays cached). - Tested under Node 22 and headless Chromium (Playwright's build) against a local server, cross-origin included; not in Firefox or Safari. -- Compound, reference, opaque, bitfield, time and VL-sequence datasets are - refused with an error naming the type; attributes of those types come back +- Compound (h5py's complex numbers, a compound `{r, i}`, included), + reference, opaque, bitfield, time and VL-sequence datasets are + refused with an error naming the type (HDF5 2.0 native complex datasets + read, as `[re, im]` pairs); attributes of those types come back as `value: null` with their `dtype`. (VL strings read, with h5py's values, through the same `VlResolver` as `File` and `h5rs`.) - No Zstd or SZIP (both link C): such datasets fail with @@ -669,6 +668,32 @@ netCDF4-python-written files with netCDF-C itself, and `corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)). +## HDF5 2.0 native complex numbers read as a `{r, i}` compound + +**Status:** fixed 2026-09-29 (branch `feat/complex-first-class`; PR not +yet opened). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0) +and `main` until then. Listed until now under +[HDF5 features still unsupported](#hdf5-features-still-unsupported). + +A class-11 datatype (`H5T_COMPLEX_IEEE_F64LE`, ...) parsed into the +equivalent compound `{r, i}`. Values were right (`read_complex_f64`, +numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs ls`/`dump` +and the browser reader showed a compound, so `h5rs dump` did not match +h5dump 2.x, the browser refused the dataset, and a parsed type written +back out became a compound. `Datatype::parse` now returns +`Datatype::Complex` (in compounds, arrays and VL types too), +`Dataset::dtype()` is `DType::Complex(..)`, `h5rs` prints what h5dump, +h5ls and h5diff 2.2.0 print (checked on tank, 2026-09-29, against a file +written by h5py 3.16 / libhdf5 2.0.0: +`cargo test -p clawhdf5-tools --test native_complex_dump`), and the +browser reads `[re, im]` pairs. + +**What users must do:** code that matched a native complex type as a +`Datatype::Compound` (or a `DType::Compound` of `r`/`i`) must match +`Datatype::Complex` / `DType::Complex`, or take the compound view with +`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a +`Complex` arm. Files are unchanged; nothing needs rewriting. + ## NetCDF-4: variables' dimensions are guessed from sizes diff --git a/examples/wasm-viewer/README.md b/examples/wasm-viewer/README.md index c697357..83c2fd3 100644 --- a/examples/wasm-viewer/README.md +++ b/examples/wasm-viewer/README.md @@ -123,8 +123,12 @@ else throws an `Error` naming the type. `readHyperslab`). Files of 4 GiB or more are refused at open (wasm32). Every response body is cut off past the length asked for. More in `docs/known-issues.md`. -- Compound, reference, opaque and variable-length-sequence datasets are - refused with an error. Attributes of those types are listed with +- HDF5 2.0 native complex datasets read as `[re, im]` pairs: `dtype` + `complex`, `elementShape` ending in `2`, the parts interleaved in + the typed array. +- Compound (h5py's complex numbers, a compound `{r, i}`, included), + reference, opaque and variable-length-sequence datasets are refused with + an error. Attributes of those types are listed with `value: null` and their `dtype`. - No Zstd or SZIP filters (they link C): such a dataset fails with `unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and diff --git a/examples/wasm-viewer/test/browser.sh b/examples/wasm-viewer/test/browser.sh index df31f7f..2032db6 100644 --- a/examples/wasm-viewer/test/browser.sh +++ b/examples/wasm-viewer/test/browser.sh @@ -93,6 +93,12 @@ expect $H5 "/vlen_str" '"двa"' '
vlen string
' expect $H5 "/pairs" '[2, 3]' '
array[2]<i32>
' # 3-D: leading dimension held at 0, window over the last two. expect $H5 "/cube" '
(2, 5, 6)
' '29' 'dim 0' +# HDF5 2.0 native complex (written when h5py's libhdf5 is 2.0 or later): +# each cell is the [re, im] pair. +if "$PY" -c 'import sys, h5py; sys.exit(not getattr(h5py.get_config(), "has_native_complex", False))'; then + expect $H5 "/native_c128" '
complex<f64>
' '
(3, 2)
' '[1, -1.5]' \ + '[5, -7.5]' +fi # Unsupported type: an error, not values. expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported' # Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that diff --git a/examples/wasm-viewer/test/make_fixture.py b/examples/wasm-viewer/test/make_fixture.py index 2a2af72..c884daa 100644 --- a/examples/wasm-viewer/test/make_fixture.py +++ b/examples/wasm-viewer/test/make_fixture.py @@ -80,6 +80,15 @@ with h5py.File(h5, "w") as f: **hdf5plugin.Zstd()) comp = np.zeros(2, dtype=[("x", "c8")]: + z = (np.arange(6) - 1.5j * np.arange(6)).reshape(3, 2).astype(dt) + d = h5d.create(f.id, name, t, h5s.create_simple(z.shape)) + d.write(h5s.ALL, h5s.ALL, z, mtype=t) g = f.create_group("sensors") g.attrs["location"] = "lab" g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype=" { const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { blockSize: 512, - fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })), + fetch: tamper(async (r) => withHeaders(r, { "Content-Range": `bytes 0-511/${statSync(join(fixDir, "fixture.h5")).size}` })), }); await f.read("/grid"); }, /the server sent 0-511/, "wrong range"); @@ -550,7 +550,7 @@ async function floodTests() { const range = new Headers(init.headers).get("Range"); if (range === "bytes=0-511") return fetch(url, init); const m = /^bytes=(\d+)-(\d+)$/.exec(range); - return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } }); + return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/${statSync(join(fixDir, "fixture.h5")).size}` } }); }, }).catch((e) => e); // The open itself may need a second range: flooded either way. -- 2.54.0 From e4b3a76dc7ec5c8876e7701fecb74165fc2c6860 Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 20:43:14 -0500 Subject: [PATCH 3/7] Chunk dimensions of 2^32 or more; unfiltered 4 GiB chunks written without copies - DataLayout::Chunked::chunk_dimensions is Vec (was Vec), and the chunk index readers, writers and serializers take &[u64]: layout messages of version 4/5 store each dimension in up to 8 bytes, and libhdf5 2.x writes dimensions of 2^32 or more (layout version 5). Such dimensions were refused on read (InvalidChunkDimensions) and write. A version-3 layout (4-byte dimensions) is never written for them; a chunk whose size overflows 64 bits is refused when opened. - The file writer lays chunked datasets out as pieces referring to the chunks instead of copying them into one buffer per pass, and an unfiltered chunk that is a contiguous run of the dataset's data (a dataset stored as one chunk of its shape, row blocks) borrows it; a filtered one is compressed straight from it. Contiguous datasets are not copied either. FileWriter::finish_with streams the file to a callback; FileBuilder::write uses it, so the file is never assembled in memory. DatasetBuilder::with_u8_data_owned takes the data without a copy. Peak RSS writing one unfiltered 1 GiB chunk (with_u8_data_owned + write): 5.0 GiB before, 1.0 GiB after; with_u8_data: 6.0 -> 2.0 GiB. - extract_chunk no longer panics on data shorter than the shape. - Tests: huge_chunk_dims.h5 fixture (libhdf5 2.0.0 via h5py 3.16, u8 chunks of 2^32 + 7), and opt-in end-to-end tests of chunk dims >= 2^32 (filtered and unfiltered, read and written, h5py and h5dump 2.2.0), of an unfiltered 4 GiB+ chunk written by clawhdf5, and of LZ4/Zstd chunks of that size; example write_one_chunk for memory measurements. Co-Authored-By: Claude Opus 5.5 (1M context) --- crates/clawhdf5-format/src/chunked_read.rs | 72 +-- crates/clawhdf5-format/src/chunked_write.rs | 453 ++++++++++++++---- crates/clawhdf5-format/src/data_layout.rs | 42 +- crates/clawhdf5-format/src/ea_writer.rs | 2 +- .../clawhdf5-format/src/extensible_array.rs | 16 +- crates/clawhdf5-format/src/file_writer.rs | 291 ++++++++--- crates/clawhdf5-format/src/fixed_array.rs | 18 +- crates/clawhdf5-format/src/type_builders.rs | 8 +- crates/clawhdf5-py/src/node.rs | 7 +- crates/clawhdf5-tools/src/check.rs | 8 +- crates/clawhdf5-tools/src/ls.rs | 2 +- crates/clawhdf5/examples/write_one_chunk.rs | 46 ++ crates/clawhdf5/src/edit/mod.rs | 5 +- crates/clawhdf5/src/writer.rs | 34 +- .../tests/fixtures/gen_huge_chunks.py | 59 ++- .../tests/fixtures/huge_chunk_dims.h5 | Bin 0 -> 31199 bytes crates/clawhdf5/tests/huge_chunks_interop.rs | 349 +++++++++++++- 17 files changed, 1156 insertions(+), 256 deletions(-) create mode 100644 crates/clawhdf5/examples/write_one_chunk.rs create mode 100644 crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5 diff --git a/crates/clawhdf5-format/src/chunked_read.rs b/crates/clawhdf5-format/src/chunked_read.rs index 5f0528a..593be30 100644 --- a/crates/clawhdf5-format/src/chunked_read.rs +++ b/crates/clawhdf5-format/src/chunked_read.rs @@ -555,7 +555,7 @@ pub(crate) fn checked_byte_len(elements: u64, elem_size: usize) -> Result 0, dim = {d}" ))); } - let bytes = spatial - .iter() - .fold(elem_size as u128, |acc, &c| acc * u128::from(c)); + // Saturating: 33 dimensions of up to 64 bits each can overflow u128. + let bytes = spatial.iter().fold(elem_size as u128, |acc, &c| { + acc.saturating_mul(u128::from(c)) + }); + // So the indexes can compute a chunk's size in 64 bits. + if bytes > u128::from(u64::MAX) { + return Err(FormatError::InvalidChunkDimensions(format!( + "chunk size overflows 64 bits (chunk {spatial:?} of {elem_size}-byte elements)" + ))); + } if layout_version < 4 && bytes > u128::from(u32::MAX) { return Err(FormatError::InvalidChunkDimensions(format!( "chunk size must be < 4GB with v1 b-tree index (chunk {spatial:?} of {elem_size}-byte elements)" ))); } - Ok((rank, spatial.iter().map(|&c| c as usize).collect())) + let spatial = spatial + .iter() + .map(|&c| to_usize(c)) + .collect::, _>>()?; + Ok((rank, spatial)) } /// The size of one element of `dt` as stored in the file: a @@ -625,7 +636,7 @@ pub fn check_chunk_element_size( return Ok(()); }; let expected = stored_element_size(datatype, offset_size); - if u64::from(stored) != expected { + if stored != expected { return Err(FormatError::InvalidChunkDimensions(format!( "stored datatype size in chunk layout does not match datatype description \ (layout {stored} bytes, datatype {expected})" @@ -759,7 +770,7 @@ pub fn collect_chunk_info_in( pub fn collect_chunk_info_checked( file_data: &[u8], btree_address: u64, - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], offset_size: u8, length_size: u8, ) -> Result, FormatError> { @@ -776,7 +787,7 @@ pub fn collect_chunk_info_checked( pub fn collect_chunk_info_checked_in( file_data: &S, btree_address: u64, - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], offset_size: u8, length_size: u8, ) -> Result, FormatError> { @@ -808,7 +819,7 @@ pub fn collect_chunk_info_checked_in( .iter_mut() .zip(chunk.offsets.iter().zip(chunk_dimensions)) { - *w = o / u64::from(d); + *w = o / d; } wanted[ndims - 1] = 0; if let Some(i) = root.find(&wanted) { @@ -849,9 +860,9 @@ impl ChunkNode { } } - fn scale_keys(&mut self, dims: &[u32]) { + fn scale_keys(&mut self, dims: &[u64]) { for (k, &d) in self.keys.iter_mut().zip(dims.iter().cycle()) { - *k /= u64::from(d); + *k /= d; } if let ChunkChildren::Nodes(nodes) = &mut self.children { for n in nodes { @@ -919,9 +930,9 @@ fn btree_cmp3(lt: &[u64], scaled: &[u64], rt: &[u64]) -> core::cmp::Ordering { /// Check one v1 B-tree chunk key's offsets (see /// [`collect_chunk_info_checked`]). -fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u32]) -> Result<(), FormatError> { +fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u64]) -> Result<(), FormatError> { for (&offset, &dim) in offsets.iter().zip(chunk_dimensions) { - if dim == 0 || offset % u64::from(dim) != 0 { + if dim == 0 || offset % dim != 0 { return Err(FormatError::ChunkedReadError(format!( "bad coordinate offset {offsets:?} for chunk dimensions {chunk_dimensions:?}" ))); @@ -937,7 +948,7 @@ fn read_key_offsets( file_data: &[u8], pos: usize, ndims: usize, - chunk_dimensions: Option<&[u32]>, + chunk_dimensions: Option<&[u64]>, out: &mut Vec, ) -> Result<(), FormatError> { let start = out.len(); @@ -966,7 +977,7 @@ fn parse_chunk_node( file_data: &S, btree_address: u64, ndims: usize, - chunk_dimensions: Option<&[u32]>, + chunk_dimensions: Option<&[u64]>, offset_size: u8, depth: usize, stored: &mut Vec, @@ -1075,7 +1086,7 @@ fn parse_chunk_node( pub fn generate_implicit_chunks( base_address: u64, dataset_dims: &[u64], - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], element_size: u32, ) -> Vec { generate_implicit_chunks_in_grid( @@ -1097,17 +1108,18 @@ pub fn generate_implicit_chunks_in_grid( base_address: u64, dataset_dims: &[u64], max_dims: &[u64], - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], element_size: u32, ) -> Vec { let rank = chunk_dimensions.len(); - let chunk_byte_size: u64 = - chunk_dimensions.iter().map(|&d| d as u64).product::() * element_size as u64; + let chunk_byte_size: u64 = chunk_dimensions + .iter() + .fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d)); let mut num_chunks_per_dim = Vec::with_capacity(rank); let mut grid_per_dim = Vec::with_capacity(rank); for d in 0..rank { - let ch = chunk_dimensions[d] as u64; + let ch = chunk_dimensions[d]; let n = dataset_dims[d].div_ceil(ch); num_chunks_per_dim.push(n); grid_per_dim.push(max_dims.get(d).map_or(n, |m| m.div_ceil(ch)).max(n)); @@ -1125,7 +1137,7 @@ pub fn generate_implicit_chunks_in_grid( let nchunks = num_chunks_per_dim[d]; let chunk_idx = remaining % nchunks; remaining /= nchunks; - offsets[d] = chunk_idx * chunk_dimensions[d] as u64; + offsets[d] = chunk_idx * chunk_dimensions[d]; grid_idx = grid_idx.saturating_add(chunk_idx.saturating_mul(down)); down = down.saturating_mul(grid_per_dim[d]); } @@ -1350,7 +1362,7 @@ pub fn list_chunks_in( } (4, Some(2)) => { // Implicit index — use spatial chunk dims only - let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank]; + let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank]; generate_implicit_chunks_in_grid( addr, &dataspace.dimensions, @@ -1364,7 +1376,7 @@ pub fn list_chunks_in( } (4, Some(3)) => { // Fixed Array — use spatial chunk dims only - let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank]; + let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank]; let header = FixedArrayHeader::parse_in( file_data, checked_addr(addr)?, @@ -1384,7 +1396,7 @@ pub fn list_chunks_in( } (4, Some(4)) => { // Extensible Array — use spatial chunk dims only - let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank]; + let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank]; let header = ExtensibleArrayHeader::parse_in( file_data, checked_addr(addr)?, @@ -2435,12 +2447,12 @@ mod tests { /// A leaf whose final key is the one libhdf5 writes past the last chunk: /// each coordinate of the last chunk plus its chunk dimension (the /// element size last). - fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u32], offset_size: u8) -> Vec { + fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u64], offset_size: u8) -> Vec { let last = &chunks.last().expect("a chunk").offsets; let end: Vec = dims .iter() .enumerate() - .map(|(d, &c)| last.get(d).copied().unwrap_or(0) + u64::from(c)) + .map(|(d, &c)| last.get(d).copied().unwrap_or(0) + c) .collect(); build_chunk_btree_leaf_to(chunks, &end, offset_size) } @@ -2771,7 +2783,7 @@ mod tests { } // Build B-tree at offset 0x100 - let dims = [chunk_size_elems as u32, elem_size as u32]; + let dims = [chunk_size_elems as u64, elem_size as u64]; let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os); let btree_addr = 0x100usize; file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree); @@ -3145,7 +3157,7 @@ mod tests { data_offset += compressed.len() + 16; // some padding } - let dims = [chunk_elems as u32, elem_size as u32]; + let dims = [chunk_elems as u64, elem_size as u64]; let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os); let btree_addr = 0x100usize; file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree); @@ -3228,7 +3240,7 @@ mod tests { } } - let dims = [chunk_dims[0] as u32, chunk_dims[1] as u32, elem_size as u32]; + let dims = [chunk_dims[0] as u64, chunk_dims[1] as u64, elem_size as u64]; let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os); let btree_addr = 0x100usize; file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree); @@ -3408,7 +3420,7 @@ mod tests { } let layout = DataLayout::Chunked { - chunk_dimensions: vec![chunk_elems as u32, elem_size as u32], + chunk_dimensions: vec![chunk_elems as u64, elem_size as u64], btree_address: Some(data_addr as u64), version: 4, chunk_index_type: Some(1), diff --git a/crates/clawhdf5-format/src/chunked_write.rs b/crates/clawhdf5-format/src/chunked_write.rs index d750fd8..0c61aaf 100644 --- a/crates/clawhdf5-format/src/chunked_write.rs +++ b/crates/clawhdf5-format/src/chunked_write.rs @@ -5,12 +5,14 @@ extern crate alloc; use crate::addr::saturating_usize; #[cfg(not(feature = "std"))] -use alloc::{format, vec, vec::Vec}; +use alloc::{borrow::Cow, format, vec, vec::Vec}; +#[cfg(feature = "std")] +use std::borrow::Cow; use crate::btree_v1_write; use crate::btree_v2_write::{BTreeV2Params, build_btree_v2}; use crate::checksum::jenkins_lookup3; -use crate::chunk_cache::{CACHE_LINE_SIZE, align_to_cache_line}; +use crate::chunk_cache::CACHE_LINE_SIZE; use crate::chunk_grid::ChunkGrid; use crate::ea_writer; use crate::error::FormatError; @@ -464,7 +466,9 @@ fn extract_chunk( let dst: usize = (0..rank).map(|d| idx[d] * chunk_strides[d]).sum::() * element_size; // Whole elements only, as far as `raw_data` reaches. let n = row.min(raw_data.len().saturating_sub(src)) / element_size * element_size; - chunk[dst..dst + n].copy_from_slice(&raw_data[src..src + n]); + if n > 0 { + chunk[dst..dst + n].copy_from_slice(&raw_data[src..src + n]); + } // Next row: advance every dimension but the last. let mut d = rank - 1; loop { @@ -481,6 +485,66 @@ fn extract_chunk( } } +/// The bytes of the `linear_idx`-th chunk (see [`extract_chunk`]), +/// borrowed from `raw_data` when the chunk is one contiguous run of it: a +/// chunk inside the dataset (not past its edge) whose dimensions after the +/// first one above 1 equal the dataset's. A dataset stored as one chunk of +/// its own shape is such a run, so its chunk is never copied. Other chunks +/// are copied out, as [`extract_chunk`] does. +fn chunk_bytes_of<'a>( + raw_data: &'a [u8], + shape: &[u64], + chunk_dims: &[u64], + element_size: usize, + linear_idx: u64, +) -> Cow<'a, [u8]> { + if shape.is_empty() { + return Cow::Borrowed(raw_data); + } + if let Some(range) = contiguous_chunk(shape, chunk_dims, element_size, linear_idx) + && let (Ok(start), Ok(end)) = (usize::try_from(range.start), usize::try_from(range.end)) + && end <= raw_data.len() + { + return Cow::Borrowed(&raw_data[start..end]); + } + Cow::Owned(extract_chunk(raw_data, shape, chunk_dims, element_size, linear_idx).1) +} + +/// The byte range of the `linear_idx`-th chunk in the row-major dataset, +/// when the chunk is a contiguous run of it (see [`chunk_bytes_of`]). +fn contiguous_chunk( + shape: &[u64], + chunk_dims: &[u64], + element_size: usize, + linear_idx: u64, +) -> Option> { + let rank = shape.len(); + // Past the first chunk dimension above 1, every chunk dimension must + // span the dataset's. + let first_wide = chunk_dims.iter().position(|&c| c > 1).unwrap_or(rank); + if (first_wide + 1..rank).any(|d| chunk_dims[d] != shape[d]) { + return None; + } + let mut remaining = linear_idx; + let mut start = 0u64; + let mut stride = element_size as u64; + for d in (0..rank).rev() { + let n = shape[d].div_ceil(chunk_dims[d]); + let offset = (remaining % n).checked_mul(chunk_dims[d])?; + remaining /= n; + // A chunk past the dataset's edge is padded: not a run of it. + if offset.checked_add(chunk_dims[d])? > shape[d] { + return None; + } + start = start.checked_add(offset.checked_mul(stride)?)?; + stride = stride.checked_mul(shape[d])?; + } + let len = chunk_dims + .iter() + .try_fold(element_size as u64, |acc, &c| acc.checked_mul(c))?; + Some(start..start.checked_add(len)?) +} + /// Parallel compression threshold: use rayon when chunk count exceeds this. /// /// Lowered to 2 to enable parallel compression for typical 4-chunk workloads @@ -508,25 +572,21 @@ const PARALLEL_COMPRESS_MAX_CHUNK_BYTES: u64 = 64 << 20; /// at most [`PARALLEL_COMPRESS_MAX_CHUNK_BYTES`], compression runs across /// rayon threads; otherwise it is sequential. Output order matches chunk /// order, so per-chunk bytes are identical to the sequential path. -fn compress_all_chunks( - raw_data: &[u8], +fn compress_all_chunks<'a>( + raw_data: &'a [u8], shape: &[u64], chunk_dims: &[u64], element_size: usize, chunk_bytes: u64, pipeline: &Option, -) -> Result, u32)>, FormatError> { - let one = |i: u64| -> Result<(u64, Vec, u32), FormatError> { - let (_, raw) = if shape.is_empty() { - (Vec::new(), raw_data.to_vec()) - } else { - extract_chunk(raw_data, shape, chunk_dims, element_size, i) - }; +) -> Result>, FormatError> { + let one = |i: u64| -> Result, FormatError> { + let raw = chunk_bytes_of(raw_data, shape, chunk_dims, element_size, i); let raw_size = raw.len() as u64; match pipeline { Some(pl) => { let (stored, mask) = compress_chunk_masked(&raw, pl, element_size as u32)?; - Ok((raw_size, stored, mask)) + Ok((raw_size, Cow::Owned(stored), mask)) } None => Ok((raw_size, raw, 0)), } @@ -557,7 +617,7 @@ fn compress_all_chunks( /// layout/pipeline messages. `base_address` is where the blob will be placed in the file. /// Serialize a v4 single chunk layout message (public for OH size estimation). pub fn serialize_v4_single_chunk_pub( - chunk_dims: &[u32], + chunk_dims: &[u64], chunk_address: u64, filtered_size: Option, filter_mask: Option, @@ -578,7 +638,7 @@ pub fn serialize_v4_single_chunk_pub( /// Serialize a v4 (or, for a chunk of 4 GiB or more, v5) single chunk /// layout message. fn serialize_v4_single_chunk( - chunk_dims: &[u32], + chunk_dims: &[u64], chunk_address: u64, filtered_size: Option, filter_mask: Option, @@ -622,7 +682,7 @@ fn serialize_v4_single_chunk( /// Serialize a v4 Fixed Array layout message. fn serialize_v4_fixed_array( - chunk_dims: &[u32], + chunk_dims: &[u64], fixed_array_address: u64, offset_size: u8, element_size: u32, @@ -653,7 +713,8 @@ fn serialize_v4_fixed_array( /// dimensions, then the element size). Each takes the fewest bytes that hold /// the largest, as libhdf5 computes it (`H5D__chunk_set_sizes`: /// `(log2(dim) + 8) / 8`); HDF5 2.0.0 refuses any other width. -pub(crate) fn push_v4_chunk_dims(buf: &mut Vec, chunk_dims: &[u32], element_size: u32) { +pub(crate) fn push_v4_chunk_dims(buf: &mut Vec, chunk_dims: &[u64], element_size: u32) { + let element_size = u64::from(element_size); let max_dim = chunk_dims .iter() .copied() @@ -661,14 +722,14 @@ pub(crate) fn push_v4_chunk_dims(buf: &mut Vec, chunk_dims: &[u32], element_ .max() .unwrap_or(1) .max(1); - let width = (32 - max_dim.leading_zeros()).div_ceil(8) as usize; + let width = (64 - max_dim.leading_zeros()).div_ceil(8) as usize; buf.push(width as u8); for &d in chunk_dims.iter().chain(core::iter::once(&element_size)) { buf.extend_from_slice(&d.to_le_bytes()[..width]); } } -fn layout_v4_chunked_prefix(chunk_dims: &[u32], element_size: u32, version: u8) -> Vec { +fn layout_v4_chunked_prefix(chunk_dims: &[u64], element_size: u32, version: u8) -> Vec { let mut buf = Vec::new(); buf.push(version); buf.push(2); // class = chunked @@ -884,7 +945,66 @@ pub fn precompress_chunks( element_size: usize, options: &ChunkOptions, ) -> Result { - let (_, chunk_bytes) = checked_chunk_dims(chunk_dims, element_size)?; + let p = prepare_chunks(raw_data, shape, chunk_dims, element_size, options)?; + Ok(PrecompressedChunks { + chunks: p + .chunks + .into_iter() + .map(|(raw, stored, mask)| (raw, stored.into_owned(), mask)) + .collect(), + has_filters: p.has_filters, + element_size: p.element_size, + shape: p.shape, + chunk_dims: p.chunk_dims, + pipeline_message: p.pipeline_message, + }) +} + +/// One chunk of [`PreparedChunks`]: its raw size, its stored bytes and its +/// filter mask. +pub(crate) type PreparedChunk<'a> = (u64, Cow<'a, [u8]>, u32); + +/// [`PrecompressedChunks`] whose stored bytes may borrow the dataset's data: +/// an unfiltered chunk that is a contiguous run of it (a dataset stored as +/// one chunk, say) is not copied. What the file writer lays out. +pub(crate) struct PreparedChunks<'a> { + pub(crate) chunks: Vec>, + pub(crate) has_filters: bool, + pub(crate) element_size: usize, + pub(crate) shape: Vec, + pub(crate) chunk_dims: Vec, + pub(crate) pipeline_message: Option>, +} + +impl PrecompressedChunks { + /// A view of these chunks as [`PreparedChunks`] (their bytes borrowed). + fn prepared(&self) -> PreparedChunks<'_> { + PreparedChunks { + chunks: self + .chunks + .iter() + .map(|(raw, stored, mask)| (*raw, Cow::Borrowed(&stored[..]), *mask)) + .collect(), + has_filters: self.has_filters, + element_size: self.element_size, + shape: self.shape.clone(), + chunk_dims: self.chunk_dims.clone(), + pipeline_message: self.pipeline_message.clone(), + } + } +} + +/// [`precompress_chunks`], borrowing what it can (see [`PreparedChunks`]). +/// A filtered chunk that is a contiguous run of `raw_data` is compressed +/// straight from it, without a copy of the raw chunk. +pub(crate) fn prepare_chunks<'a>( + raw_data: &'a [u8], + shape: &[u64], + chunk_dims: &[u64], + element_size: usize, + options: &ChunkOptions, +) -> Result, FormatError> { + let chunk_bytes = checked_chunk_dims(chunk_dims, element_size)?; if chunk_bytes > MAX_V4_CHUNK_BYTES { check_huge_chunk_filters(options, chunk_bytes)?; } @@ -902,7 +1022,7 @@ pub fn precompress_chunks( &pipeline, )?; - Ok(PrecompressedChunks { + Ok(PreparedChunks { chunks, has_filters, element_size, @@ -912,26 +1032,87 @@ pub fn precompress_chunks( }) } -/// The chunk dimensions as the layout message stores them (each below -/// 2^32), and one chunk's size in bytes. A chunk dimension of 2^32 or more -/// (which HDF5 2.0 can store, in wider fields) is refused, as it is when -/// read: clawhdf5 holds chunk dimensions as `u32`. So is a chunk whose size -/// overflows 64 bits, or that this platform cannot hold in memory (a chunk -/// of 4 GiB or more on a 32-bit target). -fn checked_chunk_dims( - chunk_dims: &[u64], - element_size: usize, -) -> Result<(Vec, u64), FormatError> { - let dims = chunk_dims - .iter() - .map(|&d| { - u32::try_from(d).map_err(|_| { - FormatError::InvalidChunkDimensions(format!( - "chunk dimension {d} is 2^32 or more, which clawhdf5 does not support" - )) - }) - }) - .collect::, _>>()?; +/// A piece of a chunked dataset's data blob as laid out by +/// [`place_chunked_data`]: zero padding, a chunk's stored bytes (by index +/// into [`PreparedChunks::chunks`]), or index structures. +#[derive(Debug)] +pub(crate) enum Piece { + Zeros(usize), + Chunk(usize), + Bytes(Vec), +} + +/// A chunked dataset laid out at an address, without its chunk bytes +/// copied: the pieces of its data blob, in order, and its messages. +pub(crate) struct ChunkedPlacement { + pub(crate) pieces: Vec, + /// Length of the data blob in bytes. + pub(crate) len: u64, + pub(crate) layout_message: Vec, + pub(crate) pipeline_message: Option>, +} + +impl ChunkedPlacement { + /// The data blob as one buffer (the chunks copied in). + fn into_result(self, pre: &PreparedChunks<'_>) -> ChunkedDataResult { + let mut data_bytes = Vec::with_capacity(saturating_usize(self.len)); + for piece in &self.pieces { + match piece { + Piece::Zeros(n) => data_bytes.resize(data_bytes.len() + n, 0), + Piece::Chunk(i) => data_bytes.extend_from_slice(&pre.chunks[*i].1), + Piece::Bytes(b) => data_bytes.extend_from_slice(b), + } + } + ChunkedDataResult { + data_bytes, + layout_message: self.layout_message, + pipeline_message: self.pipeline_message, + } + } +} + +/// Builds the pieces of a data blob, tracking its length. +#[derive(Default)] +struct Pieces { + pieces: Vec, + len: u64, +} + +impl Pieces { + /// Pad to the next cache line. + fn align(&mut self) { + let aligned = align_chunk_offset(self.len); + if aligned > self.len { + self.pieces + .push(Piece::Zeros(saturating_usize(aligned - self.len))); + self.len = aligned; + } + } + + fn chunk(&mut self, i: usize, len: usize) { + self.pieces.push(Piece::Chunk(i)); + self.len += len as u64; + } + + fn bytes(&mut self, b: Vec) { + self.len += b.len() as u64; + self.pieces.push(Piece::Bytes(b)); + } +} + +/// One chunk's size in bytes. A chunk dimension of 0 is refused, and so is +/// a chunk whose size overflows 64 bits, or that this platform cannot hold +/// in memory (a chunk of 4 GiB or more on a 32-bit target). Chunk +/// dimensions of 2^32 or more are allowed, as in libhdf5 2.x: such a chunk +/// is over 4 GiB, so it takes layout message version 5, whose dimension +/// fields are up to 8 bytes wide (a version-3 layout, with 4-byte fields, +/// is never written for it). +fn checked_chunk_dims(chunk_dims: &[u64], element_size: usize) -> Result { + if chunk_dims.contains(&0) { + return Err(FormatError::InvalidChunkDimensions(format!( + "chunk dimensions {chunk_dims:?} include 0" + ))); + } let bytes = chunk_dims .iter() .try_fold(element_size as u64, |acc, &d| acc.checked_mul(d)) @@ -941,7 +1122,7 @@ fn checked_chunk_dims( "a chunk of {chunk_dims:?} x {element_size} bytes exceeds this platform's address space" )) })?; - Ok((dims, bytes)) + Ok(bytes) } /// The filters clawhdf5 can apply to a chunk of more than `u32::MAX` bytes: @@ -1010,7 +1191,22 @@ pub fn build_chunked_data_from_precompressed_libver( low: LibVer, high: LibVer, ) -> Result { - let (_, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?; + let pre = pre.prepared(); + let placement = place_chunked_data(&pre, base_address, maxshape, low, high)?; + Ok(placement.into_result(&pre)) +} + +/// [`build_chunked_data_from_precompressed_libver`] without copying the +/// chunks into one buffer: the data blob as [`Piece`]s referring to +/// `pre`'s chunks. The file writer writes the pieces out one by one. +pub(crate) fn place_chunked_data( + pre: &PreparedChunks<'_>, + base_address: u64, + maxshape: Option<&[u64]>, + low: LibVer, + high: LibVer, +) -> Result { + let chunk_bytes = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?; if chunk_bytes > MAX_V4_CHUNK_BYTES { if high < LibVer::V200 { return Err(FormatError::LibverBound { @@ -1020,7 +1216,7 @@ pub fn build_chunked_data_from_precompressed_libver( }); } } else if low < LibVer::V110 { - return build_btree_v1_chunked_data(pre, base_address, maxshape); + return place_btree_v1_chunked_data(pre, base_address, maxshape); } let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?; let offset_size: u8 = 8; @@ -1028,17 +1224,14 @@ pub fn build_chunked_data_from_precompressed_libver( let num_chunks = pre.chunks.len(); let element_size = pre.element_size; - let mut data_buf = Vec::new(); + let mut blob = Pieces::default(); let mut written_chunks = Vec::with_capacity(num_chunks); - for (raw_size, compressed, filter_mask) in &pre.chunks { - let aligned_offset = align_to_cache_line(data_buf.len()); - if aligned_offset > data_buf.len() { - data_buf.resize(aligned_offset, 0u8); - } - let address = base_address + data_buf.len() as u64; + for (i, (raw_size, compressed, filter_mask)) in pre.chunks.iter().enumerate() { + blob.align(); + let address = base_address + blob.len; let compressed_size = compressed.len() as u64; - data_buf.extend_from_slice(compressed); + blob.chunk(i, compressed.len()); written_chunks.push(WrittenChunk { address, compressed_size, @@ -1047,28 +1240,22 @@ pub fn build_chunked_data_from_precompressed_libver( }); } - let (chunk_dims_u32, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, element_size)?; let version = layout_version_for(chunk_bytes); - - let aligned_idx = align_to_cache_line(data_buf.len()); - if aligned_idx > data_buf.len() { - data_buf.resize(aligned_idx, 0u8); - } + blob.align(); let layout_message = match &index { ChunkIndexPlan::ExtensibleArray(grid) => { - let ea_address = base_address + data_buf.len() as u64; + let ea_address = base_address + blob.len; let slots = index_slots(grid, &pre.shape, &pre.chunk_dims, &written_chunks, None)?; - let ea_bytes = ea_writer::build_extensible_array_at( + blob.bytes(ea_writer::build_extensible_array_at( &slots, offset_size, length_size, pre.has_filters, ea_address, - ); - data_buf.extend_from_slice(&ea_bytes); + )); ea_writer::serialize_v4_extensible_array( - &chunk_dims_u32, + &pre.chunk_dims, ea_address, offset_size, element_size as u32, @@ -1084,7 +1271,7 @@ pub fn build_chunked_data_from_precompressed_libver( }; let filter_mask = pre.has_filters.then_some(written_chunks[0].filter_mask); serialize_v4_single_chunk( - &chunk_dims_u32, + &pre.chunk_dims, chunk_addr, filtered_size, filter_mask, @@ -1094,7 +1281,7 @@ pub fn build_chunked_data_from_precompressed_libver( ) } ChunkIndexPlan::FixedArray(grid, nslots) => { - let fa_address = base_address + data_buf.len() as u64; + let fa_address = base_address + blob.len; let slots = index_slots( grid, &pre.shape, @@ -1102,16 +1289,15 @@ pub fn build_chunked_data_from_precompressed_libver( &written_chunks, Some(*nslots), )?; - let fa_bytes = build_fixed_array_at( + blob.bytes(build_fixed_array_at( &slots, offset_size, length_size, pre.has_filters, fa_address, - ); - data_buf.extend_from_slice(&fa_bytes); + )); serialize_v4_fixed_array( - &chunk_dims_u32, + &pre.chunk_dims, fa_address, offset_size, element_size as u32, @@ -1120,7 +1306,7 @@ pub fn build_chunked_data_from_precompressed_libver( ) } ChunkIndexPlan::BTreeV2 => { - let bt_address = base_address + data_buf.len() as u64; + let bt_address = base_address + blob.len; let records: Vec<(Vec, &WrittenChunk)> = written_chunks .iter() .enumerate() @@ -1134,9 +1320,9 @@ pub fn build_chunked_data_from_precompressed_libver( pre.has_filters, bt_address, )?; - data_buf.extend_from_slice(&bt_bytes); + blob.bytes(bt_bytes); serialize_v4_btree_v2( - &chunk_dims_u32, + &pre.chunk_dims, bt_address, offset_size, element_size as u32, @@ -1146,8 +1332,9 @@ pub fn build_chunked_data_from_precompressed_libver( } }; - Ok(ChunkedDataResult { - data_bytes: data_buf, + Ok(ChunkedPlacement { + pieces: blob.pieces, + len: blob.len, layout_message, pipeline_message: pre.pipeline_message.clone(), }) @@ -1156,11 +1343,11 @@ pub fn build_chunked_data_from_precompressed_libver( /// Lay out precompressed chunks at `base_address` followed by a version-1 /// B-tree chunk index, with a version-3 layout message: what libhdf5 writes /// for a chunked dataset under a low bound of 1.8. -fn build_btree_v1_chunked_data( - pre: &PrecompressedChunks, +fn place_btree_v1_chunked_data( + pre: &PreparedChunks<'_>, base_address: u64, maxshape: Option<&[u64]>, -) -> Result { +) -> Result { if let Some(ms) = maxshape { let bad = |what: &str| FormatError::ChunkedReadError(format!("maxshape: {what}")); if ms.len() != pre.shape.len() { @@ -1171,20 +1358,17 @@ fn build_btree_v1_chunked_data( } } let offset_size: u8 = 8; - let mut data_buf = Vec::new(); + let mut blob = Pieces::default(); let mut entries = Vec::with_capacity(pre.chunks.len()); for (i, (_raw_size, stored, filter_mask)) in pre.chunks.iter().enumerate() { - let aligned_offset = align_to_cache_line(data_buf.len()); - if aligned_offset > data_buf.len() { - data_buf.resize(aligned_offset, 0u8); - } + blob.align(); entries.push(btree_v1_write::ChunkEntry { scaled: scaled_coords(&pre.shape, &pre.chunk_dims, i), nbytes: stored.len() as u64, filter_mask: *filter_mask, - address: base_address + data_buf.len() as u64, + address: base_address + blob.len, }); - data_buf.extend_from_slice(stored); + blob.chunk(i, stored.len()); } let element_size = u32::try_from(pre.element_size) .map_err(|_| FormatError::Overflow("element size".into()))?; @@ -1193,11 +1377,8 @@ fn build_btree_v1_chunked_data( let btree_address = if entries.is_empty() { u64::MAX } else { - let aligned_idx = align_to_cache_line(data_buf.len()); - if aligned_idx > data_buf.len() { - data_buf.resize(aligned_idx, 0u8); - } - let addr = base_address + data_buf.len() as u64; + blob.align(); + let addr = base_address + blob.len; let tree = btree_v1_write::build_chunk_btree_v1_at( &entries, &pre.chunk_dims, @@ -1205,13 +1386,14 @@ fn build_btree_v1_chunked_data( addr, offset_size, )?; - data_buf.extend_from_slice(&tree); + blob.bytes(tree); addr }; let layout_message = serialize_v3_chunked(&pre.chunk_dims, btree_address, offset_size, element_size)?; - Ok(ChunkedDataResult { - data_bytes: data_buf, + Ok(ChunkedPlacement { + pieces: blob.pieces, + len: blob.len, layout_message, pipeline_message: pre.pipeline_message.clone(), }) @@ -1424,7 +1606,7 @@ fn build_btree_v2_chunk_index_at( /// Serialize a v4 layout message for a version-2 B-tree chunk index. fn serialize_v4_btree_v2( - chunk_dims: &[u32], + chunk_dims: &[u64], btree_address: u64, offset_size: u8, element_size: u32, @@ -2184,7 +2366,7 @@ mod tests { /// filtered index element stores the chunk's size in 8 bytes. #[test] fn huge_chunk_layout_messages_and_index_elements() { - let dims = [HUGE_DIM as u32]; + let dims = [HUGE_DIM]; let parsed = |msg: &[u8]| { assert_eq!(msg[0], 5, "layout message version"); // Class chunked, then (after the flags) 2 dimensions of 4 bytes. @@ -2197,7 +2379,7 @@ mod tests { chunk_index_type, .. } => { - assert_eq!(chunk_dimensions, vec![HUGE_DIM as u32, 8]); + assert_eq!(chunk_dimensions, vec![HUGE_DIM, 8]); chunk_index_type.unwrap() } other => panic!("{other:?}"), @@ -2245,22 +2427,89 @@ mod tests { assert_eq!(u16::from_le_bytes([bt[10], bt[11]]), 28); } + /// A chunk dimension of 2^32 or more (a 1-byte element, 2^32 + 7 per + /// chunk): the dimensions take 5 bytes each, as libhdf5 2.x writes + /// them (`(log2(dim) + 8) / 8`), and parse back whole. + #[test] + fn chunk_dimensions_of_2_pow_32_or_more_take_5_bytes() { + const DIM: u64 = (1 << 32) + 7; + let msg = serialize_v4_single_chunk(&[DIM], 0x800, None, None, 8, 1, 5); + assert_eq!(&msg[..5], &[5, 2, 0, 2, 5]); + assert_eq!(&msg[5..10], &[7, 0, 0, 0, 1]); + assert_eq!(&msg[10..15], &[1, 0, 0, 0, 0]); + assert_eq!(msg[15], 1, "Single Chunk index"); + match DataLayout::parse(&msg, 8, 8).unwrap() { + DataLayout::Chunked { + chunk_dimensions, .. + } => assert_eq!(chunk_dimensions, [DIM, 1]), + other => panic!("{other:?}"), + } + // The largest dimension takes all 8 bytes. + let mut buf = Vec::new(); + push_v4_chunk_dims(&mut buf, &[u64::MAX, 3], 4); + assert_eq!(buf[0], 8); + assert_eq!(buf.len(), 1 + 3 * 8); + // A version-3 layout has 4-byte dimensions: refused, not truncated. + assert!(serialize_v3_chunked(&[DIM], 0x800, 8, 1).is_err()); + } + + /// Which chunks are contiguous runs of the dataset's data (borrowed, + /// not copied, when unfiltered). + #[test] + fn contiguous_chunks_are_borrowed() { + // Rows of a 2-D dataset: chunk (2, 5) of shape (5, 5). + assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 0), Some(0..40)); + assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 1), Some(40..80)); + // The last row chunk runs past the edge: padded, so copied. + assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 2), None); + // Columns split: not contiguous. + assert_eq!(contiguous_chunk(&[4, 6], &[2, 3], 1, 0), None); + // Leading dimensions of 1, then a partial row: contiguous. + assert_eq!(contiguous_chunk(&[2, 3, 8], &[1, 1, 4], 2, 5), Some(40..48)); + // The whole dataset as one chunk. + assert_eq!(contiguous_chunk(&[3, 4], &[3, 4], 8, 0), Some(0..96)); + + let raw: Vec = (0..100).collect(); + let whole = chunk_bytes_of(&raw, &[10, 10], &[10, 10], 1, 0); + assert!(matches!(whole, Cow::Borrowed(b) if b.as_ptr() == raw.as_ptr())); + let rows = chunk_bytes_of(&raw, &[10, 10], &[3, 10], 1, 1); + assert!(matches!(rows, Cow::Borrowed(b) if b == &raw[30..60])); + // Data shorter than the shape is copied (and zero-padded) as before. + let short = chunk_bytes_of(&raw[..50], &[10, 10], &[10, 10], 1, 0); + assert!(matches!(&short, Cow::Owned(v) if v.len() == 100)); + // Borrowed or copied, the bytes are what extract_chunk gives. + for (shape, chunk) in [ + (&[10u64, 10][..], &[3u64, 10][..]), + (&[10, 10], &[3, 4]), + (&[4, 5, 5], &[1, 5, 5]), + (&[4, 5, 5], &[1, 2, 5]), + ] { + for i in 0..chunk_count(shape, chunk) { + let got = chunk_bytes_of(&raw, shape, chunk, 1, i); + assert_eq!(&got[..], &extract_chunk(&raw, shape, chunk, 1, i).1[..]); + } + } + } + #[test] fn huge_chunks_refused_where_unsupported() { - // A chunk dimension of 2^32 or more. + // A chunk dimension of 0. assert!(matches!( - checked_chunk_dims(&[1 << 32], 1), - Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32") + checked_chunk_dims(&[4, 0], 1), + Err(FormatError::InvalidChunkDimensions(_)) )); + // A chunk dimension of 2^32 or more is fine (on a 64-bit target). + #[cfg(target_pointer_width = "64")] + assert_eq!( + checked_chunk_dims(&[(1 << 32) + 5], 1).unwrap(), + (1 << 32) + 5 + ); // A chunk size that overflows u64. assert!(matches!( checked_chunk_dims(&[u32::MAX.into(), u32::MAX.into(), 2], 8), Err(FormatError::Overflow(_)) )); - assert_eq!( - checked_chunk_dims(&[HUGE_DIM], 8).unwrap(), - (vec![HUGE_DIM as u32], HUGE_BYTES) - ); + assert_eq!(checked_chunk_dims(&[HUGE_DIM], 8).unwrap(), HUGE_BYTES); // Filters that cannot take a chunk that large. for plugin in [ PluginFilter::Lzf, diff --git a/crates/clawhdf5-format/src/data_layout.rs b/crates/clawhdf5-format/src/data_layout.rs index cf439d6..8310cd4 100644 --- a/crates/clawhdf5-format/src/data_layout.rs +++ b/crates/clawhdf5-format/src/data_layout.rs @@ -34,7 +34,7 @@ const MAX_LAYOUT_NDIMS: usize = 33; /// (`H5O__layout_decode`): at most [`MAX_LAYOUT_NDIMS`], no dimension 0, and /// before version 4 at least one dataspace dimension plus the element size. /// A zero chunk dimension used to read the dataset as all fill values. -fn check_chunk_dims(dims: Vec, layout_version: u8) -> Result, FormatError> { +fn check_chunk_dims(dims: Vec, layout_version: u8) -> Result, FormatError> { if dims.len() > MAX_LAYOUT_NDIMS { return Err(FormatError::InvalidChunkDimensions( "dimensionality is too large".into(), @@ -72,7 +72,7 @@ pub enum DataLayout { /// Chunked: data stored in chunks via a B-tree. Chunked { /// Chunk dimension sizes. - chunk_dimensions: Vec, + chunk_dimensions: Vec, /// B-tree address, or `None` if undefined. btree_address: Option, /// Layout version (3 or 4). Version 1/2 messages (HDF5 1.4/1.6-era) @@ -404,11 +404,11 @@ impl DataLayout { _ => return Err(FormatError::InvalidLayoutClass(layout_class)), }; ensure_len(data, p, dimensionality * 4)?; - let dims: Vec = data[p..p + dimensionality * 4] + let dims: Vec = data[p..p + dimensionality * 4] .as_chunks::<4>() .0 .iter() - .map(|c| u32::from_le_bytes(*c)) + .map(|c| u64::from(u32::from_le_bytes(*c))) .collect(); p += dimensionality * 4; match layout_class { @@ -424,7 +424,7 @@ impl DataLayout { 1 => { let size = dims .iter() - .try_fold(1u64, |acc, &d| acc.checked_mul(d as u64)) + .try_fold(1u64, |acc, &d| acc.checked_mul(d)) .ok_or_else(|| { FormatError::Overflow(format!("contiguous layout size {dims:?}")) })?; @@ -489,7 +489,12 @@ impl DataLayout { ensure_len(data, p, dimensionality * 4)?; let mut chunk_dimensions = Vec::with_capacity(dimensionality); for _ in 0..dimensionality { - let dim = u32::from_le_bytes([data[p], data[p + 1], data[p + 2], data[p + 3]]); + let dim = u64::from(u32::from_le_bytes([ + data[p], + data[p + 1], + data[p + 2], + data[p + 3], + ])); chunk_dimensions.push(dim); p += 4; } @@ -565,14 +570,8 @@ impl DataLayout { .iter() .rev() .fold(0u64, |acc, &b| (acc << 8) | u64::from(b)); - // Chunk dimensions are held as u32; HDF5 2.0 can write - // larger ones (layout version 5), which are refused - // rather than truncated. - let val = u32::try_from(val).map_err(|_| { - FormatError::InvalidChunkDimensions(format!( - "chunk dimension {val} is larger than 2^32 - 1, which is not supported" - )) - })?; + // HDF5 2.0 writes dimensions of 2^32 or more (layout + // version 5, chunks over 4 GiB) in 5 to 8 bytes. chunk_dimensions.push(val); p += dim_size_encoded_length; } @@ -858,7 +857,7 @@ mod tests { .unwrap_or_else(|e| panic!("width {width}: {e:?}")); assert!( matches!(&layout, DataLayout::Chunked { chunk_dimensions, .. } - if chunk_dimensions.iter().map(|&d| u64::from(d)).eq(dims)), + if chunk_dimensions[..] == dims), "width {width}: {layout:?}" ); } @@ -871,12 +870,13 @@ mod tests { ) ); } - // A dimension past u32 cannot be represented and is refused, not - // truncated. - assert!(matches!( - DataLayout::parse(&v4_chunked_msg(5, &[1 << 32, 8]), 8, 8), - Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32") - )); + // HDF5 2.0 writes dimensions of 2^32 or more (up to 8 bytes each). + for (width, dim) in [(5u8, (1u64 << 32) + 3), (8, u64::MAX)] { + assert!(matches!( + DataLayout::parse(&v4_chunked_msg(width, &[dim, 1]), 8, 8), + Ok(DataLayout::Chunked { chunk_dimensions, .. }) if chunk_dimensions == [dim, 1] + )); + } } #[test] diff --git a/crates/clawhdf5-format/src/ea_writer.rs b/crates/clawhdf5-format/src/ea_writer.rs index 5951c50..04a996c 100644 --- a/crates/clawhdf5-format/src/ea_writer.rs +++ b/crates/clawhdf5-format/src/ea_writer.rs @@ -14,7 +14,7 @@ use crate::chunked_write::{ /// Serialize a v4 Extensible Array layout message. pub(crate) fn serialize_v4_extensible_array( - chunk_dims: &[u32], + chunk_dims: &[u64], ea_address: u64, offset_size: u8, element_size: u32, diff --git a/crates/clawhdf5-format/src/extensible_array.rs b/crates/clawhdf5-format/src/extensible_array.rs index 889f2cf..45f8d82 100644 --- a/crates/clawhdf5-format/src/extensible_array.rs +++ b/crates/clawhdf5-format/src/extensible_array.rs @@ -454,7 +454,7 @@ pub fn read_extensible_array_chunks( header: &ExtensibleArrayHeader, dataset_dims: &[u64], max_dims: Option<&[u64]>, - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], element_size: u32, offset_size: u8, length_size: u8, @@ -480,7 +480,7 @@ pub fn read_extensible_array_chunks_in( header: &ExtensibleArrayHeader, dataset_dims: &[u64], max_dims: Option<&[u64]>, - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], element_size: u32, offset_size: u8, _length_size: u8, @@ -489,12 +489,12 @@ pub fn read_extensible_array_chunks_in( // Linear indexes follow the maximum dimensions, with the unlimited // dimension swizzled to the slowest position (see `chunk_grid`). - let dims_u64: Vec = chunk_dimensions.iter().map(|&d| d as u64).collect(); - let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, &dims_u64)?; + let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, chunk_dimensions)?; let grid = &grid; - let chunk_byte_size: u64 = - chunk_dimensions.iter().map(|&d| d as u64).product::() * element_size as u64; + let chunk_byte_size: u64 = chunk_dimensions + .iter() + .fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d)); // Parse index block (EAIB): signature(4) + version(1) + client_id(1) // + header address(offset_size), then the inline elements, then the @@ -930,7 +930,7 @@ mod tests { let header = ExtensibleArrayHeader::parse(&file_data, aehd_offset, os, ls).unwrap(); let ds_dims = vec![40u64]; // 2 chunks × 20 elements - let chunk_dims = vec![20u32]; + let chunk_dims = vec![20u64]; let chunks = read_extensible_array_chunks( &file_data, &header, @@ -1057,7 +1057,7 @@ mod tests { let file_data = build_inline_plus_data_blocks(); let header = ExtensibleArrayHeader::parse(&file_data, 0x100, os, ls).unwrap(); let ds_dims = vec![40u64]; - let chunk_dims = vec![10u32]; + let chunk_dims = vec![10u64]; let chunks = read_extensible_array_chunks( &file_data, &header, diff --git a/crates/clawhdf5-format/src/file_writer.rs b/crates/clawhdf5-format/src/file_writer.rs index 4eb59ee..a1025ad 100644 --- a/crates/clawhdf5-format/src/file_writer.rs +++ b/crates/clawhdf5-format/src/file_writer.rs @@ -3,15 +3,14 @@ //! Produces valid HDF5 files with v3 superblock, v2 object headers, //! link messages, contiguous datasets, inline and dense attributes. -use crate::addr::saturating_usize; +use crate::addr::{saturating_usize, to_usize}; #[cfg(not(feature = "std"))] use alloc::{format, string::String, vec, vec::Vec}; use crate::attribute::AttributeMessage; use crate::btree_v2_write::{BTreeV2Params, build_btree_v2}; use crate::chunked_write::{ - ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed_libver, - precompress_chunks, + ChunkOptions, Piece, PreparedChunks, place_chunked_data, prepare_chunks, }; use crate::data_layout::VdsMapping; use crate::dataspace::{Dataspace, DataspaceType}; @@ -1612,7 +1611,30 @@ impl FileWriter { self } + /// The file, in memory. Its datasets' data is copied into the buffer + /// once (a dataset stored as one chunk of its own shape, or unfiltered + /// chunks that are contiguous runs of its data, are not copied before); + /// to write a file without holding it all, use [`Self::finish_with`]. pub fn finish(self) -> Result, FormatError> { + let mut buf = Vec::new(); + self.finish_into(&mut buf)?; + Ok(buf) + } + + /// The file, handed to `put` in order, a piece at a time, without + /// assembling it in memory: an unfiltered chunk or contiguous dataset + /// is passed straight from the data the dataset was given. Writing a + /// dataset stored as one unfiltered chunk this way needs little memory + /// beyond the dataset's own data. `put`'s first error stops the write + /// and is returned. + pub fn finish_with>( + self, + put: impl FnMut(&[u8]) -> Result<(), E>, + ) -> Result<(), E> { + self.finish_into(&mut FnSink(put)) + } + + fn finish_into(self, sink: &mut S) -> Result<(), S::Error> { let page_size = self.page_size; if let Some(ps) = page_size && !(MIN_FILE_SPACE_PAGE_SIZE..=MAX_FILE_SPACE_PAGE_SIZE).contains(&ps) @@ -1620,7 +1642,8 @@ impl FileWriter { return Err(FormatError::SerializationError(format!( "file space page size {ps} is outside libhdf5's \ {MIN_FILE_SPACE_PAGE_SIZE}..={MAX_FILE_SPACE_PAGE_SIZE} bytes" - ))); + )) + .into()); } let (low, high) = (self.low, self.high); @@ -1774,15 +1797,16 @@ impl FileWriter { }) .collect::>()?; - struct DataBlob { + struct DataBlob<'a> { data: Vec, oh_bytes: Vec, /// Cached compressed chunks for chunked datasets; reused in Pass 2 - /// to avoid re-compressing the same data. - precompressed: Option, + /// to avoid re-compressing the same data. Unfiltered chunks + /// borrow the dataset's data where they can. + precompressed: Option>, } - let mut dummy_blobs: Vec = Vec::new(); + let mut dummy_blobs: Vec> = Vec::new(); let mut dummy_cursor = 0u64; for (i, d) in all_ds.iter().enumerate() { let dense_blob = ds_dense[i] @@ -1820,21 +1844,16 @@ impl FileWriter { .resolve_chunk_dims_for(&d.ds.dimensions, elem_size); // Compress once in Pass 1; cache the result so Pass 2 can skip // re-compression and just rebuild the index with real addresses. - let pre = precompress_chunks( + let pre = prepare_chunks( &d.raw, &d.ds.dimensions, &chunk_dims, elem_size, &d.chunk_options, )?; - let result = build_chunked_data_from_precompressed_libver( - &pre, - dummy_cursor, - d.maxshape.as_deref(), - low, - high, - )?; - dummy_cursor += result.data_bytes.len() as u64; + let result = + place_chunked_data(&pre, dummy_cursor, d.maxshape.as_deref(), low, high)?; + dummy_cursor += result.len; let oh = build_chunked_dataset_oh( &d.dt, &d.ds, @@ -1849,7 +1868,7 @@ impl FileWriter { d.refcount, )?; dummy_blobs.push(DataBlob { - data: result.data_bytes, + data: Vec::new(), oh_bytes: oh, precompressed: Some(pre), }); @@ -1955,7 +1974,7 @@ impl FileWriter { }) .collect::>()?; - let mut ds_blobs2: Vec = Vec::new(); + let mut ds_blobs2: Vec> = Vec::new(); let global_align_threshold = self.alignment_threshold; let global_align_bytes = self.alignment_bytes; for (i, d) in all_ds.iter().enumerate() { @@ -1977,26 +1996,21 @@ impl FileWriter { &d.fill_message, d.refcount, )?; - ds_blobs2.push(DataBlob { - data: gcol_bytes.clone(), + ds_blobs2.push(OutBlob { + data: vec![Out::Slice(gcol_bytes)], oh_bytes: oh, - precompressed: None, }); } else if is_chunked[i] { let base_address = cursor2 as u64; // Reuse precompressed chunks from Pass 1 — avoids re-compressing // the same data a second time. - let result = build_chunked_data_from_precompressed_libver( - dummy_blobs[i] - .precompressed - .as_ref() - .expect("chunked dataset missing precompressed cache"), - base_address, - d.maxshape.as_deref(), - low, - high, - )?; - cursor2 += result.data_bytes.len(); + let pre = dummy_blobs[i] + .precompressed + .as_ref() + .expect("chunked dataset missing precompressed cache"); + let result = + place_chunked_data(pre, base_address, d.maxshape.as_deref(), low, high)?; + cursor2 += to_usize(result.len)?; let oh = build_chunked_dataset_oh( &d.dt, &d.ds, @@ -2010,11 +2024,16 @@ impl FileWriter { &d.fill_message, d.refcount, )?; - ds_blobs2.push(DataBlob { - data: result.data_bytes, - oh_bytes: oh, - precompressed: None, - }); + let data = result + .pieces + .into_iter() + .map(|p| match p { + Piece::Zeros(n) => Out::Zeros(n), + Piece::Chunk(c) => Out::Slice(&pre.chunks[c].1), + Piece::Bytes(b) => Out::Bytes(b), + }) + .collect(); + ds_blobs2.push(OutBlob { data, oh_bytes: oh }); } else if is_compact[i] { // Compact: data is inline in the object header, no external blob let oh = build_compact_dataset_oh( @@ -2030,10 +2049,9 @@ impl FileWriter { d.refcount, layout_version, )?; - ds_blobs2.push(DataBlob { + ds_blobs2.push(OutBlob { data: vec![], oh_bytes: oh, - precompressed: None, }); } else { // Determine alignment: per-dataset overrides global @@ -2060,13 +2078,10 @@ impl FileWriter { d.refcount, layout_version, )?; - let mut data = vec![0u8; padding]; - data.extend_from_slice(&d.raw); cursor2 += d.raw.len(); - ds_blobs2.push(DataBlob { - data, + ds_blobs2.push(OutBlob { + data: vec![Out::Zeros(padding), Out::Slice(&d.raw)], oh_bytes: oh, - precompressed: None, }); } } @@ -2080,7 +2095,8 @@ impl FileWriter { cursor2 = cursor2.next_multiple_of(ps as usize); } let eof_addr2 = cursor2 as u64; - let mut buf = Vec::with_capacity(cursor2); + sink.reserve(cursor2); + let mut out = Counted { sink, written: 0 }; let sb = Superblock { version: superblock_version, @@ -2103,9 +2119,9 @@ impl FileWriter { checksum: None, page_size: None, }; - buf.extend_from_slice(&sb.serialize()); + out.put(&sb.serialize())?; if let Some(ref ext) = sb_ext { - buf.extend_from_slice(ext); + out.put(ext)?; } // Group OHs + dense blobs (link blob, then attr blob, matching pass 2) @@ -2132,32 +2148,120 @@ impl FileWriter { g.refcount, )?; debug_assert_eq!(oh.len(), group_oh_sizes[gi]); - debug_assert_eq!(buf.len() as u64, group_addrs2[gi]); - buf.extend_from_slice(&oh); + debug_assert_eq!(out.written as u64, group_addrs2[gi]); + out.put(&oh)?; if let Some(ref b) = link_blob { - buf.extend_from_slice(&b.blob); + out.put(&b.blob)?; } if let Some(ref blob) = group_dense_blobs[gi] { - buf.extend_from_slice(&blob.blob); + out.put(&blob.blob)?; } } // Dataset OHs + dense blobs for (i, blob) in ds_blobs2.iter().enumerate() { - buf.extend_from_slice(&blob.oh_bytes); + out.put(&blob.oh_bytes)?; if let Some(ref dense) = ds_dense_blobs[i] { - buf.extend_from_slice(&dense.blob); + out.put(&dense.blob)?; } } // Data for blob in &ds_blobs2 { - buf.extend_from_slice(&blob.data); + for piece in &blob.data { + match piece { + Out::Bytes(b) => out.put(b)?, + Out::Slice(s) => out.put(s)?, + Out::Zeros(n) => out.zeros(*n)?, + } + } } - debug_assert_eq!(buf.len(), data_end); - buf.resize(cursor2, 0); - Ok(buf) + debug_assert_eq!(out.written, data_end); + out.zeros(cursor2 - data_end)?; + Ok(()) + } +} + +/// A piece of a dataset's data as [`FileWriter`] writes it out. +enum Out<'a> { + Bytes(Vec), + /// Borrowed from the dataset's data or its prepared chunks. + Slice(&'a [u8]), + Zeros(usize), +} + +/// A dataset's object header and the pieces of its data. +struct OutBlob<'a> { + oh_bytes: Vec, + data: Vec>, +} + +/// Where [`FileWriter`] puts a file's bytes, in order. +trait Sink { + type Error: From; + + /// The file will be `_len` bytes long. + fn reserve(&mut self, _len: usize) {} + + fn put(&mut self, bytes: &[u8]) -> Result<(), Self::Error>; + + fn zeros(&mut self, n: usize) -> Result<(), Self::Error> { + const ZEROS: [u8; 4096] = [0; 4096]; + let mut left = n; + while left > 0 { + let k = left.min(ZEROS.len()); + self.put(&ZEROS[..k])?; + left -= k; + } + Ok(()) + } +} + +impl Sink for Vec { + type Error = FormatError; + + fn reserve(&mut self, len: usize) { + self.reserve_exact(len.saturating_sub(self.len())); + } + + fn put(&mut self, bytes: &[u8]) -> Result<(), FormatError> { + self.extend_from_slice(bytes); + Ok(()) + } + + fn zeros(&mut self, n: usize) -> Result<(), FormatError> { + self.resize(self.len() + n, 0); + Ok(()) + } +} + +/// A [`Sink`] calling a function ([`FileWriter::finish_with`]). +struct FnSink(F); + +impl, F: FnMut(&[u8]) -> Result<(), E>> Sink for FnSink { + type Error = E; + + fn put(&mut self, bytes: &[u8]) -> Result<(), E> { + (self.0)(bytes) + } +} + +/// A [`Sink`] and the number of bytes put into it. +struct Counted<'s, S> { + sink: &'s mut S, + written: usize, +} + +impl Counted<'_, S> { + fn put(&mut self, bytes: &[u8]) -> Result<(), S::Error> { + self.written += bytes.len(); + self.sink.put(bytes) + } + + fn zeros(&mut self, n: usize) -> Result<(), S::Error> { + self.written += n; + self.sink.zeros(n) } } @@ -3015,4 +3119,75 @@ mod tests { assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 3); assert_eq!(layout_of(&bytes, "x")[0], 3); } + + /// `finish_with` hands out exactly the bytes `finish` returns, for every + /// storage (contiguous, compact, chunked filtered and not, chunks + /// borrowed and copied, a version-1 B-tree, a paged file), and stops at + /// the sink's first error. + #[test] + fn finish_with_matches_finish() { + let build = |low: LibVer, paged: bool| { + let mut fw = FileWriter::new(); + fw.libver_bounds(low, LibVer::Latest); + if paged { + fw.with_page_size(4096); + } + let v: Vec = (0..300).map(f64::from).collect(); + fw.create_dataset("contig").with_f64_data(&v); + fw.create_dataset("compact") + .with_f64_data(&[1.0, 2.0]) + .compact(); + fw.create_dataset("one_chunk") + .with_f64_data(&v) + .with_shape(&[30, 10]) + .with_chunks(&[30, 10]); + fw.create_dataset("rows") + .with_f64_data(&v) + .with_shape(&[30, 10]) + .with_chunks(&[7, 10]); + fw.create_dataset("tiles") + .with_f64_data(&v) + .with_shape(&[30, 10]) + .with_chunks(&[7, 3]); + fw.create_dataset("deflated") + .with_f64_data(&v) + .with_chunks(&[64]) + .with_shuffle() + .with_deflate(4); + fw.create_dataset("grows") + .with_u8_data_owned((0..=255).collect()) + .with_chunks(&[100]) + .with_maxshape(&[u64::MAX]); + fw + }; + for (low, paged) in [ + (LibVer::V18, false), + (LibVer::V110, false), + (LibVer::Latest, true), + ] { + let bytes = build(low, paged).finish().unwrap(); + let mut streamed = Vec::new(); + build(low, paged) + .finish_with(|b: &[u8]| { + streamed.extend_from_slice(b); + Ok::<(), FormatError>(()) + }) + .unwrap(); + assert_eq!(streamed, bytes, "{low:?} paged {paged}"); + } + + let mut calls = 0; + let err = build(LibVer::V110, false) + .finish_with(|_| { + calls += 1; + if calls == 3 { + Err(FormatError::SerializationError("full".into())) + } else { + Ok(()) + } + }) + .unwrap_err(); + assert_eq!(err, FormatError::SerializationError("full".into())); + assert_eq!(calls, 3); + } } diff --git a/crates/clawhdf5-format/src/fixed_array.rs b/crates/clawhdf5-format/src/fixed_array.rs index dbe00ae..b28f60c 100644 --- a/crates/clawhdf5-format/src/fixed_array.rs +++ b/crates/clawhdf5-format/src/fixed_array.rs @@ -155,7 +155,7 @@ pub fn read_fixed_array_chunks( header: &FixedArrayHeader, dataset_dims: &[u64], max_dims: Option<&[u64]>, - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], element_size: u32, offset_size: u8, length_size: u8, @@ -180,7 +180,7 @@ pub fn read_fixed_array_chunks_in( header: &FixedArrayHeader, dataset_dims: &[u64], max_dims: Option<&[u64]>, - chunk_dimensions: &[u32], + chunk_dimensions: &[u64], element_size: u32, offset_size: u8, _length_size: u8, @@ -227,11 +227,11 @@ pub fn read_fixed_array_chunks_in( // The index is laid out over the chunk grid of the *maximum* dimensions // (row-major), so a dataset smaller than its maxshape has gaps. - let dims_u64: Vec = chunk_dimensions.iter().map(|&d| d as u64).collect(); - let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, &dims_u64)?; + let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, chunk_dimensions)?; - let chunk_byte_size: u64 = - chunk_dimensions.iter().map(|&d| d as u64).product::() * element_size as u64; + let chunk_byte_size: u64 = chunk_dimensions + .iter() + .fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d)); let mut chunks = Vec::new(); // `rel` is relative to the data block, whose bytes are in `w`. @@ -670,7 +670,7 @@ mod tests { let header = FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap(); let ds_dims = vec![100u64]; - let chunk_dims = vec![20u32]; + let chunk_dims = vec![20u64]; let chunks = read_fixed_array_chunks( &file_data, &header, @@ -747,7 +747,7 @@ mod tests { let header = FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap(); let ds_dims = vec![60u64]; - let chunk_dims = vec![20u32]; + let chunk_dims = vec![20u64]; let chunks = read_fixed_array_chunks( &file_data, &header, @@ -848,7 +848,7 @@ mod tests { assert_eq!(header.num_elements, 11); let ds_dims = vec![11u64 * 20]; - let chunk_dims = vec![20u32]; + let chunk_dims = vec![20u64]; let chunks = read_fixed_array_chunks( &file_data, &header, diff --git a/crates/clawhdf5-format/src/type_builders.rs b/crates/clawhdf5-format/src/type_builders.rs index ba8c292..f9776b9 100644 --- a/crates/clawhdf5-format/src/type_builders.rs +++ b/crates/clawhdf5-format/src/type_builders.rs @@ -663,11 +663,17 @@ impl DatasetBuilder { } pub fn with_u8_data(&mut self, data: &[u8]) -> &mut Self { + self.with_u8_data_owned(data.to_vec()) + } + + /// [`Self::with_u8_data`] without the copy: the builder takes the + /// vector. For large data this halves the memory a write needs. + pub fn with_u8_data_owned(&mut self, data: Vec) -> &mut Self { self.datatype = Some(make_u8_type()); - self.data = Some(data.to_vec()); if self.shape.is_none() { self.shape = Some(vec![data.len() as u64]); } + self.data = Some(data); self } diff --git a/crates/clawhdf5-py/src/node.rs b/crates/clawhdf5-py/src/node.rs index 35e5dec..9ef9290 100644 --- a/crates/clawhdf5-py/src/node.rs +++ b/crates/clawhdf5-py/src/node.rs @@ -189,12 +189,7 @@ pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Optio { clawhdf5_format::data_layout::DataLayout::Chunked { chunk_dimensions, .. - } if chunk_dimensions.len() >= rank => Some( - chunk_dimensions[..rank] - .iter() - .map(|&d| u64::from(d)) - .collect(), - ), + } if chunk_dimensions.len() >= rank => Some(chunk_dimensions[..rank].to_vec()), _ => None, } } diff --git a/crates/clawhdf5-tools/src/check.rs b/crates/clawhdf5-tools/src/check.rs index 1ba8239..aa7729d 100644 --- a/crates/clawhdf5-tools/src/check.rs +++ b/crates/clawhdf5-tools/src/check.rs @@ -873,9 +873,9 @@ impl Checker<'_> { self.counts.chunk_index_checksummed += 1; } let filtered = matches!(&info.filters, Ok(Some(p)) if !p.filters.is_empty()); - let chunk_bytes = cdims.iter().try_fold(u64::from(dt.type_size()), |a, &d| { - a.checked_mul(u64::from(d)) - }); + let chunk_bytes = cdims + .iter() + .try_fold(u64::from(dt.type_size()), |a, &d| a.checked_mul(d)); let max = ds.max_dimensions.clone(); let mut seen: HashSet> = HashSet::with_capacity(chunks.len().min(1 << 20)); let mut reported = 0usize; @@ -893,7 +893,7 @@ impl Checker<'_> { )); } else { for (i, (&o, &cd)) in c.offsets.iter().zip(cdims).enumerate() { - if o % u64::from(cd) != 0 { + if o % cd != 0 { bad.push(format!( "offset {o} in dimension {i} is not a multiple of the chunk size {cd}" )); diff --git a/crates/clawhdf5-tools/src/ls.rs b/crates/clawhdf5-tools/src/ls.rs index 203cce3..8a466a2 100644 --- a/crates/clawhdf5-tools/src/ls.rs +++ b/crates/clawhdf5-tools/src/ls.rs @@ -292,7 +292,7 @@ impl Ls<'_> { .unwrap_or(0); let bytes = dims .iter() - .try_fold(esize, |a, &d| a.checked_mul(u64::from(d))) + .try_fold(esize, |a, &d| a.checked_mul(d)) .map(|b| b.to_string()) .unwrap_or_else(|| "?".into()); writeln!( diff --git a/crates/clawhdf5/examples/write_one_chunk.rs b/crates/clawhdf5/examples/write_one_chunk.rs new file mode 100644 index 0000000..2761e8b --- /dev/null +++ b/crates/clawhdf5/examples/write_one_chunk.rs @@ -0,0 +1,46 @@ +//! Write one dataset of `u8` stored as a single chunk, for measuring the +//! writer's peak memory (`/usr/bin/time -v`): +//! +//! ```text +//! cargo run --release -p clawhdf5 --example write_one_chunk -- BYTES OUT.h5 [slice|owned] [deflate|lz4|zstd] +//! ``` +//! +//! `slice` passes the data by reference (`with_u8_data`, which copies it), +//! `owned` hands the builder the vector (`with_u8_data_owned`). The data is +//! `i % 251`, one byte per element; with no filter the chunk is stored raw. +fn main() { + let args: Vec = std::env::args().collect(); + let n: u64 = args[1].parse().expect("BYTES"); + let out = &args[2]; + let mode = args.get(3).map_or("slice", String::as_str); + let filter = args.get(4).map(String::as_str); + let data: Vec = (0..n).map(|i| (i % 251) as u8).collect(); + let mut b = clawhdf5::FileBuilder::new(); + let ds = b.create_dataset("x"); + match mode { + "slice" => { + ds.with_u8_data(&data); + } + "owned" => { + ds.with_u8_data_owned(data); + } + m => panic!("mode {m}"), + } + ds.with_chunks(&[n]); + match filter { + None => {} + Some("deflate") => { + ds.with_deflate(1); + } + #[cfg(feature = "lz4")] + Some("lz4") => { + ds.with_lz4(); + } + #[cfg(feature = "zstd")] + Some("zstd") => { + ds.with_zstd(1); + } + Some(f) => panic!("filter {f} (not built in?)"), + } + b.write(out).unwrap(); +} diff --git a/crates/clawhdf5/src/edit/mod.rs b/crates/clawhdf5/src/edit/mod.rs index 169c6d3..a550edc 100644 --- a/crates/clawhdf5/src/edit/mod.rs +++ b/crates/clawhdf5/src/edit/mod.rs @@ -408,10 +408,7 @@ impl<'t> ChunkedEdit<'t> { "chunk rank differs from the dataspace".into(), )); } - let cd: Vec = chunk_dimensions[..rank] - .iter() - .map(|&c| u64::from(c)) - .collect(); + let cd: Vec = chunk_dimensions[..rank].to_vec(); // Chunks of 4 GiB or more (HDF5 2.0 writes them with layout message // version 5) are read, but not rewritten: the editor holds a chunk // it rewrites in memory, compressed and not. Writing values, and a diff --git a/crates/clawhdf5/src/writer.rs b/crates/clawhdf5/src/writer.rs index 9fcf6fc..eb05b97 100644 --- a/crates/clawhdf5/src/writer.rs +++ b/crates/clawhdf5/src/writer.rs @@ -137,10 +137,15 @@ impl FileBuilder { Ok(self.writer.finish()?) } - /// Serialize and write the file to the given path. + /// Serialize and write the file to the given path. The file is written + /// as it is serialized, not assembled in memory first: unfiltered + /// chunks and contiguous data go straight from the datasets' data. pub fn write>(self, path: P) -> Result<(), Error> { - let bytes = self.finish()?; - write_file_atomically(path.as_ref(), &bytes).map_err(Error::Io) + use std::io::Write; + write_file_atomically_with(path.as_ref(), |f| { + self.writer + .finish_with(|bytes| f.write_all(bytes).map_err(Error::Io)) + }) } } @@ -259,8 +264,19 @@ pub fn create_datasets_parallel(specs: Vec) -> Result, Erro /// file or the complete new one — never a truncated mix. `std::fs::write` /// truncates the destination first, so dying mid-write used to destroy the /// existing file. +#[cfg(test)] fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Result<()> { use std::io::Write; + write_file_atomically_with(path, |f| f.write_all(bytes)) +} + +/// [`write_file_atomically`] with the contents written by `write` (into a +/// buffered writer over the temporary file). +fn write_file_atomically_with>( + path: &std::path::Path, + write: impl FnOnce(&mut std::io::BufWriter) -> Result<(), E>, +) -> Result<(), E> { + use std::io::Write; // Same directory as the target, so the rename stays on one filesystem. let mut tmp_name = path @@ -272,11 +288,13 @@ fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Resul tmp_name.push(format!(".tmp-{}", std::process::id())); let tmp_path = path.with_file_name(tmp_name); - let result = (|| { - let mut f = std::fs::File::create(&tmp_path)?; - f.write_all(bytes)?; - f.sync_all()?; - std::fs::rename(&tmp_path, path) + let result = (|| -> Result<(), E> { + let mut f = std::io::BufWriter::with_capacity(1 << 20, std::fs::File::create(&tmp_path)?); + write(&mut f)?; + f.flush()?; + f.get_ref().sync_all()?; + std::fs::rename(&tmp_path, path)?; + Ok(()) })(); if result.is_err() { let _ = std::fs::remove_file(&tmp_path); diff --git a/crates/clawhdf5/tests/fixtures/gen_huge_chunks.py b/crates/clawhdf5/tests/fixtures/gen_huge_chunks.py index f901fd5..6cdf311 100644 --- a/crates/clawhdf5/tests/fixtures/gen_huge_chunks.py +++ b/crates/clawhdf5/tests/fixtures/gen_huge_chunks.py @@ -2,6 +2,12 @@ python gen_huge_chunks.py filtered OUT.h5 # huge_chunks_filtered.h5 python gen_huge_chunks.py unfiltered OUT.h5 # generated at test time + python gen_huge_chunks.py dims OUT.h5 # huge_chunk_dims.h5 + python gen_huge_chunks.py dims-unfiltered OUT.h5 # generated at test time + +`dims` and `dims-unfiltered` (see `make_dims`) hold `u1` datasets whose +chunk dimension is N8 = 2**32 + 7: a chunk dimension of 2**32 or more, +which a layout message of version 5 stores in 5 bytes. libhdf5 2.x writes a chunk of more than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`: "chunk size > 4GB requires @@ -38,7 +44,8 @@ allocated early (see `make`). libhdf5 holds a whole filtered chunk in memory while it writes it, so the `filtered` run needs about 4 GiB of memory and 100 s (tank). -Generated 2026-09-28 with h5py 3.16.0 (libhdf5 2.0.0). +huge_chunks_filtered.h5 was generated 2026-09-28 and huge_chunk_dims.h5 +2026-09-29, both with h5py 3.16.0 (libhdf5 2.0.0). """ import sys @@ -92,12 +99,60 @@ def make(fid, name, index, filtered): dsid.close() +N8 = 2**32 + 7 + + +def make_dims(fid, name, index, filtered): + """A `u1` dataset whose chunks are N8 elements (a chunk dimension of + 2**32 or more). `d[0:10] = 0..9` and `d[N8-10:N8] = 100..109`; where + there is a second chunk (`farray`, `earray`: shape N8 + 10), + `d[N8:N8+10] = 200..209`. + + filtered (`dims`, committed): deflate level 9 twice, fill value 7; + `single` (Single Chunk) and `earray` (Extensible Array, maxshape None). + Writing it holds a 4 GiB chunk: 4.1 GiB peak resident memory and about + a minute (tank, 2026-09-29). + unfiltered (`dims-unfiltered`): fill time "never", sparse; `single` + (allocated early, as in `make`) and `farray` (Fixed Array). + """ + dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE) + if index == "single": + shape, maxshape = (N8,), (N8,) + elif index == "farray": + shape, maxshape = (N8 + 10,), (N8 + 10,) + elif index == "earray": + shape, maxshape = (N8 + 10,), (UNLIM,) + dcpl.set_chunk((N8,)) + if filtered: + dcpl.set_deflate(9) + dcpl.set_deflate(9) + dcpl.set_fill_value(np.array(7, dtype="u1")) + else: + dcpl.set_fill_time(h5py.h5d.FILL_TIME_NEVER) + if index == "single": + dcpl.set_alloc_time(h5py.h5d.ALLOC_TIME_EARLY) + space = h5py.h5s.create_simple(shape, maxshape) + dsid = h5py.h5d.create(fid, name.encode(), h5py.h5t.STD_U8LE, space, dcpl=dcpl) + ds = h5py.Dataset(dsid) + ds[0:10] = np.arange(10, dtype="u1") + ds[N8 - 10 : N8] = np.arange(100, 110, dtype="u1") + if shape[0] > N8: + ds[N8 : N8 + 10] = np.arange(200, 210, dtype="u1") + dsid.close() + + def main(): mode, out = sys.argv[1], sys.argv[2] - filtered = {"filtered": True, "unfiltered": False}[mode] fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS) fapl.set_libver_bounds(h5py.h5f.LIBVER_EARLIEST, h5py.h5f.LIBVER_V200) fid = h5py.h5f.create(out.encode(), h5py.h5f.ACC_TRUNC, fapl=fapl) + if mode in ("dims", "dims-unfiltered"): + filtered = mode == "dims" + for index in ["single", "earray"] if filtered else ["single", "farray"]: + make_dims(fid, index, index, filtered) + fid.close() + return + filtered = {"filtered": True, "unfiltered": False}[mode] indexes = ["single", "farray", "earray", "btree2"] if not filtered: indexes.insert(1, "implicit") diff --git a/crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5 b/crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5 new file mode 100644 index 0000000000000000000000000000000000000000..34d97d5a00c7d994d66ae61a244d564f15c09950 GIT binary patch literal 31199 zcmeHPZA?>F7(TbOSlJX9j8#M})66Av8~bo0fG`$tkP&gZh(cUJsq!Ti3bX=+Ml)c6 zg=L1!=#VYT+@fZH7{G}@B@Ep{wgep_8gz<8ffZ%C)k3>_?uWZ7Vd2-7@LZC6&Uw#A z+vh&-^PbZC-jAXpR?lc-6pn#iyhUt_{fR<_5y~>5{8q+S7()LD zQ3-ZWVf{zPy}vFplCUGi^~F+{Mgrt~O(_~s&ME|=BE#3x54=(M)FPkG2s{(P(5Gf@ z&(IR0P0G$rVmT}F@-1<|3`>S6`89-v7=g(Q!yD$X>(QE6y0HAFZR$n*2%fPD!7g)= zaWm<7ddP+K#WkKFf!sXif&b9gXGpSkTSiijmdxd+v6iXrO7c86LLrmM6x^HQ-UU$?#EG5~8g$fmTU*lp@l?rpNp=02*a(Km1TQAUsE~9{Q59cG?J`B@EH?ccuRK|7 z=rs{_@W1Dkn!8#G8Ikfzu9nL^+{#?JK6-kbPwEGK(^VqFqaq07t(JSp`B*f~VA17; zkU)k*_`yLtHKEf?L)R1t7%z@q@=vc~R3%rfs*uvEN%PWc9Sf54%C)KI>(bubSaEuH z&rA2cDxP(j=a!!7T5btl8y@rOhE7GLvFHbbCN6(`Va?94py06T?vkumD(p2cFIl3q zE?#6hSzPbOzMeTcc;Z;@j`~(dVe{l4D_KiG?!r0lU3Y!P^tp{=^I5~bmg@1I<-;#t z>AYpR_kxp7I$IE2qp`IqMjWQ`k*^J`sr|ai`$^j6ZDG$O#gl8OqzxZw13knH41fW3 z8SnrYfCq>eh?>|rV2@a)H$3ttdrk!#+8iG^norc*-1Yl2^SEE1GJ6D4^{<9$yze)62YELwipaKYc>wY|%{>U;y1`7zOFKJ`~xR#tEJDeUylw(Lpy?4HW` z=dN?#gxS2T{dcdWg)Htq^i^NqWv8KS&5hn(cb#ML$CS0<1x|bC?VRGtgx2^%`U7;y z?_+niG!Dod8NKVRbcChA*8KfN9aGaZv|KYuhqzkp$=P-NvKt2u#A;f2%Z329gRGO& z))$dnV~eyAFKwWQ5MTfdsLOx{zyLfz#6Z-<&H;NQf_)+Z2CNc*2f%=Ci@*b503Pt$ zWB3Am0lolV5Wgxy51SyA z6GqHc8;+ptB1D=PnX2J!d6HyFF#(V7;m_0A#J<^5o-l8>NwGm42mk>f00e-*{|f=6 zIsA~X@22u*|@Kog};6Pu-RVMeLW>F2Wix^4%zBD2R3i}<#@oa zD|Xi%QybNxl~qZbPOV5k;`(0K>060D3*P}b-&KyBRpLES{V)s&N^Z&XztT;;c@ z*w$n)*@j1d8ob%mF;Hq>%JMPsLg=N&xKwLmv vKwLmvKwNk@{}r(jvGMWPIGEfzp3|U+(e)m&Hh5~SU7BYSdYoGCqaJ?)U%R8< literal 0 HcmV?d00001 diff --git a/crates/clawhdf5/tests/huge_chunks_interop.rs b/crates/clawhdf5/tests/huge_chunks_interop.rs index 52885e7..4a3b5cf 100644 --- a/crates/clawhdf5/tests/huge_chunks_interop.rs +++ b/crates/clawhdf5/tests/huge_chunks_interop.rs @@ -81,8 +81,15 @@ fn layout_of( let lm = msg(MessageType::DataLayout); let layout = DataLayout::parse(&lm.data, sb.offset_size, sb.length_size).unwrap(); let space = Dataspace::parse(&msg(MessageType::Dataspace).data, sb.length_size).unwrap(); + // The element size is the layout's last dimension. + let elem = match &layout { + DataLayout::Chunked { + chunk_dimensions, .. + } => *chunk_dimensions.last().unwrap() as usize, + _ => 8, + }; let (chunks, _) = - list_chunks(data, &layout, &space, 8, sb.offset_size, sb.length_size).unwrap(); + list_chunks(data, &layout, &space, elem, sb.offset_size, sb.length_size).unwrap(); (lm.data[0], layout, space, chunks) } @@ -516,3 +523,343 @@ print("ok") } } } + +// ---- Chunk dimensions of 2^32 or more ---- + +/// Elements per chunk of the `u8` datasets below: a chunk dimension of +/// 2^32 or more (4 GiB + 7 bytes), which a version-5 layout message stores +/// in 5 bytes. +const N8: u64 = (1 << 32) + 7; + +fn dims_fixture() -> PathBuf { + Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures/huge_chunk_dims.h5") +} + +/// `u8` values of `sel` in dataset `name` of `file`. +fn sel_u8(file: &File, name: &str, start: u64, count: u64) -> Vec { + let ds = file.dataset(name).unwrap(); + let s = Selection::Hyperslab { + start: vec![start], + stride: vec![1], + count: vec![count], + block: vec![1], + }; + ds.read_selection(&s).unwrap() +} + +fn u8s(range: std::ops::Range) -> Vec { + range.collect() +} + +/// `fixtures/huge_chunk_dims.h5` (libhdf5 2.0.0 through h5py 3.16, +/// `gen_huge_chunks.py dims`): `u8` datasets chunked by N8 = 2^32 + 7, in a +/// Single Chunk and an Extensible Array index. Their layouts parse to the +/// whole dimension (it was refused: chunk dimensions were `u32`) and their +/// chunks list at the right origins. +#[test] +fn huge_chunk_dims_list() { + let data = std::fs::read(dims_fixture()).unwrap(); + for (name, index, shape, origins) in [ + ("single", 1u8, N8, &[0u64][..]), + ("earray", 4, N8 + 10, &[0, N8]), + ] { + let (version, layout, space, mut chunks) = layout_of(&data, name); + assert_eq!(version, 5, "{name}"); + let DataLayout::Chunked { + chunk_dimensions, + chunk_index_type, + .. + } = &layout + else { + panic!("{name}: {layout:?}"); + }; + assert_eq!(chunk_dimensions, &[N8, 1], "{name}"); + assert_eq!(*chunk_index_type, Some(index), "{name}"); + assert_eq!(space.dimensions, [shape], "{name}"); + chunks.sort_by_key(|c| c.offsets[0]); + let got: Vec = chunks.iter().map(|c| c.offsets[0]).collect(); + assert_eq!(got, origins, "{name}"); + } +} + +/// Decoding the fixture's chunks (4 GiB each, one at a time): written +/// values, the fill value (7) around them, and the second chunk. +#[test] +fn huge_chunk_dims_read() { + if !heavy() { + return; + } + let file = File::open(dims_fixture()).unwrap(); + let mut first = u8s(0..10); + first.extend([7, 7]); + let mut last = vec![7, 7]; + last.extend(u8s(100..110)); + for name in ["single", "earray"] { + assert_eq!(sel_u8(&file, name, 0, 12), first, "{name}"); + assert_eq!(sel_u8(&file, name, N8 - 12, 12), last, "{name}"); + } + let mut across = u8s(108..110); + across.extend(u8s(200..210)); + assert_eq!(sel_u8(&file, "earray", N8 - 2, 12), across); +} + +/// Unfiltered chunks of N8 bytes that libhdf5 2.0.0 wrote (a sparse file, +/// `gen_huge_chunks.py dims-unfiltered`), in a Single Chunk and a Fixed +/// Array index: read through a memory map, and through positioned reads +/// that fetch only the rows selected. +#[test] +fn huge_chunk_dims_unfiltered_read() { + if !heavy() { + return; + } + let dir = scratch(); + let Some(path) = generate(dir.path(), "dims-unfiltered") else { + return; + }; + let check = |file: &File| { + let mut first = u8s(0..10); + first.extend([0, 0]); + let mut last = vec![0, 0]; + last.extend(u8s(100..110)); + for name in ["single", "farray"] { + assert_eq!(sel_u8(file, name, 0, 12), first, "{name}"); + assert_eq!(sel_u8(file, name, N8 - 12, 12), last, "{name}"); + } + let mut across = u8s(108..110); + across.extend(u8s(200..210)); + assert_eq!(sel_u8(file, "farray", N8 - 2, 12), across); + }; + check(&File::open(&path).unwrap()); + let file = std::fs::File::open(&path).unwrap(); + let len = file.metadata().unwrap().len(); + let storage = std::sync::Arc::new(Counting { + file, + len, + read: 0.into(), + }); + check(&File::open_storage(storage.clone()).unwrap()); + let read = storage.read.load(std::sync::atomic::Ordering::Relaxed); + assert!(read < 1 << 20, "{read} bytes read"); +} + +/// Run `script` under h5py with `path` as `sys.argv[1]`; skipped without +/// h5py (a failure under `CLAWHDF5_REQUIRE_INTEROP=1`). +fn h5py_check(script: &str, path: &Path) { + if !python_available() { + assert!(!interop_required(), "h5py not available"); + eprintln!("SKIP: python3 with h5py not available"); + return; + } + let out = Command::new(python()) + .arg("-c") + .arg(script) + .arg(path) + .output() + .unwrap(); + assert!( + out.status.success(), + "h5py:\n{}", + String::from_utf8_lossy(&out.stderr) + ); +} + +/// `h5dump` 2.x's output for `args` on `path`, checked to succeed without +/// errors; `None` without `CLAWHDF5_H5DUMP2`. +fn h5dump2_text(args: &[&str], path: &Path) -> Option { + let h5dump = h5dump2()?; + let out = Command::new(&h5dump).args(args).arg(path).output().unwrap(); + let text = format!( + "{}{}", + String::from_utf8_lossy(&out.stdout), + String::from_utf8_lossy(&out.stderr) + ); + assert!(out.status.success(), "h5dump {args:?}:\n{text}"); + assert!(!text.contains("rror"), "h5dump {args:?}:\n{text}"); + Some(text) +} + +/// clawhdf5 writes deflated `u8` chunks of N8 elements (a chunk dimension +/// of 2^32 or more) in a Single Chunk and an Extensible Array index; it, +/// h5py (libhdf5 2.0.0) and h5dump 2.2.0 (layouts only: it cannot inflate +/// chunks over 4 GiB, see known-issues) read them. +#[test] +fn writer_huge_chunk_dims_round_trip() { + if !heavy() { + return; + } + let dir = scratch(); + let path = dir.path().join("dims.h5"); + let values = u8s(0..10); + let mut b = clawhdf5::FileBuilder::new(); + for (name, max) in [("single", N8), ("earray", u64::MAX)] { + b.create_dataset(name) + .with_u8_data(&values) + .with_maxshape(&[max]) + .with_chunks(&[N8]) + .with_deflate(1); + } + b.write(&path).unwrap(); + + let data = std::fs::read(&path).unwrap(); + // Deflate level 1 leaves about 40 MiB of each 4 GiB chunk of zeros. + assert!(data.len() < 256 << 20, "{} bytes", data.len()); + for (name, index) in [("single", 1), ("earray", 4)] { + let (version, layout, _, chunks) = layout_of(&data, name); + assert_eq!(version, 5, "{name}"); + assert!( + matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(t), .. } + if *t == index && chunk_dimensions == &[N8, 1]), + "{name}: {layout:?}" + ); + assert_eq!(chunks.len(), 1, "{name}"); + } + let file = File::open(&path).unwrap(); + for name in ["single", "earray"] { + assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}"); + } + drop(file); + + h5py_check( + r#" +import sys, h5py +with h5py.File(sys.argv[1], "r") as f: + for name in ("single", "earray"): + d = f[name] + assert d.chunks == (2**32 + 7,), (name, d.chunks) + assert list(d[:]) == list(range(10)), (name, d[:]) +"#, + &path, + ); + for name in ["single", "earray"] { + if let Some(text) = h5dump2_text(&["-H", "-p", "-d", name], &path) { + assert!(text.contains("CHUNKED ( 4294967303 )"), "{name}:\n{text}"); + } + } +} + +/// An unfiltered chunk of N8 bytes (a chunk dimension of 2^32 or more) +/// written by clawhdf5 from data it was handed (`with_u8_data_owned`): +/// `FileBuilder::write` streams it to the file without copying it, so the +/// write needs about the data's own 4 GiB. Read back by clawhdf5, h5py +/// (a few elements) and h5dump 2.2.0. +#[test] +fn writer_unfiltered_huge_chunk_round_trip() { + if !heavy() { + return; + } + let dir = scratch(); + let path = dir.path().join("raw.h5"); + let value = |i: u64| (i % 251) as u8; + let mut b = clawhdf5::FileBuilder::new(); + b.create_dataset("single") + .with_u8_data_owned((0..N8).map(value).collect()) + .with_chunks(&[N8]); + b.write(&path).unwrap(); + + let file = File::open(&path).unwrap(); + let data = file.as_bytes(); + let (version, layout, space, chunks) = layout_of(data, "single"); + assert_eq!(version, 5); + assert!( + matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(1), .. } + if chunk_dimensions == &[N8, 1]), + "{layout:?}" + ); + assert_eq!(space.dimensions, [N8]); + assert_eq!(chunks.len(), 1); + assert_eq!(chunks[0].chunk_size, N8); + for start in [0, (1 << 32) - 3, N8 - 6] { + let want: Vec = (start..start + 6).map(value).collect(); + assert_eq!(sel_u8(&file, "single", start, 6), want, "at {start}"); + } + drop(file); + + h5py_check( + r#" +import sys, h5py, numpy as np +n = 2**32 + 7 +with h5py.File(sys.argv[1], "r") as f: + d = f["single"] + assert d.chunks == (n,) and d.shape == (n,) and d.compression is None + for start in (0, 2**32 - 3, n - 6): + want = [i % 251 for i in range(start, start + 6)] + assert list(d[start:start + 6]) == want, (start, d[start:start + 6]) +"#, + &path, + ); + let start = (1u64 << 32) - 3; + if let Some(text) = h5dump2_text( + &["-d", "single", "-s", &start.to_string(), "-c", "6"], + &path, + ) { + let want: Vec = (start..start + 6).map(|i| value(i).to_string()).collect(); + assert!(text.contains(&want.join(", ")), "{text}"); + } +} + +/// LZ4 and Zstandard chunks of 4 GiB or more: written by clawhdf5, read +/// back by it and by h5py (through `hdf5plugin`'s LZ4 and Zstandard +/// filters when installed). +#[cfg(any(feature = "lz4", feature = "zstd"))] +#[test] +fn writer_huge_chunks_lz4_zstd_round_trip() { + if !heavy() { + return; + } + let dir = scratch(); + let path = dir.path().join("codecs.h5"); + let values = u8s(0..10); + let mut b = clawhdf5::FileBuilder::new(); + let mut names = Vec::new(); + #[cfg(feature = "lz4")] + { + b.create_dataset("lz4") + .with_u8_data(&values) + .with_maxshape(&[N8]) + .with_chunks(&[N8]) + .with_lz4(); + names.push("lz4"); + } + #[cfg(feature = "zstd")] + { + b.create_dataset("zstd") + .with_u8_data(&values) + .with_maxshape(&[N8]) + .with_chunks(&[N8]) + .with_zstd(3); + names.push("zstd"); + } + b.write(&path).unwrap(); + let data = std::fs::read(&path).unwrap(); + assert!(data.len() < 256 << 20, "{} bytes", data.len()); + for name in &names { + let (version, _, _, chunks) = layout_of(&data, name); + assert_eq!(version, 5, "{name}"); + assert_eq!(chunks.len(), 1, "{name}"); + } + drop(data); + let file = File::open(&path).unwrap(); + for name in &names { + assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}"); + } + drop(file); + h5py_check( + &format!( + r#" +import sys, h5py +try: + import hdf5plugin +except ImportError: + print("SKIP: no hdf5plugin") + sys.exit(0) +with h5py.File(sys.argv[1], "r") as f: + for name in {names:?}: + d = f[name] + assert d.chunks == (2**32 + 7,), (name, d.chunks) + assert list(d[0:10]) == list(range(10)), (name, d[0:10]) +"#, + names = names + ), + &path, + ); +} -- 2.54.0 From 728ffeef16e6cf38d48e4135081db46c77c481fb Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 20:46:39 -0500 Subject: [PATCH 4/7] docs: chunk dimensions of 2^32 or more and copy-free 4 GiB writes known-issues: two fixed entries (chunk dimensions of 2^32 or more were refused; a 4 GiB unfiltered chunk was held several times when written), the open "Chunks of 4 GiB or more: limits" updated (LZ4/Zstd tested end to end, memory figures of 2026-09-29) and the open-issues table in sync. CHANGELOG (Unreleased) entry with the public API changes and the peak RSS measurements; README capability table. The wasm package test lists the new fixture and expects an Overflow error reading its chunks. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 59 ++++++++++++++++ README.md | 5 +- docs/known-issues.md | 104 +++++++++++++++++++++++------ examples/wasm-viewer/test/test.mjs | 9 +++ 4 files changed, 156 insertions(+), 21 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2d9e329..7831789 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -119,6 +119,65 @@ opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer hard-codes the fixture's length. +### Chunk dimensions of 2^32 or more; 4 GiB chunks written without copies (2026-09-29) +- **Chunk dimensions of 2^32 or more** (libhdf5 2.x writes them with + layout message version 5, up to 8 bytes per dimension; a 4 GiB chunk of + 1-byte elements needs one) are read and written. They were refused + (`InvalidChunkDimensions`). **Public API changes:** + `DataLayout::Chunked::chunk_dimensions` is `Vec` (was `Vec`), + and the `chunk_dimensions` argument of `chunked_read::{collect_chunk_info_checked, + collect_chunk_info_checked_in, generate_implicit_chunks_in_grid}`, + `fixed_array::{read_fixed_array_chunks, read_fixed_array_chunks_in}`, + `extensible_array::{read_extensible_array_chunks, + read_extensible_array_chunks_in}`, and the `chunk_dims` argument of + `chunked_write::serialize_v4_single_chunk_pub`, are `&[u64]` (were + `&[u32]`). A chunk whose size overflows 64 bits is refused when a dataset is read + (`InvalidChunkDimensions`), and the writer refuses a chunk dimension of + 0 up front. A version-3 layout (4-byte dimensions) is never written for + such a chunk: over 4 GiB, it always takes version 5, as in libhdf5. +- **Writing without copies:** `FileWriter` no longer copies chunks into a + buffer per pass and then into the file's buffer. A chunk that is a + contiguous run of the dataset's data (a dataset stored as one chunk of + its shape, blocks of whole rows) is borrowed when unfiltered and + compressed straight from it when filtered, and contiguous datasets are + not copied. New `FileWriter::finish_with(put)` hands the file out piece + by piece; `FileBuilder::write` uses it (through a 1 MiB `BufWriter`, + still to a temporary file renamed into place), so a file is never + assembled in memory. New `DatasetBuilder::with_u8_data_owned(Vec)` + takes the data without the copy `with_u8_data` makes. Peak resident + memory of one unfiltered 1 GiB chunk of `u8` + (`cargo run --release -p clawhdf5 --example write_one_chunk -- + 1073741824 OUT owned|slice`, tank, 2026-09-29): 5.0 GiB before (the + writer of `4260af4`), 1.0 GiB after with `owned`; 6.0 and 2.0 GiB with + `slice`; 4 GiB + 7 bytes `owned`, 4.00 GiB. Same bytes written. Write + speed was not re-measured (the machine was not idle): the A/B of + `FileBuilder::write` and `finish` on small and large files is still to + be run. +- `chunked_write::extract_chunk` (behind `split_into_chunks`) no longer + panics when the data is shorter than the shape; the missing part is + zeros, as it was meant to be. +- Tests: `crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5` (31 KB, + libhdf5 2.0.0 via h5py 3.16, `gen_huge_chunks.py dims`: `u8` chunks of + 2^32 + 7 deflated twice, Single Chunk and Extensible Array), and in + `huge_chunks_interop`: `huge_chunk_dims_list` (always), and with + `CLAWHDF5_HUGE_CHUNKS=1` `huge_chunk_dims_read`, + `huge_chunk_dims_unfiltered_read` (a sparse file h5py writes, Single + Chunk and Fixed Array, mmap and positioned reads), + `writer_huge_chunk_dims_round_trip` (deflated, read back by clawhdf5, + h5py 3.16 and — layouts — h5dump 2.2.0), + `writer_unfiltered_huge_chunk_round_trip` (an unfiltered 4 GiB + 7 byte + chunk: clawhdf5, h5py, and h5dump 2.2.0 printing values past 2^32) and + `writer_huge_chunks_lz4_zstd_round_trip` (with the `lz4`/`zstd` + features; h5py through hdf5plugin 7.1.0). The whole opt-in suite + (`cargo test --release -p clawhdf5 --features lz4,zstd --test + huge_chunks_interop -- --test-threads=1`, h5py and h5dump 2.2.0 + included) passed with a peak of 10.04 GiB resident (tank, 2026-09-29, + commit `0065c6b`). Unit tests: 5-byte dimension encoding, contiguous + chunk detection, `finish_with` against `finish`. The wasm package test + lists the new fixture and gets `FormatError::Overflow` reading it. + `docs/known-issues.md`: two fixed entries; the open "Chunks of 4 GiB or + more: limits" keeps the refused filters, the editor and memory. + ### NetCDF-4: variables' dimensions come from the file (2026-09-28) - `clawhdf5-netcdf4` gave each variable the first unused dimension of equal size (else an anonymous `dim_`), so a variable on an unlimited diff --git a/README.md b/README.md index 9acde11..567e230 100644 --- a/README.md +++ b/README.md @@ -102,7 +102,8 @@ writes 85.2 µs vs 877 µs (10.3x); 64 group creates 130 µs vs 1.37 ms (10.6x); a 512×512 `f32` chunked deflate-6 write 1.44 ms vs 65.0 ms (re-measured 2026-09-23 with the pure-Rust deflate: 1.46 ms vs 51.4 ms, 35x); a 100K `f32` sequential write is a tie. The writer (`FileBuilder`) -assembles a file in memory and writes it once, which is part of that +lays out the whole file before writing it (when these were measured, in +one buffer in memory), which is part of that difference; read the caveats in [BENCHMARKS.md](BENCHMARKS.md#caveats) before quoting these. @@ -116,7 +117,7 @@ Limits and open issues, with dates, are in | **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) | | **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | | **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 | -| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more | +| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more and chunk dimensions of 2^32 or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error) | | **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | | **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) | | **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | diff --git a/docs/known-issues.md b/docs/known-issues.md index 9c17945..2ef9b02 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -17,6 +17,11 @@ Checked against `main` at `9b5803f` on 2026-09-28. |---|---|---| | [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 | | [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 | + + +| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 | +| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 | +| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 | | [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 | | [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | | [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | @@ -241,7 +246,10 @@ Values are correct in every case; this is cost only. Selections other than ## Chunks of 4 GiB or more: limits -**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more +**Status:** open (documented 2026-09-28; updated 2026-09-29, when +[chunk dimensions of 2^32 or more](#chunk-dimensions-of-232-or-more-were-refused) +and [unfiltered writes](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it) +were fixed). HDF5 2.0 writes a chunk of more than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct` in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which also makes a filtered chunk index element store the chunk's size in "size @@ -253,34 +261,40 @@ The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by `crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`). What remains: -- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under - `H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`) - and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec`. - A 4 GiB chunk of 1-byte elements needs one. - **Filters the writer refuses for such a chunk** (`FilterError`, before anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5 filters record sizes or block lengths in 32 bits, or cannot take such a buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are - written; only shuffle + deflate is tested end to end (LZ4's framing and - the 256 MiB limit it had are unit-tested; Zstd is not tested at this - size). -- **Unfiltered chunks this size are written but not tested end to end:** - `FileBuilder` assembles the whole file in memory, several copies of the - chunk (well over the 12 GiB the tests may use). Their layout messages - are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`). + written and tested end to end (`writer_huge_chunks_round_trip`, + `writer_huge_chunk_dims_round_trip`, + `writer_huge_chunks_lz4_zstd_round_trip`: read back by clawhdf5 and + h5py 3.16, LZ4 and Zstd through hdf5plugin 7.1.0). - **`FileEditor`** refuses to rewrite such chunks (see [its limits](#in-place-modification-fileeditor-limits)). - **Memory:** decoding a filtered chunk holds all of it (4 GiB and more), twice when it is shuffled (the inflated and the unshuffled copy, as - libhdf5's filters do too); writing one holds the chunk and its shuffled - copy. A selection decodes only the chunks it touches; one of an - unfiltered chunk reads only the rows it selects, through a memory map or + libhdf5's filters do too). Writing a filtered chunk holds its compressed + copy and, where it is not a contiguous run of the dataset's data (a + chunk past the dataset's edge is zero-padded), the raw chunk; shuffled, + the shuffled copy too. An unfiltered chunk is not copied when it is such + a run (a dataset stored as one chunk of its own shape): writing one + unfiltered chunk of 2^32 + 7 bytes with `with_u8_data_owned` and + `FileBuilder::write` peaked at 4.00 GiB resident, the data itself + (`write_one_chunk 4294967303 OUT owned`, tank, 2026-09-29, commit + `0065c6b`; see [the fix](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)). + `FileWriter::finish` (a `Vec`) holds the data and the file. + A selection decodes only the chunks it touches; one of an unfiltered + chunk reads only the rows it selects, through a memory map or positioned reads (`File::open_storage`): every selection of `unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident - memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop - -- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py - included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection - reads of the double-deflated fixture alone, 4.0 GiB. + memory of `cargo test --release -p clawhdf5 --features lz4,zstd --test + huge_chunks_interop -- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` + (all twelve tests, h5py and h5dump included) was 10.04 GiB (tank, + 2026-09-29, commit `0065c6b`), in `writer_huge_chunks_lz4_zstd_round_trip` + (the other tests, 8.2 GiB or less; the first run of the suite, six + tests, 8.06 GiB on 2026-09-28). Which step of that test holds 10 GiB + (clawhdf5's LZ4 or Zstd write or read, or hdf5plugin's) was not + measured separately. - **32-bit targets (wasm32):** such a chunk cannot be held in memory: reading or writing one is `FormatError::Overflow` ("exceeds the addressable size" / "address space"); the file still opens and lists @@ -694,6 +708,58 @@ browser reads `[re, im]` pairs. `Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a `Complex` arm. Files are unchanged; nothing needs rewriting. +## Chunk dimensions of 2^32 or more were refused + +**Status:** fixed 2026-09-29 (branch `feat/huge-chunk-dims`, PR not yet +opened); affected every release (v2.1.0 to v2.7.0). An error, never wrong +data. Nothing for users to do but upgrade; code that matches on +`DataLayout::Chunked::chunk_dimensions` gets `u64`s now. + +libhdf5 2.x writes a chunk larger than 4 GiB with layout message version +5, whose dimensions take up to 8 bytes each, so a chunk dimension can be +2^32 or more (a 4 GiB chunk of 1-byte elements needs one). +`DataLayout::Chunked::chunk_dimensions` was `Vec`: such a file was +refused when opened for reading (`InvalidChunkDimensions`, "larger than +2^32 - 1"), and the writer refused such a chunk. Chunk dimensions are +`u64` throughout (format crate, facade, editor, `h5rs`, Python bindings); +a version-3 layout, whose dimensions are 4 bytes, is never written for +them (a chunk that large always takes version 5, as in libhdf5). Tests: +`huge_chunks_interop` (`huge_chunk_dims_list` always, over +`fixtures/huge_chunk_dims.h5`, which libhdf5 2.0.0 wrote: `u8` chunks of +2^32 + 7 in a Single Chunk and an Extensible Array index; with +`CLAWHDF5_HUGE_CHUNKS=1`, reads of it and of an unfiltered sparse file +h5py writes, and clawhdf5's own such chunks read back by h5py 3.16 and +h5dump 2.2.0), unit tests of the 5-byte encoding, and the wasm package +test (the file lists; reading such a chunk on wasm32 is +`FormatError::Overflow`). + +## Writing a 4 GiB unfiltered chunk held several copies of it + +**Status:** fixed 2026-09-29 (branch `feat/huge-chunk-dims`, PR not yet +opened); affected every release (v2.1.0 to v2.7.0). Memory only. To write +large data with the least memory, hand the builder the data +(`DatasetBuilder::with_u8_data_owned`) and write with +`FileBuilder::write` (or `FileWriter::finish_with`). + +`FileWriter::finish` copied every chunk into a per-chunk buffer, laid the +chunks out into one buffer per pass (two passes, the first kept), then +copied everything into the file's buffer, and `FileBuilder::write` wrote +that buffer out: about five times a dataset's size for one unfiltered +chunk, six with `with_u8_data` (which copies its argument). Measured +before and after on tank, 2026-09-29, peak resident memory of +`write_one_chunk 1073741824 OUT ` (`crates/clawhdf5/examples`, one +unfiltered 1 GiB chunk of `u8`): `owned` 5.0 GiB before (the writer at +`4260af4`), 1.0 GiB after (commit `0065c6b`); `slice` 6.0 GiB before, +2.0 GiB after. With 4 GiB + 7 bytes and `owned`: 4.00 GiB after. The +files written are byte for byte the same. Now a chunk that is a contiguous +run of the dataset's data is borrowed (and a filtered one compressed +straight from it), the layout refers to chunks instead of copying them, +and `FileBuilder::write` streams the file to disk +(`FileWriter::finish_with`); contiguous datasets are not copied either. +Test: `writer_unfiltered_huge_chunk_round_trip` (with +`CLAWHDF5_HUGE_CHUNKS=1`: 4 GiB + 7 bytes written and read back by +clawhdf5, h5py 3.16 and h5dump 2.2.0; the test process peaked at 4.00 GiB). + ## NetCDF-4: variables' dimensions are guessed from sizes diff --git a/examples/wasm-viewer/test/test.mjs b/examples/wasm-viewer/test/test.mjs index 4a2f26c..eccbb5d 100644 --- a/examples/wasm-viewer/test/test.mjs +++ b/examples/wasm-viewer/test/test.mjs @@ -514,6 +514,15 @@ async function limitTests() { `huge chunk: ${name}`); } hc.free(); + // Likewise chunks whose dimension is 2^32 or more (u8 chunks of 2^32 + 7). + const dims = pkg.open(new Uint8Array(readFileSync(join(import.meta.dirname, "..", "..", "..", + "crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5")))); + eq(dims.list("/").map((e) => e.name), ["earray", "single"], "huge chunk dims: list"); + for (const name of ["single", "earray"]) { + await fails(() => dims.readHyperslab(`/${name}`, [0], [4]), /exceeds this platform.s address space/, + `huge chunk dims: ${name}`); + } + dims.free(); } // A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()` -- 2.54.0 From 549e442aff3ff418ca6636797d3c2b673f3953c2 Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 21:16:01 -0500 Subject: [PATCH 5/7] FileBuilder::write: 64 KiB write buffer, not 1 MiB A 1 MiB BufWriter sits above glibc's 128 KiB mmap threshold; depending on allocator history it was mapped afresh by every write, and its page faults (1.6 M vs 13 K over the criterion write_2d_chunked group) made 1 MiB chunked writes 1.35x-1.83x slower than main. With 64 KiB they match main. Co-Authored-By: Claude Opus 5.5 (1M context) --- crates/clawhdf5/src/writer.rs | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/crates/clawhdf5/src/writer.rs b/crates/clawhdf5/src/writer.rs index eb05b97..7949d13 100644 --- a/crates/clawhdf5/src/writer.rs +++ b/crates/clawhdf5/src/writer.rs @@ -289,7 +289,12 @@ fn write_file_atomically_with>( let tmp_path = path.with_file_name(tmp_name); let result = (|| -> Result<(), E> { - let mut f = std::io::BufWriter::with_capacity(1 << 20, std::fs::File::create(&tmp_path)?); + // 64 KiB, below glibc's 128 KiB mmap threshold: a 1 MiB buffer was + // mapped afresh by each write once the threshold had not risen, and + // its page faults made a 1 MiB chunked write 1.6x slower (criterion + // `write_2d_chunked/512x512` after the smaller cases, 2026-09-29). + // Writes larger than the buffer go straight to the file. + let mut f = std::io::BufWriter::with_capacity(64 << 10, std::fs::File::create(&tmp_path)?); write(&mut f)?; f.flush()?; f.get_ref().sync_all()?; -- 2.54.0 From 7fba2f39969f16261a1046596080f5d783a49ad9 Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 21:33:37 -0500 Subject: [PATCH 6/7] BENCHMARKS: streaming FileBuilder::write A/B (the 1 MiB buffer regression and the fix) Co-Authored-By: Claude Opus 5.5 (1M context) --- BENCHMARKS.md | 36 ++++++++++++++++++++++++++++++++++++ 1 file changed, 36 insertions(+) diff --git a/BENCHMARKS.md b/BENCHMARKS.md index c81101a..f0ec068 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -601,6 +601,42 @@ about 150 per dataset here) against a Fixed Array's few blocks. Writing costs th same; files grow by about 36 bytes per chunk. The default therefore stays the 1.10 format; the 1.8 format is opt-in. +### Streaming `FileBuilder::write` (2026-09-29, tank) + +`FileBuilder::write` now streams the file through a buffered writer +instead of building it in memory first (`feat/huge-chunk-dims`, the change +that removes copies of 4 GiB unfiltered chunks). Measured 2026-09-29 on +tank (AMD Ryzen 7 7800X3D), idle (1-minute load average 1.66 to 1.83 at the +start of each run): `main` at `4260af4` against the stacked branch at +`549e442`, alternating, three full runs each; median (min-max) of +criterion's estimate, ms. + +> **Run:** `cargo bench -p clawhdf5-bench --bench h5bench_write -- --noplot --warm-up-time 1 --measurement-time 3` + +The first candidate used a 1 MiB write buffer and was 1.35x to 1.83x slower +on every 512 x 512 (1 MiB) chunked case when the full suite ran (e.g. +`write_2d_chunked/512x512` 0.454 -> 0.754 ms): the buffer is above glibc's +128 KiB mmap threshold, and after the smaller cases it was mapped afresh on +each write (1 629 368 minor page faults over the `write_2d_chunked` group +against 12 888 on `main`). Run alone, the case showed no difference. With a +64 KiB buffer (`549e442`): + +| benchmark | main | candidate | ratio | +|---|---:|---:|---:| +| write_1d_contiguous/100000 | 0.216 (0.215-0.218) | 0.206 (0.204-0.207) | 0.96x | +| write_2d_chunked/32x32 | 0.062 (0.061-0.063) | 0.061 (0.061-0.064) | 0.99x | +| write_2d_chunked/128x128 | 0.102 (0.100-0.103) | 0.101 (0.099-0.102) | 0.98x | +| write_2d_chunked/512x512 | 0.460 (0.457-0.463) | 0.466 (0.463-0.477) | 1.01x | +| write_2d_chunked_zstd/deflate-6/512x512 | 0.457 (0.450-0.465) | 0.465 (0.458-0.477) | 1.02x | +| write_2d_chunked_zstd/zstd-3/512x512 | 0.376 (0.369-0.381) | 0.387 (0.379-0.390) | 1.03x | +| write_2d_chunked_pcodec/pcodec/512x512 | 0.906 (0.895-0.908) | 0.901 (0.898-0.914) | 0.99x | +| write_multi_dataset/64 | 0.136 (0.135-0.138) | 0.136 (0.135-0.136) | 1.00x | +| write_with_attrs/64 | 0.063 (0.062-0.063) | 0.063 (0.062-0.063) | 1.00x | + +The other 18 cases are within 2%. Minor page faults per full run: 59 602 to +64 486 on both sides. The write path is as fast as before; what changes is +peak memory for large unfiltered chunks (see CHANGELOG, 2026-09-29). + ### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank) Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load -- 2.54.0 From 4cc31111513033a553e320c4b668158d3fb6450f Mon Sep 17 00:00:00 2001 From: osobh Date: Tue, 29 Sep 2026 21:33:48 -0500 Subject: [PATCH 7/7] =?UTF-8?q?docs:=20known-issues=20=E2=80=94=20cite=20P?= =?UTF-8?q?Rs=20#30,=20#31,=20#32?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/known-issues.md | 11 ++++------- 1 file changed, 4 insertions(+), 7 deletions(-) diff --git a/docs/known-issues.md b/docs/known-issues.md index 2ef9b02..b66cec1 100644 --- a/docs/known-issues.md +++ b/docs/known-issues.md @@ -629,7 +629,7 @@ given. ## NetCDF-4: files without dimension scales, types and order differed from netCDF-C -**Status:** fixed 2026-09-29 (branch `fix/netcdf-phony-dims`). Affected +**Status:** fixed 2026-09-29 (#30). Affected every release (v2.1.0 to v2.7.0) and `main` after #27. Wrong metadata only (names, order and types of dimensions and variables); values were read right. Users: the dimensions of a file without dimension scales are now @@ -684,8 +684,7 @@ netCDF4-python-written files with netCDF-C itself, and ## HDF5 2.0 native complex numbers read as a `{r, i}` compound -**Status:** fixed 2026-09-29 (branch `feat/complex-first-class`; PR not -yet opened). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0) +**Status:** fixed 2026-09-29 (#31). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0) and `main` until then. Listed until now under [HDF5 features still unsupported](#hdf5-features-still-unsupported). @@ -710,8 +709,7 @@ browser reads `[re, im]` pairs. ## Chunk dimensions of 2^32 or more were refused -**Status:** fixed 2026-09-29 (branch `feat/huge-chunk-dims`, PR not yet -opened); affected every release (v2.1.0 to v2.7.0). An error, never wrong +**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). An error, never wrong data. Nothing for users to do but upgrade; code that matches on `DataLayout::Chunked::chunk_dimensions` gets `u64`s now. @@ -735,8 +733,7 @@ test (the file lists; reading such a chunk on wasm32 is ## Writing a 4 GiB unfiltered chunk held several copies of it -**Status:** fixed 2026-09-29 (branch `feat/huge-chunk-dims`, PR not yet -opened); affected every release (v2.1.0 to v2.7.0). Memory only. To write +**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). Memory only. To write large data with the least memory, hand the builder the data (`DatasetBuilder::with_u8_data_owned`) and write with `FileBuilder::write` (or `FileWriter::finish_with`). -- 2.54.0