clawhdf5-netcdf4: an unlimited dimension reports its current length
CI / test-arm64 (pull_request) Successful in 1m34s
CI / test (pull_request) Successful in 30m38s

Dimension::size of an unlimited dimension was its dimension scale's
extent, which netCDF-C leaves at 0, so it read 0 for a dimension holding
records. It is now what netCDF-C reports (nc4_find_dim_len): the largest
current extent of the variables using it in any group, found through the
scale's REFERENCE_LIST, a coordinate variable's own extent included; 0
when nothing has been written.

interop_tests::unlimited_dimension_lengths_match_netcdf4_python compares
with netCDF4-python (variables of different lengths, one in a subgroup,
an unwritten dimension, a coordinate variable shorter than another
variable on its dimension, a subgroup's own dimension); before the fix it
got time 0/6, rec 3/5, srec 0/1.

The known-issues entry moves to Fixed; the crate README's warning goes.
A new open entry records a related bug found meanwhile: variables'
dimensions are matched by size, not DIMENSION_LIST.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 21:24:59 -05:00
co-authored by Claude Opus 5.5
parent 3eca5d8334
commit 31efac1ae1
5 changed files with 236 additions and 21 deletions
+91 -6
View File
@@ -2,7 +2,8 @@
//!
//! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE
//! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited).
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited); their length
//! is the largest extent of the variables attached to them.
use std::collections::HashMap;
@@ -23,9 +24,10 @@ pub struct Dimension {
/// Extract dimensions from an HDF5 group (root or subgroup).
///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. The dimension
/// size is the dataset's first (and typically only) shape extent. Unlimited dimensions
/// have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace.
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// dimension's size is the dataset's first (and typically only) shape extent.
/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`.
pub(crate) fn extract_dimensions(
file: &clawhdf5::File,
group: &clawhdf5::Group<'_>,
@@ -44,9 +46,12 @@ pub(crate) fn extract_dimensions(
}
let shape = ds.shape()?;
let size = shape.first().copied().unwrap_or(0);
let is_unlimited = check_unlimited(file, group, ds_name);
let size = if is_unlimited {
unlimited_len(file, &attrs, &shape)
} else {
shape.first().copied().unwrap_or(0)
};
let dimid = get_dimid(&attrs);
@@ -86,6 +91,86 @@ pub(crate) fn extract_dimensions(
Ok(dims)
}
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that
/// is only a dimension, not also a (coordinate) variable.
const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable";
/// The current length of an unlimited dimension, as netCDF-C reports it
/// (`NC4_inq_dim` → `nc4_find_dim_len`): the largest current extent, along
/// the dimension, of the variables that use it, in any group; 0 when none
/// has been written. netCDF-C does not extend a dimension scale that is not
/// also a variable, so such a scale's own extent (0) is not counted; a
/// coordinate variable's is. The variables are the scale's attachments,
/// listed with the axis they use in its `REFERENCE_LIST` attribute (the
/// mirror of each variable's `DIMENSION_LIST`). Attachments that cannot be
/// read are skipped; without a readable `REFERENCE_LIST` the length is the
/// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0);
let is_variable = !matches!(
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
);
let Some(refs) = reference_list(file, attrs) else {
return own;
};
refs.into_iter()
.filter_map(|(address, axis)| {
let shape = file.dataset_at(address).ok()?.shape().ok()?;
shape.get(usize::try_from(axis).ok()?).copied()
})
.chain(is_variable.then_some(own))
.max()
.unwrap_or(0)
}
/// The `(dataset address, axis)` pairs of a dimension scale's
/// `REFERENCE_LIST` attribute (HDF5 dimension scales: a compound of an
/// object reference `dataset` and an integer `dimension`), or `None` when it
/// is missing or not in that form.
fn reference_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<(u64, u64)>> {
use clawhdf5_format::data_read::{read_compound_field, read_object_references};
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("REFERENCE_LIST") else {
return None;
};
let dataset = read_compound_field(data, datatype, "dataset").ok()?;
let addresses = read_object_references(
&dataset.raw_data,
&dataset.datatype,
file.superblock().offset_size,
)
.ok()?;
let dimension = read_compound_field(data, datatype, "dimension").ok()?;
let Datatype::FixedPoint {
size, byte_order, ..
} = dimension.datatype
else {
return None;
};
let size = usize::try_from(size).ok().filter(|s| (1..=8).contains(s))?;
let axes = dimension.raw_data.chunks_exact(size).map(|b| {
let mut v = [0u8; 8];
match byte_order {
DatatypeByteOrder::BigEndian => {
v[8 - size..].copy_from_slice(b);
u64::from_be_bytes(v)
}
_ => {
v[..size].copy_from_slice(b);
u64::from_le_bytes(v)
}
}
});
if axes.len() != addresses.len() {
return None;
}
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
}
/// Check if a dataset's attributes mark it as a dimension scale.
fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
if let Some(AttrValue::String(class)) = attrs.get("CLASS") {