clawhdf5-netcdf4: an unlimited dimension reports its current length
CI / test-arm64 (pull_request) Successful in 1m34s
CI / test (pull_request) Successful in 30m38s

Dimension::size of an unlimited dimension was its dimension scale's
extent, which netCDF-C leaves at 0, so it read 0 for a dimension holding
records. It is now what netCDF-C reports (nc4_find_dim_len): the largest
current extent of the variables using it in any group, found through the
scale's REFERENCE_LIST, a coordinate variable's own extent included; 0
when nothing has been written.

interop_tests::unlimited_dimension_lengths_match_netcdf4_python compares
with netCDF4-python (variables of different lengths, one in a subgroup,
an unwritten dimension, a coordinate variable shorter than another
variable on its dimension, a subgroup's own dimension); before the fix it
got time 0/6, rec 3/5, srec 0/1.

The known-issues entry moves to Fixed; the crate README's warning goes.
A new open entry records a related bug found meanwhile: variables'
dimensions are matched by size, not DIMENSION_LIST.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 21:24:59 -05:00
co-authored by Claude Opus 5.5
parent 3eca5d8334
commit 31efac1ae1
5 changed files with 236 additions and 21 deletions
+15
View File
@@ -15,6 +15,21 @@
after. The agent store's lock file unlocks on drop the same way (its
250 ms retry on open had hidden the race). `docs/known-issues.md`.
### NetCDF-4: unlimited dimensions report their length (2026-09-28)
- `clawhdf5-netcdf4`'s `Dimension::size` of an unlimited dimension was the
extent of its dimension scale, which netCDF-C leaves at 0 (or, for a
coordinate variable, that variable's own length), so it read 0 for a
dimension holding records. It is now what netCDF-C reports
(`nc4_find_dim_len`): the largest current extent of the variables using
the dimension in any group (the scale's `REFERENCE_LIST`), a coordinate
variable included; 0 when nothing has been written. Test:
`interop_tests::unlimited_dimension_lengths_match_netcdf4_python`
(variables of different lengths, one in a subgroup, an unwritten
dimension, a coordinate variable shorter than another variable on its
dimension, a subgroup's own dimension), compared with netCDF4-python.
Affected v2.1.0 to v2.7.0. `docs/known-issues.md` also gains an open
entry found meanwhile: variables' dimensions are matched by size.
### `ObjectHeader::parse` back at its pre-M2/M3 speed (2026-09-27)
- Parsing a version-1 object header was 4% slower than before range-read
M2/M3 (`docs/known-issues.md`). The cause was the call to the per-chunk
+1 -1
View File
@@ -39,7 +39,7 @@ let values: Vec<f64> = temp.read_f64()?;
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is wrongly 0 when it holds records; use the variables' shapes — [known issue](../../docs/known-issues.md#netcdf-4-an-unlimited-dimension-reports-size-0)) |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable |
+91 -6
View File
@@ -2,7 +2,8 @@
//!
//! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE
//! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited).
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited); their length
//! is the largest extent of the variables attached to them.
use std::collections::HashMap;
@@ -23,9 +24,10 @@ pub struct Dimension {
/// Extract dimensions from an HDF5 group (root or subgroup).
///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. The dimension
/// size is the dataset's first (and typically only) shape extent. Unlimited dimensions
/// have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace.
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// dimension's size is the dataset's first (and typically only) shape extent.
/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`.
pub(crate) fn extract_dimensions(
file: &clawhdf5::File,
group: &clawhdf5::Group<'_>,
@@ -44,9 +46,12 @@ pub(crate) fn extract_dimensions(
}
let shape = ds.shape()?;
let size = shape.first().copied().unwrap_or(0);
let is_unlimited = check_unlimited(file, group, ds_name);
let size = if is_unlimited {
unlimited_len(file, &attrs, &shape)
} else {
shape.first().copied().unwrap_or(0)
};
let dimid = get_dimid(&attrs);
@@ -86,6 +91,86 @@ pub(crate) fn extract_dimensions(
Ok(dims)
}
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that
/// is only a dimension, not also a (coordinate) variable.
const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable";
/// The current length of an unlimited dimension, as netCDF-C reports it
/// (`NC4_inq_dim` → `nc4_find_dim_len`): the largest current extent, along
/// the dimension, of the variables that use it, in any group; 0 when none
/// has been written. netCDF-C does not extend a dimension scale that is not
/// also a variable, so such a scale's own extent (0) is not counted; a
/// coordinate variable's is. The variables are the scale's attachments,
/// listed with the axis they use in its `REFERENCE_LIST` attribute (the
/// mirror of each variable's `DIMENSION_LIST`). Attachments that cannot be
/// read are skipped; without a readable `REFERENCE_LIST` the length is the
/// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0);
let is_variable = !matches!(
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
);
let Some(refs) = reference_list(file, attrs) else {
return own;
};
refs.into_iter()
.filter_map(|(address, axis)| {
let shape = file.dataset_at(address).ok()?.shape().ok()?;
shape.get(usize::try_from(axis).ok()?).copied()
})
.chain(is_variable.then_some(own))
.max()
.unwrap_or(0)
}
/// The `(dataset address, axis)` pairs of a dimension scale's
/// `REFERENCE_LIST` attribute (HDF5 dimension scales: a compound of an
/// object reference `dataset` and an integer `dimension`), or `None` when it
/// is missing or not in that form.
fn reference_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<(u64, u64)>> {
use clawhdf5_format::data_read::{read_compound_field, read_object_references};
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("REFERENCE_LIST") else {
return None;
};
let dataset = read_compound_field(data, datatype, "dataset").ok()?;
let addresses = read_object_references(
&dataset.raw_data,
&dataset.datatype,
file.superblock().offset_size,
)
.ok()?;
let dimension = read_compound_field(data, datatype, "dimension").ok()?;
let Datatype::FixedPoint {
size, byte_order, ..
} = dimension.datatype
else {
return None;
};
let size = usize::try_from(size).ok().filter(|s| (1..=8).contains(s))?;
let axes = dimension.raw_data.chunks_exact(size).map(|b| {
let mut v = [0u8; 8];
match byte_order {
DatatypeByteOrder::BigEndian => {
v[8 - size..].copy_from_slice(b);
u64::from_be_bytes(v)
}
_ => {
v[..size].copy_from_slice(b);
u64::from_le_bytes(v)
}
}
});
if axes.len() != addresses.len() {
return None;
}
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
}
/// Check if a dataset's attributes mark it as a dimension scale.
fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
if let Some(AttrValue::String(class)) = attrs.get("CLASS") {
@@ -377,3 +377,87 @@ ds.close()
let names = file.variable("name").unwrap().read_string().unwrap();
assert_eq!(names, vec!["Oslo", "", "São Paulo", "x"]);
}
// ===========================================================================
// Unlimited dimensions: the length netCDF-C reports
// ===========================================================================
/// An unlimited dimension's length is the largest extent of the variables
/// using it, in any group (netCDF-C's `nc4_find_dim_len`), not its dimension
/// scale's extent (which netCDF-C leaves at 0): variables of different
/// lengths, one in a subgroup, a dimension no variable has written, a
/// coordinate variable, a subgroup's own unlimited dimension. Compared with
/// what netCDF4-python reports for the same file.
#[test]
fn unlimited_dimension_lengths_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("unlimited.nc");
let script = format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w", format="NETCDF4") as f:
f.createDimension("time", None)
f.createDimension("empty", None)
f.createDimension("rec", None)
f.createDimension("x", 3)
f.createVariable("t", "f8", ("time", "x"))[0:2, :] = np.ones((2, 3))
f.createVariable("a", "i4", ("time",))[0:4] = np.arange(4)
f.createVariable("e", "i4", ("empty",))
f.createVariable("rec", "f4", ("rec",))[0:3] = [1, 2, 3]
f.createVariable("r", "f4", ("x", "rec"))[:, 0:5] = np.ones((3, 5))
g = f.createGroup("sub")
g.createVariable("c", "i4", ("time",))[0:6] = np.arange(6)
g.createDimension("srec", None)
g.createVariable("s", "i4", ("srec", "x"))[0:1, :] = np.ones((1, 3))
with nc.Dataset({path:?}) as f:
for grp in (f, f.groups["sub"]):
for name, d in grp.dimensions.items():
print(grp.path, name, len(d), d.isunlimited())
"#,
path = path.display().to_string()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.expect("failed to run python3");
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let expected: Vec<String> = String::from_utf8(out.stdout)
.unwrap()
.lines()
.map(str::to_string)
.collect();
// time: t has 2 records, a 4 and sub/c 6; rec: the coordinate variable
// has 3, r 5; empty: nothing written.
for want in [
"/ time 6 True",
"/ empty 0 True",
"/ rec 5 True",
"/ x 3 False",
"/sub srec 1 True",
] {
assert!(
expected.iter().any(|l| l == want),
"netCDF4 reports {expected:?}"
);
}
let file = NetCDF4File::open(&path).unwrap();
let sub = file.group("sub").unwrap();
let mut got = Vec::new();
for (path, dims) in [
("/", file.dimensions().unwrap()),
("/sub", sub.dimensions().unwrap()),
] {
for d in dims {
let unlimited = if d.is_unlimited { "True" } else { "False" };
got.push(format!("{path} {} {} {unlimited}", d.name, d.size));
}
}
assert_eq!(got, expected);
}
+45 -14
View File
@@ -15,7 +15,7 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since |
|---|---|---|
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
| [Selection reads that decode more than the selection](#selection-reads-that-decode-more-than-the-selection) | speed only | 2026-09-26 |
@@ -401,20 +401,24 @@ always expose it. A fix belongs in `conformance/ref_bugs.py` (more or more
varied reads for this object) or in documenting the file as a known
refusal; neither is done.
## NetCDF-4: an unlimited dimension reports size 0
## NetCDF-4: variables' dimensions are guessed from sizes
**Status:** open (found 2026-09-28 while verifying the README refresh).
`clawhdf5-netcdf4`'s `NetCDF4File::dimensions()` reports an unlimited
dimension's `size` as 0 when variables along it hold records. Reproducer: with
netCDF4-python, create dimension `time` (unlimited) and `x` (3), a variable
`t(time, x)`, and write 2 records; netCDF4 reports `time` = 2 and `t` shape
(2, 3). clawhdf5-netcdf4 reports `dim time size 0 unlimited true`, while
`variable("t").shape()` correctly gives `[2, 3]`. In NetCDF-4 an unlimited
dimension's length is the largest extent of the variables that use it (its
dimension-scale dataset is not extended by netCDF-C), and the size is read
from the dimension scale instead. Variable shapes and values are correct;
only `Dimension::size` of unlimited dimensions is wrong. Workaround: use the
variables' shapes.
**Status:** open (found 2026-09-28 while fixing unlimited dimension sizes).
`clawhdf5-netcdf4` gives each variable the dimensions it finds by size
(`match_dimensions_to_variable`: the first unused dimension of equal size,
else an anonymous `dim_<n>`), not the ones its `DIMENSION_LIST` names, and
`variables()` also lists the dimension scales that are only dimensions
(netCDF-C hides them). With netCDF4-python: unlimited `time` and `empty`,
`x` (3), `a(time)` with 2 records, `b(time, x)` with 5, `e(empty)` — netCDF4
reports `time` = 5, `a` on `time`, and variables `a`, `b`, `e` and a `c(x)`;
clawhdf5-netcdf4 reports `time` = 5 (right) but `a` on `dim_2`, and also
variables `time` and `empty` (the scales, put on `empty`). Two dimensions of
one size can be swapped the same way. Related: netCDF4 gives `a` the shape
(5,) (a variable along an unlimited dimension has the dimension's length,
unwritten records read as fill); `Variable::shape` is the HDF5 extent,
`[2]`, and reads return those 2 values. `dimensions()` is right.
Workaround: read the variable's `_Netcdf4Coordinates` attribute (dimension
ids, matching each scale's `_Netcdf4Dimid`).
## The Node.js package (`packages/clawhdf5-node`) does not work
@@ -477,6 +481,33 @@ helped: it is inherited across `fork` the same way, and on Linux it does
not conflict with `flock`, so libhdf5 (h5py), which locks with `flock`,
would no longer be refused while an editor holds the file.
## NetCDF-4: an unlimited dimension reports size 0
**Status:** fixed 2026-09-28 (`fix/netcdf-unlimited-dim-size`). Affected
every release (v2.1.0 to v2.7.0). Wrong metadata only: variable values and
HDF5 extents were always right. Users who worked around it with the variables'
shapes can use `Dimension::size` again.
Found 2026-09-28 while verifying the README refresh.
`clawhdf5-netcdf4`'s `NetCDF4File::dimensions()` reported an unlimited
dimension's `size` as 0 when variables along it hold records. Reproducer: with
netCDF4-python, create dimension `time` (unlimited) and `x` (3), a variable
`t(time, x)`, and write 2 records; netCDF4 reports `time` = 2 and `t` shape
(2, 3). clawhdf5-netcdf4 reported `dim time size 0 unlimited true`, while
`variable("t").shape()` correctly gave `[2, 3]`. The size was read from the
dimension scale's extent, which netCDF-C does not extend (it stays 0; a
coordinate variable is extended only by its own writes).
The size is now what netCDF-C reports (`NC4_inq_dim` → `nc4_find_dim_len`):
the largest current extent, along the dimension, of the variables that use
it in any group, found through the scale's `REFERENCE_LIST`, plus a
coordinate variable's own extent; 0 when nothing has been written.
`interop_tests::unlimited_dimension_lengths_match_netcdf4_python` compares
with netCDF4-python for variables of different lengths (one in a subgroup),
an unwritten dimension, a coordinate variable shorter than a variable on its
dimension, and a subgroup's own unlimited dimension; before the fix it got
`time` 0 (want 6), `rec` 3 (want 5), `srec` 0 (want 1).
## `ObjectHeader::parse` 4% slower after range-read M2/M3
**Status:** fixed 2026-09-27 (`96086ad`, PR #21), before any release.