Merge pull request 'clawhdf5-netcdf4: unlimited dimensions report their length' (#24) from fix/netcdf-unlimited-dim-size into main
CI / test-arm64 (push) Successful in 1m34s
CI / test (push) Successful in 23m33s

Reviewed-on: #24
This commit was merged in pull request #24.
This commit is contained in:
2026-09-29 03:10:56 +00:00
5 changed files with 236 additions and 21 deletions
+15
View File
@@ -15,6 +15,21 @@
after. The agent store's lock file unlocks on drop the same way (its after. The agent store's lock file unlocks on drop the same way (its
250 ms retry on open had hidden the race). `docs/known-issues.md`. 250 ms retry on open had hidden the race). `docs/known-issues.md`.
### NetCDF-4: unlimited dimensions report their length (2026-09-28)
- `clawhdf5-netcdf4`'s `Dimension::size` of an unlimited dimension was the
extent of its dimension scale, which netCDF-C leaves at 0 (or, for a
coordinate variable, that variable's own length), so it read 0 for a
dimension holding records. It is now what netCDF-C reports
(`nc4_find_dim_len`): the largest current extent of the variables using
the dimension in any group (the scale's `REFERENCE_LIST`), a coordinate
variable included; 0 when nothing has been written. Test:
`interop_tests::unlimited_dimension_lengths_match_netcdf4_python`
(variables of different lengths, one in a subgroup, an unwritten
dimension, a coordinate variable shorter than another variable on its
dimension, a subgroup's own dimension), compared with netCDF4-python.
Affected v2.1.0 to v2.7.0. `docs/known-issues.md` also gains an open
entry found meanwhile: variables' dimensions are matched by size.
### `ObjectHeader::parse` back at its pre-M2/M3 speed (2026-09-27) ### `ObjectHeader::parse` back at its pre-M2/M3 speed (2026-09-27)
- Parsing a version-1 object header was 4% slower than before range-read - Parsing a version-1 object header was 4% slower than before range-read
M2/M3 (`docs/known-issues.md`). The cause was the call to the per-chunk M2/M3 (`docs/known-issues.md`). The cause was the call to the per-chunk
+1 -1
View File
@@ -39,7 +39,7 @@ let values: Vec<f64> = temp.read_f64()?;
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | | `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) | | `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` | | `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is wrongly 0 when it holds records; use the variables' shapes — [known issue](../../docs/known-issues.md#netcdf-4-an-unlimited-dimension-reports-size-0)) | | `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` | | `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable | | `NcType` | the NetCDF type of a variable |
+91 -6
View File
@@ -2,7 +2,8 @@
//! //!
//! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE //! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE
//! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via //! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited). //! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited); their length
//! is the largest extent of the variables attached to them.
use std::collections::HashMap; use std::collections::HashMap;
@@ -23,9 +24,10 @@ pub struct Dimension {
/// Extract dimensions from an HDF5 group (root or subgroup). /// Extract dimensions from an HDF5 group (root or subgroup).
/// ///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. The dimension /// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// size is the dataset's first (and typically only) shape extent. Unlimited dimensions /// dimension's size is the dataset's first (and typically only) shape extent.
/// have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace. /// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`.
pub(crate) fn extract_dimensions( pub(crate) fn extract_dimensions(
file: &clawhdf5::File, file: &clawhdf5::File,
group: &clawhdf5::Group<'_>, group: &clawhdf5::Group<'_>,
@@ -44,9 +46,12 @@ pub(crate) fn extract_dimensions(
} }
let shape = ds.shape()?; let shape = ds.shape()?;
let size = shape.first().copied().unwrap_or(0);
let is_unlimited = check_unlimited(file, group, ds_name); let is_unlimited = check_unlimited(file, group, ds_name);
let size = if is_unlimited {
unlimited_len(file, &attrs, &shape)
} else {
shape.first().copied().unwrap_or(0)
};
let dimid = get_dimid(&attrs); let dimid = get_dimid(&attrs);
@@ -86,6 +91,86 @@ pub(crate) fn extract_dimensions(
Ok(dims) Ok(dims)
} }
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that
/// is only a dimension, not also a (coordinate) variable.
const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable";
/// The current length of an unlimited dimension, as netCDF-C reports it
/// (`NC4_inq_dim` → `nc4_find_dim_len`): the largest current extent, along
/// the dimension, of the variables that use it, in any group; 0 when none
/// has been written. netCDF-C does not extend a dimension scale that is not
/// also a variable, so such a scale's own extent (0) is not counted; a
/// coordinate variable's is. The variables are the scale's attachments,
/// listed with the axis they use in its `REFERENCE_LIST` attribute (the
/// mirror of each variable's `DIMENSION_LIST`). Attachments that cannot be
/// read are skipped; without a readable `REFERENCE_LIST` the length is the
/// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0);
let is_variable = !matches!(
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
);
let Some(refs) = reference_list(file, attrs) else {
return own;
};
refs.into_iter()
.filter_map(|(address, axis)| {
let shape = file.dataset_at(address).ok()?.shape().ok()?;
shape.get(usize::try_from(axis).ok()?).copied()
})
.chain(is_variable.then_some(own))
.max()
.unwrap_or(0)
}
/// The `(dataset address, axis)` pairs of a dimension scale's
/// `REFERENCE_LIST` attribute (HDF5 dimension scales: a compound of an
/// object reference `dataset` and an integer `dimension`), or `None` when it
/// is missing or not in that form.
fn reference_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<(u64, u64)>> {
use clawhdf5_format::data_read::{read_compound_field, read_object_references};
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("REFERENCE_LIST") else {
return None;
};
let dataset = read_compound_field(data, datatype, "dataset").ok()?;
let addresses = read_object_references(
&dataset.raw_data,
&dataset.datatype,
file.superblock().offset_size,
)
.ok()?;
let dimension = read_compound_field(data, datatype, "dimension").ok()?;
let Datatype::FixedPoint {
size, byte_order, ..
} = dimension.datatype
else {
return None;
};
let size = usize::try_from(size).ok().filter(|s| (1..=8).contains(s))?;
let axes = dimension.raw_data.chunks_exact(size).map(|b| {
let mut v = [0u8; 8];
match byte_order {
DatatypeByteOrder::BigEndian => {
v[8 - size..].copy_from_slice(b);
u64::from_be_bytes(v)
}
_ => {
v[..size].copy_from_slice(b);
u64::from_le_bytes(v)
}
}
});
if axes.len() != addresses.len() {
return None;
}
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
}
/// Check if a dataset's attributes mark it as a dimension scale. /// Check if a dataset's attributes mark it as a dimension scale.
fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool { fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
if let Some(AttrValue::String(class)) = attrs.get("CLASS") { if let Some(AttrValue::String(class)) = attrs.get("CLASS") {
@@ -377,3 +377,87 @@ ds.close()
let names = file.variable("name").unwrap().read_string().unwrap(); let names = file.variable("name").unwrap().read_string().unwrap();
assert_eq!(names, vec!["Oslo", "", "São Paulo", "x"]); assert_eq!(names, vec!["Oslo", "", "São Paulo", "x"]);
} }
// ===========================================================================
// Unlimited dimensions: the length netCDF-C reports
// ===========================================================================
/// An unlimited dimension's length is the largest extent of the variables
/// using it, in any group (netCDF-C's `nc4_find_dim_len`), not its dimension
/// scale's extent (which netCDF-C leaves at 0): variables of different
/// lengths, one in a subgroup, a dimension no variable has written, a
/// coordinate variable, a subgroup's own unlimited dimension. Compared with
/// what netCDF4-python reports for the same file.
#[test]
fn unlimited_dimension_lengths_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("unlimited.nc");
let script = format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w", format="NETCDF4") as f:
f.createDimension("time", None)
f.createDimension("empty", None)
f.createDimension("rec", None)
f.createDimension("x", 3)
f.createVariable("t", "f8", ("time", "x"))[0:2, :] = np.ones((2, 3))
f.createVariable("a", "i4", ("time",))[0:4] = np.arange(4)
f.createVariable("e", "i4", ("empty",))
f.createVariable("rec", "f4", ("rec",))[0:3] = [1, 2, 3]
f.createVariable("r", "f4", ("x", "rec"))[:, 0:5] = np.ones((3, 5))
g = f.createGroup("sub")
g.createVariable("c", "i4", ("time",))[0:6] = np.arange(6)
g.createDimension("srec", None)
g.createVariable("s", "i4", ("srec", "x"))[0:1, :] = np.ones((1, 3))
with nc.Dataset({path:?}) as f:
for grp in (f, f.groups["sub"]):
for name, d in grp.dimensions.items():
print(grp.path, name, len(d), d.isunlimited())
"#,
path = path.display().to_string()
);
let out = Command::new(python())
.args(["-c", &script])
.output()
.expect("failed to run python3");
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let expected: Vec<String> = String::from_utf8(out.stdout)
.unwrap()
.lines()
.map(str::to_string)
.collect();
// time: t has 2 records, a 4 and sub/c 6; rec: the coordinate variable
// has 3, r 5; empty: nothing written.
for want in [
"/ time 6 True",
"/ empty 0 True",
"/ rec 5 True",
"/ x 3 False",
"/sub srec 1 True",
] {
assert!(
expected.iter().any(|l| l == want),
"netCDF4 reports {expected:?}"
);
}
let file = NetCDF4File::open(&path).unwrap();
let sub = file.group("sub").unwrap();
let mut got = Vec::new();
for (path, dims) in [
("/", file.dimensions().unwrap()),
("/sub", sub.dimensions().unwrap()),
] {
for d in dims {
let unlimited = if d.is_unlimited { "True" } else { "False" };
got.push(format!("{path} {} {} {unlimited}", d.name, d.size));
}
}
assert_eq!(got, expected);
}
+45 -14
View File
@@ -15,7 +15,7 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since | | Issue | Kind | Since |
|---|---|---| |---|---|---|
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 | | [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | | [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | | [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
| [Selection reads that decode more than the selection](#selection-reads-that-decode-more-than-the-selection) | speed only | 2026-09-26 | | [Selection reads that decode more than the selection](#selection-reads-that-decode-more-than-the-selection) | speed only | 2026-09-26 |
@@ -401,20 +401,24 @@ always expose it. A fix belongs in `conformance/ref_bugs.py` (more or more
varied reads for this object) or in documenting the file as a known varied reads for this object) or in documenting the file as a known
refusal; neither is done. refusal; neither is done.
## NetCDF-4: an unlimited dimension reports size 0 ## NetCDF-4: variables' dimensions are guessed from sizes
**Status:** open (found 2026-09-28 while verifying the README refresh). **Status:** open (found 2026-09-28 while fixing unlimited dimension sizes).
`clawhdf5-netcdf4`'s `NetCDF4File::dimensions()` reports an unlimited `clawhdf5-netcdf4` gives each variable the dimensions it finds by size
dimension's `size` as 0 when variables along it hold records. Reproducer: with (`match_dimensions_to_variable`: the first unused dimension of equal size,
netCDF4-python, create dimension `time` (unlimited) and `x` (3), a variable else an anonymous `dim_<n>`), not the ones its `DIMENSION_LIST` names, and
`t(time, x)`, and write 2 records; netCDF4 reports `time` = 2 and `t` shape `variables()` also lists the dimension scales that are only dimensions
(2, 3). clawhdf5-netcdf4 reports `dim time size 0 unlimited true`, while (netCDF-C hides them). With netCDF4-python: unlimited `time` and `empty`,
`variable("t").shape()` correctly gives `[2, 3]`. In NetCDF-4 an unlimited `x` (3), `a(time)` with 2 records, `b(time, x)` with 5, `e(empty)` — netCDF4
dimension's length is the largest extent of the variables that use it (its reports `time` = 5, `a` on `time`, and variables `a`, `b`, `e` and a `c(x)`;
dimension-scale dataset is not extended by netCDF-C), and the size is read clawhdf5-netcdf4 reports `time` = 5 (right) but `a` on `dim_2`, and also
from the dimension scale instead. Variable shapes and values are correct; variables `time` and `empty` (the scales, put on `empty`). Two dimensions of
only `Dimension::size` of unlimited dimensions is wrong. Workaround: use the one size can be swapped the same way. Related: netCDF4 gives `a` the shape
variables' shapes. (5,) (a variable along an unlimited dimension has the dimension's length,
unwritten records read as fill); `Variable::shape` is the HDF5 extent,
`[2]`, and reads return those 2 values. `dimensions()` is right.
Workaround: read the variable's `_Netcdf4Coordinates` attribute (dimension
ids, matching each scale's `_Netcdf4Dimid`).
## The Node.js package (`packages/clawhdf5-node`) does not work ## The Node.js package (`packages/clawhdf5-node`) does not work
@@ -477,6 +481,33 @@ helped: it is inherited across `fork` the same way, and on Linux it does
not conflict with `flock`, so libhdf5 (h5py), which locks with `flock`, not conflict with `flock`, so libhdf5 (h5py), which locks with `flock`,
would no longer be refused while an editor holds the file. would no longer be refused while an editor holds the file.
## NetCDF-4: an unlimited dimension reports size 0
**Status:** fixed 2026-09-28 (`fix/netcdf-unlimited-dim-size`). Affected
every release (v2.1.0 to v2.7.0). Wrong metadata only: variable values and
HDF5 extents were always right. Users who worked around it with the variables'
shapes can use `Dimension::size` again.
Found 2026-09-28 while verifying the README refresh.
`clawhdf5-netcdf4`'s `NetCDF4File::dimensions()` reported an unlimited
dimension's `size` as 0 when variables along it hold records. Reproducer: with
netCDF4-python, create dimension `time` (unlimited) and `x` (3), a variable
`t(time, x)`, and write 2 records; netCDF4 reports `time` = 2 and `t` shape
(2, 3). clawhdf5-netcdf4 reported `dim time size 0 unlimited true`, while
`variable("t").shape()` correctly gave `[2, 3]`. The size was read from the
dimension scale's extent, which netCDF-C does not extend (it stays 0; a
coordinate variable is extended only by its own writes).
The size is now what netCDF-C reports (`NC4_inq_dim` → `nc4_find_dim_len`):
the largest current extent, along the dimension, of the variables that use
it in any group, found through the scale's `REFERENCE_LIST`, plus a
coordinate variable's own extent; 0 when nothing has been written.
`interop_tests::unlimited_dimension_lengths_match_netcdf4_python` compares
with netCDF4-python for variables of different lengths (one in a subgroup),
an unwritten dimension, a coordinate variable shorter than a variable on its
dimension, and a subgroup's own unlimited dimension; before the fix it got
`time` 0 (want 6), `rec` 3 (want 5), `srec` 0 (want 1).
## `ObjectHeader::parse` 4% slower after range-read M2/M3 ## `ObjectHeader::parse` 4% slower after range-read M2/M3
**Status:** fixed 2026-09-27 (`96086ad`, PR #21), before any release. **Status:** fixed 2026-09-27 (`96086ad`, PR #21), before any release.