clawhdf5-netcdf4: variables' dimensions come from the file
CI / test-arm64 (pull_request) Successful in 1m33s
CI / test (pull_request) Successful in 18m24s

Variables got the first unused dimension of equal size, so a variable on
an unlimited dimension with fewer records got an anonymous dim_<n>, and
dimensions of one size could be swapped. Resolve them as netCDF-C does
(libhdf5/hdf5open.c): _Netcdf4Coordinates ids, else the scales
DIMENSION_LIST references (the last one attached to an axis), searched in
the variable's group and its parents; a coordinate variable is on its own
scale. Size matching remains only for axes the file names nothing for.

variables()/variable_names() leave out dimension scales that are only
dimensions, and _nc4_non_coord_<name> is the variable <name>.
Variable::shape is the netCDF shape (an unlimited dimension's length) and
the reads pad unwritten records with the fill value (_FillValue, else
NC_FILL_*; NaN from read_f64); Variable::stored_shape is the HDF5 extent.
New NetCDF4File::variable_names.

Tests compare with netCDF4-python variable by variable: the known-issues
reproducer, equal sizes, (p, p), scalars, inherited dimensions, unwritten
records, h5py dimension scales, h5netcdf and xarray files. CI installs
h5netcdf. known-issues entry moved to Fixed (history); stale open-table
row for the unlimited-size fix removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This commit is contained in:
osobh
2026-09-28 22:52:04 -05:00
co-authored by Claude Opus 5.5
parent f9edf4d6ad
commit 00b6f76ee0
12 changed files with 1191 additions and 237 deletions
+1 -1
View File
@@ -40,7 +40,7 @@ jobs:
python3 -m venv /opt/interop python3 -m venv /opt/interop
# maturin + pytest: ci-test.sh builds the Python package # maturin + pytest: ci-test.sh builds the Python package
# (crates/clawhdf5-py) and runs its tests against h5py. # (crates/clawhdf5-py) and runs its tests against h5py.
/opt/interop/bin/pip install --no-cache-dir h5py numpy netCDF4 xarray hdf5plugin maturin pytest /opt/interop/bin/pip install --no-cache-dir h5py numpy netCDF4 xarray h5netcdf hdf5plugin maturin pytest
echo "/opt/interop/bin" >> "$GITHUB_PATH" echo "/opt/interop/bin" >> "$GITHUB_PATH"
- name: Show interop library versions - name: Show interop library versions
# h5dump's version too: the h5rs dump test requires its exact output # h5dump's version too: the h5rs dump test requires its exact output
+53
View File
@@ -2,6 +2,59 @@
## Unreleased ## Unreleased
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
dimension it had written fewer records of got `dim_<n>`, and dimensions
of one size could be swapped. It now resolves them as netCDF-C does
(`libhdf5/hdf5open.c`): the ids in the variable's `_Netcdf4Coordinates`
(each scale's `_Netcdf4Dimid`), else the scales its `DIMENSION_LIST`
references (object references, also the revised `H5T_STD_REF`; of
several scales attached to one axis, the last, as netCDF-C's
`dimscale_visitor` keeps), looked up in its group and then each parent
group; a coordinate variable is on
its own scale. Only an axis with neither (a file not written by a netCDF
library) is still matched by size. A variable may use one dimension
twice (`(p, p)`).
- `variables()` and `variable_names()` leave out the dimension scales that
are only dimensions (`NAME` "This is a netCDF dimension but not a netCDF
variable."), as netCDF-C does, and `variable()` refuses them
(`VariableNotFound`). A variable stored as `_nc4_non_coord_<name>`
(netCDF-C's name for a variable sharing a dimension's name without being
its coordinate variable) is listed and found as `<name>`.
`NetCDF4File::variable_names` is new; `NetCDF4Group::variable_names` used
to list every dataset.
- A variable along an unlimited dimension has the dimension's length, as in
netCDF: `Variable::shape` is that length and the reads return that many
values, the records it has not written as the fill value (`_FillValue`,
else netCDF's `NC_FILL_*` for the type; `""` for strings; NaN from
`read_f64`). It was the HDF5 extent. `Variable::stored_shape` (new) is the
HDF5 extent. Where the unlimited dimension is not a variable's first,
unwritten values are placed per row, as netCDF-C's element and row reads
return them; a whole-variable read through netCDF-C 4.9.3
(netCDF4-python 1.7.4) instead returns the written values first, then the
fill.
- A variable's attributes are read when it is opened (they were read on
first use).
- Tests, compared with netCDF4-python 1.7.4 (netCDF-C 4.9.3) variable by
variable (dimensions, shape, every value): the known-issues reproducer;
two dimensions of one size in either order, one dimension used twice, a
scalar, a non-coordinate variable named like a dimension, a subgroup and
a sub-subgroup on their ancestors' dimensions; unwritten records of
`i1`/`i4`/`u8`/`f4`/`f8`/string variables, with and without
`_FillValue`; a file of h5py dimension scales (no netCDF attributes);
h5netcdf 1.8.1 and xarray 2026.7.0 files (engines netcdf4 and h5netcdf).
The h5netcdf cases skip when h5netcdf is not installed; CI now installs
it. A one-off comparison over the 78 conformance-corpus files
netCDF4-python opens (tank, 2026-09-28; a throwaway test, not committed)
found 52 files with the same variables, dimensions and shapes, and the
same values in every numeric variable of up to 5000 elements (enum
variables were not compared). The other 26 have no dimension scales
(netCDF-C names their axes `phony_dim_<n>`; this crate still matches by
size or names them `dim_<size>`) or hold datasets of types netCDF-C
skips (opaque, references). Affected v2.1.0 to v2.7.0.
`docs/known-issues.md`.
### A dropped `FileEditor` releases its lock at once (2026-09-28) ### A dropped `FileEditor` releases its lock at once (2026-09-28)
- `FileEditor`'s `flock` could outlive the editor for a moment when - `FileEditor`'s `flock` could outlive the editor for a moment when
another thread forked to spawn a process: the child shared the locked another thread forked to spawn a process: the child shared the locked
+2 -1
View File
@@ -411,7 +411,8 @@ scripts/ci-test.sh # what CI runs: fmt, clippy matrix, tests,
conformance/run.sh # the conformance report (needs h5py, hdf5plugin, h5dump) conformance/run.sh # the conformance report (needs h5py, hdf5plugin, h5dump)
``` ```
The interop suites need a Python with h5py (and netCDF4, xarray); on a The interop suites need a Python with h5py (and netCDF4, xarray; the
NetCDF-4 tests of h5netcdf-written files skip without h5netcdf); on a
PEP 668 system that has to be a virtualenv, which `ci-test.sh` finds as PEP 668 system that has to be a virtualenv, which `ci-test.sh` finds as
`.venv` or through `CLAWHDF5_PYTHON`. Without one they skip; set `.venv` or through `CLAWHDF5_PYTHON`. Without one they skip; set
`CLAWHDF5_REQUIRE_INTEROP=1` to make that a failure, as CI does: `CLAWHDF5_REQUIRE_INTEROP=1` to make that a failure, as CI does:
+25 -7
View File
@@ -36,17 +36,35 @@ let values: Vec<f64> = temp.read_f64()?;
| Item | What | | Item | What |
|---|---| |---|---|
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | | `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) | | `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `variable_names`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` | | `Variable` | `name`, `shape`, `stored_shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) | | `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` | | `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable | | `NcType` | the NetCDF type of a variable |
No cargo features. Tests compare against files written by netCDF4-python Variables and dimensions follow netCDF-C:
(`tests/interop_tests.rs`; the CI job requires them with
`CLAWHDF5_REQUIRE_INTEROP=1`). What the HDF5 reader underneath cannot - A variable's dimensions are the ones the file names: the ids in its
read is listed in [`docs/known-issues.md`](../../docs/known-issues.md). `_Netcdf4Coordinates` attribute, else the dimension scales its
`DIMENSION_LIST` references, found in its group or a parent group. Only
an axis the file names no dimension for (an HDF5 file not written by a
netCDF library) gets the first dimension of the group of the same size,
else an anonymous `dim_<size>`.
- Dimension scales that are only dimensions are not variables; a dataset
`_nc4_non_coord_<name>` is the variable `<name>`.
- A variable along an unlimited dimension has the dimension's length:
`shape` is that length, and the reads return that many values, the
records the variable has not written as its fill value (`_FillValue`,
else netCDF's default for the type; NaN from `read_f64`).
`stored_shape` is the HDF5 dataset's extent.
No cargo features. Tests compare against files written by netCDF4-python,
h5py dimension scales, h5netcdf and xarray, variable by variable with what
netCDF4-python reads (`tests/interop_tests.rs`; the CI job requires them
with `CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is
not installed). What the HDF5 reader underneath cannot read is listed in
[`docs/known-issues.md`](../../docs/known-issues.md).
## License ## License
+144 -94
View File
@@ -22,73 +22,123 @@ pub struct Dimension {
pub is_unlimited: bool, pub is_unlimited: bool,
} }
/// Extract dimensions from an HDF5 group (root or subgroup). /// A dimension scale of one group: the dataset that defines a dimension.
#[derive(Debug, Clone)]
pub(crate) struct Scale {
/// Object header address of the scale's dataset (what a variable's
/// `DIMENSION_LIST` references).
pub address: u64,
/// Its `_Netcdf4Dimid` (what a variable's `_Netcdf4Coordinates` lists).
pub dimid: Option<i64>,
/// Index of its dimension in [`GroupDims::dims`].
pub dim: usize,
}
/// The dimensions a group defines, with the scales that define them.
#[derive(Debug, Clone, Default)]
pub(crate) struct GroupDims {
/// The group's dimensions, in `_Netcdf4Dimid` order (then discovery order).
pub dims: Vec<Dimension>,
/// The dimension scales behind `dims`; empty when the group has no
/// dimension scale and `dims` were inferred from 1-D datasets.
pub scales: Vec<Scale>,
}
impl GroupDims {
/// The dimension defined by the scale at `address`.
pub fn by_address(&self, address: u64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.address == address)
.map(|s| &self.dims[s.dim])
}
/// The dimension whose scale has `_Netcdf4Dimid` `id`.
pub fn by_dimid(&self, id: i64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.dimid == Some(id))
.map(|s| &self.dims[s.dim])
}
}
/// The dimensions of an HDF5 group (root or subgroup).
/// ///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed /// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// dimension's size is the dataset's first (and typically only) shape extent. /// dimension's size is the dataset's first (and typically only) shape extent.
/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace; /// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`. /// their size is computed by `unlimited_len`. A group with no dimension
pub(crate) fn extract_dimensions( /// scale at all (not written by a netCDF library) gets one dimension per
/// 1-D dataset instead.
pub(crate) fn group_dims(
file: &clawhdf5::File, file: &clawhdf5::File,
group: &clawhdf5::Group<'_>, group: &clawhdf5::Group<'_>,
) -> Result<Vec<Dimension>, Error> { ) -> Result<GroupDims, Error> {
let addresses: HashMap<String, u64> = group.entries()?.into_iter().collect();
let dataset_names = group.datasets()?; let dataset_names = group.datasets()?;
let mut dims = Vec::new(); // (dimid, dimension, scale address), in discovery order.
let mut seen_dimids: HashMap<i64, usize> = HashMap::new(); let mut found: Vec<(Option<i64>, Dimension, u64)> = Vec::new();
for ds_name in &dataset_names { for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?; let ds = group.dataset(ds_name)?;
let attrs = ds.attrs()?; let attrs = ds.attrs()?;
// Check if this is a dimension scale
if !is_dimension_scale(&attrs) { if !is_dimension_scale(&attrs) {
continue; continue;
} }
let Some(&address) = addresses.get(ds_name) else {
continue;
};
let shape = ds.shape()?; let shape = ds.shape()?;
let is_unlimited = check_unlimited(file, group, ds_name); let is_unlimited = is_unlimited(&ds);
let size = if is_unlimited { let size = if is_unlimited {
unlimited_len(file, &attrs, &shape) unlimited_len(file, &attrs, &shape)
} else { } else {
shape.first().copied().unwrap_or(0) shape.first().copied().unwrap_or(0)
}; };
let dimid = get_dimid(&attrs);
let dim = Dimension { let dim = Dimension {
name: ds_name.clone(), name: ds_name.clone(),
size, size,
is_unlimited, is_unlimited,
}; };
found.push((get_dimid(&attrs), dim, address));
if let Some(id) = dimid {
seen_dimids.insert(id, dims.len());
}
dims.push(dim);
} }
// Sort by dimid if available, otherwise keep discovery order if found.is_empty() {
if !seen_dimids.is_empty() { // Fallback: infer dimensions from dataset shapes and names.
let mut pairs: Vec<(i64, Dimension)> = Vec::new(); // In NetCDF-4, coordinate variables are datasets whose name matches
let mut unordered = Vec::new(); // a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
for (i, dim) in dims.into_iter().enumerate() { let mut dims = Vec::new();
let id = seen_dimids for ds_name in &dataset_names {
.iter() let ds = group.dataset(ds_name)?;
.find(|(_, idx)| **idx == i) let shape = ds.shape()?;
.map(|(k, _)| *k); if shape.len() == 1 {
if let Some(id) = id { dims.push(Dimension {
pairs.push((id, dim)); name: ds_name.clone(),
} else { size: shape[0],
unordered.push(dim); is_unlimited: is_unlimited(&ds),
});
} }
} }
pairs.sort_by_key(|(id, _)| *id); return Ok(GroupDims {
dims = pairs.into_iter().map(|(_, d)| d).collect(); dims,
dims.extend(unordered); scales: Vec::new(),
});
} }
Ok(dims) // By dimid; scales without one keep their discovery order after those
// with one (the sort is stable).
found.sort_by_key(|(id, ..)| (id.is_none(), id.unwrap_or(0)));
let mut out = GroupDims::default();
for (i, (dimid, dim, address)) in found.into_iter().enumerate() {
out.dims.push(dim);
out.scales.push(Scale {
address,
dimid,
dim: i,
});
}
Ok(out)
} }
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that /// The start of the `NAME` attribute netCDF-C gives a dimension scale that
@@ -107,10 +157,7 @@ const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF v
/// scale's own extent, as before. /// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 { fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0); let own = shape.first().copied().unwrap_or(0);
let is_variable = !matches!( let is_variable = !is_pure_dimension(attrs);
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
);
let Some(refs) = reference_list(file, attrs) else { let Some(refs) = reference_list(file, attrs) else {
return own; return own;
}; };
@@ -171,8 +218,59 @@ fn reference_list(
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect()) Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
} }
/// The dimension scale attached to each axis of a variable, from its
/// `DIMENSION_LIST` attribute (HDF5 dimension scales: one variable-length
/// sequence of object references per axis) — the address of the scale
/// netCDF-C takes for the axis, or `None` for an axis with none. netCDF-C's
/// `dimscale_visitor` lets `H5DSiterate_scales` visit every scale attached
/// to the axis and keeps the last, so with several (h5py's `attach_scale`
/// twice) the last one is the axis's dimension. `None` overall when the
/// attribute is missing or not in that form.
pub(crate) fn dimension_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<Option<u64>>> {
use clawhdf5_format::data_read::read_object_references;
use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::vl_data::VlResolver;
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("DIMENSION_LIST") else {
return None;
};
let Datatype::VariableLength {
is_string: false,
base_type,
..
} = datatype
else {
return None;
};
let sb = file.superblock();
let base_size = usize::try_from(base_type.type_size()).ok()?;
let sequences = VlResolver::new_in(file.storage(), sb.offset_size, sb.length_size)
.sequences(data, base_size)
.ok()?;
sequences
.iter()
.map(|refs| {
let refs = read_object_references(refs, base_type, sb.offset_size).ok()?;
Some(refs.iter().rev().find(|r| !r.is_null()).map(|r| r.address))
})
.collect()
}
/// Whether a dataset is a dimension scale that is only a dimension, not a
/// netCDF variable: netCDF-C and h5netcdf give it this `NAME`, and netCDF-C
/// does not list it among the variables.
pub(crate) fn is_pure_dimension(attrs: &HashMap<String, AttrValue>) -> bool {
is_dimension_scale(attrs)
&& matches!(
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
)
}
/// Check if a dataset's attributes mark it as a dimension scale. /// Check if a dataset's attributes mark it as a dimension scale.
fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool { pub(crate) fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
if let Some(AttrValue::String(class)) = attrs.get("CLASS") { if let Some(AttrValue::String(class)) = attrs.get("CLASS") {
return class == "DIMENSION_SCALE"; return class == "DIMENSION_SCALE";
} }
@@ -180,7 +278,7 @@ fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
} }
/// Get the _Netcdf4Dimid attribute value if present. /// Get the _Netcdf4Dimid attribute value if present.
fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> { pub(crate) fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
match attrs.get("_Netcdf4Dimid") { match attrs.get("_Netcdf4Dimid") {
Some(AttrValue::I64(id)) => Some(*id), Some(AttrValue::I64(id)) => Some(*id),
Some(AttrValue::U64(id)) => Some(*id as i64), Some(AttrValue::U64(id)) => Some(*id as i64),
@@ -188,56 +286,8 @@ fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
} }
} }
/// Check if a dimension is unlimited by inspecting the HDF5 dataspace max_dimensions. /// Whether a dataset's first axis is unlimited (`max_dimensions[0] ==
/// /// u64::MAX` in its dataspace).
/// A dimension is unlimited when `max_dimensions[0] == u64::MAX` in the HDF5 dataspace. fn is_unlimited(ds: &clawhdf5::Dataset<'_>) -> bool {
fn check_unlimited(_file: &clawhdf5::File, group: &clawhdf5::Group<'_>, ds_name: &str) -> bool { matches!(ds.max_dimensions(), Ok(Some(max_dims)) if max_dims.first() == Some(&u64::MAX))
let ds = match group.dataset(ds_name) {
Ok(ds) => ds,
Err(_) => return false,
};
match ds.max_dimensions() {
Ok(Some(max_dims)) => max_dims.first().copied() == Some(u64::MAX),
_ => false,
}
}
/// Extract dimensions from an HDF5 group using both dimension scale attributes
/// and variable DIMENSION_LIST references.
///
/// This is a more robust approach that also discovers dimensions from variables
/// that reference them, even when dimension scales aren't explicitly set.
pub(crate) fn extract_dimensions_from_datasets(
group: &clawhdf5::Group<'_>,
file: &clawhdf5::File,
) -> Result<Vec<Dimension>, Error> {
// First try the standard approach with DIMENSION_SCALE
let mut dims = extract_dimensions(file, group)?;
// If we found dimensions, return them
if !dims.is_empty() {
return Ok(dims);
}
// Fallback: infer dimensions from dataset shapes and names.
// In NetCDF-4, coordinate variables are datasets whose name matches
// a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
let dataset_names = group.datasets()?;
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
if shape.len() == 1 {
// This 1-D dataset could be a coordinate variable / dimension
let is_unlimited = check_unlimited(file, group, ds_name);
dims.push(Dimension {
name: ds_name.clone(),
size: shape[0],
is_unlimited,
});
}
}
Ok(dims)
} }
+23 -17
View File
@@ -9,12 +9,15 @@ use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension}; use crate::dimension::{self, Dimension};
use crate::error::Error; use crate::error::Error;
use crate::variable::{self, Variable}; use crate::scope::{self, Scope};
use crate::variable::Variable;
/// A NetCDF-4 group corresponding to an HDF5 group. /// A NetCDF-4 group corresponding to an HDF5 group.
pub struct NetCDF4Group<'f> { pub struct NetCDF4Group<'f> {
/// Group name. /// Group name.
name: String, name: String,
/// Path of the group from the root (`/`-separated).
path: String,
/// Underlying HDF5 file. /// Underlying HDF5 file.
file: &'f clawhdf5::File, file: &'f clawhdf5::File,
/// Underlying HDF5 group. /// Underlying HDF5 group.
@@ -25,11 +28,13 @@ impl<'f> NetCDF4Group<'f> {
/// Create a new NetCDF4Group from an HDF5 group. /// Create a new NetCDF4Group from an HDF5 group.
pub(crate) fn new( pub(crate) fn new(
name: String, name: String,
path: String,
file: &'f clawhdf5::File, file: &'f clawhdf5::File,
hdf5_group: clawhdf5::Group<'f>, hdf5_group: clawhdf5::Group<'f>,
) -> Self { ) -> Self {
Self { Self {
name, name,
path,
file, file,
hdf5_group, hdf5_group,
} }
@@ -40,27 +45,22 @@ impl<'f> NetCDF4Group<'f> {
&self.name &self.name
} }
/// List dimensions defined in this group. /// List dimensions defined in this group (not those of its parent
/// groups, which its variables can also use).
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> { pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
dimension::extract_dimensions_from_datasets(&self.hdf5_group, self.file) Ok(dimension::group_dims(self.file, &self.hdf5_group)?.dims)
} }
/// List variables in this group. /// List variables in this group: its datasets, except the dimension
/// scales that are only dimensions. Their dimensions can be defined in
/// this group or a parent group.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> { pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let dims = self.dimensions()?; Scope::new(self.file, &self.path)?.variables()
variable::build_variables(&self.hdf5_group, &dims)
} }
/// Get a specific variable by name. /// Get a specific variable by name.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> { pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let dims = self.dimensions()?; scope::variable_at(self.file, &self.path, name)
let ds = self
.hdf5_group
.dataset(name)
.map_err(|_| Error::VariableNotFound(name.to_string()))?;
let shape = ds.shape()?;
let var_dims = crate::variable::match_dimensions_to_variable(&shape, &dims);
Ok(Variable::new(name.to_string(), ds, var_dims))
} }
/// Read all attributes of this group. /// Read all attributes of this group.
@@ -79,12 +79,18 @@ impl<'f> NetCDF4Group<'f> {
.hdf5_group .hdf5_group
.group(name) .group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?; .map_err(|_| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(name.to_string(), self.file, hdf5_group)) Ok(NetCDF4Group::new(
name.to_string(),
format!("{}/{name}", self.path),
self.file,
hdf5_group,
))
} }
/// List dataset (variable) names in this group. /// The names of this group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> { pub fn variable_names(&self) -> Result<Vec<String>, Error> {
Ok(self.hdf5_group.datasets()?) Scope::new(self.file, &self.path)?.variable_names()
} }
} }
+19 -13
View File
@@ -27,6 +27,7 @@ pub mod cf;
pub mod dimension; pub mod dimension;
pub mod error; pub mod error;
pub mod group; pub mod group;
mod scope;
pub mod types; pub mod types;
pub mod variable; pub mod variable;
@@ -75,25 +76,25 @@ impl NetCDF4File {
/// List dimensions defined in the root group. /// List dimensions defined in the root group.
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> { pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
dimension::extract_dimensions_from_datasets(&self.hdf5.root(), &self.hdf5) Ok(dimension::group_dims(&self.hdf5, &self.hdf5.root())?.dims)
} }
/// List all variables in the root group. /// List all variables in the root group: its datasets, except the
/// dimension scales that are only dimensions (netCDF-C does not list
/// them either).
pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> { pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> {
let dims = self.dimensions()?; scope::Scope::new(&self.hdf5, "/")?.variables()
variable::build_variables(&self.hdf5.root(), &dims) }
/// The names of the root group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
scope::Scope::new(&self.hdf5, "/")?.variable_names()
} }
/// Get a specific variable by name from the root group. /// Get a specific variable by name from the root group.
pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> { pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> {
let dims = self.dimensions()?; scope::variable_at(&self.hdf5, "", name)
let ds = self
.hdf5
.dataset(name)
.map_err(|_| Error::VariableNotFound(name.to_string()))?;
let shape = ds.shape()?;
let var_dims = variable::match_dimensions_to_variable(&shape, &dims);
Ok(Variable::new(name.to_string(), ds, var_dims))
} }
/// Read all global (root group) attributes. /// Read all global (root group) attributes.
@@ -112,7 +113,12 @@ impl NetCDF4File {
.hdf5 .hdf5
.group(name) .group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?; .map_err(|_| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(name.to_string(), &self.hdf5, hdf5_group)) Ok(NetCDF4Group::new(
name.to_string(),
name.to_string(),
&self.hdf5,
hdf5_group,
))
} }
/// Access the underlying HDF5 file for advanced operations. /// Access the underlying HDF5 file for advanced operations.
+231
View File
@@ -0,0 +1,231 @@
//! A group's variables and the dimensions they are defined on.
//!
//! netCDF-C (`libhdf5/hdf5open.c`) gives a variable its dimensions from the
//! file, never by size: the dimension ids in its `_Netcdf4Coordinates`
//! attribute (each dimension scale's `_Netcdf4Dimid`), else the dimension
//! scales its `DIMENSION_LIST` attribute references, looked up in the
//! variable's group and then each parent group up to the root. Only an axis
//! with neither (a file not written by a netCDF library) gets a dimension
//! by size. Dimension scales that are only dimensions are not variables, and
//! a variable stored as `_nc4_non_coord_<name>` (a variable sharing a
//! dimension's name without being its coordinate variable) is `<name>`.
use std::collections::{HashMap, HashSet};
use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension, GroupDims};
use crate::error::Error;
use crate::variable::Variable;
/// The prefix netCDF-C gives the dataset of a variable that has a
/// dimension's name but is not that dimension's coordinate variable (the
/// dimension's scale holds the name).
const NON_COORD_PREFIX: &str = "_nc4_non_coord_";
/// A group, with the dimensions visible from it.
pub(crate) struct Scope<'f> {
file: &'f clawhdf5::File,
group: clawhdf5::Group<'f>,
/// This group's dimensions, then its parent's, and so on to the root's.
levels: Vec<GroupDims>,
}
impl<'f> Scope<'f> {
/// The group at `path` (`/`-separated from the root; `""` or `"/"` is
/// the root).
pub fn new(file: &'f clawhdf5::File, path: &str) -> Result<Self, Error> {
let parts: Vec<&str> = path.split('/').filter(|p| !p.is_empty()).collect();
let mut levels = Vec::with_capacity(parts.len() + 1);
for n in (0..=parts.len()).rev() {
let group = file.group(&parts[..n].join("/"))?;
levels.push(dimension::group_dims(file, &group)?);
}
let group = file.group(&parts.join("/"))?;
Ok(Self {
file,
group,
levels,
})
}
/// The group's datasets, as `(dataset name, object header address)` in
/// listing order.
fn datasets(&self) -> Result<Vec<(String, u64)>, Error> {
let datasets: HashSet<String> = self.group.datasets()?.into_iter().collect();
Ok(self
.group
.entries()?
.into_iter()
.filter(|(name, _)| datasets.contains(name))
.collect())
}
/// The group's variables: every dataset but the dimension scales that
/// are only dimensions.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let mut variables = Vec::new();
for (ds_name, address) in self.datasets()? {
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
continue;
}
variables.push(self.variable_from(nc_name(&ds_name), address, ds, attrs)?);
}
Ok(variables)
}
/// The names of the group's variables.
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
let mut names = Vec::new();
for (ds_name, address) in self.datasets()? {
let attrs = self.file.dataset_at(address)?.attrs()?;
if !dimension::is_pure_dimension(&attrs) {
names.push(nc_name(&ds_name));
}
}
Ok(names)
}
/// The variable called `name`: the dataset `_nc4_non_coord_<name>` if
/// there is one, else the dataset `<name>` unless it is only a
/// dimension.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let not_found = || Error::VariableNotFound(name.to_string());
let datasets = self.datasets()?;
let prefixed = format!("{NON_COORD_PREFIX}{name}");
let address = datasets
.iter()
.find(|(n, _)| *n == prefixed)
.or_else(|| datasets.iter().find(|(n, _)| n == name))
.map(|&(_, address)| address)
.ok_or_else(not_found)?;
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
return Err(not_found());
}
self.variable_from(nc_name(name), address, ds, attrs)
}
fn variable_from(
&self,
name: String,
address: u64,
ds: clawhdf5::Dataset<'f>,
attrs: HashMap<String, AttrValue>,
) -> Result<Variable<'f>, Error> {
let shape = ds.shape()?;
let dims = self.variable_dims(address, &attrs, &shape);
Ok(Variable::new(name, ds, dims, attrs))
}
/// The first dimension, searching this group and then its ancestors,
/// that `find` picks.
fn find<'a>(
&'a self,
find: impl Fn(&'a GroupDims) -> Option<&'a Dimension>,
) -> Option<Dimension> {
self.levels.iter().find_map(find).cloned()
}
/// The dimensions of the dataset at `address`, one per axis of `shape`,
/// as netCDF-C resolves them (see the module docs).
fn variable_dims(
&self,
address: u64,
attrs: &HashMap<String, AttrValue>,
shape: &[u64],
) -> Vec<Dimension> {
let rank = shape.len();
let mut dims: Vec<Option<Dimension>> = vec![None; rank];
if rank == 0 {
return Vec::new();
}
// A coordinate variable is the scale of its (first) dimension.
dims[0] = self.levels[0].by_address(address).cloned();
if let Some(ids) = coordinates(attrs).filter(|ids| ids.len() == rank) {
for (slot, id) in dims.iter_mut().zip(ids) {
if slot.is_none() {
*slot = self.find(|level| level.by_dimid(id));
}
}
}
if dims.iter().any(Option::is_none)
&& let Some(scales) =
dimension::dimension_list(self.file, attrs).filter(|s| s.len() == rank)
{
for (slot, scale) in dims.iter_mut().zip(scales) {
if slot.is_none()
&& let Some(scale) = scale
{
*slot = self.find(|level| level.by_address(scale));
}
}
}
// Neither: the first dimension of this group of the same size not
// already taken by another such axis, else an anonymous one.
let own = &self.levels[0].dims;
let mut used = vec![false; own.len()];
dims.into_iter()
.zip(shape)
.map(|(dim, &size)| {
dim.unwrap_or_else(|| {
match own
.iter()
.enumerate()
.find(|&(i, d)| !used[i] && d.size == size)
{
Some((i, d)) => {
used[i] = true;
d.clone()
}
None => Dimension {
name: format!("dim_{size}"),
size,
is_unlimited: false,
},
}
})
})
.collect()
}
}
/// The variable `name` of the group at `group_path`; `name` may itself be
/// a path (`"sub/var"`), relative to that group.
pub(crate) fn variable_at<'f>(
file: &'f clawhdf5::File,
group_path: &str,
name: &str,
) -> Result<Variable<'f>, Error> {
match name.trim_start_matches('/').rsplit_once('/') {
Some((dir, leaf)) => Scope::new(file, &format!("{group_path}/{dir}"))
.map_err(|_| Error::VariableNotFound(name.to_string()))?
.variable(leaf),
None => Scope::new(file, group_path)?.variable(name.trim_start_matches('/')),
}
}
/// The netCDF name of the dataset `ds_name`.
fn nc_name(ds_name: &str) -> String {
ds_name
.strip_prefix(NON_COORD_PREFIX)
.unwrap_or(ds_name)
.to_string()
}
/// A variable's `_Netcdf4Coordinates`: the `_Netcdf4Dimid` of the dimension
/// of each axis.
fn coordinates(attrs: &HashMap<String, AttrValue>) -> Option<Vec<i64>> {
match attrs.get("_Netcdf4Coordinates")? {
AttrValue::I64Array(ids) => Some(ids.clone()),
AttrValue::I64(id) => Some(vec![*id]),
AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(),
AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]),
_ => None,
}
}
+208 -81
View File
@@ -2,12 +2,18 @@
//! //!
//! Variables in NetCDF-4 are HDF5 datasets. This module wraps them with //! Variables in NetCDF-4 are HDF5 datasets. This module wraps them with
//! dimension associations and CF attribute support. //! dimension associations and CF attribute support.
//!
//! A variable along an unlimited dimension has that dimension's length in
//! netCDF, even when fewer records of it have been written (its HDF5 dataset
//! is shorter): [`Variable::shape`] is the netCDF shape and the reads return
//! that many values, the unwritten ones as the fill value, as netCDF-C does.
//! [`Variable::stored_shape`] is the dataset's extent.
use std::collections::HashMap; use std::collections::HashMap;
use clawhdf5::AttrValue; use clawhdf5::AttrValue;
use crate::cf::{self, CfAttributes}; use crate::cf::{self, CfAttributes, FillValue};
use crate::dimension::Dimension; use crate::dimension::Dimension;
use crate::error::Error; use crate::error::Error;
use crate::types::{NcType, dtype_to_nctype}; use crate::types::{NcType, dtype_to_nctype};
@@ -20,18 +26,23 @@ pub struct Variable<'f> {
dataset: clawhdf5::Dataset<'f>, dataset: clawhdf5::Dataset<'f>,
/// Dimensions associated with this variable. /// Dimensions associated with this variable.
dims: Vec<Dimension>, dims: Vec<Dimension>,
/// Cached attributes. /// The dataset's attributes.
attrs_cache: Option<HashMap<String, AttrValue>>, attrs: HashMap<String, AttrValue>,
} }
impl<'f> Variable<'f> { impl<'f> Variable<'f> {
/// Create a new Variable wrapping an HDF5 dataset. /// Create a new Variable wrapping an HDF5 dataset.
pub(crate) fn new(name: String, dataset: clawhdf5::Dataset<'f>, dims: Vec<Dimension>) -> Self { pub(crate) fn new(
name: String,
dataset: clawhdf5::Dataset<'f>,
dims: Vec<Dimension>,
attrs: HashMap<String, AttrValue>,
) -> Self {
Self { Self {
name, name,
dataset, dataset,
dims, dims,
attrs_cache: None, attrs,
} }
} }
@@ -40,13 +51,27 @@ impl<'f> Variable<'f> {
&self.name &self.name
} }
/// The dimensions of this variable. /// The dimensions of this variable, one per axis: the ones the file
/// gives it (`_Netcdf4Coordinates`, else `DIMENSION_LIST`), found in its
/// group or a parent group. An axis the file gives no dimension (a file
/// not written by a netCDF library) gets the first dimension of the
/// variable's group of the same size, else an anonymous `dim_<size>`.
pub fn dimensions(&self) -> &[Dimension] { pub fn dimensions(&self) -> &[Dimension] {
&self.dims &self.dims
} }
/// The shape of this variable (dimension sizes). /// The shape of this variable as netCDF reports it: along an unlimited
/// dimension, the dimension's current length (the longest variable on
/// it), even if fewer records of this variable have been written;
/// otherwise the dataset's extent. The reads return this many values.
pub fn shape(&self) -> Result<Vec<u64>, Error> { pub fn shape(&self) -> Result<Vec<u64>, Error> {
Ok(nc_shape(&self.dataset.shape()?, &self.dims))
}
/// The extent of the HDF5 dataset: what has been written. It differs
/// from [`shape`](Self::shape) only along an unlimited dimension that
/// another variable has more records of.
pub fn stored_shape(&self) -> Result<Vec<u64>, Error> {
Ok(self.dataset.shape()?) Ok(self.dataset.shape()?)
} }
@@ -58,59 +83,88 @@ impl<'f> Variable<'f> {
/// Read all attributes as a HashMap. /// Read all attributes as a HashMap.
pub fn attrs(&mut self) -> Result<&HashMap<String, AttrValue>, Error> { pub fn attrs(&mut self) -> Result<&HashMap<String, AttrValue>, Error> {
if self.attrs_cache.is_none() { Ok(&self.attrs)
self.attrs_cache = Some(self.dataset.attrs()?);
}
Ok(self
.attrs_cache
.as_ref()
.expect("invariant: attrs_cache is Some after initialization"))
} }
/// Extract CF convention attributes. /// Extract CF convention attributes.
pub fn cf_attributes(&mut self) -> Result<CfAttributes, Error> { pub fn cf_attributes(&mut self) -> Result<CfAttributes, Error> {
let attrs = self.attrs()?; Ok(cf::extract_cf_attributes(&self.attrs))
Ok(cf::extract_cf_attributes(attrs))
} }
/// Read data as f64 with scale_factor/add_offset applied. /// Read data as f64 with scale_factor/add_offset applied.
/// ///
/// Missing values (matching `_FillValue` or `missing_value`) become NaN. /// Missing values (matching `_FillValue` or `missing_value`) become NaN.
/// If no scale_factor or add_offset attributes exist, returns the raw f64 data. /// If no scale_factor or add_offset attributes exist, returns the raw f64 data.
/// Records along an unlimited dimension that this variable has not
/// written (see [`shape`](Self::shape)) are NaN.
pub fn read_f64(&mut self) -> Result<Vec<f64>, Error> { pub fn read_f64(&mut self) -> Result<Vec<f64>, Error> {
let raw = self.dataset.read_f64()?; let raw = self.dataset.read_f64()?;
let cf = self.cf_attributes()?; let cf = cf::extract_cf_attributes(&self.attrs);
Ok(cf::apply_scale_offset(&raw, &cf)) self.padded(cf::apply_scale_offset(&raw, &cf), || Ok(f64::NAN))
} }
/// Read raw data as f64 without any scale/offset transformation. /// Read raw data as f64 without any scale/offset transformation.
///
/// Unwritten records along an unlimited dimension read as the fill
/// value (`_FillValue`, else netCDF's default for the type), as in the
/// other `read_raw_*` methods and [`read_string`](Self::read_string).
pub fn read_raw_f64(&self) -> Result<Vec<f64>, Error> { pub fn read_raw_f64(&self) -> Result<Vec<f64>, Error> {
Ok(self.dataset.read_f64()?) self.padded_read(self.dataset.read_f64()?)
} }
/// Read raw data as f32 without any scale/offset transformation. /// Read raw data as f32 without any scale/offset transformation.
pub fn read_raw_f32(&self) -> Result<Vec<f32>, Error> { pub fn read_raw_f32(&self) -> Result<Vec<f32>, Error> {
Ok(self.dataset.read_f32()?) self.padded_read(self.dataset.read_f32()?)
} }
/// Read raw data as i32 without any scale/offset transformation. /// Read raw data as i32 without any scale/offset transformation.
pub fn read_raw_i32(&self) -> Result<Vec<i32>, Error> { pub fn read_raw_i32(&self) -> Result<Vec<i32>, Error> {
Ok(self.dataset.read_i32()?) self.padded_read(self.dataset.read_i32()?)
} }
/// Read raw data as i64 without any scale/offset transformation. /// Read raw data as i64 without any scale/offset transformation.
pub fn read_raw_i64(&self) -> Result<Vec<i64>, Error> { pub fn read_raw_i64(&self) -> Result<Vec<i64>, Error> {
Ok(self.dataset.read_i64()?) self.padded_read(self.dataset.read_i64()?)
} }
/// Read raw data as u64 without any scale/offset transformation. /// Read raw data as u64 without any scale/offset transformation.
pub fn read_raw_u64(&self) -> Result<Vec<u64>, Error> { pub fn read_raw_u64(&self) -> Result<Vec<u64>, Error> {
Ok(self.dataset.read_u64()?) self.padded_read(self.dataset.read_u64()?)
} }
/// Read raw data as strings. /// Read raw data as strings.
pub fn read_string(&self) -> Result<Vec<String>, Error> { pub fn read_string(&self) -> Result<Vec<String>, Error> {
Ok(self.dataset.read_string()?) self.padded_read(self.dataset.read_string()?)
}
/// The fill value netCDF-C gives the variable's unwritten values: its
/// `_FillValue`, else the default fill value of its type (`NC_FILL_*`).
fn fill_value(&self) -> Result<FillValue, Error> {
if let Some(fill) = cf::extract_cf_attributes(&self.attrs).fill_value {
return Ok(fill);
}
Ok(default_fill(self.nc_type()?))
}
/// `data`, read in the dataset's extent, laid out in the variable's
/// netCDF shape with the fill value in the positions not written.
fn padded_read<T: Clone + FromFill>(&self, data: Vec<T>) -> Result<Vec<T>, Error> {
self.padded(data, || Ok(T::from_fill(&self.fill_value()?)))
}
/// Like [`padded_read`](Self::padded_read), padding with what `fill`
/// returns (called only when there is something to pad).
fn padded<T: Clone>(
&self,
data: Vec<T>,
fill: impl FnOnce() -> Result<T, Error>,
) -> Result<Vec<T>, Error> {
let extent = self.dataset.shape()?;
let shape = nc_shape(&extent, &self.dims);
if shape == extent {
return Ok(data);
}
pad(data, &extent, &shape, fill()?)
} }
/// Read raw bytes without any type conversion. /// Read raw bytes without any type conversion.
@@ -122,23 +176,23 @@ impl<'f> Variable<'f> {
let dtype = self.dataset.dtype()?; let dtype = self.dataset.dtype()?;
match dtype { match dtype {
clawhdf5::DType::F64 => { clawhdf5::DType::F64 => {
let vals = self.dataset.read_f64()?; let vals = self.read_raw_f64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::F32 => { clawhdf5::DType::F32 => {
let vals = self.dataset.read_f32()?; let vals = self.read_raw_f32()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::I32 => { clawhdf5::DType::I32 => {
let vals = self.dataset.read_i32()?; let vals = self.read_raw_i32()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::I64 => { clawhdf5::DType::I64 => {
let vals = self.dataset.read_i64()?; let vals = self.read_raw_i64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::U64 => { clawhdf5::DType::U64 => {
let vals = self.dataset.read_u64()?; let vals = self.read_raw_u64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
other => { other => {
@@ -169,64 +223,137 @@ impl std::fmt::Debug for Variable<'_> {
} }
} }
/// Build variables from a group's datasets and associated dimensions. /// The netCDF shape of a variable whose dataset has `extent`: along an
pub(crate) fn build_variables<'f>( /// unlimited dimension the dimension's length, which is at least the extent.
group: &clawhdf5::Group<'f>, fn nc_shape(extent: &[u64], dims: &[Dimension]) -> Vec<u64> {
available_dims: &[Dimension], if dims.len() != extent.len() {
) -> Result<Vec<Variable<'f>>, Error> { return extent.to_vec();
let dataset_names = group.datasets()?;
let mut variables = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
// Associate dimensions with this variable.
// First try DIMENSION_LIST attribute, then fall back to shape matching.
let var_dims = match_dimensions_to_variable(&shape, available_dims);
variables.push(Variable::new(ds_name.clone(), ds, var_dims));
} }
extent
Ok(variables) .iter()
.zip(dims)
.map(|(&e, d)| if d.is_unlimited { e.max(d.size) } else { e })
.collect()
} }
/// Match dimensions to a variable based on shape. /// `data`, row-major in `extent`, placed in a row-major array of `shape`
/// /// (as many axes, each at least as long) filled with `fill`.
/// For each axis of the variable, find a dimension with matching size. fn pad<T: Clone>(data: Vec<T>, extent: &[u64], shape: &[u64], fill: T) -> Result<Vec<T>, Error> {
/// If multiple dimensions have the same size, prefer exact name matching let too_big = || Error::TypeError(format!("variable of shape {shape:?} is too large"));
/// from the convention order. let to_usize = |dims: &[u64]| -> Result<Vec<usize>, Error> {
pub(crate) fn match_dimensions_to_variable( dims.iter()
shape: &[u64], .map(|&d| usize::try_from(d).map_err(|_| too_big()))
available_dims: &[Dimension], .collect()
) -> Vec<Dimension> { };
let mut result = Vec::with_capacity(shape.len()); let (extent, shape) = (to_usize(extent)?, to_usize(shape)?);
let total = shape
// Track which dimensions have been used to avoid duplicates .iter()
let mut used = vec![false; available_dims.len()]; .try_fold(1usize, |n, &d| n.checked_mul(d))
.ok_or_else(too_big)?;
for &dim_size in shape { if extent.len() != shape.len()
let mut matched = false; || extent.iter().zip(&shape).any(|(e, s)| e > s)
|| extent.iter().product::<usize>() != data.len()
// Find a dimension with matching size that hasn't been used yet {
for (i, dim) in available_dims.iter().enumerate() { return Err(Error::TypeError(format!(
if !used[i] && dim.size == dim_size { "{} values of extent {extent:?} do not fit shape {shape:?}",
result.push(dim.clone()); data.len()
used[i] = true; )));
matched = true; }
let (Some((&row, outer)), Some(&row_stride)) = (extent.split_last(), shape.last()) else {
return Ok(data);
};
let mut out = vec![fill; total];
if row == 0 {
return Ok(out);
}
// The position of the current row along each outer axis.
let mut index = vec![0usize; outer.len()];
for chunk in data.chunks_exact(row) {
let offset = index.iter().zip(&shape).fold(0, |o, (&i, &n)| o * n + i);
out[offset * row_stride..][..row].clone_from_slice(chunk);
for (i, &n) in index.iter_mut().zip(outer).rev() {
*i += 1;
if *i < n {
break; break;
} }
} *i = 0;
if !matched {
// Create an anonymous dimension for unmatched sizes
result.push(Dimension {
name: format!("dim_{dim_size}"),
size: dim_size,
is_unlimited: false,
});
} }
} }
Ok(out)
}
result /// netCDF's default fill value for a type (`NC_FILL_*` in `netcdf.h`).
fn default_fill(nc_type: NcType) -> FillValue {
match nc_type {
NcType::Byte => FillValue::Int(-127),
NcType::UByte => FillValue::UInt(255),
NcType::Short => FillValue::Int(-32767),
NcType::UShort => FillValue::UInt(65535),
NcType::Int => FillValue::Int(-2_147_483_647),
NcType::UInt => FillValue::UInt(4_294_967_295),
NcType::Int64 => FillValue::Int(-9_223_372_036_854_775_806),
NcType::UInt64 => FillValue::UInt(18_446_744_073_709_551_614),
NcType::Float => FillValue::Float(f64::from(9.969_21e36_f32)),
NcType::Double => FillValue::Float(9.969_209_968_386_869e36),
NcType::String => FillValue::String(String::new()),
NcType::Char => FillValue::Int(0),
}
}
/// A fill value converted to the element type of a read, as the read
/// converts the stored values.
trait FromFill {
fn from_fill(fill: &FillValue) -> Self;
}
macro_rules! numeric_from_fill {
($($t:ty),*) => {$(
impl FromFill for $t {
fn from_fill(fill: &FillValue) -> Self {
match fill {
FillValue::Float(v) => *v as $t,
FillValue::Int(v) => *v as $t,
FillValue::UInt(v) => *v as $t,
FillValue::String(_) => <$t>::default(),
}
}
}
)*};
}
numeric_from_fill!(f64, f32, i32, i64, u64);
impl FromFill for String {
fn from_fill(fill: &FillValue) -> Self {
match fill {
FillValue::String(s) => s.clone(),
_ => String::new(),
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn pad_places_rows() {
// (2, 1) written of (2, 4): each row padded, not the tail.
let out = pad(vec![1, 2], &[2, 1], &[2, 4], 0).unwrap();
assert_eq!(out, vec![1, 0, 0, 0, 2, 0, 0, 0]);
// Leading axis short.
let out = pad(vec![1, 2, 3, 4], &[2, 2], &[3, 2], -1).unwrap();
assert_eq!(out, vec![1, 2, 3, 4, -1, -1]);
// Nothing written.
let out = pad(Vec::<i32>::new(), &[0, 3], &[2, 3], 7).unwrap();
assert_eq!(out, vec![7; 6]);
// 3-D, middle axis short.
let out = pad(vec![1, 2, 3, 4], &[2, 1, 2], &[2, 2, 2], 0).unwrap();
assert_eq!(out, vec![1, 2, 0, 0, 3, 4, 0, 0]);
}
#[test]
fn pad_rejects_wrong_length() {
assert!(pad(vec![1, 2, 3], &[2, 2], &[3, 2], 0).is_err());
assert!(pad(vec![1, 2, 3, 4], &[2, 2], &[1, 4], 0).is_err());
}
} }
+382 -1
View File
@@ -4,7 +4,7 @@
use std::process::Command; use std::process::Command;
use clawhdf5_netcdf4::{AttrValue, NetCDF4File}; use clawhdf5_netcdf4::{AttrValue, NcType, NetCDF4File};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Helpers // Helpers
@@ -461,3 +461,384 @@ with nc.Dataset({path:?}) as f:
} }
assert_eq!(got, expected); assert_eq!(got, expected);
} }
// ===========================================================================
// Variables' dimensions, shapes and values as netCDF4-python reports them
// ===========================================================================
/// Whether python can import `module`.
fn python_has(module: &str) -> bool {
Command::new(python())
.args(["-c", &format!("import {module}")])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
/// h5netcdf is not in every interop environment (CI installs it; a local
/// `.venv` may not have it), so its tests skip without it even under
/// `CLAWHDF5_REQUIRE_INTEROP=1`.
macro_rules! skip_if_no_h5netcdf {
() => {
if !python_has("h5netcdf") {
eprintln!("SKIP: python3 with h5netcdf not available");
return;
}
};
}
/// Every variable of the file at `path`, in every group, as netCDF4-python
/// reports it: `"<group path> <name> (<dims>) (<shape>)"` and its values
/// (numeric variables; element by element with masking off, so unwritten
/// records are the fill value), sorted by the description.
///
/// Values are read one element at a time because netCDF-C 4.9.3 lays out a
/// whole-variable read of a variable shorter than an unlimited dimension
/// that is not its first wrongly (the written values first, then the fill);
/// element reads, and reads of one index of the leading axis, are right.
fn netcdf4_view(path: &std::path::Path) -> Vec<(String, Vec<f64>)> {
let script = r#"
import sys
import numpy as np
import netCDF4 as nc
def walk(g):
for name, v in g.variables.items():
v.set_auto_mask(False)
head = "%s %s (%s) (%s)" % (g.path, name, ",".join(v.dimensions), ",".join(map(str, v.shape)))
vals = []
if v.dtype != str and v.dtype.kind in "iuf":
vals = [repr(float(v[i])) for i in np.ndindex(v.shape)]
print(head + "|" + " ".join(vals))
for sub in g.groups.values():
walk(sub)
with nc.Dataset(sys.argv[1]) as f:
walk(f)
"#;
let out = Command::new(python())
.args(["-c", script, &path.display().to_string()])
.output()
.expect("failed to run python3");
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let mut view: Vec<(String, Vec<f64>)> = String::from_utf8(out.stdout)
.unwrap()
.lines()
.map(|line| {
let (head, vals) = line.split_once('|').unwrap();
let vals = vals
.split_whitespace()
.map(|v| v.parse().unwrap())
.collect();
(head.to_string(), vals)
})
.collect();
view.sort_by(|a, b| a.0.cmp(&b.0));
view
}
/// The same view of the file through clawhdf5-netcdf4.
fn clawhdf5_view(path: &std::path::Path) -> Vec<(String, Vec<f64>)> {
fn describe(
group_path: &str,
vars: Vec<clawhdf5_netcdf4::Variable<'_>>,
) -> Vec<(String, Vec<f64>)> {
vars.into_iter()
.map(|v| {
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
let shape: Vec<String> = v.shape().unwrap().iter().map(u64::to_string).collect();
let head = format!(
"{group_path} {} ({}) ({})",
v.name(),
dims.join(","),
shape.join(",")
);
let vals = match v.nc_type().unwrap() {
NcType::String | NcType::Char => Vec::new(),
_ => v.read_raw_f64().unwrap(),
};
(head, vals)
})
.collect()
}
fn walk(
group_path: &str,
group: &clawhdf5_netcdf4::NetCDF4Group<'_>,
out: &mut Vec<(String, Vec<f64>)>,
) {
out.extend(describe(group_path, group.variables().unwrap()));
for name in group.group_names().unwrap() {
walk(
&format!("{group_path}/{name}"),
&group.group(&name).unwrap(),
out,
);
}
}
let file = NetCDF4File::open(path).unwrap();
let mut view = describe("/", file.variables().unwrap());
for name in file.group_names().unwrap() {
walk(&format!("/{name}"), &file.group(&name).unwrap(), &mut view);
}
view.sort_by(|a, b| a.0.cmp(&b.0));
view
}
/// clawhdf5-netcdf4 reports the same variables, dimensions, shapes and
/// values (bit for bit, NaN equal to NaN) as netCDF4-python.
fn assert_same_view(path: &std::path::Path) {
let want = netcdf4_view(path);
let got = clawhdf5_view(path);
let heads = |v: &[(String, Vec<f64>)]| v.iter().map(|(h, _)| h.clone()).collect::<Vec<_>>();
assert_eq!(heads(&got), heads(&want), "variables differ from netCDF4's");
for ((head, got), (_, want)) in got.iter().zip(&want) {
let same = got.len() == want.len()
&& got
.iter()
.zip(want)
.all(|(a, b)| a.to_bits() == b.to_bits() || (a.is_nan() && b.is_nan()));
assert!(same, "{head}: got {got:?}, netCDF4 reads {want:?}");
}
}
/// The reproducer of the known-issues entry: `a` is on the unlimited `time`
/// (5 long through `b`) with 2 records, not on an anonymous `dim_2`; the
/// pure dimension scales `time` and `empty` are not variables; `a` has
/// shape (5,) and reads its 3 unwritten records as the fill value.
#[test]
fn variable_dimensions_come_from_the_file() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("repro.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("time", None)
f.createDimension("empty", None)
f.createDimension("x", 3)
f.createVariable("a", "i4", ("time",))[0:2] = [1, 2]
f.createVariable("b", "f4", ("time", "x"))[0:5, :] = np.arange(15).reshape(5, 3)
f.createVariable("e", "i4", ("empty",))
f.createVariable("c", "i4", ("x",))[:] = [7, 8, 9]
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let mut names = file.variable_names().unwrap();
names.sort();
assert_eq!(names, ["a", "b", "c", "e"]);
assert!(matches!(
file.variable("time"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
let a = file.variable("a").unwrap();
assert_eq!(a.dimensions()[0].name, "time");
assert_eq!(a.shape().unwrap(), [5]);
assert_eq!(a.stored_shape().unwrap(), [2]);
assert_eq!(
a.read_raw_i32().unwrap(),
[1, 2, -2_147_483_647, -2_147_483_647, -2_147_483_647]
);
}
/// Dimensions of one size are told apart by the file, not by order: `p`
/// and `q` are both 2 long, and `v(q, p)`, `same(p, p)` (one dimension
/// twice), a scalar, `q`'s coordinate variable, a variable called `p` that
/// is not `p`'s coordinate variable (stored as `_nc4_non_coord_p`), and
/// variables in a subgroup and a sub-subgroup on dimensions of their
/// ancestors.
#[test]
fn equal_size_and_inherited_dimensions_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("dims.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("p", 2)
f.createDimension("q", 2)
f.createVariable("v", "i4", ("q", "p"))[:] = np.array([[1, 2], [3, 4]])
f.createVariable("same", "i4", ("p", "p"))[:] = np.array([[5, 6], [7, 8]])
f.createVariable("s", "f8", ())[...] = 3.5
f.createVariable("q", "f4", ("q",))[:] = [0, 1]
f.createVariable("p", "f4", ("q", "p"))[:] = np.array([[0, 1], [2, 3]])
g = f.createGroup("g")
g.createDimension("r", 2)
g.createVariable("w", "i4", ("r", "q", "p"))[:] = np.arange(8).reshape(2, 2, 2)
h = g.createGroup("h")
h.createVariable("z", "i4", ("p", "r"))[:] = np.array([[1, 2], [3, 4]])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let v = file.variable("v").unwrap();
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
assert_eq!(dims, ["q", "p"]);
let p = file.variable("p").unwrap();
assert!(!p.is_coordinate());
assert!(file.variable("q").unwrap().is_coordinate());
let s = file.variable("s").unwrap();
assert!(s.dimensions().is_empty());
assert_eq!(s.shape().unwrap(), Vec::<u64>::new());
let z = file
.group("g")
.unwrap()
.group("h")
.unwrap()
.variable("z")
.unwrap();
let dims: Vec<&str> = z.dimensions().iter().map(|d| d.name.as_str()).collect();
assert_eq!(dims, ["p", "r"]);
}
/// Variables shorter than their unlimited dimension have its length and
/// read the fill value (`_FillValue`, else netCDF's default for the type)
/// where nothing was written — also when the unlimited dimension is not
/// the first; `read_f64` gives NaN there.
#[test]
fn unwritten_records_read_as_fill_like_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("pad.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("t", None)
f.createDimension("x", 2)
f.createVariable("a", "i4", ("t",))[0:2] = [1, 2]
f.createVariable("f", "f4", ("x", "t"), fill_value=-5.0)[:, 0:1] = np.array([[1], [2]])
f.createVariable("d", "f8", ("t",))[0:4] = [1, 2, 3, 4]
f.createVariable("u", "u8", ("t",))[0:1] = [1]
f.createVariable("b", "i1", ("t", "x"))[0:3, :] = np.ones((3, 2))
f.createVariable("st", str, ("t",))[0] = "hi"
g = f.createGroup("g")
g.createVariable("k", "f4", ("t",))[0:1] = [9]
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let mut f = file.variable("f").unwrap();
assert_eq!(f.shape().unwrap(), [2, 4]);
assert_eq!(f.stored_shape().unwrap(), [2, 1]);
assert_eq!(
f.read_raw_f32().unwrap(),
[1.0, -5.0, -5.0, -5.0, 2.0, -5.0, -5.0, -5.0]
);
let read = f.read_f64().unwrap();
assert_eq!(read[0], 1.0);
assert!(read[1].is_nan() && read[7].is_nan());
let st = file.variable("st").unwrap();
assert_eq!(st.read_string().unwrap(), ["hi", "", "", ""]);
assert_eq!(st.shape().unwrap(), [4]);
}
/// A file with HDF5 dimension scales but none of netCDF's own attributes
/// (h5py's `dims` API): the dimensions come from `DIMENSION_LIST`, so
/// `v(q, p)` is not `v(p, q)` although both are 2 long; with two scales
/// attached to one axis (`w`), netCDF-C takes the last.
#[test]
fn h5py_dimension_scales_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("scales.h5");
run_python(&format!(
r#"
import h5py
import numpy as np
with h5py.File({path:?}, "w") as f:
f["p"] = np.arange(2.0)
f["q"] = np.arange(2.0) + 10
f["p"].make_scale("p")
f["q"].make_scale("q")
f["v"] = np.arange(4).reshape(2, 2)
f["v"].dims[0].attach_scale(f["q"])
f["v"].dims[1].attach_scale(f["p"])
f["w"] = np.arange(2)
f["w"].dims[0].attach_scale(f["p"])
f["w"].dims[0].attach_scale(f["q"])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
/// Files h5netcdf writes (its own implementation of the netCDF-4
/// conventions over h5py): an unlimited dimension, equal sizes, a subgroup
/// on inherited dimensions, a scalar.
#[test]
fn h5netcdf_file_matches_netcdf4_python() {
skip_if_no_netcdf4!();
skip_if_no_h5netcdf!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("h5netcdf.nc");
run_python(&format!(
r#"
import h5netcdf
import numpy as np
with h5netcdf.File({path:?}, "w") as f:
f.dimensions = {{"p": 2, "q": 2, "t": None}}
f.create_variable("v", ("q", "p"), "i4")[...] = np.array([[1, 2], [3, 4]])
f.create_variable("q", ("q",), "f4")[...] = [0, 1]
f.create_variable("same", ("p", "p"), "i4")[...] = np.array([[5, 6], [7, 8]])
a = f.create_variable("a", ("t", "p"), "f8")
f.resize_dimension("t", 3)
a[...] = np.ones((3, 2))
f.create_variable("short", ("t",), "i4")
g = f.create_group("g")
g.dimensions = {{"r": 2}}
g.create_variable("w", ("r", "q", "p"), "i4")[...] = np.arange(8).reshape(2, 2, 2)
g.create_variable("s", (), "f8")[...] = 2.5
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
/// Files xarray writes, through netCDF4 and (when installed) h5netcdf:
/// coordinates, two dimensions of one size, an unlimited dimension.
#[test]
fn xarray_files_match_netcdf4_python() {
skip_if_no_netcdf4!();
skip_if_no_xarray!();
let dir = tempfile::tempdir().unwrap();
let mut engines = vec!["netcdf4"];
if python_has("h5netcdf") {
engines.push("h5netcdf");
} else {
eprintln!("SKIP: xarray with engine h5netcdf (h5netcdf not available)");
}
for engine in engines {
let path = dir.path().join(format!("xarray_{engine}.nc"));
run_python(&format!(
r#"
import numpy as np
import xarray as xr
ds = xr.Dataset(
{{
"temp": (("time", "lat", "lon"), np.arange(12.0).reshape(3, 2, 2)),
"grid": (("lon", "lat"), np.array([[1, 2], [3, 4]], dtype="i4")),
"scalar": ((), 1.5),
}},
coords={{"time": [0.0, 6.0, 12.0], "lat": [10.0, 20.0], "lon": [5.0, 6.0]}},
)
ds.to_netcdf({path:?}, engine={engine:?}, unlimited_dims=["time"])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
}
@@ -754,3 +754,53 @@ fn test_dimension_struct_equality() {
}; };
assert_ne!(d1, d3); assert_ne!(d1, d3);
} }
/// A dimension scale that is only a dimension (netCDF-C's `NAME`) is not a
/// variable, and `_nc4_non_coord_<name>` is the variable `<name>`, found in
/// place of the scale of the same name.
#[test]
fn test_pure_dimensions_hidden_and_non_coord_names() {
let pure = "This is a netCDF dimension but not a netCDF variable. 2";
let mut b = FileBuilder::new();
b.create_dataset("x")
.with_f32_data(&[0.0, 0.0])
.with_shape(&[2])
.set_attr("CLASS", AttrValue::String("DIMENSION_SCALE".into()))
.set_attr("NAME", AttrValue::String(pure.into()))
.set_attr("_Netcdf4Dimid", AttrValue::I64(0));
b.create_dataset("_nc4_non_coord_x")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_shape(&[3]);
b.create_dataset("v")
.with_f64_data(&[5.0, 6.0])
.with_shape(&[2]);
let file = NetCDF4File::from_bytes(b.finish().unwrap()).unwrap();
let dims = file.dimensions().unwrap();
assert_eq!(dims.len(), 1);
assert_eq!(dims[0].name, "x");
let mut names = file.variable_names().unwrap();
names.sort();
assert_eq!(names, ["v", "x"]);
let x = file.variable("x").unwrap();
assert_eq!(x.name(), "x");
assert_eq!(x.read_raw_f64().unwrap(), [1.0, 2.0, 3.0]);
assert!(!x.is_coordinate());
// No DIMENSION_LIST: `v` gets `x` by size, as before.
assert_eq!(file.variable("v").unwrap().dimensions()[0].name, "x");
}
/// `variable` still takes a path relative to the group, as it did when it
/// opened the dataset by path.
#[test]
fn test_variable_by_path() {
let file = NetCDF4File::from_bytes(make_grouped_netcdf4()).unwrap();
let pressure = file.variable("surface/pressure").unwrap();
assert_eq!(pressure.name(), "pressure");
assert_eq!(pressure.read_raw_f64().unwrap(), [1013.25, 1012.0, 1011.5]);
assert!(file.variable("/time").is_ok());
assert!(matches!(
file.variable("nowhere/pressure"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
}
+53 -22
View File
@@ -15,9 +15,6 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since | | Issue | Kind | Since |
|---|---|---| |---|---|---|
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 | | [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | | [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | | [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
@@ -413,25 +410,6 @@ always expose it. A fix belongs in `conformance/ref_bugs.py` (more or more
varied reads for this object) or in documenting the file as a known varied reads for this object) or in documenting the file as a known
refusal; neither is done. refusal; neither is done.
## NetCDF-4: variables' dimensions are guessed from sizes
**Status:** open (found 2026-09-28 while fixing unlimited dimension sizes).
`clawhdf5-netcdf4` gives each variable the dimensions it finds by size
(`match_dimensions_to_variable`: the first unused dimension of equal size,
else an anonymous `dim_<n>`), not the ones its `DIMENSION_LIST` names, and
`variables()` also lists the dimension scales that are only dimensions
(netCDF-C hides them). With netCDF4-python: unlimited `time` and `empty`,
`x` (3), `a(time)` with 2 records, `b(time, x)` with 5, `e(empty)` — netCDF4
reports `time` = 5, `a` on `time`, and variables `a`, `b`, `e` and a `c(x)`;
clawhdf5-netcdf4 reports `time` = 5 (right) but `a` on `dim_2`, and also
variables `time` and `empty` (the scales, put on `empty`). Two dimensions of
one size can be swapped the same way. Related: netCDF4 gives `a` the shape
(5,) (a variable along an unlimited dimension has the dimension's length,
unwritten records read as fill); `Variable::shape` is the HDF5 extent,
`[2]`, and reads return those 2 values. `dimensions()` is right.
Workaround: read the variable's `_Netcdf4Coordinates` attribute (dimension
ids, matching each scale's `_Netcdf4Dimid`).
## Small floats decode as libhdf5 does, not as the OCP MX specification ## Small floats decode as libhdf5 does, not as the OCP MX specification
**Status:** open, deliberate (documented 2026-09-28). HDF5 2.x predefines **Status:** open, deliberate (documented 2026-09-28). HDF5 2.x predefines
@@ -484,6 +462,59 @@ Newest first. "Before any release" means no tagged release (v2.7.0 and
earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date
given. given.
## NetCDF-4: variables' dimensions are guessed from sizes
**Status:** fixed 2026-09-28 (branch `fix/netcdf-dimension-list`). Affected
every release (v2.1.0 to v2.7.0: size matching dates from the crate's
first version). Wrong metadata only: stored values were always read right.
Users who read `_Netcdf4Coordinates` themselves can use
`Variable::dimensions` again; code that relied on `Variable::shape` being
the HDF5 extent, or on the reads returning only the written records, should
use `Variable::stored_shape` (new) — `shape` and the reads now follow
netCDF (below). `variables()` no longer lists pure dimension scales.
Found 2026-09-28 while fixing unlimited dimension sizes.
`clawhdf5-netcdf4` gave each variable the dimensions it found by size
(`match_dimensions_to_variable`: the first unused dimension of equal size,
else an anonymous `dim_<n>`), not the ones its `DIMENSION_LIST` names, and
`variables()` also listed the dimension scales that are only dimensions
(netCDF-C hides them). With netCDF4-python: unlimited `time` and `empty`,
`x` (3), `a(time)` with 2 records, `b(time, x)` with 5, `e(empty)` — netCDF4
reports `time` = 5, `a` on `time`, and variables `a`, `b`, `e` and a `c(x)`;
clawhdf5-netcdf4 reported `time` = 5 (right) but `a` on `dim_2`, and also
variables `time` and `empty` (the scales, put on `empty`). Two dimensions of
one size could be swapped the same way. Related: netCDF4 gives `a` the shape
(5,) (a variable along an unlimited dimension has the dimension's length,
unwritten records read as fill); `Variable::shape` was the HDF5 extent,
`[2]`, and reads returned those 2 values. `dimensions()` was right. The
workaround was to read the variable's `_Netcdf4Coordinates` attribute
(dimension ids, matching each scale's `_Netcdf4Dimid`).
Variables now get their dimensions as netCDF-C resolves them
(`libhdf5/hdf5open.c`): `_Netcdf4Coordinates` ids, else the scales
`DIMENSION_LIST` references, looked up in the variable's group and its
parents; size matching remains only for axes with neither (files not
written by a netCDF library). Pure dimension scales are not variables, and
`_nc4_non_coord_<name>` datasets are the variables `<name>`. A variable
along an unlimited dimension has the dimension's length, and its unwritten
records read as the fill value. One difference from netCDF-C 4.9.3 is
deliberate: when the unlimited dimension is not a variable's first
(`f(x, t)` with 1 of 4 records), a whole-variable read through netCDF-C
returns the written values first and then the fill
(`[[1, 2, fill, fill], [fill, ...]]` for rows `[1]` and `[2]`), while its
element and row reads — and clawhdf5-netcdf4 — place each row's values in
their row (`[[1, fill, fill, fill], [2, fill, fill, fill]]`).
`interop_tests` compares variables, dimensions, shapes and every value
with netCDF4-python 1.7.4 for the reproducer, dimensions of equal size,
one dimension used twice, a scalar, subgroups on their parents'
dimensions, unwritten records with and without `_FillValue`, h5py
dimension scales, and h5netcdf 1.8.1 and xarray files (tank, 2026-09-28,
`CLAWHDF5_PYTHON=<venv with h5netcdf> cargo test -p clawhdf5-netcdf4`).
Files without dimension scales still get dimensions by size (netCDF-C
gives them `phony_dim_<n>`), as before; see `CHANGELOG.md` for a
comparison over the conformance corpus's netCDF-readable files.
## A dropped `FileEditor` could keep its file locked for a moment ## A dropped `FileEditor` could keep its file locked for a moment
**Status:** fixed 2026-09-28 (#23), before any release **Status:** fixed 2026-09-28 (#23), before any release