Merge order: this is 1 of 3 (#27 → #28 → #29). The later PRs are stacked on this one.
clawhdf5-netcdf4 now takes each variable's dimensions from the file, the way netCDF-C does, instead of guessing them by size:
How a variable's dimensions are found, in order:
a coordinate variable is its own dimension scale;
otherwise the ids in _Netcdf4Coordinates are matched to each scale's _Netcdf4Dimid;
otherwise the DIMENSION_LIST references are used (old and revised references);
the search runs from the variable's group up through each parent;
size matching remains only for axes the file names nothing for.
Hidden scales: dimension scales that are only dimensions (netCDF's "not a netCDF variable" NAME) no longer appear in variables(). A dataset named _nc4_non_coord_<name> is listed as <name>.
Unlimited dimensions: a variable shorter than its unlimited dimension now has that dimension's length, as in netCDF4. Unwritten records read as the fill value: _FillValue, else netCDF's default for the type. The HDF5 extent is available as Variable::stored_shape().
One deliberate difference from netCDF-C 4.9.3: for f(x, t) with t unlimited, we put each row's values in their own row. netCDF-C's whole-variable read puts all the written values first; its single-element and single-row reads agree with ours. This is documented.
API:shape and the reads now pad as described. variable_names() and variables() omit pure dimension scales. Nothing else in the workspace uses this crate, and ClawBrainHub doesn't either.
Tests: six new interop tests, compared with netCDF4-python 1.7.4 (netCDF-C 4.9.3). They cover:
the known-issues reproducer;
swapped equal-size dimensions, (x, x), a scalar, and subgroups using their ancestors' dimensions;
unwritten records across several types;
h5py dimension scales;
h5netcdf files, and xarray files with both engines.
With the new code disabled, 6 of the 11 interop tests fail. CI now installs h5netcdf. Files with no dimension scales still use dim_<size> rather than netCDF-C's phony_dim_<n>.
**Merge order: this is 1 of 3 (#27 → #28 → #29).** The later PRs are stacked on this one.
`clawhdf5-netcdf4` now takes each variable's dimensions from the file, the way netCDF-C does, instead of guessing them by size:
- **How a variable's dimensions are found**, in order:
- a coordinate variable is its own dimension scale;
- otherwise the ids in `_Netcdf4Coordinates` are matched to each scale's `_Netcdf4Dimid`;
- otherwise the `DIMENSION_LIST` references are used (old and revised references);
- the search runs from the variable's group up through each parent;
- size matching remains only for axes the file names nothing for.
- **Hidden scales:** dimension scales that are only dimensions (netCDF's "not a netCDF variable" `NAME`) no longer appear in `variables()`. A dataset named `_nc4_non_coord_<name>` is listed as `<name>`.
- **Unlimited dimensions:** a variable shorter than its unlimited dimension now has that dimension's length, as in netCDF4. Unwritten records read as the fill value: `_FillValue`, else netCDF's default for the type. The HDF5 extent is available as `Variable::stored_shape()`.
- **One deliberate difference from netCDF-C 4.9.3:** for `f(x, t)` with `t` unlimited, we put each row's values in their own row. netCDF-C's whole-variable read puts all the written values first; its single-element and single-row reads agree with ours. This is documented.
**API:** `shape` and the reads now pad as described. `variable_names()` and `variables()` omit pure dimension scales. Nothing else in the workspace uses this crate, and ClawBrainHub doesn't either.
**Tests:** six new interop tests, compared with netCDF4-python 1.7.4 (netCDF-C 4.9.3). They cover:
- the known-issues reproducer;
- swapped equal-size dimensions, `(x, x)`, a scalar, and subgroups using their ancestors' dimensions;
- unwritten records across several types;
- h5py dimension scales;
- h5netcdf files, and xarray files with both engines.
With the new code disabled, 6 of the 11 interop tests fail. CI now installs `h5netcdf`. Files with no dimension scales still use `dim_<size>` rather than netCDF-C's `phony_dim_<n>`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Variables got the first unused dimension of equal size, so a variable on
an unlimited dimension with fewer records got an anonymous dim_<n>, and
dimensions of one size could be swapped. Resolve them as netCDF-C does
(libhdf5/hdf5open.c): _Netcdf4Coordinates ids, else the scales
DIMENSION_LIST references (the last one attached to an axis), searched in
the variable's group and its parents; a coordinate variable is on its own
scale. Size matching remains only for axes the file names nothing for.
variables()/variable_names() leave out dimension scales that are only
dimensions, and _nc4_non_coord_<name> is the variable <name>.
Variable::shape is the netCDF shape (an unlimited dimension's length) and
the reads pad unwritten records with the fill value (_FillValue, else
NC_FILL_*; NaN from read_f64); Variable::stored_shape is the HDF5 extent.
New NetCDF4File::variable_names.
Tests compare with netCDF4-python variable by variable: the known-issues
reproducer, equal sizes, (p, p), scalars, inherited dimensions, unwritten
records, h5py dimension scales, h5netcdf and xarray files. CI installs
h5netcdf. known-issues entry moved to Fixed (history); stale open-table
row for the unlimited-size fix removed.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Merge order: this is 1 of 3 (#27 → #28 → #29). The later PRs are stacked on this one.
clawhdf5-netcdf4now takes each variable's dimensions from the file, the way netCDF-C does, instead of guessing them by size:_Netcdf4Coordinatesare matched to each scale's_Netcdf4Dimid;DIMENSION_LISTreferences are used (old and revised references);NAME) no longer appear invariables(). A dataset named_nc4_non_coord_<name>is listed as<name>._FillValue, else netCDF's default for the type. The HDF5 extent is available asVariable::stored_shape().f(x, t)withtunlimited, we put each row's values in their own row. netCDF-C's whole-variable read puts all the written values first; its single-element and single-row reads agree with ours. This is documented.API:
shapeand the reads now pad as described.variable_names()andvariables()omit pure dimension scales. Nothing else in the workspace uses this crate, and ClawBrainHub doesn't either.Tests: six new interop tests, compared with netCDF4-python 1.7.4 (netCDF-C 4.9.3). They cover:
(x, x), a scalar, and subgroups using their ancestors' dimensions;With the new code disabled, 6 of the 11 interop tests fail. CI now installs
h5netcdf. Files with no dimension scales still usedim_<size>rather than netCDF-C'sphony_dim_<n>.🤖 Generated with Claude Code
View command line instructions
Checkout
From your project repository, check out a new branch and test the changes.