HDF5 2.0 native complex as a first-class type; Python libver= #31

Open
osobh wants to merge 2 commits from feat/complex-first-class into main
37 changed files with 3251 additions and 605 deletions
+117
View File
@@ -2,6 +2,123 @@
## Unreleased ## Unreleased
### NetCDF-4: phony dimensions, skipped types and order as in netCDF-C (2026-09-29)
- `clawhdf5-netcdf4` now reads a file's metadata as netCDF-C 4.9.3 does
(`libhdf5/hdf5open.c`), for the whole file on first use
(`src/model.rs`): links in creation order when the group tracks it,
else in name order; a group's datasets before its subgroups; dimension
ids file-wide (`_Netcdf4Dimid`, else the next free id); then variables'
dimensions, subgroups first: `_Netcdf4Coordinates` ids (looked up
file-wide), else the scales `DIMENSION_LIST` attaches (when the first
axis has one), else phony dimensions.
- **Files without dimension scales** get netCDF-C's phony dimensions,
`phony_dim_<id>` (`create_phony_dims`): per axis, the first dimension of
the variable's group of the same length and unlimitedness not used by an
earlier axis of the variable (real dimensions included), else a new one;
a length of 0 is always unlimited. They used to be one dimension per 1-D
dataset, named after it, and `dim_<size>` for other axes.
- **Datasets netCDF-C skips are not variables:** references, bit fields,
time and array types, and compounds, enums and VLENs whose members or
base type are not netCDF atomic types (4- and 8-byte floats only) or a
type read before — replaying netCDF-C's file-wide type list, which also
keeps a type it failed to read, so the second dataset of a compound with
a reference member is a variable, as in netCDF-C.
- `NcType` gains `Enum`, `Compound`, `VLen` and `Opaque` and is now
`#[non_exhaustive]` (breaking for exhaustive matches);
`Variable::nc_type` is the type netCDF-C gives the variable: a 1-byte
fixed-length string is `Char`, a longer one `String` (both were
`String`); user-defined types have their class (they were `Char`).
`dtype_to_nctype` maps compounds and enums to their classes.
- Groups (`group_names`), variables (`variables`, `variable_names`) and
dimensions (`dimensions`, by id) come in netCDF-C's order; a zero-length
dimension scale is unlimited; an unlimited dimension's length is the
longest extent along it of the variables in its group and below
(`nc4_find_dim_len`; was: of the variables its `REFERENCE_LIST` names).
`NetCDF4File::group` and `NetCDF4Group::group` take a path (`"a/b"`).
- Deliberate differences: floats of other than 4 or 8 bytes are
`Float`/`Double` (netCDF-C 4.9.3 on libhdf5 1.14.6 labels them
`NC_STRING`); an axis netCDF-C leaves without a dimension (where it
reads uninitialised memory) gets one by the phony rule.
- New `clawhdf5_format::group_v2::links_in_creation_order_in`: a group's
link names in creation order, or `None` when it does not track it.
- Tests compare with netCDF-C itself (`tests/netcdf_c_view.py` calls the
libnetcdf netCDF4-python bundles through ctypes, since netCDF4-python
hides variables of types it does not support): `interop_tests` cases of
h5py files without dimension scales (sharing, unlimited and zero-length
axes, subgroups, a real dimension taken by length), of every HDF5 type
class, of creation and name order in compact and dense groups, and a
netCDF4-python file; and the gated `tests/corpus_vs_netcdf_c.rs`
(`CLAWHDF5_NETCDF_CORPUS=<dir>`): over the conformance corpus 420 of the
429 files netCDF-C 4.9.3 opens match in groups, dimensions, variables,
types, shapes and numeric values (up to 5000 elements); `main` at
`4260af4` matched 68. The other 9 (external links, values the HDF5
reader refuses) are explained in `tests/corpus_known_differences.txt`
and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4).
Affected v2.1.0 to v2.7.0.
### HDF5 2.0 native complex is its own type on read; Python `libver=` (2026-09-29)
- **`Datatype::parse` returns `Datatype::Complex { size, base_type }`** for
a class-11 message, also as a compound member, array base or
variable-length base. It used to return the equivalent `{r, i}`
compound (`docs/known-issues.md`, "HDF5 2.0 native complex numbers read
as a `{r, i}` compound"). **Breaking** for code that matched the
compound view of a native complex type: `Dataset::raw_datatype()` and
attribute datatypes are now `Datatype::Complex`;
`Datatype::complex_as_compound(size, base)` still gives the `{r, i}` view,
and `data_read::read_compound_fields` accepts a `Complex` directly.
Re-serializing a parsed type (e.g. copying it to another file) now writes
class 11 again instead of a compound.
- **Facade:** `DType::Complex(Box<DType>)` (new variant, **breaking** for
exhaustive `match`es over `DType`): `Complex(F32)` is numpy `complex64`,
`Complex(F64)` `complex128`, a binary16 part `Complex(Other("float16"))`
(the same `Other` name small floats already use); `Display` prints
`complex<f64>`. h5py's own complex encoding, the compound `{r, i}`, is
still `DType::Compound`. `read_complex_f64`/`read_complex_f32` read both.
- **`h5rs`:** `dump` prints native complex as h5dump 2.2.0 does:
`H5T_COMPLEX_IEEE_F{16,32,64}{LE,BE}` (else `H5T_COMPLEX { <base> }`), in
arrays and compounds too, and values as `1.5-2i` (`%g%+gi`; for binary16
parts, which h5dump has no C type for, `1+-2i` as it prints them). `ls`
shows `complex64`, `complex128-be`, `complex32` (and `complex<part>` for
non-IEEE parts); `ls -v` shows h5ls 2.2.0's `complex number of` /
`IEEE 64-bit big-endian float`. `diff` compares complex values part by
part and prints the difference as h5diff 2.2.0 does (`1+0i`); a native
complex and an `{r, i}` compound are no longer comparable (different
classes, as in h5diff). `dump --json` keeps the `{r, i}` compound and
`[re, im]` values: hdf5-json (h5json 2.0.0) has no complex class.
- **Browser (`clawhdf5-wasm`):** native complex datasets are readable, as
`[re, im]` pairs: `info().elementShape` ends in `2`, `read()` returns the
parts interleaved in a `Float32Array`/`Float64Array` (binary16 parts
widened to f32), and `dtype` is `complex<f64>` (`complex<f32
(big-endian)>`, `array[2]<complex<f32>>`). The viewer shows each element
as `[re, im]` with no change to the page. h5py's `{r, i}` compound is
still refused, as every compound is.
- **Python:** `clawhdf5.File(path, 'w', libver=...)` takes h5py's values:
`'v108'`, `'v110'`, `'v112'`, `'v114'`, `'v200'`, `'latest'` (the low
bound; high `'latest'`, as h5py) or a `(low, high)` tuple, mapped to
`FileBuilder::libver_bounds`. `'earliest'` as the low bound writes the
1.8 format with a `UserWarning` (clawhdf5 cannot write the pre-1.8
format); as the high bound it is a `ValueError`, as are unknown names and
a low bound above the high one. It is ignored for `'r'` and refused
(`NotImplementedError`) for `'r+'`/`'a'`. Native complex already read as
numpy `complex64`/`complex128` and still does
(`test_read_native_complex_from_h5py`).
- Tests (tank, 2026-09-29): `crates/clawhdf5-tools/tests/native_complex_dump.rs`
compares `h5rs dump` with h5dump 2.2.0's output stored next to a fixture
written by h5py 3.16 / libhdf5 2.0.0
(`crates/clawhdf5/tests/fixtures/gen_native_complex.py`:
`F32LE`/`F64LE`/`F64BE`/`F16LE`, scalar, compound member, array, root
attribute), line for line outside the values and value for value within
them, and `ls`/`diff` with h5ls/h5diff 2.2.0;
`integration_tests::complex_datasets_and_attributes_round_trip`
(`DType::Complex`, `raw_datatype()`); `clawhdf5-wasm` unit tests on the
fixture, `h5py_interop` and `examples/wasm-viewer/test` (Node:
`test.mjs`, 274 + 1324 checks; the page in headless Chromium for
`/native_c128`, `/native_c64_be`, `/pairs`) with two native complex
datasets added to `make_fixture.py` (when h5py's libhdf5 is 2.0+);
`crates/clawhdf5-py/tests/test_libver.py` (every bound; `'v108'` output
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
hard-codes the fixture's length.
### NetCDF-4: variables' dimensions come from the file (2026-09-28) ### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of - `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
+3 -3
View File
@@ -115,12 +115,12 @@ Limits and open issues, with dates, are in
|---|---|---|---| |---|---|---|---|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) | | **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | | **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 | | **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more | | **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | | **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) | | **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | | **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) | | **Bindings** | Python (read, `'w'` for numeric and complex arrays with h5py's `libver=`, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`; HDF5 2.0 native complex as `[re, im]` pairs); no Zstd/SZIP/pcodec, no compound (h5py's `{r, i}` complex included), reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`, Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`,
`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure `blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure
@@ -386,7 +386,7 @@ filter).
| `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) | | `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) |
| `clawhdf5-io` | I/O helpers: mmap, async, an HSDS client, `mpi-io` (not collective I/O) | | `clawhdf5-io` | I/O helpers: mmap, async, an HSDS client, `mpi-io` (not collective I/O) |
| `clawhdf5-remote` | HTTP(S) and object-store files through a block cache | | `clawhdf5-remote` | HTTP(S) and object-store files through a block cache |
| `clawhdf5-netcdf4` | NetCDF-4 dimensions, variables, CF attributes | | `clawhdf5-netcdf4` | NetCDF-4 dimensions, variables, CF attributes, as netCDF-C reports them (phony dimensions included) |
| `clawhdf5-derive` | Derive macros for HDF5-serialisable structs | | `clawhdf5-derive` | Derive macros for HDF5-serialisable structs |
| `clawhdf5-tools` | `h5rs`: `ls`, `dump`, `stat`, `diff`, `check` | | `clawhdf5-tools` | `h5rs`: `ls`, `dump`, `stat`, `diff`, `check` |
| **Bindings** | | | **Bindings** | |
+35 -29
View File
@@ -145,13 +145,16 @@ pub enum Datatype {
/// imaginary, in rectangular form. `size` is twice the base size and the /// imaginary, in rectangular form. `size` is twice the base size and the
/// base is an IEEE float. /// base is an IEEE float.
/// ///
/// This variant exists for **writing** (see /// [`Datatype::parse`] returns this variant for every class-11 message,
/// also inside compounds, arrays and variable-length types. Readers that
/// want the member view (`{r, i}`, as h5py writes complex numbers by
/// default) use [`Datatype::complex_as_compound`]; `read_compound_fields`
/// accepts this variant directly.
///
/// Writing it is opt-in (see
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and /// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
/// newer can read class 11, so it is opt-in and h5py's compound `{r, i}` /// newer can read class 11, so h5py's compound `{r, i}` stays the
/// stays the default complex encoding. [`Datatype::parse`] still /// default complex encoding.
/// surfaces a class-11 message as that equivalent `{r, i}` compound, so
/// every compound reader handles both encodings; parsing what this
/// variant serializes therefore yields a `Compound`, not a `Complex`.
Complex { size: u32, base_type: Box<Datatype> }, Complex { size: u32, base_type: Box<Datatype> },
} }
@@ -857,9 +860,7 @@ impl Datatype {
// Complex number (HDF5 2.0, datatype version 5). The properties // Complex number (HDF5 2.0, datatype version 5). The properties
// are a single base floating-point datatype message; an element // are a single base floating-point datatype message; an element
// is two consecutive base-type values (real, imaginary). There // is two consecutive base-type values (real, imaginary). There
// is no member list. Surface it as the equivalent two-member // is no member list.
// compound `{r, i}` — the same shape h5py writes for numpy
// complex dtypes — so downstream compound readers work as-is.
if version != 5 { if version != 5 {
return Err(FormatError::InvalidDatatypeVersion { return Err(FormatError::InvalidDatatypeVersion {
class: class_id, class: class_id,
@@ -875,7 +876,13 @@ impl Datatype {
actual: size as usize, actual: size as usize,
}); });
} }
Ok((Self::complex_as_compound(size, &base_type), pos)) Ok((
Datatype::Complex {
size,
base_type: Box::new(base_type),
},
pos,
))
} }
_ => Err(FormatError::InvalidDatatypeClass(class_id)), _ => Err(FormatError::InvalidDatatypeClass(class_id)),
} }
@@ -1174,9 +1181,8 @@ impl Datatype {
/// The `{r, i}` compound equivalent to a native complex type of `size` /// The `{r, i}` compound equivalent to a native complex type of `size`
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the /// bytes over `base_type`: `r` at offset 0, `i` right after it — the
/// shape h5py writes for numpy complex dtypes, and what [`Self::parse`] /// shape h5py writes for numpy complex dtypes. Readers that want a
/// returns for a class-11 message. Readers that meet a /// [`Datatype::Complex`] as members handle it through this view.
/// [`Datatype::Complex`] handle it through this view.
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype { pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
let base_size = base_type.type_size(); let base_size = base_type.type_size();
Datatype::Compound { Datatype::Compound {
@@ -1849,21 +1855,24 @@ mod tests {
fn test_complex_v5_from_hdf5_2_0() { fn test_complex_v5_from_hdf5_2_0() {
let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap(); let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap();
assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len()); assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len());
match dt { match &dt {
Datatype::Compound { size, members } => { Datatype::Complex { size, base_type } => {
assert_eq!(size, 16); assert_eq!(*size, 16);
assert_eq!(members.len(), 2);
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
for m in &members {
assert!(matches!( assert!(matches!(
m.datatype, base_type.as_ref(),
Datatype::FloatingPoint { size: 8, .. } Datatype::FloatingPoint { size: 8, .. }
)); ));
} }
other => panic!("expected Complex, got {other:?}"),
} }
other => panic!("expected Compound, got {other:?}"), // The member view is h5py's `{r, i}` compound.
} let Datatype::Compound { members, .. } =
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
else {
unreachable!()
};
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
} }
#[test] #[test]
@@ -1887,7 +1896,7 @@ mod tests {
assert_eq!(members.len(), 2); assert_eq!(members.len(), 2);
assert!(matches!( assert!(matches!(
&members[0].datatype, &members[0].datatype,
Datatype::Compound { size: 16, members } if members.len() == 2 Datatype::Complex { size: 16, .. }
)); ));
assert_eq!( assert_eq!(
(members[1].name.as_str(), members[1].byte_offset), (members[1].name.as_str(), members[1].byte_offset),
@@ -1918,12 +1927,9 @@ mod tests {
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0); assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
assert_eq!(dt.type_size(), 16); assert_eq!(dt.type_size(), 16);
dt.check_encodable().unwrap(); dt.check_encodable().unwrap();
// Parsing surfaces class 11 as the equivalent `{r, i}` compound. // Parsing returns the native complex type unchanged.
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap(); let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
assert_eq!( assert_eq!(parsed, dt);
parsed,
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
);
let f32c = make_native_complex_f32_type().serialize(); let f32c = make_native_complex_f32_type().serialize();
assert_eq!( assert_eq!(
+43
View File
@@ -731,6 +731,49 @@ fn group_children<S: Storage + ?Sized>(
Ok(entries) Ok(entries)
} }
/// The names of the links of the group at `group_address` in creation
/// order — the order libhdf5 iterates them in with `H5_INDEX_CRT_ORDER`
/// (`H5Literate`) — when the group tracks the creation order of its links;
/// `None` when it does not (a version-1 group never does), in which case
/// libhdf5 can only iterate by name (`H5_INDEX_NAME`: byte order of the
/// names). Hard, soft and external links are listed (user-defined ones,
/// which cannot be followed, are not); links without a creation order
/// value, which a tracking group should not have, come last in the order
/// they are stored.
pub fn links_in_creation_order_in<S: Storage + ?Sized>(
file_data: &S,
superblock: &Superblock,
group_address: u64,
) -> Result<Option<Vec<String>>, FormatError> {
let os = superblock.offset_size;
let ls = superblock.length_size;
let header = ObjectHeader::parse_in(file_data, checked_addr(group_address)?, os, ls)?;
if !is_v2_group(&header) {
return Ok(None);
}
let link_info = find_link_info(&header, os)?;
if link_info.max_creation_order.is_none() {
return Ok(None);
}
let mut links: Vec<(u64, String)> = Vec::new();
let mut visit = |link: LinkMessage| {
links.push((link.creation_order.unwrap_or(u64::MAX), link.name));
};
if let Some(fh_addr) = link_info.fractal_heap_address {
for_each_dense_link(file_data, &link_info, fh_addr, os, ls, false, visit)?;
} else {
for msg in &header.messages {
if msg.msg_type == MessageType::Link
&& let Some(link) = parse_link(&msg.data, os)?
{
visit(link);
}
}
}
links.sort_by_key(|&(order, _)| order);
Ok(Some(links.into_iter().map(|(_, name)| name).collect()))
}
/// Soft links followed while resolving one path. Guards against link cycles /// Soft links followed while resolving one path. Guards against link cycles
/// (`a -> b -> a`), which are legal to create. /// (`a -> b -> a`), which are legal to create.
const MAX_SOFT_LINK_DEPTH: u8 = 16; const MAX_SOFT_LINK_DEPTH: u8 = 16;
+50 -15
View File
@@ -36,23 +36,46 @@ let values: Vec<f64> = temp.read_f64()?;
| Item | What | | Item | What |
|---|---| |---|---|
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | | `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable` (a name or a path into a subgroup), `global_attrs`, `group` (a name or a path), `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `variable_names`, `attrs`, nested `group`) | | `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `variable_names`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `stored_shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` | | `Variable` | `name`, `shape`, `stored_shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) | | `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` | | `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable | | `NcType` | the NetCDF type of a variable: an atomic type (`Byte` ... `UInt64`, `Float`, `Double`, `Char`, `String`) or the class of a user-defined one (`Enum`, `Compound`, `VLen`, `Opaque`); `#[non_exhaustive]` |
Variables and dimensions follow netCDF-C: Groups, dimensions and variables are what netCDF-C 4.9.3 reports
(`libhdf5/hdf5open.c`), names, order and types included. The first call
that needs them reads the metadata of the whole file, as `nc_open` does:
- A variable's dimensions are the ones the file names: the ids in its - Order: a group's links in creation order when it tracks it (netCDF-4
`_Netcdf4Coordinates` attribute, else the dimension scales its files do), else in name order (h5py's default); groups and variables in
`DIMENSION_LIST` references, found in its group or a parent group. Only that order, dimensions by id.
an axis the file names no dimension for (an HDF5 file not written by a - A dimension scale defines a dimension (its `_Netcdf4Dimid`, else the
netCDF library) gets the first dimension of the group of the same size, next free id; unlimited when its first axis is or its length is 0);
else an anonymous `dim_<size>`. scales that are only dimensions are not variables, and a dataset
- Dimension scales that are only dimensions are not variables; a dataset
`_nc4_non_coord_<name>` is the variable `<name>`. `_nc4_non_coord_<name>` is the variable `<name>`.
- A variable's dimensions are the ones the file names: the ids in its
`_Netcdf4Coordinates` attribute (any group), else — when its first axis
has one — the dimension scales its `DIMENSION_LIST` attaches, found in
its group or a parent group.
- Otherwise (an HDF5 file not written by a netCDF library) its axes get
netCDF-C's **phony dimensions** `phony_dim_<n>`: per axis, the first
dimension of the variable's group with the same length and
unlimitedness that an earlier axis of the variable does not use, else a
new one. The numbers run file-wide, a group's subgroups before its own
variables.
- Datasets of types netCDF-C cannot represent are not variables:
references, bit fields, time and array types, and compounds, enums and
VLENs over members or base types that are not netCDF atomic types (4-
and 8-byte floats only) or types netCDF-C has read before in the file.
Enum, compound, VLEN and opaque datasets are variables of those classes
(netCDF4-python itself leaves out opaque ones).
- Two deliberate differences: floats of other than 4 or 8 bytes (half,
bfloat16, 4/6/8-bit floats, `long double`) are `Float`/`Double` and
read as numbers, where netCDF-C 4.9.3 on libhdf5 1.14.6 labels them
`NC_STRING`; and an axis netCDF-C leaves without a dimension (an id or
scale it cannot find, or no scale on an axis after a first one that has
one, where it reads uninitialised memory) gets one by the phony rule.
- A variable along an unlimited dimension has the dimension's length: - A variable along an unlimited dimension has the dimension's length:
`shape` is that length, and the reads return that many values, the `shape` is that length, and the reads return that many values, the
records the variable has not written as its fill value (`_FillValue`, records the variable has not written as its fill value (`_FillValue`,
@@ -60,11 +83,23 @@ Variables and dimensions follow netCDF-C:
`stored_shape` is the HDF5 dataset's extent. `stored_shape` is the HDF5 dataset's extent.
No cargo features. Tests compare against files written by netCDF4-python, No cargo features. Tests compare against files written by netCDF4-python,
h5py dimension scales, h5netcdf and xarray, variable by variable with what h5py (with and without dimension scales, every HDF5 type class), h5netcdf
netCDF4-python reads (`tests/interop_tests.rs`; the CI job requires them and xarray, with what netCDF4-python reads and with what netCDF-C itself
with `CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is reports (`tests/interop_tests.rs`; `tests/netcdf_c_view.py` calls the
not installed). What the HDF5 reader underneath cannot read is listed in libnetcdf netCDF4-python bundles; the CI job requires them with
[`docs/known-issues.md`](../../docs/known-issues.md). `CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is not
installed). `tests/corpus_vs_netcdf_c.rs` compares every file of a corpus
netCDF-C opens, when `CLAWHDF5_NETCDF_CORPUS` names one: over the
conformance corpus, 420 of 429 files match (tank, 2026-09-29), and the
other 9 are explained in `tests/corpus_known_differences.txt`:
```sh
CLAWHDF5_NETCDF_CORPUS=conformance/.cache/corpus CLAWHDF5_PYTHON=$PWD/.venv/bin/python \
cargo test -p clawhdf5-netcdf4 --test corpus_vs_netcdf_c -- --nocapture
```
The differences from netCDF-C, and what the HDF5 reader underneath cannot
read, are listed in [`docs/known-issues.md`](../../docs/known-issues.md).
## License ## License
+7 -204
View File
@@ -1,16 +1,15 @@
//! NetCDF-4 dimension representation. //! NetCDF-4 dimension representation.
//! //!
//! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE //! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE
//! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via //! attribute and a `_Netcdf4Dimid` attribute; variables name theirs in
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited); their length //! `_Netcdf4Coordinates` and `DIMENSION_LIST`. A file without them gets
//! is the largest extent of the variables attached to them. //! netCDF-C's phony dimensions. How they are put together is in
//! `crate::model`.
use std::collections::HashMap; use std::collections::HashMap;
use clawhdf5::AttrValue; use clawhdf5::AttrValue;
use crate::error::Error;
/// A NetCDF-4 dimension. /// A NetCDF-4 dimension.
#[derive(Debug, Clone, PartialEq, Eq)] #[derive(Debug, Clone, PartialEq, Eq)]
pub struct Dimension { pub struct Dimension {
@@ -22,202 +21,10 @@ pub struct Dimension {
pub is_unlimited: bool, pub is_unlimited: bool,
} }
/// A dimension scale of one group: the dataset that defines a dimension.
#[derive(Debug, Clone)]
pub(crate) struct Scale {
/// Object header address of the scale's dataset (what a variable's
/// `DIMENSION_LIST` references).
pub address: u64,
/// Its `_Netcdf4Dimid` (what a variable's `_Netcdf4Coordinates` lists).
pub dimid: Option<i64>,
/// Index of its dimension in [`GroupDims::dims`].
pub dim: usize,
}
/// The dimensions a group defines, with the scales that define them.
#[derive(Debug, Clone, Default)]
pub(crate) struct GroupDims {
/// The group's dimensions, in `_Netcdf4Dimid` order (then discovery order).
pub dims: Vec<Dimension>,
/// The dimension scales behind `dims`; empty when the group has no
/// dimension scale and `dims` were inferred from 1-D datasets.
pub scales: Vec<Scale>,
}
impl GroupDims {
/// The dimension defined by the scale at `address`.
pub fn by_address(&self, address: u64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.address == address)
.map(|s| &self.dims[s.dim])
}
/// The dimension whose scale has `_Netcdf4Dimid` `id`.
pub fn by_dimid(&self, id: i64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.dimid == Some(id))
.map(|s| &self.dims[s.dim])
}
}
/// The dimensions of an HDF5 group (root or subgroup).
///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// dimension's size is the dataset's first (and typically only) shape extent.
/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`. A group with no dimension
/// scale at all (not written by a netCDF library) gets one dimension per
/// 1-D dataset instead.
pub(crate) fn group_dims(
file: &clawhdf5::File,
group: &clawhdf5::Group<'_>,
) -> Result<GroupDims, Error> {
let addresses: HashMap<String, u64> = group.entries()?.into_iter().collect();
let dataset_names = group.datasets()?;
// (dimid, dimension, scale address), in discovery order.
let mut found: Vec<(Option<i64>, Dimension, u64)> = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let attrs = ds.attrs()?;
if !is_dimension_scale(&attrs) {
continue;
}
let Some(&address) = addresses.get(ds_name) else {
continue;
};
let shape = ds.shape()?;
let is_unlimited = is_unlimited(&ds);
let size = if is_unlimited {
unlimited_len(file, &attrs, &shape)
} else {
shape.first().copied().unwrap_or(0)
};
let dim = Dimension {
name: ds_name.clone(),
size,
is_unlimited,
};
found.push((get_dimid(&attrs), dim, address));
}
if found.is_empty() {
// Fallback: infer dimensions from dataset shapes and names.
// In NetCDF-4, coordinate variables are datasets whose name matches
// a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
let mut dims = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
if shape.len() == 1 {
dims.push(Dimension {
name: ds_name.clone(),
size: shape[0],
is_unlimited: is_unlimited(&ds),
});
}
}
return Ok(GroupDims {
dims,
scales: Vec::new(),
});
}
// By dimid; scales without one keep their discovery order after those
// with one (the sort is stable).
found.sort_by_key(|(id, ..)| (id.is_none(), id.unwrap_or(0)));
let mut out = GroupDims::default();
for (i, (dimid, dim, address)) in found.into_iter().enumerate() {
out.dims.push(dim);
out.scales.push(Scale {
address,
dimid,
dim: i,
});
}
Ok(out)
}
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that /// The start of the `NAME` attribute netCDF-C gives a dimension scale that
/// is only a dimension, not also a (coordinate) variable. /// is only a dimension, not also a (coordinate) variable.
const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable"; const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable";
/// The current length of an unlimited dimension, as netCDF-C reports it
/// (`NC4_inq_dim` → `nc4_find_dim_len`): the largest current extent, along
/// the dimension, of the variables that use it, in any group; 0 when none
/// has been written. netCDF-C does not extend a dimension scale that is not
/// also a variable, so such a scale's own extent (0) is not counted; a
/// coordinate variable's is. The variables are the scale's attachments,
/// listed with the axis they use in its `REFERENCE_LIST` attribute (the
/// mirror of each variable's `DIMENSION_LIST`). Attachments that cannot be
/// read are skipped; without a readable `REFERENCE_LIST` the length is the
/// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0);
let is_variable = !is_pure_dimension(attrs);
let Some(refs) = reference_list(file, attrs) else {
return own;
};
refs.into_iter()
.filter_map(|(address, axis)| {
let shape = file.dataset_at(address).ok()?.shape().ok()?;
shape.get(usize::try_from(axis).ok()?).copied()
})
.chain(is_variable.then_some(own))
.max()
.unwrap_or(0)
}
/// The `(dataset address, axis)` pairs of a dimension scale's
/// `REFERENCE_LIST` attribute (HDF5 dimension scales: a compound of an
/// object reference `dataset` and an integer `dimension`), or `None` when it
/// is missing or not in that form.
fn reference_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<(u64, u64)>> {
use clawhdf5_format::data_read::{read_compound_field, read_object_references};
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("REFERENCE_LIST") else {
return None;
};
let dataset = read_compound_field(data, datatype, "dataset").ok()?;
let addresses = read_object_references(
&dataset.raw_data,
&dataset.datatype,
file.superblock().offset_size,
)
.ok()?;
let dimension = read_compound_field(data, datatype, "dimension").ok()?;
let Datatype::FixedPoint {
size, byte_order, ..
} = dimension.datatype
else {
return None;
};
let size = usize::try_from(size).ok().filter(|s| (1..=8).contains(s))?;
let axes = dimension.raw_data.chunks_exact(size).map(|b| {
let mut v = [0u8; 8];
match byte_order {
DatatypeByteOrder::BigEndian => {
v[8 - size..].copy_from_slice(b);
u64::from_be_bytes(v)
}
_ => {
v[..size].copy_from_slice(b);
u64::from_le_bytes(v)
}
}
});
if axes.len() != addresses.len() {
return None;
}
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
}
/// The dimension scale attached to each axis of a variable, from its /// The dimension scale attached to each axis of a variable, from its
/// `DIMENSION_LIST` attribute (HDF5 dimension scales: one variable-length /// `DIMENSION_LIST` attribute (HDF5 dimension scales: one variable-length
/// sequence of object references per axis) — the address of the scale /// sequence of object references per axis) — the address of the scale
@@ -281,13 +88,9 @@ pub(crate) fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
pub(crate) fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> { pub(crate) fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
match attrs.get("_Netcdf4Dimid") { match attrs.get("_Netcdf4Dimid") {
Some(AttrValue::I64(id)) => Some(*id), Some(AttrValue::I64(id)) => Some(*id),
Some(AttrValue::U64(id)) => Some(*id as i64), Some(AttrValue::U64(id)) => i64::try_from(*id).ok(),
Some(AttrValue::I64Array(ids)) if ids.len() == 1 => Some(ids[0]),
Some(AttrValue::U64Array(ids)) if ids.len() == 1 => i64::try_from(ids[0]).ok(),
_ => None, _ => None,
} }
} }
/// Whether a dataset's first axis is unlimited (`max_dimensions[0] ==
/// u64::MAX` in its dataspace).
fn is_unlimited(ds: &clawhdf5::Dataset<'_>) -> bool {
matches!(ds.max_dimensions(), Ok(Some(max_dims)) if max_dims.first() == Some(&u64::MAX))
}
+35 -31
View File
@@ -7,36 +7,36 @@ use std::collections::HashMap;
use clawhdf5::AttrValue; use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension}; use crate::dimension::Dimension;
use crate::error::Error; use crate::error::Error;
use crate::scope::{self, Scope}; use crate::model::Model;
use crate::variable::Variable; use crate::variable::Variable;
/// A NetCDF-4 group corresponding to an HDF5 group. /// A NetCDF-4 group corresponding to an HDF5 group.
pub struct NetCDF4Group<'f> { pub struct NetCDF4Group<'f> {
/// Group name. /// Group name.
name: String, name: String,
/// Path of the group from the root (`/`-separated).
path: String,
/// Underlying HDF5 file. /// Underlying HDF5 file.
file: &'f clawhdf5::File, file: &'f clawhdf5::File,
/// Underlying HDF5 group. /// The file's groups, dimensions and variables.
hdf5_group: clawhdf5::Group<'f>, model: &'f Model,
/// This group's index in `model`.
index: usize,
} }
impl<'f> NetCDF4Group<'f> { impl<'f> NetCDF4Group<'f> {
/// Create a new NetCDF4Group from an HDF5 group. /// The group `index` of `model`.
pub(crate) fn new( pub(crate) fn new(
name: String, name: String,
path: String,
file: &'f clawhdf5::File, file: &'f clawhdf5::File,
hdf5_group: clawhdf5::Group<'f>, model: &'f Model,
index: usize,
) -> Self { ) -> Self {
Self { Self {
name, name,
path,
file, file,
hdf5_group, model,
index,
} }
} }
@@ -45,52 +45,56 @@ impl<'f> NetCDF4Group<'f> {
&self.name &self.name
} }
/// List dimensions defined in this group (not those of its parent /// List dimensions defined in this group, in dimension-id order (not
/// groups, which its variables can also use). /// those of its parent groups, which its variables can also use).
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> { pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
Ok(dimension::group_dims(self.file, &self.hdf5_group)?.dims) Ok(self.model.dimensions(self.index))
} }
/// List variables in this group: its datasets, except the dimension /// List variables in this group, in netCDF-C's order: its datasets,
/// scales that are only dimensions. Their dimensions can be defined in /// except the dimension scales that are only dimensions and datasets
/// this group or a parent group. /// of types netCDF-C cannot represent. Their dimensions can be defined
/// in this group or a parent group.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> { pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
Scope::new(self.file, &self.path)?.variables() self.model.variables(self.file, self.index)
} }
/// Get a specific variable by name. /// Get a specific variable by name (or by path, `"sub/var"`).
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> { pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
scope::variable_at(self.file, &self.path, name) self.model.variable(self.file, self.index, name)
} }
/// Read all attributes of this group. /// Read all attributes of this group.
pub fn attrs(&self) -> Result<HashMap<String, AttrValue>, Error> { pub fn attrs(&self) -> Result<HashMap<String, AttrValue>, Error> {
Ok(self.hdf5_group.attrs()?) Ok(self
.file
.group_at(self.model.group_address(self.index))
.attrs()?)
} }
/// List subgroup names. /// List subgroup names, in netCDF-C's order.
pub fn group_names(&self) -> Result<Vec<String>, Error> { pub fn group_names(&self) -> Result<Vec<String>, Error> {
Ok(self.hdf5_group.groups()?) Ok(self.model.group_names(self.index))
} }
/// Get a subgroup by name. /// Get a subgroup by name (or by path, `"a/b"`).
pub fn group(&self, name: &str) -> Result<NetCDF4Group<'f>, Error> { pub fn group(&self, name: &str) -> Result<NetCDF4Group<'f>, Error> {
let hdf5_group = self let index = self
.hdf5_group .model
.group(name) .find_group(self.index, name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?; .ok_or_else(|| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new( Ok(NetCDF4Group::new(
name.to_string(), name.to_string(),
format!("{}/{name}", self.path),
self.file, self.file,
hdf5_group, self.model,
index,
)) ))
} }
/// The names of this group's variables (see /// The names of this group's variables (see
/// [`variables`](Self::variables)). /// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> { pub fn variable_names(&self) -> Result<Vec<String>, Error> {
Scope::new(self.file, &self.path)?.variable_names() Ok(self.model.variable_names(self.index))
} }
} }
+48 -23
View File
@@ -27,7 +27,7 @@ pub mod cf;
pub mod dimension; pub mod dimension;
pub mod error; pub mod error;
pub mod group; pub mod group;
mod scope; mod model;
pub mod types; pub mod types;
pub mod variable; pub mod variable;
@@ -40,26 +40,49 @@ pub use types::NcType;
pub use variable::Variable; pub use variable::Variable;
use std::collections::HashMap; use std::collections::HashMap;
use std::sync::OnceLock;
use model::Model;
/// A NetCDF-4 file reader. /// A NetCDF-4 file reader.
/// ///
/// Wraps a clawhdf5 File and provides NetCDF-4 semantics: dimensions, /// Wraps a clawhdf5 File and provides NetCDF-4 semantics: dimensions,
/// variables with CF attributes, groups, and type mapping. /// variables with CF attributes, groups, and type mapping.
///
/// The groups, dimensions and variables are what netCDF-C reports for the
/// file (see the crate README): the first call that needs them reads the
/// metadata of the whole file, as `nc_open` does, and keeps it; the values
/// are read when asked for.
pub struct NetCDF4File { pub struct NetCDF4File {
hdf5: clawhdf5::File, hdf5: clawhdf5::File,
model: OnceLock<Model>,
} }
impl NetCDF4File { impl NetCDF4File {
/// Open a NetCDF-4 file from a filesystem path. /// Open a NetCDF-4 file from a filesystem path.
pub fn open<P: AsRef<std::path::Path>>(path: P) -> Result<Self, Error> { pub fn open<P: AsRef<std::path::Path>>(path: P) -> Result<Self, Error> {
let hdf5 = clawhdf5::File::open(path)?; Ok(Self::from_hdf5(clawhdf5::File::open(path)?))
Ok(Self { hdf5 })
} }
/// Open a NetCDF-4 file from in-memory bytes. /// Open a NetCDF-4 file from in-memory bytes.
pub fn from_bytes(data: Vec<u8>) -> Result<Self, Error> { pub fn from_bytes(data: Vec<u8>) -> Result<Self, Error> {
let hdf5 = clawhdf5::File::from_bytes(data)?; Ok(Self::from_hdf5(clawhdf5::File::from_bytes(data)?))
Ok(Self { hdf5 }) }
fn from_hdf5(hdf5: clawhdf5::File) -> Self {
Self {
hdf5,
model: OnceLock::new(),
}
}
/// The file's groups, dimensions and variables, read on first use.
fn model(&self) -> Result<&Model, Error> {
if let Some(model) = self.model.get() {
return Ok(model);
}
let model = Model::build(&self.hdf5)?;
Ok(self.model.get_or_init(|| model))
} }
/// Get the _NCProperties root attribute, if present. /// Get the _NCProperties root attribute, if present.
@@ -74,27 +97,29 @@ impl NetCDF4File {
} }
} }
/// List dimensions defined in the root group. /// List dimensions defined in the root group, in dimension-id order.
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> { pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
Ok(dimension::group_dims(&self.hdf5, &self.hdf5.root())?.dims) Ok(self.model()?.dimensions(0))
} }
/// List all variables in the root group: its datasets, except the /// List the variables of the root group, in netCDF-C's order: its
/// dimension scales that are only dimensions (netCDF-C does not list /// datasets, except the dimension scales that are only dimensions and
/// them either). /// datasets of types netCDF-C cannot represent (see
/// [`NcType`]).
pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> { pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> {
scope::Scope::new(&self.hdf5, "/")?.variables() self.model()?.variables(&self.hdf5, 0)
} }
/// The names of the root group's variables (see /// The names of the root group's variables (see
/// [`variables`](Self::variables)). /// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> { pub fn variable_names(&self) -> Result<Vec<String>, Error> {
scope::Scope::new(&self.hdf5, "/")?.variable_names() Ok(self.model()?.variable_names(0))
} }
/// Get a specific variable by name from the root group. /// Get a specific variable by name from the root group; the name may be
/// a path into a subgroup (`"sub/var"`).
pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> { pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> {
scope::variable_at(&self.hdf5, "", name) self.model()?.variable(&self.hdf5, 0, name)
} }
/// Read all global (root group) attributes. /// Read all global (root group) attributes.
@@ -102,22 +127,22 @@ impl NetCDF4File {
Ok(self.hdf5.root().attrs()?) Ok(self.hdf5.root().attrs()?)
} }
/// List subgroup names in the root group. /// List subgroup names in the root group, in netCDF-C's order.
pub fn group_names(&self) -> Result<Vec<String>, Error> { pub fn group_names(&self) -> Result<Vec<String>, Error> {
Ok(self.hdf5.root().groups()?) Ok(self.model()?.group_names(0))
} }
/// Get a subgroup by name. /// Get a subgroup by name (or by path, `"a/b"`).
pub fn group(&self, name: &str) -> Result<NetCDF4Group<'_>, Error> { pub fn group(&self, name: &str) -> Result<NetCDF4Group<'_>, Error> {
let hdf5_group = self let model = self.model()?;
.hdf5 let index = model
.group(name) .find_group(0, name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?; .ok_or_else(|| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new( Ok(NetCDF4Group::new(
name.to_string(),
name.to_string(), name.to_string(),
&self.hdf5, &self.hdf5,
hdf5_group, model,
index,
)) ))
} }
+853
View File
@@ -0,0 +1,853 @@
//! A file's groups, dimensions and variables, as netCDF-C reads them.
//!
//! netCDF-C (`libhdf5/hdf5open.c`, 4.9.3) reads a file in two passes, and
//! the names, order and sharing of the dimensions depend on both, so this
//! module replays them for the whole file at once:
//!
//! 1. `rec_read_metadata`: each group's links in creation order when the
//! group tracks it, else in name order; a group's datasets and named
//! datatypes before its subgroups, which follow in the same order. A
//! dimension scale (`CLASS` `DIMENSION_SCALE`) defines a dimension with
//! the id in its `_Netcdf4Dimid`, else the next free id (ids are
//! file-wide), its first extent as length, unlimited when that axis is
//! (or the length is 0); it is a variable too unless its `NAME` says it
//! is only a dimension. Every other dataset is a variable, unless its
//! type is one netCDF-C cannot represent ([`VarTypes::nc_type`]); a
//! dataset `_nc4_non_coord_<name>` is the variable `<name>`.
//! 2. `rec_match_dimscales`, subgroups first, then the group's variables in
//! order: a variable gets the dimensions whose ids its
//! `_Netcdf4Coordinates` lists (looked up file-wide), else — when its
//! `DIMENSION_LIST` attaches a scale to its first axis — the scales that
//! list attaches, looked up in its group and then each parent, else
//! "phony" dimensions (`create_phony_dims`): per axis, the first
//! dimension of the variable's group of that length and unlimitedness
//! not already used by an earlier axis of the variable, else a new one
//! called `phony_dim_<id>`.
//!
//! An unlimited dimension's length is the largest extent, along it, of the
//! variables that use it in its group and the groups below
//! (`nc4_find_dim_len`).
//!
//! Where netCDF-C 4.9.3 leaves an axis without a dimension (an id or a scale
//! it cannot find, or an axis without a scale of a variable whose first
//! axis has one — there it reads uninitialised memory), netCDF4-python
//! cannot open the file; this crate gives such an axis a dimension by the
//! phony rule instead.
use std::collections::HashMap;
use clawhdf5::AttrValue;
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
use clawhdf5_format::object_header::{ObjectClass, ObjectHeader};
use crate::dimension::{self, Dimension};
use crate::error::Error;
use crate::types::NcType;
use crate::variable::Variable;
/// The prefix netCDF-C gives the dataset of a variable that has a
/// dimension's name but is not that dimension's coordinate variable (the
/// dimension's scale holds the name).
const NON_COORD_PREFIX: &str = "_nc4_non_coord_";
/// Groups read at most, a guard against files whose groups are hard-linked
/// into each other many times over (netCDF-C reads each link as its own
/// group).
const MAX_GROUPS: usize = 100_000;
/// A file as netCDF-C sees it.
#[derive(Debug)]
pub(crate) struct Model {
/// Every group; the root is the first.
groups: Vec<Group>,
/// Every dimension, in the order they were created.
dims: Vec<Dim>,
}
#[derive(Debug)]
struct Group {
/// Object header address of the HDF5 group.
address: u64,
/// Its subgroups, `(name, index in Model::groups)`, in netCDF-C's order.
children: Vec<(String, usize)>,
parent: Option<usize>,
/// The dimensions it defines (indexes in `Model::dims`), in creation
/// order.
dims: Vec<usize>,
/// Its variables, in netCDF-C's order.
vars: Vec<Var>,
}
#[derive(Debug)]
struct Dim {
id: i64,
name: String,
/// The length it was created with (for an unlimited dimension, replaced
/// by its current length once every variable has its dimensions).
len: u64,
unlimited: bool,
/// Object header address of its dimension scale; `None` for a phony
/// dimension.
scale: Option<u64>,
}
#[derive(Debug)]
struct Var {
/// The netCDF name.
name: String,
/// Whether its dataset is `_nc4_non_coord_<name>`.
non_coord: bool,
address: u64,
nc_type: NcType,
extent: Vec<u64>,
/// Whether each axis is unlimited in the dataspace.
unlimited: Vec<bool>,
/// How netCDF-C finds its dimensions.
source: DimSource,
/// Its dimensions (indexes in `Model::dims`), one per axis, once found.
dims: Vec<usize>,
}
#[derive(Debug)]
enum DimSource {
/// A one-dimensional coordinate variable: its own scale's dimension.
OwnScale(usize),
/// The ids in `_Netcdf4Coordinates`; for a multi-dimensional coordinate
/// variable, also its own dimension (for the first axis, should an id
/// not be found).
Coordinates(Vec<i64>, Option<usize>),
/// The scale `DIMENSION_LIST` attaches to each axis (the first has one).
Scales(Vec<Option<u64>>, Option<usize>),
/// None: phony dimensions.
Phony(Option<usize>),
}
impl Model {
/// Read the metadata of `file` as netCDF-C does.
pub fn build(file: &clawhdf5::File) -> Result<Self, Error> {
let mut builder = Builder {
file,
model: Model {
groups: Vec::new(),
dims: Vec::new(),
},
next_id: 0,
types: VarTypes::default(),
};
let root = file.superblock().root_group_address;
builder.model.groups.push(Group {
address: root,
children: Vec::new(),
parent: None,
dims: Vec::new(),
vars: Vec::new(),
});
builder.read_group(0, &mut vec![root])?;
builder.match_dims(0);
builder.unlimited_lengths();
Ok(builder.model)
}
/// The group at `path` (`/`-separated names), from the group `from`.
pub fn find_group(&self, from: usize, path: &str) -> Option<usize> {
path.split('/')
.filter(|p| !p.is_empty())
.try_fold(from, |g, name| {
self.groups[g]
.children
.iter()
.find(|(n, _)| n == name)
.map(|&(_, i)| i)
})
}
/// The object header address of group `g`.
pub fn group_address(&self, g: usize) -> u64 {
self.groups[g].address
}
/// The names of group `g`'s subgroups, in netCDF-C's order.
pub fn group_names(&self, g: usize) -> Vec<String> {
self.groups[g]
.children
.iter()
.map(|(n, _)| n.clone())
.collect()
}
/// The dimensions group `g` defines, in id order (as `nc_inq_dimids`).
pub fn dimensions(&self, g: usize) -> Vec<Dimension> {
let mut dims: Vec<&Dim> = self.groups[g].dims.iter().map(|&d| &self.dims[d]).collect();
dims.sort_by_key(|d| d.id);
dims.into_iter().map(Dim::dimension).collect()
}
/// The names of group `g`'s variables, in netCDF-C's order.
pub fn variable_names(&self, g: usize) -> Vec<String> {
self.groups[g].vars.iter().map(|v| v.name.clone()).collect()
}
/// Group `g`'s variables, in netCDF-C's order.
pub fn variables<'f>(
&self,
file: &'f clawhdf5::File,
g: usize,
) -> Result<Vec<Variable<'f>>, Error> {
self.groups[g]
.vars
.iter()
.map(|v| self.open(file, v))
.collect()
}
/// The variable `name` of group `g`; `name` may be a path (`"sub/var"`)
/// relative to it. Of two variables of one name (datasets `<name>` and
/// `_nc4_non_coord_<name>`), the second.
pub fn variable<'f>(
&self,
file: &'f clawhdf5::File,
g: usize,
name: &str,
) -> Result<Variable<'f>, Error> {
let not_found = || Error::VariableNotFound(name.to_string());
let trimmed = name.trim_start_matches('/');
let (g, leaf) = match trimmed.rsplit_once('/') {
Some((dir, leaf)) => (self.find_group(g, dir).ok_or_else(not_found)?, leaf),
None => (g, trimmed),
};
let vars = &self.groups[g].vars;
let var = vars
.iter()
.find(|v| v.name == leaf && v.non_coord)
.or_else(|| vars.iter().find(|v| v.name == leaf))
.ok_or_else(not_found)?;
self.open(file, var)
}
fn open<'f>(&self, file: &'f clawhdf5::File, var: &Var) -> Result<Variable<'f>, Error> {
let ds = file.dataset_at(var.address)?;
let attrs = ds.attrs().unwrap_or_default();
let dims = var.dims.iter().map(|&d| self.dims[d].dimension()).collect();
Ok(Variable::new(
var.name.clone(),
ds,
dims,
attrs,
var.nc_type,
))
}
}
impl Dim {
fn dimension(&self) -> Dimension {
Dimension {
name: self.name.clone(),
size: self.len,
is_unlimited: self.unlimited,
}
}
}
struct Builder<'f> {
file: &'f clawhdf5::File,
model: Model,
/// netCDF-C's `next_dimid`.
next_id: i64,
types: VarTypes,
}
impl Builder<'_> {
/// Pass 1 for group `g` and, after its own links, its subgroups.
/// `ancestors` holds the addresses of the groups from the root to `g`,
/// so that a group linked into itself is not read forever.
fn read_group(&mut self, g: usize, ancestors: &mut Vec<u64>) -> Result<(), Error> {
let address = self.model.groups[g].address;
let sb = self.file.superblock();
let mut subgroups = Vec::new();
for (name, child) in ordered_entries(self.file, address)? {
let Ok(header) =
ObjectHeader::parse_in(self.file.storage(), child, sb.offset_size, sb.length_size)
else {
continue;
};
match header.object_class() {
Some(ObjectClass::Dataset) => self.read_dataset(g, name, child),
Some(ObjectClass::NamedDatatype) => {
if let Some(dt) = header_datatype(&header) {
self.types.named(&dt);
}
}
_ if is_group(&header) => subgroups.push((name, child)),
_ => {}
}
}
for (name, child) in subgroups {
if ancestors.contains(&child) || self.model.groups.len() >= MAX_GROUPS {
continue;
}
let index = self.model.groups.len();
self.model.groups.push(Group {
address: child,
children: Vec::new(),
parent: Some(g),
dims: Vec::new(),
vars: Vec::new(),
});
self.model.groups[g].children.push((name, index));
ancestors.push(child);
let read = self.read_group(index, ancestors);
ancestors.pop();
read?;
}
Ok(())
}
/// `read_dataset`: a dimension for a dimension scale, a variable for
/// the rest. A dataset that cannot be opened is skipped.
fn read_dataset(&mut self, g: usize, name: String, address: u64) {
let Ok(ds) = self.file.dataset_at(address) else {
return;
};
let attrs = ds.attrs().unwrap_or_default();
let Ok(extent) = ds.shape() else {
return;
};
let max = ds.max_dimensions().ok().flatten();
let unlimited: Vec<bool> = (0..extent.len())
.map(|i| matches!(&max, Some(m) if m.get(i) == Some(&u64::MAX)))
.collect();
let mut own = None;
if dimension::is_dimension_scale(&attrs) && !extent.is_empty() {
// read_scale
let id = match dimension::get_dimid(&attrs) {
Some(id) => {
if id >= self.next_id {
self.next_id = id.saturating_add(1);
}
id
}
None => self.take_id(),
};
let len = extent[0];
own = Some(self.add_dim(
g,
Dim {
id,
name: name.clone(),
len,
unlimited: unlimited[0] || len == 0,
scale: Some(address),
},
));
if dimension::is_pure_dimension(&attrs) {
return;
}
}
// read_var: a type netCDF-C cannot represent drops the variable
// (the dimension of a scale stays).
let Some(nc_type) = ds
.raw_datatype()
.ok()
.and_then(|dt| self.types.nc_type(&dt))
else {
return;
};
let rank = extent.len();
let coordinates = coordinates(&attrs).filter(|ids| ids.len() == rank && rank > 0);
let source = match (own, coordinates) {
(Some(dim), _) if rank == 1 => DimSource::OwnScale(dim),
(_, Some(ids)) => DimSource::Coordinates(ids, own),
_ => match dimension::dimension_list(self.file, &attrs) {
Some(scales)
if scales.len() == rank && scales.first().is_some_and(Option::is_some) =>
{
DimSource::Scales(scales, own)
}
_ => DimSource::Phony(own),
},
};
let (name, non_coord) = match name.strip_prefix(NON_COORD_PREFIX) {
Some(rest) if !rest.is_empty() => (rest.to_string(), true),
_ => (name, false),
};
self.model.groups[g].vars.push(Var {
name,
non_coord,
address,
nc_type,
extent,
unlimited,
source,
dims: Vec::new(),
});
}
fn take_id(&mut self) -> i64 {
let id = self.next_id;
self.next_id = self.next_id.saturating_add(1);
id
}
fn add_dim(&mut self, g: usize, dim: Dim) -> usize {
let index = self.model.dims.len();
self.model.dims.push(dim);
self.model.groups[g].dims.push(index);
index
}
/// Pass 2 (`rec_match_dimscales`): subgroups first, then the group's
/// variables in order.
fn match_dims(&mut self, g: usize) {
let children: Vec<usize> = self.model.groups[g]
.children
.iter()
.map(|&(_, c)| c)
.collect();
for child in children {
self.match_dims(child);
}
for v in 0..self.model.groups[g].vars.len() {
let var = &self.model.groups[g].vars[v];
let rank = var.extent.len();
let mut found: Vec<Option<usize>> = vec![None; rank];
match &var.source {
DimSource::OwnScale(dim) => found[0] = Some(*dim),
DimSource::Coordinates(ids, own) => {
for (slot, id) in found.iter_mut().zip(ids) {
*slot = self.dim_by_id(*id);
}
if found[0].is_none() {
found[0] = *own;
}
}
DimSource::Scales(scales, own) => {
for (slot, scale) in found.iter_mut().zip(scales) {
*slot = scale.and_then(|s| self.dim_by_scale(g, s));
}
if found[0].is_none() {
found[0] = *own;
}
}
DimSource::Phony(own) => {
if rank > 0 {
found[0] = *own;
}
}
}
let mut dims: Vec<usize> = Vec::with_capacity(rank);
for (axis, found) in found.into_iter().enumerate() {
let dim = match found {
Some(dim) => dim,
None => {
let var = &self.model.groups[g].vars[v];
let (len, unlimited) = (var.extent[axis], var.unlimited[axis]);
self.phony_dim(g, len, unlimited, &dims)
}
};
dims.push(dim);
}
self.model.groups[g].vars[v].dims = dims;
}
}
/// The dimension with id `id`: the last one created with it
/// (`nc4_find_dim` looks ids up in a file-wide table, where a later
/// dimension of the same id replaces an earlier one).
fn dim_by_id(&self, id: i64) -> Option<usize> {
self.model.dims.iter().rposition(|d| d.id == id)
}
/// The dimension of the scale at `address`, in group `g` or the nearest
/// parent that has it.
fn dim_by_scale(&self, g: usize, address: u64) -> Option<usize> {
let mut group = Some(g);
while let Some(i) = group {
let found = self.model.groups[i]
.dims
.iter()
.copied()
.find(|&d| self.model.dims[d].scale == Some(address));
if found.is_some() {
return found;
}
group = self.model.groups[i].parent;
}
None
}
/// `create_phony_dims` for one axis: the first dimension of group `g`
/// of length `len` and unlimitedness `unlimited` that no earlier axis
/// of the variable uses (`taken`), else a new `phony_dim_<id>`.
fn phony_dim(&mut self, g: usize, len: u64, unlimited: bool, taken: &[usize]) -> usize {
let dims = &self.model.dims;
let existing = self.model.groups[g].dims.iter().copied().find(|&d| {
let dim = &dims[d];
dim.len == len
&& dim.unlimited == unlimited
&& !taken.iter().any(|&t| dims[t].id == dim.id)
});
if let Some(dim) = existing {
return dim;
}
let id = self.take_id();
self.add_dim(
g,
Dim {
id,
name: format!("phony_dim_{id}"),
len,
// `nc4_dim_list_add`: a length of 0 is NC_UNLIMITED.
unlimited: unlimited || len == 0,
scale: None,
},
)
}
/// Each unlimited dimension's length: the largest extent along it of
/// the variables using it in its group and the groups below.
fn unlimited_lengths(&mut self) {
let mut owner = vec![0usize; self.model.dims.len()];
for (g, group) in self.model.groups.iter().enumerate() {
for &d in &group.dims {
owner[d] = g;
}
}
let mut lens = vec![0u64; self.model.dims.len()];
for (g, group) in self.model.groups.iter().enumerate() {
for var in &group.vars {
for (&d, &e) in var.dims.iter().zip(&var.extent) {
if self.model.dims[d].unlimited && self.is_within(g, owner[d]) {
lens[d] = lens[d].max(e);
}
}
}
}
for (dim, len) in self.model.dims.iter_mut().zip(lens) {
if dim.unlimited {
dim.len = len;
}
}
}
/// Whether group `g` is `ancestor` or below it.
fn is_within(&self, g: usize, ancestor: usize) -> bool {
let mut group = Some(g);
while let Some(i) = group {
if i == ancestor {
return true;
}
group = self.model.groups[i].parent;
}
false
}
}
/// A group's entries in netCDF-C's order: creation order when the group
/// tracks it, else the byte order of the names.
fn ordered_entries(file: &clawhdf5::File, address: u64) -> Result<Vec<(String, u64)>, Error> {
let mut entries = file.group_at(address).entries()?;
let order = clawhdf5_format::group_v2::links_in_creation_order_in(
file.storage(),
file.superblock(),
address,
)
.ok()
.flatten();
match order {
Some(names) => {
let position: HashMap<&str, usize> = names
.iter()
.enumerate()
.map(|(i, n)| (n.as_str(), i))
.collect();
entries.sort_by_key(|(n, _)| position.get(n.as_str()).copied().unwrap_or(usize::MAX));
}
None => entries.sort_by(|a, b| a.0.as_bytes().cmp(b.0.as_bytes())),
}
Ok(entries)
}
fn is_group(header: &ObjectHeader) -> bool {
use clawhdf5_format::message_type::MessageType;
header.messages.iter().any(|m| {
matches!(
m.msg_type,
MessageType::LinkInfo | MessageType::Link | MessageType::SymbolTable
)
})
}
/// The datatype a named datatype's header holds.
fn header_datatype(header: &ObjectHeader) -> Option<Datatype> {
use clawhdf5_format::message_type::MessageType;
let msg = header
.messages
.iter()
.find(|m| m.msg_type == MessageType::Datatype)?;
Datatype::parse_in_header(&msg.data, header.version)
.ok()
.map(|(dt, _)| dt)
}
/// A variable's `_Netcdf4Coordinates`: the id of the dimension of each
/// axis.
fn coordinates(attrs: &HashMap<String, AttrValue>) -> Option<Vec<i64>> {
match attrs.get("_Netcdf4Coordinates")? {
AttrValue::I64Array(ids) => Some(ids.clone()),
AttrValue::I64(id) => Some(vec![*id]),
AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(),
AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]),
_ => None,
}
}
/// The user-defined types netCDF-C has read so far, and the netCDF type of
/// a dataset.
///
/// netCDF-C keeps every type `read_type` meets in a file-wide list and
/// finds one again by `H5Tequal` on the native types — also a type it
/// failed to read: `read_type` adds the type (with its class, for a
/// compound, enum, variable-length or opaque type) before it looks at the
/// members or base type, and does not remove it when one of those fails.
/// So a dataset of a type netCDF-C skipped once becomes a variable the
/// second time (a compound with a reference member, say), and a compound
/// member of a type it skipped (a bit field) is accepted. This replays
/// that, except that a dataset whose skipped type has no class (a bit
/// field, time, array or complex type) stays hidden: netCDF-C lists the
/// second such dataset with an invalid type (class 0).
#[derive(Debug, Default)]
pub(crate) struct VarTypes {
/// Every type read, with its class when it has one netCDF knows.
known: Vec<(Datatype, Option<NcType>)>,
}
impl VarTypes {
/// The netCDF type netCDF-C gives a dataset of type `dt`
/// (`get_type_info2`), or `None` when it skips the dataset
/// (`NC_EBADTYPID`): a reference, bit field, time or array type, a
/// compound with a member, or an enum or variable-length type with a
/// base type, that is not a netCDF atomic type or a type read before
/// (see the type docs).
///
/// Floats other than 4 and 8 bytes (half floats, bfloat16, the 4-, 6-
/// and 8-bit floats, `long double`) are `Float` (up to 4 bytes) and
/// `Double` here; netCDF-C 4.9.3 on libhdf5 1.14.6 labels them
/// `NC_STRING` (their native type matches none of its own).
pub fn nc_type(&mut self, dt: &Datatype) -> Option<NcType> {
match dt {
Datatype::FixedPoint { size, signed, .. } => int_type(*size, *signed),
Datatype::FloatingPoint { size, .. } => Some(if *size <= 4 {
NcType::Float
} else {
NcType::Double
}),
Datatype::String { size, .. } => Some(if *size > 1 {
NcType::String
} else {
NcType::Char
}),
Datatype::VariableLength {
is_string: true, ..
} => Some(NcType::String),
_ => match self.find(dt) {
Some(class) => class,
None => self.read_type(dt),
},
}
}
/// A named datatype of the file (`read_type` on a committed type).
pub fn named(&mut self, dt: &Datatype) {
if self.find(dt).is_none() {
self.read_type(dt);
}
}
/// The type read before that is `dt`, by its class.
fn find(&self, dt: &Datatype) -> Option<Option<NcType>> {
self.known
.iter()
.find(|(k, _)| native_eq(k, dt))
.map(|&(_, class)| class)
}
/// `read_type`: remember `dt` (not a reference type), then its class if
/// netCDF-C can represent its members or base type.
fn read_type(&mut self, dt: &Datatype) -> Option<NcType> {
let class = match dt {
Datatype::Reference { .. } => return None,
Datatype::Compound { .. } => Some(NcType::Compound),
Datatype::VariableLength { .. } => Some(NcType::VLen),
Datatype::Opaque { .. } => Some(NcType::Opaque),
Datatype::Enumeration { .. } => Some(NcType::Enum),
_ => None,
};
self.known.push((dt.clone(), class));
let parts_ok = match dt {
Datatype::Compound { members, .. } => members.iter().all(|m| match &m.datatype {
Datatype::Array { base_type, .. } => self.is_member_type(base_type),
other => self.is_member_type(other),
}),
Datatype::VariableLength { base_type, .. }
| Datatype::Enumeration { base_type, .. } => self.is_member_type(base_type),
_ => true,
};
class.filter(|_| parts_ok)
}
/// `get_netcdf_type`: whether netCDF-C takes `dt` as the type of a
/// compound member or the base of an enum or variable-length type — an
/// atomic type (4- and 8-byte floats only) or a type read before.
fn is_member_type(&self, dt: &Datatype) -> bool {
match dt {
Datatype::FixedPoint { size, signed, .. } => int_type(*size, *signed).is_some(),
Datatype::FloatingPoint { size: 4 | 8, .. }
| Datatype::String { .. }
| Datatype::VariableLength {
is_string: true, ..
} => true,
_ => self.find(dt).is_some(),
}
}
}
/// The netCDF integer type of an HDF5 integer: the native integer libhdf5
/// converts it to (`H5Tget_native_type`: the smallest at least as wide).
fn int_type(size: u32, signed: bool) -> Option<NcType> {
Some(match (size, signed) {
(1, true) => NcType::Byte,
(1, false) => NcType::UByte,
(2, true) => NcType::Short,
(2, false) => NcType::UShort,
(3..=4, true) => NcType::Int,
(3..=4, false) => NcType::UInt,
(5..=8, true) => NcType::Int64,
(5..=8, false) => NcType::UInt64,
_ => return None,
})
}
/// Whether two datatypes have the same native type (`H5Tequal` after
/// `H5Tget_native_type`): byte order, padding and compound member offsets
/// do not count.
fn native_eq(a: &Datatype, b: &Datatype) -> bool {
use Datatype as D;
match (a, b) {
(
D::FixedPoint {
size: s1,
signed: g1,
..
},
D::FixedPoint {
size: s2,
signed: g2,
..
},
) => g1 == g2 && int_type(*s1, *g1) == int_type(*s2, *g2),
(D::FloatingPoint { size: s1, .. }, D::FloatingPoint { size: s2, .. }) => s1 == s2,
(
D::String {
size: s1,
padding: p1,
charset: c1,
},
D::String {
size: s2,
padding: p2,
charset: c2,
},
) => s1 == s2 && p1 == p2 && c1 == c2,
(
D::VariableLength {
is_string: i1,
base_type: b1,
charset: c1,
..
},
D::VariableLength {
is_string: i2,
base_type: b2,
charset: c2,
..
},
) => i1 == i2 && if *i1 { c1 == c2 } else { native_eq(b1, b2) },
(D::Compound { members: m1, .. }, D::Compound { members: m2, .. }) => {
m1.len() == m2.len()
&& m1
.iter()
.zip(m2)
.all(|(x, y)| x.name == y.name && native_eq(&x.datatype, &y.datatype))
}
(
D::Enumeration {
base_type: b1,
members: m1,
..
},
D::Enumeration {
base_type: b2,
members: m2,
..
},
) => {
native_eq(b1, b2)
&& m1.len() == m2.len()
&& m1.iter().zip(m2).all(|(x, y)| {
x.name == y.name && enum_value(&x.value, b1) == enum_value(&y.value, b2)
})
}
(D::Opaque { size: s1, tag: t1 }, D::Opaque { size: s2, tag: t2 }) => s1 == s2 && t1 == t2,
(
D::Array {
base_type: b1,
dimensions: d1,
},
D::Array {
base_type: b2,
dimensions: d2,
},
) => d1 == d2 && native_eq(b1, b2),
(
D::Reference {
size: s1,
ref_type: r1,
},
D::Reference {
size: s2,
ref_type: r2,
},
) => s1 == s2 && r1 == r2,
(D::BitField { size: s1, .. }, D::BitField { size: s2, .. })
| (D::Time { size: s1, .. }, D::Time { size: s2, .. }) => s1 == s2,
(
D::Complex {
size: s1,
base_type: b1,
},
D::Complex {
size: s2,
base_type: b2,
},
) => s1 == s2 && native_eq(b1, b2),
_ => false,
}
}
/// An enum member's value, in the order of its base type's bytes.
fn enum_value(value: &[u8], base: &Datatype) -> Vec<u8> {
let big = matches!(
base,
Datatype::FixedPoint {
byte_order: DatatypeByteOrder::BigEndian,
..
}
);
let mut v = value.to_vec();
if big {
v.reverse();
}
v
}
-231
View File
@@ -1,231 +0,0 @@
//! A group's variables and the dimensions they are defined on.
//!
//! netCDF-C (`libhdf5/hdf5open.c`) gives a variable its dimensions from the
//! file, never by size: the dimension ids in its `_Netcdf4Coordinates`
//! attribute (each dimension scale's `_Netcdf4Dimid`), else the dimension
//! scales its `DIMENSION_LIST` attribute references, looked up in the
//! variable's group and then each parent group up to the root. Only an axis
//! with neither (a file not written by a netCDF library) gets a dimension
//! by size. Dimension scales that are only dimensions are not variables, and
//! a variable stored as `_nc4_non_coord_<name>` (a variable sharing a
//! dimension's name without being its coordinate variable) is `<name>`.
use std::collections::{HashMap, HashSet};
use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension, GroupDims};
use crate::error::Error;
use crate::variable::Variable;
/// The prefix netCDF-C gives the dataset of a variable that has a
/// dimension's name but is not that dimension's coordinate variable (the
/// dimension's scale holds the name).
const NON_COORD_PREFIX: &str = "_nc4_non_coord_";
/// A group, with the dimensions visible from it.
pub(crate) struct Scope<'f> {
file: &'f clawhdf5::File,
group: clawhdf5::Group<'f>,
/// This group's dimensions, then its parent's, and so on to the root's.
levels: Vec<GroupDims>,
}
impl<'f> Scope<'f> {
/// The group at `path` (`/`-separated from the root; `""` or `"/"` is
/// the root).
pub fn new(file: &'f clawhdf5::File, path: &str) -> Result<Self, Error> {
let parts: Vec<&str> = path.split('/').filter(|p| !p.is_empty()).collect();
let mut levels = Vec::with_capacity(parts.len() + 1);
for n in (0..=parts.len()).rev() {
let group = file.group(&parts[..n].join("/"))?;
levels.push(dimension::group_dims(file, &group)?);
}
let group = file.group(&parts.join("/"))?;
Ok(Self {
file,
group,
levels,
})
}
/// The group's datasets, as `(dataset name, object header address)` in
/// listing order.
fn datasets(&self) -> Result<Vec<(String, u64)>, Error> {
let datasets: HashSet<String> = self.group.datasets()?.into_iter().collect();
Ok(self
.group
.entries()?
.into_iter()
.filter(|(name, _)| datasets.contains(name))
.collect())
}
/// The group's variables: every dataset but the dimension scales that
/// are only dimensions.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let mut variables = Vec::new();
for (ds_name, address) in self.datasets()? {
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
continue;
}
variables.push(self.variable_from(nc_name(&ds_name), address, ds, attrs)?);
}
Ok(variables)
}
/// The names of the group's variables.
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
let mut names = Vec::new();
for (ds_name, address) in self.datasets()? {
let attrs = self.file.dataset_at(address)?.attrs()?;
if !dimension::is_pure_dimension(&attrs) {
names.push(nc_name(&ds_name));
}
}
Ok(names)
}
/// The variable called `name`: the dataset `_nc4_non_coord_<name>` if
/// there is one, else the dataset `<name>` unless it is only a
/// dimension.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let not_found = || Error::VariableNotFound(name.to_string());
let datasets = self.datasets()?;
let prefixed = format!("{NON_COORD_PREFIX}{name}");
let address = datasets
.iter()
.find(|(n, _)| *n == prefixed)
.or_else(|| datasets.iter().find(|(n, _)| n == name))
.map(|&(_, address)| address)
.ok_or_else(not_found)?;
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
return Err(not_found());
}
self.variable_from(nc_name(name), address, ds, attrs)
}
fn variable_from(
&self,
name: String,
address: u64,
ds: clawhdf5::Dataset<'f>,
attrs: HashMap<String, AttrValue>,
) -> Result<Variable<'f>, Error> {
let shape = ds.shape()?;
let dims = self.variable_dims(address, &attrs, &shape);
Ok(Variable::new(name, ds, dims, attrs))
}
/// The first dimension, searching this group and then its ancestors,
/// that `find` picks.
fn find<'a>(
&'a self,
find: impl Fn(&'a GroupDims) -> Option<&'a Dimension>,
) -> Option<Dimension> {
self.levels.iter().find_map(find).cloned()
}
/// The dimensions of the dataset at `address`, one per axis of `shape`,
/// as netCDF-C resolves them (see the module docs).
fn variable_dims(
&self,
address: u64,
attrs: &HashMap<String, AttrValue>,
shape: &[u64],
) -> Vec<Dimension> {
let rank = shape.len();
let mut dims: Vec<Option<Dimension>> = vec![None; rank];
if rank == 0 {
return Vec::new();
}
// A coordinate variable is the scale of its (first) dimension.
dims[0] = self.levels[0].by_address(address).cloned();
if let Some(ids) = coordinates(attrs).filter(|ids| ids.len() == rank) {
for (slot, id) in dims.iter_mut().zip(ids) {
if slot.is_none() {
*slot = self.find(|level| level.by_dimid(id));
}
}
}
if dims.iter().any(Option::is_none)
&& let Some(scales) =
dimension::dimension_list(self.file, attrs).filter(|s| s.len() == rank)
{
for (slot, scale) in dims.iter_mut().zip(scales) {
if slot.is_none()
&& let Some(scale) = scale
{
*slot = self.find(|level| level.by_address(scale));
}
}
}
// Neither: the first dimension of this group of the same size not
// already taken by another such axis, else an anonymous one.
let own = &self.levels[0].dims;
let mut used = vec![false; own.len()];
dims.into_iter()
.zip(shape)
.map(|(dim, &size)| {
dim.unwrap_or_else(|| {
match own
.iter()
.enumerate()
.find(|&(i, d)| !used[i] && d.size == size)
{
Some((i, d)) => {
used[i] = true;
d.clone()
}
None => Dimension {
name: format!("dim_{size}"),
size,
is_unlimited: false,
},
}
})
})
.collect()
}
}
/// The variable `name` of the group at `group_path`; `name` may itself be
/// a path (`"sub/var"`), relative to that group.
pub(crate) fn variable_at<'f>(
file: &'f clawhdf5::File,
group_path: &str,
name: &str,
) -> Result<Variable<'f>, Error> {
match name.trim_start_matches('/').rsplit_once('/') {
Some((dir, leaf)) => Scope::new(file, &format!("{group_path}/{dir}"))
.map_err(|_| Error::VariableNotFound(name.to_string()))?
.variable(leaf),
None => Scope::new(file, group_path)?.variable(name.trim_start_matches('/')),
}
}
/// The netCDF name of the dataset `ds_name`.
fn nc_name(ds_name: &str) -> String {
ds_name
.strip_prefix(NON_COORD_PREFIX)
.unwrap_or(ds_name)
.to_string()
}
/// A variable's `_Netcdf4Coordinates`: the `_Netcdf4Dimid` of the dimension
/// of each axis.
fn coordinates(attrs: &HashMap<String, AttrValue>) -> Option<Vec<i64>> {
match attrs.get("_Netcdf4Coordinates")? {
AttrValue::I64Array(ids) => Some(ids.clone()),
AttrValue::I64(id) => Some(vec![*id]),
AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(),
AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]),
_ => None,
}
}
+28 -3
View File
@@ -4,8 +4,16 @@
use clawhdf5::DType; use clawhdf5::DType;
/// NetCDF-4 data types corresponding to the standard NetCDF type system. /// NetCDF-4 data types corresponding to the standard NetCDF type system:
/// the atomic types, and the class of a user-defined type.
///
/// A dataset whose type netCDF-C cannot represent — a reference, bit
/// field, time or array type; a compound with such a member, a
/// half-precision float member, or a member of a user-defined type not
/// read before it; an enum or variable-length type over such a base — is
/// not a variable, as in netCDF-C.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] #[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
#[non_exhaustive]
pub enum NcType { pub enum NcType {
/// NC_BYTE: signed 8-bit integer /// NC_BYTE: signed 8-bit integer
Byte, Byte,
@@ -29,8 +37,18 @@ pub enum NcType {
Double, Double,
/// NC_STRING: variable-length string /// NC_STRING: variable-length string
String, String,
/// NC_CHAR: fixed-length string / character data /// NC_CHAR: a fixed-length string of one byte (longer ones are
/// `String`, as netCDF-C reads them)
Char, Char,
/// NC_ENUM: an enumeration (user-defined type)
Enum,
/// NC_COMPOUND: a compound (user-defined type)
Compound,
/// NC_VLEN: a variable-length sequence (user-defined type)
VLen,
/// NC_OPAQUE: an opaque type (user-defined type; netCDF4-python skips
/// such variables)
Opaque,
} }
impl std::fmt::Display for NcType { impl std::fmt::Display for NcType {
@@ -48,11 +66,16 @@ impl std::fmt::Display for NcType {
NcType::Double => write!(f, "NC_DOUBLE"), NcType::Double => write!(f, "NC_DOUBLE"),
NcType::String => write!(f, "NC_STRING"), NcType::String => write!(f, "NC_STRING"),
NcType::Char => write!(f, "NC_CHAR"), NcType::Char => write!(f, "NC_CHAR"),
NcType::Enum => write!(f, "NC_ENUM"),
NcType::Compound => write!(f, "NC_COMPOUND"),
NcType::VLen => write!(f, "NC_VLEN"),
NcType::Opaque => write!(f, "NC_OPAQUE"),
} }
} }
} }
/// Map a clawhdf5 DType to a NetCDF type. /// Map a clawhdf5 DType to a NetCDF type. A variable's own type, as
/// netCDF-C reads it, is [`Variable::nc_type`](crate::Variable::nc_type).
pub fn dtype_to_nctype(dtype: &DType) -> NcType { pub fn dtype_to_nctype(dtype: &DType) -> NcType {
match dtype { match dtype {
DType::I8 => NcType::Byte, DType::I8 => NcType::Byte,
@@ -66,6 +89,8 @@ pub fn dtype_to_nctype(dtype: &DType) -> NcType {
DType::F32 => NcType::Float, DType::F32 => NcType::Float,
DType::F64 => NcType::Double, DType::F64 => NcType::Double,
DType::String | DType::VariableLengthString => NcType::String, DType::String | DType::VariableLengthString => NcType::String,
DType::Compound(_) => NcType::Compound,
DType::Enum(_) => NcType::Enum,
_ => NcType::Char, // fallback for other types _ => NcType::Char, // fallback for other types
} }
} }
+18 -9
View File
@@ -16,7 +16,7 @@ use clawhdf5::AttrValue;
use crate::cf::{self, CfAttributes, FillValue}; use crate::cf::{self, CfAttributes, FillValue};
use crate::dimension::Dimension; use crate::dimension::Dimension;
use crate::error::Error; use crate::error::Error;
use crate::types::{NcType, dtype_to_nctype}; use crate::types::NcType;
/// A NetCDF-4 variable backed by an HDF5 dataset. /// A NetCDF-4 variable backed by an HDF5 dataset.
pub struct Variable<'f> { pub struct Variable<'f> {
@@ -28,6 +28,8 @@ pub struct Variable<'f> {
dims: Vec<Dimension>, dims: Vec<Dimension>,
/// The dataset's attributes. /// The dataset's attributes.
attrs: HashMap<String, AttrValue>, attrs: HashMap<String, AttrValue>,
/// Its netCDF type.
nc_type: NcType,
} }
impl<'f> Variable<'f> { impl<'f> Variable<'f> {
@@ -37,12 +39,14 @@ impl<'f> Variable<'f> {
dataset: clawhdf5::Dataset<'f>, dataset: clawhdf5::Dataset<'f>,
dims: Vec<Dimension>, dims: Vec<Dimension>,
attrs: HashMap<String, AttrValue>, attrs: HashMap<String, AttrValue>,
nc_type: NcType,
) -> Self { ) -> Self {
Self { Self {
name, name,
dataset, dataset,
dims, dims,
attrs, attrs,
nc_type,
} }
} }
@@ -51,11 +55,12 @@ impl<'f> Variable<'f> {
&self.name &self.name
} }
/// The dimensions of this variable, one per axis: the ones the file /// The dimensions of this variable, one per axis, as netCDF-C gives
/// gives it (`_Netcdf4Coordinates`, else `DIMENSION_LIST`), found in its /// them: the ones the file names (`_Netcdf4Coordinates`, else the
/// group or a parent group. An axis the file gives no dimension (a file /// scales `DIMENSION_LIST` attaches), found in its group or a parent
/// not written by a netCDF library) gets the first dimension of the /// group; for a dataset without them (a file not written by a netCDF
/// variable's group of the same size, else an anonymous `dim_<size>`. /// library), netCDF-C's phony dimensions `phony_dim_<n>`, shared by
/// the variables of a group by length.
pub fn dimensions(&self) -> &[Dimension] { pub fn dimensions(&self) -> &[Dimension] {
&self.dims &self.dims
} }
@@ -75,10 +80,11 @@ impl<'f> Variable<'f> {
Ok(self.dataset.shape()?) Ok(self.dataset.shape()?)
} }
/// The NetCDF data type of this variable. /// The NetCDF data type of this variable, as netCDF-C reads it (a
/// fixed-length string longer than one byte is `String`, one byte
/// long `Char`).
pub fn nc_type(&self) -> Result<NcType, Error> { pub fn nc_type(&self) -> Result<NcType, Error> {
let dtype = self.dataset.dtype()?; Ok(self.nc_type)
Ok(dtype_to_nctype(&dtype))
} }
/// Read all attributes as a HashMap. /// Read all attributes as a HashMap.
@@ -297,6 +303,9 @@ fn default_fill(nc_type: NcType) -> FillValue {
NcType::Double => FillValue::Float(9.969_209_968_386_869e36), NcType::Double => FillValue::Float(9.969_209_968_386_869e36),
NcType::String => FillValue::String(String::new()), NcType::String => FillValue::String(String::new()),
NcType::Char => FillValue::Int(0), NcType::Char => FillValue::Int(0),
// User-defined types have no default fill value in netCDF-C; their
// values are not read through these methods.
_ => FillValue::Int(0),
} }
} }
+281
View File
@@ -0,0 +1,281 @@
//! clawhdf5-netcdf4 compared with netCDF-C itself: `tests/netcdf_c_view.py`
//! prints what netCDF-C (the libnetcdf netCDF4-python bundles) reports for
//! a file, and [`differences`] lists how clawhdf5-netcdf4's reading differs.
#![allow(dead_code)]
use std::collections::{BTreeMap, BTreeSet};
use std::path::{Path, PathBuf};
use std::process::Command;
use std::sync::atomic::{AtomicUsize, Ordering};
use clawhdf5_netcdf4::{NcType, NetCDF4File, NetCDF4Group, Variable};
/// The Python with netCDF4-python (`CLAWHDF5_PYTHON`, else `python3`).
pub fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
/// netCDF-C's view of one file.
#[derive(Default)]
pub struct View {
/// `G`, `D` and `V` lines, in order.
pub meta: Vec<String>,
/// `(group, variable)` → values, from the `X` lines.
pub values: BTreeMap<(String, String), Vec<f64>>,
}
/// netCDF-C's view of each file (`None`: netCDF-C cannot open it), from
/// `tests/netcdf_c_view.py`.
pub fn netcdf_c_views(files: &[PathBuf]) -> BTreeMap<PathBuf, Option<View>> {
let dir = tempfile::tempdir().unwrap();
let list = dir.path().join("files.txt");
let mut paths = String::new();
for f in files {
paths.push_str(&f.to_string_lossy());
paths.push('\n');
}
std::fs::write(&list, paths).unwrap();
let script = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/netcdf_c_view.py");
let out = Command::new(python())
.arg(&script)
.arg("--list")
.arg(&list)
.output()
.expect("failed to run python");
assert!(
out.status.success(),
"netcdf_c_view.py failed: {}",
String::from_utf8_lossy(&out.stderr)
);
parse_views(&String::from_utf8_lossy(&out.stdout))
}
/// The views in the output of `tests/netcdf_c_view.py`.
pub fn parse_views(text: &str) -> BTreeMap<PathBuf, Option<View>> {
let mut views = BTreeMap::new();
let mut current: Option<(PathBuf, Option<View>)> = None;
for line in text.lines() {
let fields: Vec<&str> = line.split('\t').collect();
match fields[0] {
"FILE" => current = Some((PathBuf::from(fields[1]), Some(View::default()))),
"END" => {
let (path, view) = current.take().expect("END without FILE");
views.insert(path, view);
}
"ERROR" => current.as_mut().expect("ERROR without FILE").1 = None,
"X" => {
let view = current.as_mut().and_then(|c| c.1.as_mut()).unwrap();
let vals = if fields[3] == "-" {
Vec::new()
} else {
fields[3]
.split(' ')
.map(|v| v.parse().expect("value"))
.collect()
};
view.values
.insert((fields[1].to_string(), fields[2].to_string()), vals);
}
_ => {
let view = current.as_mut().and_then(|c| c.1.as_mut()).unwrap();
view.meta.push(line.to_string());
}
}
}
views
}
/// Variables whose type is labelled as netCDF-C labels it rather than as
/// clawhdf5-netcdf4 does (see [`nc_type_label`]).
pub static RELABELLED: AtomicUsize = AtomicUsize::new(0);
/// Values not compared because this build of the HDF5 reader lacks a
/// filter (a cargo feature).
pub static FILTER_SKIPPED: AtomicUsize = AtomicUsize::new(0);
/// The same `G`/`D`/`V` lines from clawhdf5-netcdf4.
pub fn our_meta(file: &NetCDF4File) -> Result<Vec<String>, String> {
let e = |e: clawhdf5_netcdf4::Error| e.to_string();
let mut out = Vec::new();
out.push("G\t/".to_string());
push_group(
&mut out,
file.hdf5_file(),
"/",
file.dimensions().map_err(e)?,
file.variables().map_err(e)?,
)?;
for name in file.group_names().map_err(e)? {
walk(
&mut out,
file.hdf5_file(),
&format!("/{name}"),
&file.group(&name).map_err(e)?,
)?;
}
Ok(out)
}
fn walk(
out: &mut Vec<String>,
hdf5: &clawhdf5::File,
path: &str,
group: &NetCDF4Group<'_>,
) -> Result<(), String> {
let e = |e: clawhdf5_netcdf4::Error| e.to_string();
out.push(format!("G\t{path}"));
push_group(
out,
hdf5,
path,
group.dimensions().map_err(e)?,
group.variables().map_err(e)?,
)?;
for name in group.group_names().map_err(e)? {
walk(
out,
hdf5,
&format!("{path}/{name}"),
&group.group(&name).map_err(e)?,
)?;
}
Ok(())
}
/// The type of variable `name` of the group at `path` as the `V` line
/// shows it: its `NcType`, except for the one deliberate difference from
/// netCDF-C — a float of other than 4 or 8 bytes (half, bfloat16, the 4-,
/// 6- and 8-bit floats, `long double`), `NC_FLOAT`/`NC_DOUBLE` here, is
/// labelled `NC_STRING` by netCDF-C 4.9.3 on libhdf5 1.14.6; such a
/// variable is shown as netCDF-C shows it, and counted.
fn nc_type_label(hdf5: &clawhdf5::File, path: &str, var: &Variable<'_>) -> Result<String, String> {
let nc_type = var.nc_type().map_err(|e| e.to_string())?;
if matches!(nc_type, NcType::Float | NcType::Double) {
let dataset = format!("{}/{}", path.trim_end_matches('/'), var.name());
if let Ok(ds) = hdf5.dataset(&dataset)
&& let Ok(clawhdf5_format::datatype::Datatype::FloatingPoint { size, .. }) =
ds.raw_datatype()
&& size != 4
&& size != 8
{
RELABELLED.fetch_add(1, Ordering::Relaxed);
return Ok("NC_STRING".to_string());
}
}
Ok(nc_type.to_string())
}
fn push_group(
out: &mut Vec<String>,
hdf5: &clawhdf5::File,
path: &str,
dims: Vec<clawhdf5_netcdf4::Dimension>,
vars: Vec<Variable<'_>>,
) -> Result<(), String> {
for d in dims {
out.push(format!(
"D\t{path}\t{}\t{}\t{}",
d.name,
d.size,
u8::from(d.is_unlimited)
));
}
for v in vars {
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
let shape: Vec<String> = v
.shape()
.map_err(|e| e.to_string())?
.iter()
.map(u64::to_string)
.collect();
let or_dash = |s: String| if s.is_empty() { "-".to_string() } else { s };
out.push(format!(
"V\t{path}\t{}\t{}\t{}\t{}",
v.name(),
nc_type_label(hdf5, path, &v)?,
or_dash(dims.join(",")),
or_dash(shape.join(","))
));
}
Ok(())
}
/// The values of variable `name` of group `group`.
fn our_values(file: &NetCDF4File, group: &str, name: &str) -> Result<Vec<f64>, String> {
let var = if group == "/" {
file.variable(name)
} else {
file.variable(&format!("{}/{name}", group.trim_start_matches('/')))
}
.map_err(|e| e.to_string())?;
match var.nc_type().map_err(|e| e.to_string())? {
NcType::String | NcType::Char => Err("not numeric".into()),
_ => var.read_raw_f64().map_err(|e| e.to_string()),
}
}
/// How clawhdf5-netcdf4's reading of `path` differs from `want`; empty
/// when it does not.
pub fn differences(path: &Path, want: &View) -> Vec<String> {
let file = match NetCDF4File::open(path) {
Ok(f) => f,
Err(e) => return vec![format!("netCDF-C opens it, clawhdf5-netcdf4 does not: {e}")],
};
let meta = match std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| our_meta(&file))) {
Ok(Ok(meta)) => meta,
Ok(Err(e)) => return vec![format!("error: {e}")],
Err(_) => return vec!["panic".to_string()],
};
let mut diffs = Vec::new();
if meta != want.meta {
let ours: BTreeSet<&String> = meta.iter().collect();
let theirs: BTreeSet<&String> = want.meta.iter().collect();
for line in want.meta.iter().filter(|l| !ours.contains(l)) {
diffs.push(format!("netCDF-C: {line}"));
}
for line in meta.iter().filter(|l| !theirs.contains(l)) {
diffs.push(format!("ours: {line}"));
}
if diffs.is_empty() {
diffs.push("same lines, different order".to_string());
}
}
for ((group, name), want) in &want.values {
let got = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
our_values(&file, group, name)
}))
.unwrap_or_else(|_| Err("panic".to_string()));
match got {
Ok(got) => {
let same = got.len() == want.len()
&& got
.iter()
.zip(want)
.all(|(a, b)| a.to_bits() == b.to_bits() || (a.is_nan() && b.is_nan()));
if !same {
diffs.push(format!("values of {group} {name} differ"));
}
}
Err(e) if e.contains("this build lacks the") => {
FILTER_SKIPPED.fetch_add(1, Ordering::Relaxed);
}
Err(e) => diffs.push(format!("values of {group} {name}: {e}")),
}
}
diffs
}
/// clawhdf5-netcdf4 reads the file at `path` as netCDF-C does: the same
/// groups, dimensions, variables and values (see [`differences`]).
pub fn assert_matches_netcdf_c(path: &Path) {
let views = netcdf_c_views(&[path.to_path_buf()]);
let Some(Some(want)) = views.get(path) else {
panic!("netCDF-C cannot open {}", path.display());
};
let diffs = differences(path, want);
assert!(
diffs.is_empty(),
"{} differs from netCDF-C:\n{}",
path.display(),
diffs.join("\n")
);
}
@@ -0,0 +1,26 @@
# Files of the conformance corpus (conformance/.cache/corpus) that netCDF-C
# opens and clawhdf5-netcdf4 reads differently, with the reason.
# <path relative to the corpus root><TAB><reason>
# Read by tests/corpus_vs_netcdf_c.rs; see docs/known-issues.md
# ("NetCDF-4: differences from netCDF-C").
#
# External links: libhdf5 follows them into the other file (present in the
# corpus); clawhdf5 does not follow external links (known-issues: "External
# links and external raw data are not followed"), so the linked groups and
# datasets are missing and, in files without dimension scales, the phony
# dimension numbers after them shift.
hdf5/test/testfiles/be_extlink1.h5 external link not followed
hdf5/test/testfiles/le_extlink1.h5 external link not followed
hdf5/tools/test/testfiles/h5diff_ext2softlink_src.h5 external link not followed
hdf5/tools/test/testfiles/h5diff_grp_recurse_ext2-1.h5 external link not followed
hdf5/tools/test/testfiles/h5diff_grp_recurse_ext2-2.h5 external link not followed
#
# Values the HDF5 reader refuses and libhdf5 1.14.6 (the netCDF4-python
# wheel's) returns; metadata matches. The first three are the scale-offset
# and N-Bit ref-bug / our-error files of CONFORMANCE.md (libhdf5 reads past
# the stored data); bad_nbit_decompress.h5 is not in the conformance run and
# is not investigated yet.
cve_hdf5/cvefiles/cve-2025-2308.h5 values: scale-offset chunk shorter than its values (libhdf5 over-read)
cve_hdf5/cvefiles/cve-2025-44904.h5 values: unfiltered chunks shorter than a chunk (libhdf5 over-read)
hdf5/test/testfiles/bad_nbit_parms_walk.h5 values: N-Bit parameters too short (conformance our-error/ref-bug)
hdf5/test/testfiles/bad_nbit_decompress.h5 values: N-Bit chunk refused ("element count exceeds chunk size"); not investigated
@@ -0,0 +1,173 @@
//! Every file of a corpus that netCDF-C opens, read by clawhdf5-netcdf4 and
//! compared with what netCDF-C reports: groups (order), dimensions (names,
//! lengths, unlimited, order), variables (names, order, types, dimensions,
//! shapes) and the values of numeric variables of at most 5000 elements.
//!
//! Gated: set `CLAWHDF5_NETCDF_CORPUS` to a directory (relative to the
//! workspace root, or absolute) — the conformance corpus
//! (`conformance/.cache/corpus`) or any tree of HDF5/netCDF-4 files — and
//! `CLAWHDF5_PYTHON` to a Python with netCDF4-python (whose bundled
//! libnetcdf `tests/netcdf_c_view.py` calls). Without the variable the test
//! does nothing. netCDF-C reads each file in its own process (4 at a time,
//! `CLAWHDF5_NETCDF_JOBS`), under a 60 s timeout and a 4 GiB address-space
//! limit: 40 s to 3 minutes for the conformance corpus on tank.
//! `CLAWHDF5_NETCDF_VIEW` may name the saved output of an earlier
//! `netcdf_c_view.py --list <file of paths>` run over the same (absolute)
//! paths, to skip that.
//!
//! A file whose differences are explained is listed, with the reason, in
//! `tests/corpus_known_differences.txt` (and in `docs/known-issues.md`); a
//! listed file that matches fails the test too, so the list stays true.
//!
//! ```sh
//! CLAWHDF5_NETCDF_CORPUS=conformance/.cache/corpus CLAWHDF5_PYTHON=$PWD/.venv/bin/python \
//! cargo test -p clawhdf5-netcdf4 --test corpus_vs_netcdf_c -- --nocapture
//! ```
mod common;
use std::collections::BTreeMap;
use std::path::{Path, PathBuf};
use std::sync::atomic::Ordering;
use common::{FILTER_SKIPPED, RELABELLED, differences, netcdf_c_views, parse_views};
/// The file extensions `conformance/list_files.py` sweeps.
const EXTS: [&str; 7] = ["h5", "hdf5", "he5", "nc", "nc4", "hdf", "h5f"];
/// The files of the corpus: those with an HDF5/netCDF-4 extension, except
/// netCDF classic files (magic `CDF`), following directory symlinks, in
/// byte order of their paths.
fn corpus_files(root: &Path) -> Vec<PathBuf> {
fn walk(dir: &Path, out: &mut Vec<PathBuf>, depth: usize) {
let Ok(entries) = std::fs::read_dir(dir) else {
return;
};
for entry in entries.flatten() {
let path = entry.path();
if path.file_name().is_some_and(|n| n == ".git") {
continue;
}
let Ok(meta) = std::fs::metadata(&path) else {
continue;
};
if meta.is_dir() && depth < 32 {
walk(&path, out, depth + 1);
} else if meta.is_file()
&& path
.extension()
.and_then(|e| e.to_str())
.is_some_and(|e| EXTS.contains(&e.to_ascii_lowercase().as_str()))
{
let mut magic = [0u8; 3];
let classic = std::fs::File::open(&path)
.and_then(|mut f| std::io::Read::read_exact(&mut f, &mut magic))
.is_ok()
&& &magic == b"CDF";
if !classic {
out.push(path);
}
}
}
}
let mut out = Vec::new();
walk(root, &mut out, 0);
out.sort_by(|a, b| {
a.as_os_str()
.as_encoded_bytes()
.cmp(b.as_os_str().as_encoded_bytes())
});
out
}
/// The files `tests/corpus_known_differences.txt` explains: path relative
/// to the corpus root → reason.
fn known_differences() -> BTreeMap<String, String> {
let list = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/corpus_known_differences.txt");
std::fs::read_to_string(list)
.expect("tests/corpus_known_differences.txt")
.lines()
.filter(|l| !l.trim().is_empty() && !l.starts_with('#'))
.map(|l| {
let (path, reason) = l.split_once('\t').unwrap_or((l, ""));
(path.trim().to_string(), reason.trim().to_string())
})
.collect()
}
#[test]
fn corpus_matches_netcdf_c() {
let Ok(root) = std::env::var("CLAWHDF5_NETCDF_CORPUS") else {
eprintln!("SKIP: set CLAWHDF5_NETCDF_CORPUS to a corpus directory");
return;
};
// A relative path is from the workspace root (cargo runs the test in
// the crate's directory).
let root = Path::new(env!("CARGO_MANIFEST_DIR"))
.join("../..")
.join(root);
let root = std::fs::canonicalize(&root).unwrap_or(root);
let files = corpus_files(&root);
assert!(!files.is_empty(), "no files under {}", root.display());
// `CLAWHDF5_NETCDF_VIEW`: the output of an earlier run of
// `tests/netcdf_c_view.py` over the same files, which takes a few
// minutes over the conformance corpus.
let views = match std::env::var("CLAWHDF5_NETCDF_VIEW") {
Ok(saved) => parse_views(&std::fs::read_to_string(saved).expect("CLAWHDF5_NETCDF_VIEW")),
Err(_) => netcdf_c_views(&files),
};
let known = known_differences();
let (mut opened, mut matched) = (0, 0);
let mut explained = Vec::new();
let mut unexplained = Vec::new();
let mut stale = Vec::new();
for path in &files {
let Some(Some(want)) = views.get(path) else {
continue;
};
opened += 1;
let rel = path
.strip_prefix(&root)
.unwrap_or(path)
.to_string_lossy()
.into_owned();
let diffs = differences(path, want);
match (diffs.is_empty(), known.get(&rel)) {
(true, None) => matched += 1,
(true, Some(_)) => stale.push(rel),
(false, Some(reason)) => explained.push((rel, reason.clone(), diffs)),
(false, None) => unexplained.push((rel, diffs)),
}
}
eprintln!(
"{} files; netCDF-C opens {opened}; {matched} match; {} differ as explained; \
{} differ unexplained; {} listed but match; {} variables relabelled \
(floats of other than 4 or 8 bytes); {} variables' values not compared \
(filter not in this build)",
files.len(),
explained.len(),
unexplained.len(),
stale.len(),
RELABELLED.load(Ordering::Relaxed),
FILTER_SKIPPED.load(Ordering::Relaxed),
);
for (rel, reason, diffs) in &explained {
eprintln!("explained: {rel} ({reason}): {} differences", diffs.len());
}
for (rel, diffs) in &unexplained {
eprintln!("DIFFERS: {rel}");
for d in diffs.iter().take(20) {
eprintln!(" {d}");
}
}
assert!(
unexplained.is_empty(),
"{} files differ from netCDF-C unexplained",
unexplained.len()
);
assert!(
stale.is_empty(),
"listed in corpus_known_differences.txt but match: {stale:?}"
);
}
@@ -2,6 +2,8 @@
//! //!
//! Tests are skipped if python3 or netCDF4/xarray Python packages are not available. //! Tests are skipped if python3 or netCDF4/xarray Python packages are not available.
mod common;
use std::process::Command; use std::process::Command;
use clawhdf5_netcdf4::{AttrValue, NcType, NetCDF4File}; use clawhdf5_netcdf4::{AttrValue, NcType, NetCDF4File};
@@ -842,3 +844,267 @@ ds.to_netcdf({path:?}, engine={engine:?}, unlimited_dims=["time"])
assert_same_view(&path); assert_same_view(&path);
} }
} }
// ===========================================================================
// Compared with netCDF-C itself (tests/netcdf_c_view.py): groups,
// dimensions, variables, types and values, in netCDF-C's order
// ===========================================================================
/// The dimensions of the variable `name` (a path) of `file`.
fn dim_names(file: &NetCDF4File, name: &str) -> Vec<String> {
file.variable(name)
.unwrap()
.dimensions()
.iter()
.map(|d| d.name.clone())
.collect()
}
/// An h5py file without dimension scales: netCDF-C's phony dimensions,
/// numbered file-wide (subgroups before their parent's variables, each
/// group's datasets in name order, as h5py does not track creation order),
/// shared by length within a group — but not between two axes of one
/// variable, nor between a fixed and an unlimited axis — with a length of 0
/// always unlimited and never shared with a fixed axis, and the real
/// dimension of a scale taken by length too.
#[test]
fn phony_dimensions_match_netcdf_c() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("phony.h5");
run_python(&format!(
r#"
import h5py
import numpy as np
with h5py.File({path:?}, "w") as f:
f["zz"] = np.arange(12.0).reshape(2, 2, 3)
f["aa"] = np.arange(6, dtype="i4").reshape(3, 2)
f.create_dataset("un", data=np.ones((2, 3), "f4"), maxshape=(None, 3))
f.create_dataset("un2", data=np.ones(2, "f4"), maxshape=(None,))
f["zero"] = np.zeros((0,))
f.create_dataset("zero_un", (0,), "f4", maxshape=(None,))
f["zero2"] = np.zeros((0, 2))
f["s"] = 1.5
g = f.create_group("g")
g["x"] = np.arange(5, dtype="i2")
g["y"] = np.arange(2, dtype="u1")
g.create_group("h")["q"] = np.arange(7.0)
f.create_group("b")["w"] = np.arange(18, dtype="i8").reshape(2, 9)
f["sc"] = np.arange(6.0)
f["sc"].make_scale("sc")
f["second_only"] = np.zeros((2, 6))
f["second_only"].dims[1].attach_scale(f["sc"])
"#,
path = path.display().to_string()
));
common::assert_matches_netcdf_c(&path);
let file = NetCDF4File::open(&path).unwrap();
// `sc` is dimension 0; /b gets 1 and 2, /g/h 3, /g 4 and 5, / from 6.
assert_eq!(dim_names(&file, "b/w"), ["phony_dim_1", "phony_dim_2"]);
assert_eq!(dim_names(&file, "g/h/q"), ["phony_dim_3"]);
let h = file.group("g/h").unwrap();
assert_eq!(h.variable_names().unwrap(), ["q"]);
assert_eq!(h.dimensions().unwrap()[0].name, "phony_dim_3");
assert_eq!(dim_names(&file, "aa"), ["phony_dim_6", "phony_dim_7"]);
assert_eq!(
dim_names(&file, "zz"),
["phony_dim_7", "phony_dim_11", "phony_dim_6"]
);
assert_eq!(dim_names(&file, "second_only"), ["phony_dim_7", "sc"]);
assert_eq!(dim_names(&file, "un"), ["phony_dim_8", "phony_dim_6"]);
assert_eq!(dim_names(&file, "un2"), ["phony_dim_8"]);
assert_eq!(dim_names(&file, "zero"), ["phony_dim_9"]);
assert_eq!(dim_names(&file, "zero_un"), ["phony_dim_9"]);
assert_eq!(dim_names(&file, "zero2"), ["phony_dim_10", "phony_dim_7"]);
let root: Vec<(String, u64, bool)> = file
.dimensions()
.unwrap()
.into_iter()
.map(|d| (d.name, d.size, d.is_unlimited))
.collect();
assert_eq!(root[0], ("sc".to_string(), 6, false));
assert!(root.contains(&("phony_dim_8".to_string(), 2, true)));
assert!(root.contains(&("phony_dim_9".to_string(), 0, true)));
assert!(root.contains(&("phony_dim_10".to_string(), 0, true)));
assert_eq!(
file.variable_names().unwrap(),
[
"aa",
"s",
"sc",
"second_only",
"un",
"un2",
"zero",
"zero2",
"zero_un",
"zz"
]
);
}
/// Datasets of types netCDF-C cannot represent are not variables:
/// references, bit fields, array types, a compound with a reference member
/// or a half-float member, a compound nesting a compound not seen before;
/// enum, compound, variable-length and opaque types are variables of those
/// classes. netCDF-C also remembers a type it failed to read, so the second
/// dataset of a compound with a reference member is a variable, and a
/// nested compound is accepted once a dataset of the inner type was read.
#[test]
fn types_netcdf_c_skips_are_not_variables() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("types.h5");
run_python(&format!(
r#"
import h5py
import numpy as np
inner = np.dtype([("x", "i2"), ("y", "f4")])
with h5py.File({path:?}, "w") as f:
f["i4"] = np.arange(3, dtype="i4")
f["f2"] = np.arange(3, dtype="f2")
f["s1"] = np.array([b"a", b"b"], dtype="S1")
f["s5"] = np.array([b"abc", b"b"], dtype="S5")
f["vs"] = np.array(["x", "yy"], dtype=h5py.string_dtype())
f["bool"] = np.array([True, False])
f["cmp"] = np.array([(1, 2.0)], dtype=[("a", "i4"), ("b", "f8")])
f["en"] = np.array([0, 1], dtype=h5py.enum_dtype({{"A": 0, "B": 1}}, basetype="i1"))
d = f.create_dataset("vl", (2,), dtype=h5py.vlen_dtype("i4"))
d[0] = [1, 2]
d[1] = [3]
f["op"] = np.array([b"ab", b"cd"], dtype="V2")
f.create_dataset("ref", (1,), dtype=h5py.ref_dtype)[0] = f["i4"].ref
f["cref1"] = np.array([(1, f["i4"].ref)], dtype=[("a", "i4"), ("r", h5py.ref_dtype)])
f["cref2"] = np.array([(2, f["i4"].ref)], dtype=[("a", "i4"), ("r", h5py.ref_dtype)])
f["cmp_f2"] = np.zeros(2, dtype=[("a", "f2")])
f["in1"] = np.zeros(2, dtype=inner)
f["nested"] = np.zeros(2, dtype=[("a", "i4"), ("in", inner)])
f["a_nested"] = np.zeros(2, dtype=[("b", "i4"), ("in", np.dtype([("p", "i1")]))])
sid = h5py.h5s.create_simple((2,))
h5py.h5d.create(f.id, b"bitf", h5py.h5t.STD_B8LE.copy(), sid)
h5py.h5d.create(f.id, b"arr", h5py.h5t.array_create(h5py.h5t.NATIVE_INT32, (3,)), sid)
"#,
path = path.display().to_string()
));
common::assert_matches_netcdf_c(&path);
let file = NetCDF4File::open(&path).unwrap();
assert_eq!(
file.variable_names().unwrap(),
[
"bool", "cmp", "cref2", "en", "f2", "i4", "in1", "nested", "op", "s1", "s5", "vl", "vs"
]
);
let nc_type = |name: &str| file.variable(name).unwrap().nc_type().unwrap();
assert_eq!(nc_type("bool"), NcType::Enum);
assert_eq!(nc_type("cmp"), NcType::Compound);
assert_eq!(nc_type("vl"), NcType::VLen);
assert_eq!(nc_type("op"), NcType::Opaque);
assert_eq!(nc_type("s1"), NcType::Char);
assert_eq!(nc_type("s5"), NcType::String);
// netCDF-C 4.9.3 on libhdf5 1.14.6 says NC_STRING (deliberate
// difference, see the README).
assert_eq!(nc_type("f2"), NcType::Float);
assert_eq!(
file.variable("f2").unwrap().read_raw_f32().unwrap(),
[0.0, 1.0, 2.0]
);
assert!(matches!(
file.variable("ref"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
}
/// The order of groups and variables: creation order where the group
/// tracks it (every netCDF-4 file; also h5py with `track_order`), compact
/// or dense (more than 8 links), else name order (h5py by default).
#[test]
fn group_and_variable_order_match_netcdf_c() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let nc = dir.path().join("order.nc");
let h5 = dir.path().join("order.h5");
run_python(&format!(
r#"
import h5py
import netCDF4 as nc
import numpy as np
names = ["zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c"]
with nc.Dataset({nc:?}, "w") as f:
f.createDimension("x", 2)
for i, n in enumerate(names):
f.createVariable(n, "i4", ("x",))[:] = [i, i + 1]
for n in ["gz", "ga", "gm"]:
f.createGroup(n).createVariable("v", "f8", ("x",))[:] = [1, 2]
small = f.createGroup("small")
for n in ["q", "b", "p"]:
small.createVariable(n, "i2", ("x",))[:] = [3, 4]
with h5py.File({h5:?}, "w") as f:
for i, n in enumerate(names):
f[n] = np.arange(i + 1)
t = f.create_group("tracked", track_order=True)
for i, n in enumerate(names):
t[n] = np.arange(3, dtype="i1")
for n in ["gz", "ga"]:
f.create_group(n)["v"] = np.arange(4.0)
"#,
nc = nc.display().to_string(),
h5 = h5.display().to_string()
));
common::assert_matches_netcdf_c(&nc);
common::assert_matches_netcdf_c(&h5);
let file = NetCDF4File::open(&nc).unwrap();
assert_eq!(
file.variable_names().unwrap(),
[
"zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c"
]
);
assert_eq!(file.group_names().unwrap(), ["gz", "ga", "gm", "small"]);
let file = NetCDF4File::open(&h5).unwrap();
assert_eq!(file.group_names().unwrap(), ["ga", "gz", "tracked"]);
assert_eq!(
file.variable_names().unwrap(),
[
"a", "alpha", "beta", "c", "gamma", "k", "mid", "omega", "zeta", "zz"
]
);
assert_eq!(
file.group("tracked").unwrap().variable_names().unwrap(),
[
"zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c"
]
);
}
/// The files of the earlier tests, compared with netCDF-C itself too:
/// dimension scales of h5py, netCDF4-python files with groups, unlimited
/// dimensions and non-coordinate variables named like a dimension.
#[test]
fn netcdf4_python_files_match_netcdf_c() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("mixed.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("time", None)
f.createDimension("p", 2)
f.createDimension("q", 2)
f.createVariable("a", "i4", ("time",))[0:2] = [1, 2]
f.createVariable("b", "f4", ("time", "q"))[0:5, :] = np.arange(10).reshape(5, 2)
f.createVariable("q", "f4", ("q",))[:] = [0, 1]
f.createVariable("p", "f4", ("q", "p"))[:] = np.array([[0, 1], [2, 3]])
g = f.createGroup("g")
g.createDimension("r", 3)
g.createVariable("w", "i4", ("r", "time"))[:, 0:1] = np.ones((3, 1))
g.createGroup("h").createVariable("z", "i8", ("p", "r"))[:] = np.arange(6).reshape(2, 3)
"#,
path = path.display().to_string()
));
common::assert_matches_netcdf_c(&path);
}
@@ -776,9 +776,12 @@ fn test_pure_dimensions_hidden_and_non_coord_names() {
.with_shape(&[2]); .with_shape(&[2]);
let file = NetCDF4File::from_bytes(b.finish().unwrap()).unwrap(); let file = NetCDF4File::from_bytes(b.finish().unwrap()).unwrap();
// `x`, and the phony dimension netCDF-C gives the 3 values of
// `_nc4_non_coord_x` (no dimension scale attached).
let dims = file.dimensions().unwrap(); let dims = file.dimensions().unwrap();
assert_eq!(dims.len(), 1); let names: Vec<&str> = dims.iter().map(|d| d.name.as_str()).collect();
assert_eq!(dims[0].name, "x"); assert_eq!(names, ["x", "phony_dim_1"]);
assert_eq!(dims[1].size, 3);
let mut names = file.variable_names().unwrap(); let mut names = file.variable_names().unwrap();
names.sort(); names.sort();
assert_eq!(names, ["v", "x"]); assert_eq!(names, ["v", "x"]);
@@ -786,7 +789,8 @@ fn test_pure_dimensions_hidden_and_non_coord_names() {
assert_eq!(x.name(), "x"); assert_eq!(x.name(), "x");
assert_eq!(x.read_raw_f64().unwrap(), [1.0, 2.0, 3.0]); assert_eq!(x.read_raw_f64().unwrap(), [1.0, 2.0, 3.0]);
assert!(!x.is_coordinate()); assert!(!x.is_coordinate());
// No DIMENSION_LIST: `v` gets `x` by size, as before. // No DIMENSION_LIST: `v` gets the group's first dimension of its length
// (netCDF-C's phony rule, which also takes real dimensions).
assert_eq!(file.variable("v").unwrap().dimensions()[0].name, "x"); assert_eq!(file.variable("v").unwrap().dimensions()[0].name, "x");
} }
@@ -0,0 +1,238 @@
#!/usr/bin/env python3
"""netcdf_c_view.py FILE...: print each file as netCDF-C sees it.
The metadata comes from netCDF-C itself (the libnetcdf that netCDF4-python
bundles, called through ctypes), not from netCDF4-python's objects, which
leave out variables of types netCDF-C supports but netCDF4-python does not
(opaque, compounds of vlen strings, ...). Values come from netCDF4-python.
Each file is read in its own process under a timeout, so a file that crashes
or hangs libnetcdf only costs that file.
Output, per file, fields separated by tabs:
FILE <path>
ERROR <message> netCDF-C cannot open it; nothing else
G <group> every group, pre-order, children in
netCDF-C's order ("/" is the root)
D <group> <name> <length> <0|1> its dimensions in dimension-id order
(1: unlimited)
V <group> <name> <type> <dims> <shape>
its variables in netCDF-C's order;
<dims> and <shape> comma-separated,
"-" when there are none; <type> as
clawhdf5_netcdf4::NcType prints it
X <group> <name> <values> the values of a numeric variable of
at most MAX_VALUES elements, as
space-separated Python float reprs
("-" when there are none)
END
Used by tests/corpus_vs_netcdf_c.rs (CLAWHDF5_NETCDF_CORPUS).
"""
import concurrent.futures
import ctypes
import glob
import os
import subprocess
import sys
MAX_VALUES = 5000
TIMEOUT = 60
MEMORY = 4 << 30 # address space of each file's process
ATOMIC = {
1: "NC_BYTE", 2: "NC_CHAR", 3: "NC_SHORT", 4: "NC_INT", 5: "NC_FLOAT",
6: "NC_DOUBLE", 7: "NC_UBYTE", 8: "NC_USHORT", 9: "NC_UINT",
10: "NC_INT64", 11: "NC_UINT64", 12: "NC_STRING",
}
USER_CLASS = {13: "NC_VLEN", 14: "NC_OPAQUE", 15: "NC_ENUM", 16: "NC_COMPOUND"}
NUMERIC = {"NC_BYTE", "NC_SHORT", "NC_INT", "NC_FLOAT", "NC_DOUBLE",
"NC_UBYTE", "NC_USHORT", "NC_UINT", "NC_INT64", "NC_UINT64"}
NAME = 257 # NC_MAX_NAME + 1
MAXDIMS = 1024
def libnetcdf():
"""The libnetcdf netCDF4-python is linked with (same process, same copy)."""
import netCDF4
site = os.path.dirname(os.path.dirname(netCDF4.__file__))
found = glob.glob(os.path.join(site, "netcdf4.libs", "libnetcdf*.so*")) + \
glob.glob(os.path.join(site, "netCDF4.libs", "libnetcdf*.so*"))
if found:
return ctypes.CDLL(found[0])
return ctypes.CDLL("libnetcdf.so")
def view(path, out):
nc = libnetcdf()
ncid = ctypes.c_int()
rc = nc.nc_open(path.encode(), 0, ctypes.byref(ncid)) # NC_NOWRITE
if rc != 0:
nc.nc_strerror.restype = ctypes.c_char_p
out.append("ERROR\t" + nc.nc_strerror(rc).decode(errors="replace"))
return
numeric = [] # (group path, variable name)
try:
walk(nc, ncid.value, "/", out, numeric)
finally:
nc.nc_close(ncid)
values(path, numeric, out)
def check(rc, what):
if rc != 0:
raise RuntimeError(f"{what} failed: {rc}")
def name_of(fn, *args):
buf = ctypes.create_string_buffer(NAME)
check(fn(*args, buf), fn.__name__)
return buf.value.decode(errors="replace")
def walk(nc, gid, path, out, numeric):
out.append(f"G\t{path}")
n = ctypes.c_int()
ids = (ctypes.c_int * MAXDIMS)()
check(nc.nc_inq_dimids(gid, ctypes.byref(n), ids, 0), "nc_inq_dimids")
for dimid in ids[: n.value]:
length = ctypes.c_size_t()
dname = ctypes.create_string_buffer(NAME)
check(nc.nc_inq_dim(gid, dimid, dname, ctypes.byref(length)), "nc_inq_dim")
out.append(f"D\t{path}\t{dname.value.decode(errors='replace')}\t{length.value}\t"
f"{int(is_unlimited(nc, gid, dimid))}")
nvars = ctypes.c_int()
varids = (ctypes.c_int * 65536)()
check(nc.nc_inq_varids(gid, ctypes.byref(nvars), varids), "nc_inq_varids")
for varid in varids[: nvars.value]:
vname = ctypes.create_string_buffer(NAME)
xtype = ctypes.c_int()
ndims = ctypes.c_int()
dimids = (ctypes.c_int * MAXDIMS)()
natts = ctypes.c_int()
check(nc.nc_inq_var(gid, varid, vname, ctypes.byref(xtype), ctypes.byref(ndims),
dimids, ctypes.byref(natts)), "nc_inq_var")
names, shape = [], []
for dimid in dimids[: ndims.value]:
dname = ctypes.create_string_buffer(NAME)
length = ctypes.c_size_t()
if nc.nc_inq_dim(gid, dimid, dname, ctypes.byref(length)) != 0:
names.append("?")
shape.append("?")
continue
names.append(dname.value.decode(errors="replace"))
shape.append(str(length.value))
tname = type_name(nc, gid, xtype.value)
vn = vname.value.decode(errors="replace")
out.append(f"V\t{path}\t{vn}\t{tname}\t{','.join(names) or '-'}\t{','.join(shape) or '-'}")
if tname in NUMERIC and "?" not in shape:
count = 1
for s in shape:
count *= int(s)
if count <= MAX_VALUES:
numeric.append((path, vn))
ngrps = ctypes.c_int()
grps = (ctypes.c_int * 65536)()
check(nc.nc_inq_grps(gid, ctypes.byref(ngrps), grps), "nc_inq_grps")
for child in grps[: ngrps.value]:
cname = ctypes.create_string_buffer(NAME)
check(nc.nc_inq_grpname(child, cname), "nc_inq_grpname")
cpath = path.rstrip("/") + "/" + cname.value.decode(errors="replace")
walk(nc, child, cpath, out, numeric)
def is_unlimited(nc, gid, dimid):
n = ctypes.c_int()
ids = (ctypes.c_int * MAXDIMS)()
# Unlimited dimensions visible from this group include its parents'.
if nc.nc_inq_unlimdims(gid, ctypes.byref(n), ids) != 0:
return False
return dimid in ids[: n.value]
def type_name(nc, gid, xtype):
if xtype in ATOMIC:
return ATOMIC[xtype]
size = ctypes.c_size_t()
base = ctypes.c_int()
nfields = ctypes.c_size_t()
klass = ctypes.c_int()
tname = ctypes.create_string_buffer(NAME)
if nc.nc_inq_user_type(gid, xtype, tname, ctypes.byref(size), ctypes.byref(base),
ctypes.byref(nfields), ctypes.byref(klass)) != 0:
return f"type{xtype}"
return USER_CLASS.get(klass.value, f"class{klass.value}")
def values(path, numeric, out):
if not numeric:
return
import numpy as np
import netCDF4
try:
ds = netCDF4.Dataset(path)
except Exception: # noqa: BLE001 - netCDF4-python refuses some files netCDF-C opens
return
with ds:
for gpath, name in numeric:
try:
group = ds if gpath == "/" else ds[gpath]
var = group.variables[name]
var.set_auto_maskandscale(False)
# A whole-variable read through netCDF-C 4.9.3 lays out a
# variable shorter than an unlimited dimension that is not its
# first wrongly (written values first); reads of one index of
# the leading axis are right.
if var.ndim >= 2:
data = np.stack([np.asarray(var[i]) for i in range(var.shape[0])]) \
if var.shape[0] else np.zeros(var.shape)
else:
data = np.asarray(var[...])
flat = np.asarray(data, dtype=np.float64).ravel()
except Exception: # noqa: BLE001
continue
vals = " ".join(repr(float(v)) for v in flat) or "-"
out.append(f"X\t{gpath}\t{name}\t{vals}")
def limit_memory():
import resource
resource.setrlimit(resource.RLIMIT_AS, (MEMORY, MEMORY))
def one(path):
"""Run view() for one file in a child process."""
try:
proc = subprocess.run([sys.executable, __file__, "--one", path],
capture_output=True, timeout=TIMEOUT, text=True,
preexec_fn=limit_memory)
except subprocess.TimeoutExpired:
return [f"FILE\t{path}", "ERROR\ttimeout", "END"]
lines = proc.stdout.splitlines()
if proc.returncode != 0 or not lines or lines[-1] != "END":
err = (proc.stderr.strip().splitlines() or [f"exit {proc.returncode}"])[-1]
return [f"FILE\t{path}", f"ERROR\tchild failed: {err}", "END"]
return lines
def main(argv):
if argv and argv[0] == "--one":
out = [f"FILE\t{argv[1]}"]
try:
view(argv[1], out)
except Exception as e: # noqa: BLE001
out = [f"FILE\t{argv[1]}", f"ERROR\t{e}"]
out.append("END")
sys.stdout.write("\n".join(out) + "\n")
return
if argv and argv[0] == "--list":
with open(argv[1]) as fh:
argv = [line.rstrip("\n") for line in fh if line.strip()]
jobs = int(os.environ.get("CLAWHDF5_NETCDF_JOBS", "4"))
with concurrent.futures.ThreadPoolExecutor(jobs) as pool:
for lines in pool.map(one, argv):
sys.stdout.write("\n".join(lines) + "\n")
if __name__ == "__main__":
main(sys.argv[1:])
+16 -1
View File
@@ -104,7 +104,22 @@ writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and
`complex128` arrays; the file is written on `close()`. Complex arrays are `complex128` arrays; the file is written on `close()`. Complex arrays are
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the (not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
Rust API writes that on request). Rust API writes that on request). Both forms read back as numpy
`complex64`/`complex128`.
`libver=` sets the library version bounds as in h5py: `'v108'`,
`'v110'`, `'v112'`, `'v114'`, `'v200'` or `'latest'` (the low bound; the
high bound is then `'latest'`), or a `(low, high)` tuple. With `'v108'`
the file is written in the HDF5 1.8 format (version-2 superblock,
version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the
HDF5 1.10 format. `'earliest'` as the low bound writes the 1.8 format too,
with a `UserWarning`: clawhdf5 cannot write the pre-1.8 format. The
argument is ignored when reading and refused for `'r+'`/`'a'`.
```python
with clawhdf5.File("old.h5", "w", libver="v108") as f: # HDF5 1.8 reads it
f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip")
```
## Editing a file in place ## Editing a file in place
+90 -2
View File
@@ -13,10 +13,13 @@ use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group}; use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle; use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err}; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
use clawhdf5_rs::LibVer;
/// Internal state for write mode. /// Internal state for write mode.
struct WriteState { struct WriteState {
path: PathBuf, path: PathBuf,
/// `libver=` as (low, high); `None` keeps the writer's default.
libver: Option<(LibVer, LibVer)>,
root_datasets: Vec<DatasetSpec>, root_datasets: Vec<DatasetSpec>,
root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>, root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
groups: Vec<Arc<Mutex<WriteGroupState>>>, groups: Vec<Arc<Mutex<WriteGroupState>>>,
@@ -80,10 +83,32 @@ impl PyFile {
/// `az://`; which schemes work depends on how the wheel was built) /// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`) /// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write /// mode: 'r' for read (default), 'w' for write
/// libver: library version bounds for mode 'w', as h5py's: one of
/// 'earliest', 'v108', 'v110', 'v112', 'v114', 'v200', 'latest' (the
/// low bound; the high bound is then 'latest') or a (low, high)
/// tuple of them. The low bound is the oldest HDF5 release whose
/// format the file uses ('v108': HDF5 1.8 can read it); the high
/// bound the newest whose features it may use. clawhdf5 cannot write
/// the pre-1.8 format, so a low bound of 'earliest' writes the 1.8
/// format (with a warning) and a high bound of 'earliest' is an
/// error. Default (None): the HDF5 1.10 format clawhdf5 has always
/// written. Ignored for reading; not supported with 'r+' / 'a'.
#[new] #[new]
#[pyo3(signature = (path, mode="r"))] #[pyo3(signature = (path, mode="r", libver=None))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> { fn new(
py: Python<'_>,
path: &str,
mode: &str,
libver: Option<&Bound<'_, PyAny>>,
) -> PyResult<Self> {
let filename = path.to_string(); let filename = path.to_string();
let libver = libver.map(|v| parse_libver(py, v)).transpose()?;
if libver.is_some() && matches!(mode, "r+" | "a") {
return Err(PyNotImplementedError::new_err(format!(
"libver with mode '{mode}': clawhdf5's in-place editor keeps the format \
versions the file already uses"
)));
}
if is_url(path) { if is_url(path) {
if mode != "r" { if mode != "r" {
return Err(PyValueError::new_err(format!( return Err(PyValueError::new_err(format!(
@@ -113,6 +138,7 @@ impl PyFile {
// Absolute now: the file is written at close, possibly // Absolute now: the file is written at close, possibly
// after the working directory changed. // after the working directory changed.
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)), path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
libver,
root_datasets: Vec::new(), root_datasets: Vec::new(),
root_attrs: Arc::new(Mutex::new(Vec::new())), root_attrs: Arc::new(Mutex::new(Vec::new())),
groups: Vec::new(), groups: Vec::new(),
@@ -460,10 +486,71 @@ fn parse_compression(
} }
} }
/// One of h5py's `libver` names as a bound; `high` says which end it is.
/// `Ok(None)` is 'earliest' as the low bound: the pre-1.8 format, which
/// clawhdf5 cannot write.
fn libver_name(name: &str, high: bool) -> PyResult<Option<LibVer>> {
Ok(Some(match name {
"earliest" if high => {
return Err(PyValueError::new_err(
"libver high bound 'earliest' (the pre-1.8 format) cannot be written by \
clawhdf5; the oldest format it writes is 'v108'",
));
}
"earliest" => return Ok(None),
"v108" => LibVer::V18,
"v110" => LibVer::V110,
"v112" => LibVer::V112,
"v114" => LibVer::V114,
"v200" => LibVer::V200,
"latest" => LibVer::Latest,
other => {
return Err(PyValueError::new_err(format!(
"unknown libver '{other}'; expected 'earliest', 'v108', 'v110', 'v112', \
'v114', 'v200' or 'latest'"
)));
}
}))
}
/// h5py's `libver=`: a name (the low bound, high bound 'latest') or a
/// `(low, high)` pair.
fn parse_libver(py: Python<'_>, v: &Bound<'_, PyAny>) -> PyResult<(LibVer, LibVer)> {
let (low, high): (String, String) = match v.extract::<String>() {
Ok(name) => (name, "latest".into()),
Err(_) => v.extract().map_err(|_| {
PyValueError::new_err("libver must be a string or a (low, high) tuple of strings")
})?,
};
let Some(high) = libver_name(&high, true)? else {
unreachable!("a high bound is never None")
};
let low = match libver_name(&low, false)? {
Some(low) => low,
None => {
PyModule::import(py, "warnings")?.getattr("warn")?.call1((
"libver 'earliest': clawhdf5 cannot write the pre-1.8 format; the file \
is written in the HDF5 1.8 format ('v108') instead",
py.get_type::<pyo3::exceptions::PyUserWarning>(),
))?;
LibVer::V18
}
};
if low > high {
return Err(PyValueError::new_err(format!(
"libver low bound {low} is newer than the high bound {high}"
)));
}
Ok((low, high))
}
/// Build and write the HDF5 file from accumulated write state. /// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> { fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| { crate::no_panic(|| {
let mut builder = clawhdf5_rs::FileBuilder::new(); let mut builder = clawhdf5_rs::FileBuilder::new();
if let Some((low, high)) = state.libver {
builder.libver_bounds(low, high);
}
// Root attributes // Root attributes
let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner()); let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner());
@@ -520,6 +607,7 @@ mod tests {
let state = WriteState { let state = WriteState {
path: path.clone(), path: path.clone(),
libver: None,
root_datasets: vec![DatasetSpec { root_datasets: vec![DatasetSpec {
name: "data".into(), name: "data".into(),
data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]), data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]),
+143
View File
@@ -0,0 +1,143 @@
"""`clawhdf5.File(path, 'w', libver=...)`: h5py's library version bounds.
A file written with libver='v108' must open in HDF5 1.8. Its h5dump is
found through CLAWHDF5_H5DUMP18 or at ~/.cache/hdf5-1.8.23/bin/h5dump
(scripts/build-hdf5-1.8.sh builds it); without it that check is skipped.
"""
import os
import subprocess
import numpy as np
import pytest
import clawhdf5
def superblock_version(path):
with open(path, "rb") as f:
head = f.read(9)
assert head[:8] == b"\x89HDF\r\n\x1a\n"
return head[8]
def h5dump18():
path = os.environ.get("CLAWHDF5_H5DUMP18") or os.path.expanduser(
"~/.cache/hdf5-1.8.23/bin/h5dump"
)
try:
out = subprocess.run([path, "--version"], capture_output=True, text=True)
except OSError:
return None
return path if "1.8." in out.stdout else None
def write_sample(path, libver):
with clawhdf5.File(path, "w", libver=libver) as f:
f.create_dataset("x", data=np.arange(10, dtype=np.float64))
f.create_dataset(
"chunked",
data=np.arange(100, dtype=np.int32),
chunks=(30,),
compression="gzip",
)
f.create_dataset("z", data=np.array([1 + 2j, -3j], dtype=np.complex128))
g = f.create_group("g")
g.create_dataset("y", data=np.ones(3, dtype=np.float32))
f.attrs["version"] = 1
@pytest.mark.parametrize(
"libver, sb",
[
(None, 3),
("v108", 2),
(("v108", "v108"), 2),
(("v108", "latest"), 2),
("v110", 3),
("v112", 3),
("v114", 3),
("v200", 3),
("latest", 3),
(("v110", "v200"), 3),
],
)
def test_libver_bounds(h5py, tmp_path, libver, sb):
path = str(tmp_path / "libver.h5")
write_sample(path, libver)
# v108 low bound: HDF5 1.8's version-2 superblock; 1.10 and later: 3.
assert superblock_version(path) == sb
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
np.testing.assert_array_equal(f["z"][:], [1 + 2j, -3j])
np.testing.assert_array_equal(f["g/y"][:], np.ones(3))
assert f.attrs["version"] == 1
with clawhdf5.File(path, "r") as f:
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
def test_libver_v108_opens_in_hdf5_1_8(tmp_path):
h5dump = h5dump18()
if h5dump is None:
pytest.skip("no HDF5 1.8 h5dump (CLAWHDF5_H5DUMP18)")
path = str(tmp_path / "v108.h5")
write_sample(path, "v108")
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode == 0, out.stderr
assert "h5dump error" not in out.stderr
assert 'DATASET "y"' in out.stdout
# The chunked, deflated dataset (a version-1 B-tree) reads in full.
out = subprocess.run(
[h5dump, "-w", "0", "-d", "/chunked", path], capture_output=True, text=True
)
assert out.returncode == 0, out.stderr
assert "(0): " + ", ".join(str(k) for k in range(100)) + "\n" in out.stdout
# The default (1.10 format) does not open in 1.8: the bound matters.
path = str(tmp_path / "default.h5")
write_sample(path, None)
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode != 0
def test_libver_earliest_writes_v108_with_a_warning(h5py, tmp_path):
path = str(tmp_path / "earliest.h5")
with pytest.warns(UserWarning, match="pre-1.8"):
write_sample(path, "earliest")
assert superblock_version(path) == 2
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
@pytest.mark.parametrize(
"libver, match",
[
(("v108", "earliest"), "high bound 'earliest'"),
("v109", "unknown libver 'v109'"),
(("latest", "v108"), "newer than the high bound"),
(("v108",), "tuple"),
(3, "tuple"),
],
)
def test_libver_errors(tmp_path, libver, match):
with pytest.raises(ValueError, match=match):
clawhdf5.File(str(tmp_path / "bad.h5"), "w", libver=libver)
def test_libver_v108_high_bound_writes_everything(tmp_path):
# Native complex numbers need HDF5 2.0; the Python writer stores complex
# as h5py's {r, i} compound, which 1.8 reads, so the 1.8 high bound is
# fine for everything clawhdf5.File writes.
path = str(tmp_path / "v18only.h5")
write_sample(path, ("v108", "v108"))
assert superblock_version(path) == 2
def test_libver_not_for_editing(tmp_path):
path = str(tmp_path / "e.h5")
write_sample(path, None)
with pytest.raises(NotImplementedError, match="libver"):
clawhdf5.File(path, "r+", libver="latest")
# Reading ignores it, as the bounds only affect what is written.
with clawhdf5.File(path, "r", libver="v108") as f:
assert f["x"].shape == (10,)
+22 -1
View File
@@ -603,7 +603,22 @@ impl Diff {
let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) { let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) {
(Some(x), Some(y), ..) => x.abs_diff(y).to_string(), (Some(x), Some(y), ..) => x.abs_diff(y).to_string(),
(_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64), (_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64),
// Each part's absolute difference, as h5diff 2.x prints
// it (`1+0i`).
_ => match (&va, &vb) {
(
Value::Complex { re: ar, im: ai, .. },
Value::Complex { re: br, im: bi, .. },
) => match (number(ar), number(ai), number(br), number(bi)) {
(Some(ar), Some(ai), Some(br), Some(bi)) => format!(
"{}+{}i",
value::fmt_float((ar - br).abs(), 64),
value::fmt_float((ai - bi).abs(), 64)
),
_ => String::new(), _ => String::new(),
},
_ => String::new(),
},
}; };
rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}")); rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}"));
} }
@@ -713,6 +728,9 @@ impl Diff {
(Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => { (Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => {
p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v)) p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v))
} }
(Value::Complex { re: pr, im: pi, .. }, Value::Complex { re: qr, im: qi, .. }) => {
self.equal(a, b, pr, qr) && self.equal(a, b, pi, qi)
}
(Value::Ref(None), Value::Ref(None)) => true, (Value::Ref(None), Value::Ref(None)) => true,
(Value::Ref(Some(p)), Value::Ref(Some(q))) => { (Value::Ref(Some(p)), Value::Ref(Some(q))) => {
// Addresses mean nothing across files: compare the paths the // Addresses mean nothing across files: compare the paths the
@@ -846,7 +864,10 @@ fn types_comparable(a: &Datatype, b: &Datatype) -> bool {
( (
Datatype::Enumeration { base_type: x, .. }, Datatype::Enumeration { base_type: x, .. },
Datatype::Enumeration { base_type: y, .. }, Datatype::Enumeration { base_type: y, .. },
) => types_comparable(x, y), )
| (Datatype::Complex { base_type: x, .. }, Datatype::Complex { base_type: y, .. }) => {
types_comparable(x, y)
}
_ => true, _ => true,
} }
} }
+34 -6
View File
@@ -135,9 +135,16 @@ pub fn short(dt: &Datatype) -> String {
return f.short.into(); return f.short.into();
} }
match dt { match dt {
Datatype::Complex { size, base_type } => { // numpy's names (`complex64` is two `float32`) for IEEE parts,
short(&Datatype::complex_as_compound(*size, base_type)) // `complex<part>` otherwise.
} Datatype::Complex { size, base_type } => match base_type.as_ref() {
Datatype::FloatingPoint { byte_order, .. } if is_ieee(base_type) => format!(
"complex{}{}",
u64::from(*size) * 8,
if be(byte_order) { "-be" } else { "" }
),
_ => format!("complex<{}>", short(base_type)),
},
Datatype::FixedPoint { Datatype::FixedPoint {
size, size,
signed, signed,
@@ -216,8 +223,9 @@ pub fn long(dt: &Datatype) -> String {
return f.long.into(); return f.long.into();
} }
match dt { match dt {
Datatype::Complex { size, base_type } => { // As h5ls 2.x prints a complex type it has no native name for.
long(&Datatype::complex_as_compound(*size, base_type)) Datatype::Complex { base_type, .. } => {
format!("complex number of\n {}", long(base_type))
} }
Datatype::FixedPoint { Datatype::FixedPoint {
size, size,
@@ -361,6 +369,19 @@ fn atomic_ddl(dt: &Datatype) -> Option<String> {
)), )),
// As h5dump 2.x names them (checked against h5dump 2.2.0). // As h5dump 2.x names them (checked against h5dump 2.2.0).
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()), Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
// h5dump 2.x's predefined complex names: IEEE binary16/32/64 parts.
Datatype::Complex { base_type, .. } => match base_type.as_ref() {
Datatype::FloatingPoint {
size, byte_order, ..
} if is_ieee(base_type) && !matches!(byte_order, DatatypeByteOrder::Vax) => {
Some(format!(
"H5T_COMPLEX_IEEE_F{}{}",
u64::from(*size) * 8,
order_suffix(byte_order)
))
}
_ => None,
},
Datatype::BitField { Datatype::BitField {
size, byte_order, .. size, byte_order, ..
} => Some(format!( } => Some(format!(
@@ -418,6 +439,10 @@ pub fn ddl(dt: &Datatype, ind: usize) -> String {
Datatype::VariableLength { base_type, .. } => { Datatype::VariableLength { base_type, .. } => {
format!("H5T_VLEN {{ {} }}", ddl(base_type, ind)) format!("H5T_VLEN {{ {} }}", ddl(base_type, ind))
} }
// Parts h5dump has no predefined complex name for.
Datatype::Complex { base_type, .. } => {
format!("H5T_COMPLEX {{ {} }}", ddl(base_type, ind))
}
Datatype::Opaque { size, tag } => { Datatype::Opaque { size, tag } => {
let tag = String::from_utf8_lossy(tag); let tag = String::from_utf8_lossy(tag);
format!( format!(
@@ -498,6 +523,9 @@ fn string_ddl(
/// hdf5-json type object. /// hdf5-json type object.
pub fn json(dt: &Datatype) -> J { pub fn json(dt: &Datatype) -> J {
match dt { match dt {
// hdf5-json (h5json 2.0.0) has no complex class: a native complex
// is written as the `{r, i}` compound h5py uses for complex numbers,
// its values as `[re, im]` pairs.
Datatype::Complex { size, base_type } => { Datatype::Complex { size, base_type } => {
json(&Datatype::complex_as_compound(*size, base_type)) json(&Datatype::complex_as_compound(*size, base_type))
} }
@@ -598,7 +626,7 @@ fn pad_json(p: &StringPadding) -> &'static str {
/// Class name used to decide whether two datatypes can be compared. /// Class name used to decide whether two datatypes can be compared.
pub fn class(dt: &Datatype) -> &'static str { pub fn class(dt: &Datatype) -> &'static str {
match dt { match dt {
Datatype::Complex { .. } => "compound", Datatype::Complex { .. } => "complex",
Datatype::FixedPoint { .. } => "integer", Datatype::FixedPoint { .. } => "integer",
Datatype::FloatingPoint { .. } => "float", Datatype::FloatingPoint { .. } => "float",
Datatype::Time { .. } => "time", Datatype::Time { .. } => "time",
+31 -5
View File
@@ -27,6 +27,15 @@ pub enum Value {
/// An enum member (name, when the value matches one) and its value. /// An enum member (name, when the value matches one) and its value.
Enum(Option<String>, i128), Enum(Option<String>, i128),
Compound(Vec<(String, Value)>), Compound(Vec<(String, Value)>),
/// A native complex number (HDF5 2.0 class 11): real and imaginary
/// part. `sign_free` is set when h5dump joins the parts with a bare `+`
/// (parts it has no native C type for: binary16, non-IEEE), giving
/// `1+-2i`; otherwise the imaginary part carries its sign (`1-2i`).
Complex {
re: Box<Value>,
im: Box<Value>,
sign_free: bool,
},
Array(Vec<Value>), Array(Vec<Value>),
/// A variable-length sequence. /// A variable-length sequence.
Seq(Vec<Value>), Seq(Vec<Value>),
@@ -183,11 +192,18 @@ impl<'a> Decoder<'a> {
return Value::Error("short element".into()); return Value::Error("short element".into());
}; };
match dt { match dt {
Datatype::Complex { size, base_type } => self.decode( Datatype::Complex { base_type, .. } => {
&Datatype::complex_as_compound(*size, base_type), let part = base_type.type_size() as usize;
b, let (re, im) = b.split_at(part.min(b.len()));
depth + 1, Value::Complex {
), re: Box::new(self.decode(base_type, re, depth + 1)),
im: Box::new(self.decode(base_type, im, depth + 1)),
// h5dump prints `float`/`double` complex (whatever the
// byte order) with `%g%+gi`, anything else as
// `<re>+<im>i`.
sign_free: !(dtype::is_ieee(base_type) && matches!(part, 4 | 8)),
}
}
Datatype::FixedPoint { .. } => match decode_int(dt, b) { Datatype::FixedPoint { .. } => match decode_int(dt, b) {
Some(v) => Value::Int(v), Some(v) => Value::Int(v),
None => Value::Bytes(b.to_vec()), None => Value::Bytes(b.to_vec()),
@@ -354,6 +370,15 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
.collect::<Vec<_>>() .collect::<Vec<_>>()
.join(", ") .join(", ")
), ),
Value::Complex { re, im, sign_free } => {
let (re, im) = (text(re, h5paths), text(im, h5paths));
let plus = if *sign_free || !im.starts_with('-') {
"+"
} else {
""
};
format!("{re}{plus}{im}i")
}
Value::Array(vs) => format!( Value::Array(vs) => format!(
"[ {} ]", "[ {} ]",
vs.iter() vs.iter()
@@ -408,6 +433,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)), Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths), Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()), Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
Value::Complex { re, im, .. } => J::Array(vec![to_json(re, h5paths), to_json(im, h5paths)]),
Value::Array(vs) | Value::Seq(vs) => { Value::Array(vs) | Value::Seq(vs) => {
J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect()) J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect())
} }
@@ -0,0 +1,202 @@
//! `h5rs dump` and `h5rs ls` of HDF5 2.0 native complex numbers (datatype
//! class 11), against the output of h5dump 2.2.0 and h5ls 2.2.0 (the Debian
//! h5dump CI installs, 1.14.x, predates the type). The h5dump output is
//! stored next to the fixture; see
//! `crates/clawhdf5/tests/fixtures/gen_native_complex.py`.
use std::path::PathBuf;
use std::process::Command;
fn fixture(name: &str) -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name)
}
fn h5rs(args: &[&str]) -> String {
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(args)
.output()
.unwrap();
assert!(out.status.success(), "h5rs {args:?}: {out:?}");
String::from_utf8(out.stdout).unwrap()
}
/// Split a dump into its lines, with the data lines (`(i): ...`) replaced
/// by their index, and the data values in order.
fn split(ddl: &str) -> (Vec<String>, Vec<String>) {
let (mut frame, mut data) = (Vec::new(), Vec::new());
for line in ddl.lines() {
match line.split_once("):") {
Some((idx, body)) if idx.trim_start().starts_with('(') => {
frame.push(format!("{idx}):"));
// Values are split at `, `; array elements come bracketed.
data.extend(
body.split(", ")
.map(|s| s.trim().trim_matches(|c| c == '[' || c == ']' || c == ','))
.map(str::trim)
.filter(|s| !s.is_empty() && *s != "{")
.map(String::from),
);
}
_ => {
let t = line.trim().trim_end_matches(',');
if t.ends_with('i') || t.parse::<f64>().is_ok() {
// A compound member's value on its own line.
data.push(t.to_string());
} else {
frame.push(line.to_string());
}
}
}
}
(frame, data)
}
/// `a+bi`, `a-bi` or `a+-bi` as its two parts (a real number as one).
fn parts(v: &str) -> Vec<f64> {
let Some(z) = v.strip_suffix('i') else {
return vec![v.parse().unwrap()];
};
let b = z.as_bytes();
let at = (1..b.len())
.find(|&k| (b[k] == b'+' || b[k] == b'-') && !matches!(b[k - 1], b'e' | b'E' | b'+'))
.unwrap_or_else(|| panic!("not a complex value: {v}"));
let (re, im) = z.split_at(at);
let im = im.strip_prefix('+').unwrap_or(im);
vec![re.parse().unwrap(), im.parse().unwrap()]
}
#[test]
fn dump_matches_h5dump_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
let ours = h5rs(&["dump", path.to_str().unwrap()]);
let reference = std::fs::read_to_string(fixture("native_complex_hdf5_2.ddl")).unwrap();
let (mut our_frame, our_data) = split(&ours);
let (mut ref_frame, ref_data) = split(&reference);
// The first line names the file as it was given.
our_frame.remove(0);
ref_frame.remove(0);
// Everything but the values is h5dump's byte for byte: the types print
// as H5T_COMPLEX_IEEE_F64LE, H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }, ...
assert_eq!(our_frame, ref_frame);
assert_eq!(our_data.len(), ref_data.len(), "{our_data:?}\n{ref_data:?}");
assert_eq!(our_data.len(), 36);
// Values: h5dump prints `%g` (6 digits), inf/nan, and the binary16
// parts it has no C type for as `<re>+<im>i` (`1+-2i`); h5rs prints
// the shortest round-trip string and Inf/NaN, joined the same way.
for (i, (o, r)) in our_data.iter().zip(&ref_data).enumerate() {
assert_eq!(o.contains("+-"), r.contains("+-"), "[{i}] {o} vs {r}");
let (op, rp) = (parts(o), parts(r));
assert_eq!(op.len(), rp.len(), "[{i}] {o} vs {r}");
for (x, y) in op.iter().zip(&rp) {
let same = if y.is_nan() || y.is_infinite() {
x.is_nan() == y.is_nan() && (y.is_nan() || x == y)
} else {
*x as f32 == *y as f32 || (x - y).abs() <= 5e-6 * y.abs()
};
assert!(same, "[{i}] h5rs {o}, h5dump {r}");
assert_eq!(
x.is_sign_negative(),
y.is_sign_negative(),
"[{i}] {o} vs {r}"
);
}
}
}
#[test]
fn ls_describes_complex_like_h5ls_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
// h5ls 2.2.0 -v, `Type:` of the types it has no native C name for (it
// prints `native double _Complex` for the others, as it prints `native
// double` where h5rs prints `IEEE 64-bit little-endian float`).
for (name, long) in [
(
"c128be",
"complex number of\n IEEE 64-bit big-endian float",
),
(
"c128",
"complex number of\n IEEE 64-bit little-endian float",
),
(
"c32",
"complex number of\n IEEE 16-bit little-endian float",
),
(
"array",
"[2] complex number of\n IEEE 32-bit little-endian float",
),
] {
let target = format!("{}/{name}", path.display());
let verbose = h5rs(&["ls", "-v", &target]);
assert!(
verbose.contains(&format!(" Type: {long}\n")),
"{name}:\n{verbose}"
);
}
let listing = h5rs(&["ls", path.to_str().unwrap()]);
for (name, short) in [
("c64", "complex64"),
("c128", "complex128"),
("c128be", "complex128-be"),
("c32", "complex32"),
("scalar", "complex128"),
("array", "array[2]<complex64>"),
("compound", "compound{z: complex128, k: int64}"),
] {
assert!(
listing
.lines()
.any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))),
"{name}:\n{listing}"
);
}
}
#[test]
fn json_dump_uses_the_r_i_compound() {
// hdf5-json (h5json 2.0.0) has no complex class: the type is written as
// the `{r, i}` compound h5py uses, the values as `[re, im]` pairs.
let path = fixture("native_complex_hdf5_2.h5");
let j: serde_json::Value =
serde_json::from_str(&h5rs(&["dump", "--json", path.to_str().unwrap()])).unwrap();
let ds = j["datasets"]
.as_object()
.unwrap()
.values()
.find(|d| d["alias"][0] == "/scalar")
.unwrap();
assert_eq!(ds["type"]["class"], "H5T_COMPOUND");
assert_eq!(ds["type"]["fields"][0]["name"], "r");
assert_eq!(ds["type"]["fields"][1]["type"]["base"], "H5T_IEEE_F64LE");
assert_eq!(ds["value"], serde_json::json!([2.5, -0.5]));
}
#[test]
fn diff_compares_complex_values() {
let path = fixture("native_complex_hdf5_2.h5");
let p = path.to_str().unwrap();
// Same values, different byte order: no differences.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", p, p, "/c128", "/c128be"])
.output()
.unwrap();
assert!(out.status.success(), "{out:?}");
// One element differs in its imaginary part: one difference, exit
// status 1, as h5diff 2.2.0 reports it.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", "-r", p, p, "/c128", "/c128x"])
.output()
.unwrap();
assert_eq!(out.status.code(), Some(1), "{out:?}");
let text = String::from_utf8_lossy(&out.stdout);
assert!(text.contains("1 difference(s) found"), "{text}");
// h5diff 2.2.0: `[ 1 1 ] 0+1i 1+1i 1+0i`.
let row = text.lines().find(|l| l.starts_with("[ 1 1 ]")).unwrap();
assert_eq!(
row.split_whitespace().collect::<Vec<_>>(),
["[", "1", "1", "]", "0+1i", "1+1i", "1+0i"]
);
}
+85 -7
View File
@@ -312,11 +312,14 @@ impl Reader {
let base = array_base(dt); let base = array_base(dt);
let is_array = !std::ptr::eq(base, dt); let is_array = !std::ptr::eq(base, dt);
Ok(match base { Ok(match base {
// Read as the part type: an array or complex element is its
// parts stored one after another (a complex number's real, then
// imaginary part).
Datatype::FloatingPoint { size, .. } if *size <= 4 => { Datatype::FloatingPoint { size, .. } if *size <= 4 => {
Data::F32(data_read::read_as_f32(raw, dt).map_err(err)?) Data::F32(data_read::read_as_f32(raw, base).map_err(err)?)
} }
Datatype::FloatingPoint { .. } => { Datatype::FloatingPoint { .. } => {
Data::F64(data_read::read_as_f64(raw, dt).map_err(err)?) Data::F64(data_read::read_as_f64(raw, base).map_err(err)?)
} }
Datatype::FixedPoint { size, signed, .. } => { Datatype::FixedPoint { size, signed, .. } => {
let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err); let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err);
@@ -393,10 +396,13 @@ fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T
.collect() .collect()
} }
/// Innermost element type of (possibly nested) array datatypes. /// Innermost element type of (possibly nested) array datatypes; for a
/// native complex type, its part type (each element holds two).
fn array_base(dt: &Datatype) -> &Datatype { fn array_base(dt: &Datatype) -> &Datatype {
match dt { match dt {
Datatype::Array { base_type, .. } => array_base(base_type), Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => {
array_base(base_type)
}
_ => dt, _ => dt,
} }
} }
@@ -411,6 +417,12 @@ fn element_shape(dt: &Datatype) -> Vec<u64> {
dims.extend(element_shape(base_type)); dims.extend(element_shape(base_type));
dims dims
} }
// `[re, im]`: a complex element reads as its two parts.
Datatype::Complex { base_type, .. } => {
let mut dims = vec![2];
dims.extend(element_shape(base_type));
dims
}
_ => Vec::new(), _ => Vec::new(),
} }
} }
@@ -487,9 +499,7 @@ pub fn describe(dt: &Datatype) -> String {
} }
} }
match dt { match dt {
Datatype::Complex { size, base_type } => { Datatype::Complex { base_type, .. } => format!("complex<{}>", describe(base_type)),
describe(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::FixedPoint { Datatype::FixedPoint {
size, size,
signed, signed,
@@ -681,6 +691,74 @@ mod tests {
assert!(e.contains("not supported"), "{e}"); assert!(e.contains("not supported"), "{e}");
} }
#[test]
fn native_complex_reads_as_re_im_pairs() {
let z = [[1.0f64, -2.0], [0.5, 0.0], [-0.0, 3.25], [f64::MAX, 1e-300]];
let mut b = FileBuilder::new();
b.create_dataset("z")
.with_native_complex_f64_data(&z)
.with_shape(&[2, 2]);
b.create_dataset("z32")
.with_native_complex_f32_data(&[[1.5f32, -2.5]]);
let r = Reader::open(b.finish().unwrap()).unwrap();
let i = r.info("z").unwrap();
assert_eq!(
(i.shape, i.dtype, i.element_shape),
(vec![2, 2], "complex<f64>".to_string(), vec![2])
);
let all = r.read("z", None).unwrap();
assert_eq!(all.shape, vec![2, 2, 2]);
assert_eq!(all.data, Data::F64(z.iter().flatten().copied().collect()));
let slab = Hyperslab {
start: vec![1, 0],
count: vec![1, 2],
stride: None,
block: None,
};
let part = r.read("z", Some(&slab)).unwrap();
assert_eq!(part.shape, vec![1, 2, 2]);
assert_eq!(part.data, Data::F64(vec![-0.0, 3.25, f64::MAX, 1e-300]));
let z32 = r.read("z32", None).unwrap();
assert_eq!(
(z32.shape, z32.data),
(vec![1, 2], Data::F32(vec![1.5, -2.5]))
);
}
#[test]
fn native_complex_written_by_libhdf5_2() {
// Written by h5py 3.16 / libhdf5 2.0.0 (gen_native_complex.py).
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5/tests/fixtures/native_complex_hdf5_2.h5"
);
let r = Reader::open(std::fs::read(path).unwrap()).unwrap();
let be = r.read("c128be", None).unwrap();
assert_eq!(be.shape, vec![2, 3, 2]);
assert_eq!(be.data, r.read("c128", None).unwrap().data);
assert_eq!(r.info("c128be").unwrap().dtype, "complex<f64 (big-endian)>");
// binary16 parts widen to f32.
assert_eq!(
r.read("c32", None).unwrap().data,
Data::F32(vec![1.0, -2.0, 0.5, 65504.0, 0.0, -0.0])
);
// An array of complex: array dimensions, then the two parts.
let i = r.info("array").unwrap();
assert_eq!(
(i.dtype.as_str(), i.element_shape),
("array[2]<complex<f32>>", vec![2, 2])
);
assert_eq!(
r.read("array", None).unwrap().data,
Data::F32(vec![1.0, 2.0, 3.0, 4.0, -1.0, -1.0, 0.0, 0.0])
);
let s = r.read("scalar", None).unwrap();
assert_eq!((s.shape, s.data), (vec![2], Data::F64(vec![2.5, -0.5])));
// A compound holding a complex member is still refused.
let e = r.read("compound", None).unwrap_err();
assert!(e.contains("compound{z: complex<f64>, k: i64}"), "{e}");
}
#[test] #[test]
fn garbage_is_an_error() { fn garbage_is_an_error() {
assert!(Reader::open(vec![0u8; 64]).is_err()); assert!(Reader::open(vec![0u8; 64]).is_err());
+4 -1
View File
@@ -1342,7 +1342,10 @@ impl<'f> Dataset<'f> {
) -> Result<(Vec<T>, Vec<T>), Error> { ) -> Result<(Vec<T>, Vec<T>), Error> {
let dt = self.datatype()?; let dt = self.datatype()?;
let is_complex = match &dt { let is_complex = match &dt {
// Class 11 parses to this same `{r, i}` compound. Datatype::Complex { size, base_type } => {
matches!(base_type.as_ref(), Datatype::FloatingPoint { .. })
&& *size == 2 * base_type.type_size()
}
Datatype::Compound { size, members } => { Datatype::Compound { size, members } => {
matches!(members.as_slice(), [r, i] matches!(members.as_slice(), [r, i]
if r.name == "r" && i.name == "i" if r.name == "r" && i.name == "i"
+11
View File
@@ -41,6 +41,13 @@ pub enum DType {
Array(Box<DType>, Vec<u32>), Array(Box<DType>, Vec<u32>),
/// Variable-length UTF-8 string stored via a global heap reference. /// Variable-length UTF-8 string stored via a global heap reference.
VariableLengthString, VariableLengthString,
/// HDF5 2.0 native complex number (datatype class 11) over the given
/// floating-point part type: `Complex(F32)` is numpy `complex64`,
/// `Complex(F64)` `complex128`, and a half-precision part is
/// `Complex(Other("float16"))`. h5py's default complex encoding, the
/// compound `{r, i}`, stays a [`DType::Compound`]; `read_complex_f64`
/// and `read_complex_f32` read both.
Complex(Box<DType>),
/// Catch-all for HDF5 datatypes that do not map to a specific variant. /// Catch-all for HDF5 datatypes that do not map to a specific variant.
Other(String), Other(String),
} }
@@ -60,6 +67,7 @@ impl fmt::Display for DType {
DType::U64 => write!(f, "u64"), DType::U64 => write!(f, "u64"),
DType::String => write!(f, "string"), DType::String => write!(f, "string"),
DType::VariableLengthString => write!(f, "vlen_string"), DType::VariableLengthString => write!(f, "vlen_string"),
DType::Complex(base) => write!(f, "complex<{base}>"),
DType::Compound(fields) => { DType::Compound(fields) => {
write!(f, "compound{{")?; write!(f, "compound{{")?;
for (i, (name, dt)) in fields.iter().enumerate() { for (i, (name, dt)) in fields.iter().enumerate() {
@@ -148,6 +156,9 @@ pub(crate) fn classify_datatype(dt: &clawhdf5_format::datatype::Datatype) -> DTy
base_type, base_type,
dimensions, dimensions,
} => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()), } => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()),
Datatype::Complex { base_type, .. } => {
DType::Complex(Box::new(classify_datatype(base_type)))
}
_ => DType::Other(format!("{dt:?}")), _ => DType::Other(format!("{dt:?}")),
} }
} }
+75
View File
@@ -0,0 +1,75 @@
"""Generate native_complex_hdf5_2.h5: HDF5 2.0 native complex (class 11).
Datasets (every one class 11, or a type holding it):
c64 H5T_COMPLEX_IEEE_F32LE, 4 values incl. -0, inf, nan, subnormal
c128 H5T_COMPLEX_IEEE_F64LE, 2 x 3
c128be H5T_COMPLEX_IEEE_F64BE, the same values
c128x H5T_COMPLEX_IEEE_F64LE, c128 with element (1,1) changed
c32 H5T_COMPLEX_IEEE_F16LE, 3 values
scalar H5T_COMPLEX_IEEE_F64LE, scalar dataspace
compound {z: H5T_COMPLEX_IEEE_F64LE @0, k: H5T_STD_I64LE @16}
array H5T_ARRAY [2] of H5T_COMPLEX_IEEE_F32LE, 2 elements
plus the root attribute `zattr` (H5T_COMPLEX_IEEE_F64LE, 2 values).
Written with h5py 3.16 (libhdf5 2.0.0) through its low-level API, the file
type as the memory type (no conversion). Re-run only if the fixture ever
needs regenerating (generated 2026-09-29):
python gen_native_complex.py native_complex_hdf5_2.h5
The h5dump / h5ls 2.2.0 reference output the tests compare with was made
with the tools of the libhdf5 2.2.0 build described in gen_mx_floats.py:
h5dump native_complex_hdf5_2.h5 > native_complex_hdf5_2.ddl
h5ls -r -v native_complex_hdf5_2.h5 > native_complex_hdf5_2.ls
"""
import sys
import h5py
import numpy as np
from h5py import h5a, h5d, h5s, h5t
out = sys.argv[1]
f = h5py.File(out, "w", libver="latest")
def ds(name, tid, arr, shape=None):
shape = arr.shape if shape is None else shape
sp = h5s.create_simple(shape) if shape != () else h5s.create(h5s.SCALAR)
d = h5d.create(f.id, name.encode(), tid, sp)
d.write(h5s.ALL, h5s.ALL, np.ascontiguousarray(arr), mtype=tid)
nan, inf = float("nan"), float("inf")
c64 = np.array(
[1.5 - 2j, complex(0.0, -0.0), complex(inf, nan), complex(-inf, 1e-40)],
dtype=np.complex64,
)
ds("c64", h5t.COMPLEX_IEEE_F32LE, c64)
c128 = np.array(
[[1 + 2j, -3.5 + 4e300j, 0.1 + 0.2j], [5e-324 - 0j, 1j, -1]], dtype=np.complex128
)
ds("c128", h5t.COMPLEX_IEEE_F64LE, c128)
ds("c128be", h5t.COMPLEX_IEEE_F64BE, c128.astype(">c16"))
c128x = c128.copy()
c128x[1, 1] = 1 + 1j
ds("c128x", h5t.COMPLEX_IEEE_F64LE, c128x)
half = np.array([1.0, -2.0, 0.5, 65504.0, 0.0, -0.0], dtype=np.float16)
ds("c32", h5t.COMPLEX_IEEE_F16LE, half, shape=(3,))
ds("scalar", h5t.COMPLEX_IEEE_F64LE, np.array(2.5 - 0.5j), shape=())
ct = h5t.create(h5t.COMPOUND, 24)
ct.insert(b"z", 0, h5t.COMPLEX_IEEE_F64LE)
ct.insert(b"k", 16, h5t.STD_I64LE)
cd = np.array([(1 + 1j, 7), (-2.5 + 0j, -8)], dtype=[("z", "<c16"), ("k", "<i8")])
ds("compound", ct, cd)
at = h5t.array_create(h5t.COMPLEX_IEEE_F32LE, (2,))
ad = np.array([[1 + 2j, 3 + 4j], [-1 - 1j, 0j]], dtype=np.complex64)
ds("array", at, ad, shape=(2,))
a = h5a.create(f.id, b"zattr", h5t.COMPLEX_IEEE_F64LE, h5s.create_simple((2,)))
a.write(np.array([0.5 - 1.5j, 2j]), mtype=h5t.COMPLEX_IEEE_F64LE)
f.close()
@@ -0,0 +1,80 @@
HDF5 "native_complex_hdf5_2.h5" {
GROUP "/" {
ATTRIBUTE "zattr" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): 0.5-1.5i, 0+2i
}
}
DATASET "array" {
DATATYPE H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): [ 1+2i, 3+4i ], [ -1-1i, 0+0i ]
}
}
DATASET "c128" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128be" {
DATATYPE H5T_COMPLEX_IEEE_F64BE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128x" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 1+1i, -1+0i
}
}
DATASET "c32" {
DATATYPE H5T_COMPLEX_IEEE_F16LE
DATASPACE SIMPLE { ( 3 ) / ( 3 ) }
DATA {
(0): 1+-2i, 0.5+65504i, 0+-0i
}
}
DATASET "c64" {
DATATYPE H5T_COMPLEX_IEEE_F32LE
DATASPACE SIMPLE { ( 4 ) / ( 4 ) }
DATA {
(0): 1.5-2i, 0-0i, inf+nani, -inf+9.99995e-41i
}
}
DATASET "compound" {
DATATYPE H5T_COMPOUND {
H5T_COMPLEX_IEEE_F64LE "z";
H5T_STD_I64LE "k";
}
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): {
1+1i,
7
},
(1): {
-2.5+0i,
-8
}
}
}
DATASET "scalar" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SCALAR
DATA {
(0): 2.5-0.5i
}
}
}
}
Binary file not shown.
+19 -5
View File
@@ -1064,12 +1064,26 @@ fn complex_datasets_and_attributes_round_trip() {
let ds = file.dataset(name).unwrap(); let ds = file.dataset(name).unwrap();
assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}"); assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}");
assert_eq!(ds.shape().unwrap(), vec![3]); assert_eq!(ds.shape().unwrap(), vec![3]);
// Both encodings read as the same {r, i} compound. // h5py's encoding is a compound, the native one a complex type.
assert_eq!( let want = if name == "native64" {
ds.dtype().unwrap(), DType::Complex(Box::new(DType::F32))
} else {
DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)]) DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)])
); };
assert_eq!(ds.dtype().unwrap(), want, "{name}");
} }
assert_eq!(
file.dataset("native128").unwrap().dtype().unwrap(),
DType::Complex(Box::new(DType::F64))
);
assert_eq!(
file.dataset("native128").unwrap().raw_datatype().unwrap(),
clawhdf5::make_native_complex_f64_type()
);
assert_eq!(
DType::Complex(Box::new(DType::F64)).to_string(),
"complex<f64>"
);
for name in ["compound128", "native128"] { for name in ["compound128", "native128"] {
same( same(
file.dataset(name).unwrap().read_complex_f64().unwrap(), file.dataset(name).unwrap().read_complex_f64().unwrap(),
@@ -1094,7 +1108,7 @@ fn complex_datasets_and_attributes_round_trip() {
data, data,
}) => { }) => {
assert!(shape.is_empty()); assert!(shape.is_empty());
assert_eq!(datatype.type_size(), 16); assert_eq!(datatype, clawhdf5::make_native_complex_f64_type());
let fields = let fields =
clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap(); clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap();
let part = |k: usize| { let part = |k: usize| {
+171 -15
View File
@@ -15,10 +15,8 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since | | Issue | Kind | Since |
|---|---|---| |---|---|---|
| [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 |
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 | | [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 | | [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | | [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | | [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
@@ -33,6 +31,83 @@ Checked against `main` at `9b5803f` on 2026-09-28.
--- ---
## NetCDF-4: differences from netCDF-C
**Status:** open (documented 2026-09-29): deliberate differences, and the
corpus files that still read differently. `clawhdf5-netcdf4` reports a
file's groups, dimensions and variables — names, order, types, shapes — as
netCDF-C 4.9.3 does (`libhdf5/hdf5open.c`; see the crate README), and
`crates/clawhdf5-netcdf4/tests/corpus_vs_netcdf_c.rs` compares it with
netCDF-C itself (the libnetcdf netCDF4-python bundles, called through
ctypes by `tests/netcdf_c_view.py`) over a corpus. Over the conformance
corpus (tank, 2026-09-29, netCDF4-python 1.7.4: netCDF-C 4.9.3 on libhdf5
1.14.6; `CLAWHDF5_NETCDF_CORPUS=<main checkout>/conformance/.cache/corpus
CLAWHDF5_PYTHON=$PWD/.venv/bin/python cargo test -p clawhdf5-netcdf4 --test
corpus_vs_netcdf_c -- --nocapture`): of 692 files, netCDF-C opens 429;
420 match in groups, dimensions (names, lengths, unlimited, order),
variables (names, order, types, dimensions, shapes) and the values of
every numeric variable of up to 5000 elements; the other 9 are listed
below and in `tests/corpus_known_differences.txt`, which the test holds to.
Deliberate:
- **Floats of other than 4 or 8 bytes** (half, bfloat16, the 4-, 6- and
8-bit floats, `long double`) are `NcType::Float` (up to 4 bytes) or
`Double`, and read as numbers. netCDF-C 4.9.3 on libhdf5 1.14.6 labels
them `NC_STRING`: their native type (libhdf5 1.14.6 has a native
`_Float16`) matches none of its own. 27 corpus variables; the comparison
shows them as netCDF-C does and counts them.
- **An axis netCDF-C leaves without a dimension** gets one by the phony
rule (the group's first dimension of its length, else a new
`phony_dim_<n>`): a `_Netcdf4Coordinates` id or a `DIMENSION_LIST` scale
it cannot find, and an axis without a scale of a variable whose first
axis has one — there netCDF-C 4.9.3 reads uninitialised memory
(`get_attached_info` mallocs the object ids and `dimscale_visitor` never
fills that axis's): in one run it gave a 2-long axis the group's first
phony dimension, 3 long; in another netCDF4-python could not open the
file. No corpus file has such an axis.
- **Types netCDF-C remembers after failing to read them.** netCDF-C adds a
type to its list before reading its members, and keeps it when that
fails, so the second dataset of such a type is a variable. For a
compound, enum or variable-length type it has a class and is shown here
too (`interop_tests::types_netcdf_c_skips_are_not_variables`); a bit
field, time, array or complex type has none, and netCDF-C lists the
second dataset with an invalid type (class 0): it stays hidden here.
- **Files netCDF-C refuses** are read as well as they can be: a
multi-dimensional dimension scale without `_Netcdf4Coordinates`, a
`_Netcdf4Coordinates` of the wrong length, a named datatype netCDF-C
cannot represent, an unreadable dataset or attribute (netCDF-C fails
`nc_open`; this crate skips it).
- **Cost:** the first call that needs the metadata reads that of the whole
file (every group, and every dataset's attributes), as `nc_open` does;
it was one group's per call. Not measured; worth measuring on a file
with many groups and variables before a release.
- A whole-variable read of a variable shorter than an unlimited dimension
that is not its first: see [the entry of #27](#netcdf-4-variables-dimensions-are-guessed-from-sizes).
Corpus files that differ (2026-09-29):
- **External links** (`hdf5/test/testfiles/be_extlink1.h5`,
`le_extlink1.h5`, `hdf5/tools/test/testfiles/h5diff_ext2softlink_src.h5`,
`h5diff_grp_recurse_ext2-1.h5`, `h5diff_grp_recurse_ext2-2.h5`): libhdf5
follows them into the other file; clawhdf5 does not
([External links](#external-links-and-external-raw-data-are-not-followed)),
so the linked objects are missing and the phony dimension numbers after
them shift.
- **Values the HDF5 reader refuses** and libhdf5 1.14.6 returns (metadata
matches): `cve_hdf5/cvefiles/cve-2025-2308.h5`, `cve-2025-44904.h5` and
`hdf5/test/testfiles/bad_nbit_parms_walk.h5`, the scale-offset and N-Bit
files of CONFORMANCE.md's ref-bug/our-error list (libhdf5 reads past the
stored data); and `hdf5/test/testfiles/bad_nbit_decompress.h5`, whose
`/Nbit_float_data_le` clawhdf5 refuses ("nbit: element count exceeds
chunk size", also `h5rs dump`) while libhdf5 1.14.6 and h5py 3.16
(libhdf5 2.0.0) return values. That file is not in the conformance
report and has not been investigated.
- Not compared: the values of 18 variables whose filter (SZIP, Blosc,
bzip2) this crate's default build of the HDF5 reader lacks. Blosc and
bzip2 are pure-Rust cargo features of `clawhdf5` (`blosc`, `bzip2`) that
a dependent can turn on; SZIP needs libaec.
## In-place modification (`FileEditor`) limits ## In-place modification (`FileEditor`) limits
**Status:** open (documented 2026-09-26, updated when version-2 B-tree **Status:** open (documented 2026-09-26, updated when version-2 B-tree
@@ -240,15 +315,12 @@ wrong data.
decoded, and external references are an error (object references decoded, and external references are an error (object references
decode). A multi-dimensional numeric attribute is returned as a flat decode). A multi-dimensional numeric attribute is returned as a flat
array (its shape is not reported; `AttrValue::Raw` carries the shape). array (its shape is not reported; `AttrValue::Raw` carries the shape).
HDF5 2.0's native complex type (class 11) is read as the equivalent HDF5 2.0's native complex type (class 11) is read as its own type since
`{r, i}` compound: values are right (`Dataset::read_complex_f64`, and 2026-09-29 ([fixed](#hdf5-20-native-complex-numbers-read-as-a-r-i-compound));
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs writing it is opt-in (`with_native_complex_f64_data`,
ls`/`dump` and the browser reader show a compound where h5dump 2.x `make_native_complex_f64_type`), and the Python bindings write complex
prints `H5T_COMPLEX_IEEE_F64LE`, so `h5rs dump` of such a file does not arrays as h5py's compound only. `h5rs dump --json` writes it as the
match h5dump 2.x (h5dump 1.14 cannot read it at all). Writing class 11 `{r, i}` compound: hdf5-json (h5json 2.0.0) has no complex class.
(added 2026-09-28) is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`); the Python bindings write complex
arrays as h5py's compound only.
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in - **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
that: libhdf5 fails only the first metadata read of an image it cannot that: libhdf5 fails only the first metadata read of an image it cannot
load and then reads the file's own (possibly stale) metadata, where we load and then reads the file's own (possibly stale) metadata, where we
@@ -458,8 +530,10 @@ browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27.
again retries it; what was fetched stays cached). again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against - Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included; not in Firefox or Safari. a local server, cross-origin included; not in Firefox or Safari.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are - Compound (h5py's complex numbers, a compound `{r, i}`, included),
refused with an error naming the type; attributes of those types come back reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type (HDF5 2.0 native complex datasets
read, as `[re, im]` pairs); attributes of those types come back
as `value: null` with their `dtype`. (VL strings read, with h5py's as `value: null` with their `dtype`. (VL strings read, with h5py's
values, through the same `VlResolver` as `File` and `h5rs`.) values, through the same `VlResolver` as `File` and `h5rs`.)
- No Zstd or SZIP (both link C): such datasets fail with - No Zstd or SZIP (both link C): such datasets fail with
@@ -539,6 +613,87 @@ Newest first. "Before any release" means no tagged release (v2.7.0 and
earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date
given. given.
## NetCDF-4: files without dimension scales, types and order differed from netCDF-C
**Status:** fixed 2026-09-29 (branch `fix/netcdf-phony-dims`). Affected
every release (v2.1.0 to v2.7.0) and `main` after #27. Wrong metadata only
(names, order and types of dimensions and variables); values were read
right. Users: the dimensions of a file without dimension scales are now
netCDF-C's `phony_dim_<n>`, so code that matched `dim_<size>` or took a
1-D dataset's name for a dimension must use the new names; datasets of
types netCDF-C skips (references, bit fields, array types, compounds of
those, ...) are no longer variables; lists come in netCDF-C's order.
`NcType` has four new variants (`Enum`, `Compound`, `VLen`, `Opaque`) and
is `#[non_exhaustive]`: an exhaustive `match` needs a wildcard arm.
Found 2026-09-28 by the one-off corpus comparison of #27 (52 of the 78
files netCDF4-python opened matched; netCDF4-python itself hides variables
of opaque and other types it does not support, so the comparison now asks
netCDF-C through its C API). Under that comparison `main` at `4260af4`
matched 68 of the 429 corpus files netCDF-C opens (tank, 2026-09-29, the
command above run on a copy of `4260af4` with the new test files). What
differed:
- **Files without dimension scales.** netCDF-C gives every axis a phony
dimension `phony_dim_<id>` (`create_phony_dims`), ids file-wide after the
real dimensions, subgroups numbered before their parent's variables,
shared by length and unlimitedness within a group but not between two
axes of one variable, a length of 0 always unlimited. The crate reported
one dimension per 1-D dataset, named after it (an h5py file with `x(5)`
and `v(5)` had dimensions `x` and `v`, and `v` on `x`), and other axes as
`dim_<size>`.
- **Datasets netCDF-C skips** (`NC_EBADTYPID`: references, bit fields,
time and array types, a compound with such a member or a half-float
member, an enum or VLEN over one) were variables, with `NcType::Char`;
so were enums, compounds, VLENs and opaque types, and 1-byte strings
were `String` where netCDF-C says `NC_CHAR` (longer ones `NC_STRING`).
- **Order.** Groups and variables came in the order of the HDF5 links as
stored (hash order in a dense group); netCDF-C uses creation order when
the group tracks it, else name order. Dimensions without
`_Netcdf4Dimid` came after the others instead of taking the next id in
reading order; a zero-length dimension scale was not unlimited.
- `_Netcdf4Coordinates` ids were looked up only in the variable's group
and its parents (netCDF-C: file-wide), and an unlimited dimension's
length came from its scale's `REFERENCE_LIST` in any group
(`nc4_find_dim_len`: the variables in its group and below).
The crate now reads the metadata of the whole file in netCDF-C's two
passes (`crates/clawhdf5-netcdf4/src/model.rs`), with a new
`clawhdf5_format::group_v2::links_in_creation_order_in` for the order.
Tests: `interop_tests::phony_dimensions_match_netcdf_c`,
`types_netcdf_c_skips_are_not_variables`,
`group_and_variable_order_match_netcdf_c` and
`netcdf4_python_files_match_netcdf_c` compare h5py- and
netCDF4-python-written files with netCDF-C itself, and
`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in
[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)).
## HDF5 2.0 native complex numbers read as a `{r, i}` compound
**Status:** fixed 2026-09-29 (branch `feat/complex-first-class`; PR not
yet opened). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0)
and `main` until then. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
A class-11 datatype (`H5T_COMPLEX_IEEE_F64LE`, ...) parsed into the
equivalent compound `{r, i}`. Values were right (`read_complex_f64`,
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs ls`/`dump`
and the browser reader showed a compound, so `h5rs dump` did not match
h5dump 2.x, the browser refused the dataset, and a parsed type written
back out became a compound. `Datatype::parse` now returns
`Datatype::Complex` (in compounds, arrays and VL types too),
`Dataset::dtype()` is `DType::Complex(..)`, `h5rs` prints what h5dump,
h5ls and h5diff 2.2.0 print (checked on tank, 2026-09-29, against a file
written by h5py 3.16 / libhdf5 2.0.0:
`cargo test -p clawhdf5-tools --test native_complex_dump`), and the
browser reads `[re, im]` pairs.
**What users must do:** code that matched a native complex type as a
`Datatype::Compound` (or a `DType::Compound` of `r`/`i`) must match
`Datatype::Complex` / `DType::Complex`, or take the compound view with
`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a
`Complex` arm. Files are unchanged; nothing needs rewriting.
## NetCDF-4: variables' dimensions are guessed from sizes ## NetCDF-4: variables' dimensions are guessed from sizes
@@ -590,7 +745,8 @@ dimension scales, and h5netcdf 1.8.1 and xarray files (tank, 2026-09-28,
`CLAWHDF5_PYTHON=<venv with h5netcdf> cargo test -p clawhdf5-netcdf4`). `CLAWHDF5_PYTHON=<venv with h5netcdf> cargo test -p clawhdf5-netcdf4`).
Files without dimension scales still get dimensions by size (netCDF-C Files without dimension scales still get dimensions by size (netCDF-C
gives them `phony_dim_<n>`), as before; see `CHANGELOG.md` for a gives them `phony_dim_<n>`), as before; see `CHANGELOG.md` for a
comparison over the conformance corpus's netCDF-readable files. comparison over the conformance corpus's netCDF-readable files. (Fixed
2026-09-29: see the entry above.)
## HDF5 1.8 could not read the files we wrote ## HDF5 1.8 could not read the files we wrote
+6 -2
View File
@@ -123,8 +123,12 @@ else throws an `Error` naming the type.
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32). `readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in Every response body is cut off past the length asked for. More in
`docs/known-issues.md`. `docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are - HDF5 2.0 native complex datasets read as `[re, im]` pairs: `dtype`
refused with an error. Attributes of those types are listed with `complex<f64>`, `elementShape` ending in `2`, the parts interleaved in
the typed array.
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque and variable-length-sequence datasets are refused with
an error. Attributes of those types are listed with
`value: null` and their `dtype`. `value: null` and their `dtype`.
- No Zstd or SZIP filters (they link C): such a dataset fails with - No Zstd or SZIP filters (they link C): such a dataset fails with
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and `unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
+6
View File
@@ -93,6 +93,12 @@ expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>' expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
# 3-D: leading dimension held at 0, window over the last two. # 3-D: leading dimension held at 0, window over the last two.
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0' expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
# HDF5 2.0 native complex (written when h5py's libhdf5 is 2.0 or later):
# each cell is the [re, im] pair.
if "$PY" -c 'import sys, h5py; sys.exit(not getattr(h5py.get_config(), "has_native_complex", False))'; then
expect $H5 "/native_c128" '<dd>complex&lt;f64&gt;</dd>' '<dd>(3, 2)</dd>' '<td>[1, -1.5]</td>' \
'<td>[5, -7.5]</td>'
fi
# Unsupported type: an error, not values. # Unsupported type: an error, not values.
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported' expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that # Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
+22 -3
View File
@@ -80,6 +80,15 @@ with h5py.File(h5, "w") as f:
**hdf5plugin.Zstd()) **hdf5plugin.Zstd())
comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")]) comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")])
f.create_dataset("table", data=comp) f.create_dataset("table", data=comp)
if getattr(h5py.get_config(), "has_native_complex", False):
# HDF5 2.0 native complex (class 11): read as [re, im] pairs.
from h5py import h5d, h5s, h5t
for name, t, dt in [(b"native_c128", h5t.COMPLEX_IEEE_F64LE, "<c16"),
(b"native_c64_be", h5t.COMPLEX_IEEE_F32BE, ">c8")]:
z = (np.arange(6) - 1.5j * np.arange(6)).reshape(3, 2).astype(dt)
d = h5d.create(f.id, name, t, h5s.create_simple(z.shape))
d.write(h5s.ALL, h5s.ALL, z, mtype=t)
g = f.create_group("sensors") g = f.create_group("sensors")
g.attrs["location"] = "lab" g.attrs["location"] = "lab"
g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4")) g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4"))
@@ -105,6 +114,8 @@ def kind(dt):
return "strings" return "strings"
if dt.subdtype is not None: if dt.subdtype is not None:
return kind(dt.subdtype[0]) return kind(dt.subdtype[0])
if dt.kind == "c":
return "f64" if dt.itemsize == 16 else "f32"
if dt.kind == "f": if dt.kind == "f":
return "f64" if dt.itemsize == 8 else "f32" return "f64" if dt.itemsize == 8 else "f32"
if dt.kind in "iu": if dt.kind in "iu":
@@ -120,20 +131,25 @@ def flat(a, k):
return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()] return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()]
if k.startswith(("i", "u")): if k.startswith(("i", "u")):
return [str(int(x)) for x in a.ravel()] return [str(int(x)) for x in a.ravel()]
if a.dtype.kind == "c":
# [re, im] pairs, in order.
a = np.stack([a.real, a.imag], axis=-1)
return [float(x) for x in a.ravel()] return [float(x) for x in a.ravel()]
def entry(ds, slab=None): def entry(ds, slab=None):
k = kind(ds.dtype) k = kind(ds.dtype)
data = ds[()] data = ds[()]
e = {"kind": k, "shape": list(np.shape(data)), "values": flat(data, k)} # A complex element reads as its two parts: one more dimension.
pair = [2] if ds.dtype.kind == "c" else []
e = {"kind": k, "shape": list(np.shape(data)) + pair, "values": flat(data, k)}
if slab: if slab:
start, count, stride = slab start, count, stride = slab
idx = tuple(slice(s, s + (c - 1) * st + 1, st) idx = tuple(slice(s, s + (c - 1) * st + 1, st)
for s, c, st in zip(start, count, stride)) for s, c, st in zip(start, count, stride))
part = ds[idx] part = ds[idx]
e["slab"] = {"start": start, "count": count, "stride": stride, e["slab"] = {"start": start, "count": count, "stride": stride,
"shape": list(part.shape), "values": flat(part, k)} "shape": list(part.shape) + pair, "values": flat(part, k)}
return e return e
@@ -176,7 +192,10 @@ def describe(path):
"datasets": sorted(sets)} "datasets": sorted(sets)}
for n, o in members.items(): for n, o in members.items():
walk(key.rstrip("/") + "/" + n, o) walk(key.rstrip("/") + "/" + n, o)
elif obj.dtype.names: elif obj.dtype.names or (
obj.dtype.kind == "c"
and obj.id.get_type().get_class() != getattr(h5py.h5t, "COMPLEX", None)):
# h5py's own complex numbers are a compound {r, i}: refused.
expected["errors"][key] = "compound" expected["errors"][key] = "compound"
else: else:
expected["datasets"][key] = entry(obj, slab_for(obj)) expected["datasets"][key] = entry(obj, slab_for(obj))
+2 -2
View File
@@ -399,7 +399,7 @@ async function remoteTests() {
await fails(async () => { await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512, blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })), fetch: tamper(async (r) => withHeaders(r, { "Content-Range": `bytes 0-511/${statSync(join(fixDir, "fixture.h5")).size}` })),
}); });
await f.read("/grid"); await f.read("/grid");
}, /the server sent 0-511/, "wrong range"); }, /the server sent 0-511/, "wrong range");
@@ -550,7 +550,7 @@ async function floodTests() {
const range = new Headers(init.headers).get("Range"); const range = new Headers(init.headers).get("Range");
if (range === "bytes=0-511") return fetch(url, init); if (range === "bytes=0-511") return fetch(url, init);
const m = /^bytes=(\d+)-(\d+)$/.exec(range); const m = /^bytes=(\d+)-(\d+)$/.exec(range);
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } }); return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/${statSync(join(fixDir, "fixture.h5")).size}` } });
}, },
}).catch((e) => e); }).catch((e) => e);
// The open itself may need a second range: flooded either way. // The open itself may need a second range: flooded either way.