Chunk dimensions of 2^32 or more; copy-free unfiltered 4 GiB writes #32

Open
osobh wants to merge 7 commits from feat/huge-chunk-dims into main
22 changed files with 988 additions and 81 deletions
Showing only changes of commit e4ba09946f - Show all commits
+63
View File
@@ -56,6 +56,69 @@
and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4).
Affected v2.1.0 to v2.7.0.
### HDF5 2.0 native complex is its own type on read; Python `libver=` (2026-09-29)
- **`Datatype::parse` returns `Datatype::Complex { size, base_type }`** for
a class-11 message, also as a compound member, array base or
variable-length base. It used to return the equivalent `{r, i}`
compound (`docs/known-issues.md`, "HDF5 2.0 native complex numbers read
as a `{r, i}` compound"). **Breaking** for code that matched the
compound view of a native complex type: `Dataset::raw_datatype()` and
attribute datatypes are now `Datatype::Complex`;
`Datatype::complex_as_compound(size, base)` still gives the `{r, i}` view,
and `data_read::read_compound_fields` accepts a `Complex` directly.
Re-serializing a parsed type (e.g. copying it to another file) now writes
class 11 again instead of a compound.
- **Facade:** `DType::Complex(Box<DType>)` (new variant, **breaking** for
exhaustive `match`es over `DType`): `Complex(F32)` is numpy `complex64`,
`Complex(F64)` `complex128`, a binary16 part `Complex(Other("float16"))`
(the same `Other` name small floats already use); `Display` prints
`complex<f64>`. h5py's own complex encoding, the compound `{r, i}`, is
still `DType::Compound`. `read_complex_f64`/`read_complex_f32` read both.
- **`h5rs`:** `dump` prints native complex as h5dump 2.2.0 does:
`H5T_COMPLEX_IEEE_F{16,32,64}{LE,BE}` (else `H5T_COMPLEX { <base> }`), in
arrays and compounds too, and values as `1.5-2i` (`%g%+gi`; for binary16
parts, which h5dump has no C type for, `1+-2i` as it prints them). `ls`
shows `complex64`, `complex128-be`, `complex32` (and `complex<part>` for
non-IEEE parts); `ls -v` shows h5ls 2.2.0's `complex number of` /
`IEEE 64-bit big-endian float`. `diff` compares complex values part by
part and prints the difference as h5diff 2.2.0 does (`1+0i`); a native
complex and an `{r, i}` compound are no longer comparable (different
classes, as in h5diff). `dump --json` keeps the `{r, i}` compound and
`[re, im]` values: hdf5-json (h5json 2.0.0) has no complex class.
- **Browser (`clawhdf5-wasm`):** native complex datasets are readable, as
`[re, im]` pairs: `info().elementShape` ends in `2`, `read()` returns the
parts interleaved in a `Float32Array`/`Float64Array` (binary16 parts
widened to f32), and `dtype` is `complex<f64>` (`complex<f32
(big-endian)>`, `array[2]<complex<f32>>`). The viewer shows each element
as `[re, im]` with no change to the page. h5py's `{r, i}` compound is
still refused, as every compound is.
- **Python:** `clawhdf5.File(path, 'w', libver=...)` takes h5py's values:
`'v108'`, `'v110'`, `'v112'`, `'v114'`, `'v200'`, `'latest'` (the low
bound; high `'latest'`, as h5py) or a `(low, high)` tuple, mapped to
`FileBuilder::libver_bounds`. `'earliest'` as the low bound writes the
1.8 format with a `UserWarning` (clawhdf5 cannot write the pre-1.8
format); as the high bound it is a `ValueError`, as are unknown names and
a low bound above the high one. It is ignored for `'r'` and refused
(`NotImplementedError`) for `'r+'`/`'a'`. Native complex already read as
numpy `complex64`/`complex128` and still does
(`test_read_native_complex_from_h5py`).
- Tests (tank, 2026-09-29): `crates/clawhdf5-tools/tests/native_complex_dump.rs`
compares `h5rs dump` with h5dump 2.2.0's output stored next to a fixture
written by h5py 3.16 / libhdf5 2.0.0
(`crates/clawhdf5/tests/fixtures/gen_native_complex.py`:
`F32LE`/`F64LE`/`F64BE`/`F16LE`, scalar, compound member, array, root
attribute), line for line outside the values and value for value within
them, and `ls`/`diff` with h5ls/h5diff 2.2.0;
`integration_tests::complex_datasets_and_attributes_round_trip`
(`DType::Complex`, `raw_datatype()`); `clawhdf5-wasm` unit tests on the
fixture, `h5py_interop` and `examples/wasm-viewer/test` (Node:
`test.mjs`, 274 + 1324 checks; the page in headless Chromium for
`/native_c128`, `/native_c64_be`, `/pairs`) with two native complex
datasets added to `make_fixture.py` (when h5py's libhdf5 is 2.0+);
`crates/clawhdf5-py/tests/test_libver.py` (every bound; `'v108'` output
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
hard-codes the fixture's length.
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
+2 -2
View File
@@ -115,12 +115,12 @@ Limits and open issues, with dates, are in
|---|---|---|---|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays with h5py's `libver=`, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`; HDF5 2.0 native complex as `[re, im]` pairs); no Zstd/SZIP/pcodec, no compound (h5py's `{r, i}` complex included), reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`,
`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure
+35 -29
View File
@@ -145,13 +145,16 @@ pub enum Datatype {
/// imaginary, in rectangular form. `size` is twice the base size and the
/// base is an IEEE float.
///
/// This variant exists for **writing** (see
/// [`Datatype::parse`] returns this variant for every class-11 message,
/// also inside compounds, arrays and variable-length types. Readers that
/// want the member view (`{r, i}`, as h5py writes complex numbers by
/// default) use [`Datatype::complex_as_compound`]; `read_compound_fields`
/// accepts this variant directly.
///
/// Writing it is opt-in (see
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
/// newer can read class 11, so it is opt-in and h5py's compound `{r, i}`
/// stays the default complex encoding. [`Datatype::parse`] still
/// surfaces a class-11 message as that equivalent `{r, i}` compound, so
/// every compound reader handles both encodings; parsing what this
/// variant serializes therefore yields a `Compound`, not a `Complex`.
/// newer can read class 11, so h5py's compound `{r, i}` stays the
/// default complex encoding.
Complex { size: u32, base_type: Box<Datatype> },
}
@@ -857,9 +860,7 @@ impl Datatype {
// Complex number (HDF5 2.0, datatype version 5). The properties
// are a single base floating-point datatype message; an element
// is two consecutive base-type values (real, imaginary). There
// is no member list. Surface it as the equivalent two-member
// compound `{r, i}` — the same shape h5py writes for numpy
// complex dtypes — so downstream compound readers work as-is.
// is no member list.
if version != 5 {
return Err(FormatError::InvalidDatatypeVersion {
class: class_id,
@@ -875,7 +876,13 @@ impl Datatype {
actual: size as usize,
});
}
Ok((Self::complex_as_compound(size, &base_type), pos))
Ok((
Datatype::Complex {
size,
base_type: Box::new(base_type),
},
pos,
))
}
_ => Err(FormatError::InvalidDatatypeClass(class_id)),
}
@@ -1174,9 +1181,8 @@ impl Datatype {
/// The `{r, i}` compound equivalent to a native complex type of `size`
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the
/// shape h5py writes for numpy complex dtypes, and what [`Self::parse`]
/// returns for a class-11 message. Readers that meet a
/// [`Datatype::Complex`] handle it through this view.
/// shape h5py writes for numpy complex dtypes. Readers that want a
/// [`Datatype::Complex`] as members handle it through this view.
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
let base_size = base_type.type_size();
Datatype::Compound {
@@ -1849,21 +1855,24 @@ mod tests {
fn test_complex_v5_from_hdf5_2_0() {
let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap();
assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len());
match dt {
Datatype::Compound { size, members } => {
assert_eq!(size, 16);
assert_eq!(members.len(), 2);
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
for m in &members {
match &dt {
Datatype::Complex { size, base_type } => {
assert_eq!(*size, 16);
assert!(matches!(
m.datatype,
base_type.as_ref(),
Datatype::FloatingPoint { size: 8, .. }
));
}
other => panic!("expected Complex, got {other:?}"),
}
other => panic!("expected Compound, got {other:?}"),
}
// The member view is h5py's `{r, i}` compound.
let Datatype::Compound { members, .. } =
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
else {
unreachable!()
};
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
}
#[test]
@@ -1887,7 +1896,7 @@ mod tests {
assert_eq!(members.len(), 2);
assert!(matches!(
&members[0].datatype,
Datatype::Compound { size: 16, members } if members.len() == 2
Datatype::Complex { size: 16, .. }
));
assert_eq!(
(members[1].name.as_str(), members[1].byte_offset),
@@ -1918,12 +1927,9 @@ mod tests {
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
assert_eq!(dt.type_size(), 16);
dt.check_encodable().unwrap();
// Parsing surfaces class 11 as the equivalent `{r, i}` compound.
// Parsing returns the native complex type unchanged.
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
assert_eq!(
parsed,
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
);
assert_eq!(parsed, dt);
let f32c = make_native_complex_f32_type().serialize();
assert_eq!(
+16 -1
View File
@@ -104,7 +104,22 @@ writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and
`complex128` arrays; the file is written on `close()`. Complex arrays are
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
Rust API writes that on request).
Rust API writes that on request). Both forms read back as numpy
`complex64`/`complex128`.
`libver=` sets the library version bounds as in h5py: `'v108'`,
`'v110'`, `'v112'`, `'v114'`, `'v200'` or `'latest'` (the low bound; the
high bound is then `'latest'`), or a `(low, high)` tuple. With `'v108'`
the file is written in the HDF5 1.8 format (version-2 superblock,
version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the
HDF5 1.10 format. `'earliest'` as the low bound writes the 1.8 format too,
with a `UserWarning`: clawhdf5 cannot write the pre-1.8 format. The
argument is ignored when reading and refused for `'r+'`/`'a'`.
```python
with clawhdf5.File("old.h5", "w", libver="v108") as f: # HDF5 1.8 reads it
f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip")
```
## Editing a file in place
+90 -2
View File
@@ -13,10 +13,13 @@ use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
use clawhdf5_rs::LibVer;
/// Internal state for write mode.
struct WriteState {
path: PathBuf,
/// `libver=` as (low, high); `None` keeps the writer's default.
libver: Option<(LibVer, LibVer)>,
root_datasets: Vec<DatasetSpec>,
root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
groups: Vec<Arc<Mutex<WriteGroupState>>>,
@@ -80,10 +83,32 @@ impl PyFile {
/// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write
/// libver: library version bounds for mode 'w', as h5py's: one of
/// 'earliest', 'v108', 'v110', 'v112', 'v114', 'v200', 'latest' (the
/// low bound; the high bound is then 'latest') or a (low, high)
/// tuple of them. The low bound is the oldest HDF5 release whose
/// format the file uses ('v108': HDF5 1.8 can read it); the high
/// bound the newest whose features it may use. clawhdf5 cannot write
/// the pre-1.8 format, so a low bound of 'earliest' writes the 1.8
/// format (with a warning) and a high bound of 'earliest' is an
/// error. Default (None): the HDF5 1.10 format clawhdf5 has always
/// written. Ignored for reading; not supported with 'r+' / 'a'.
#[new]
#[pyo3(signature = (path, mode="r"))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
#[pyo3(signature = (path, mode="r", libver=None))]
fn new(
py: Python<'_>,
path: &str,
mode: &str,
libver: Option<&Bound<'_, PyAny>>,
) -> PyResult<Self> {
let filename = path.to_string();
let libver = libver.map(|v| parse_libver(py, v)).transpose()?;
if libver.is_some() && matches!(mode, "r+" | "a") {
return Err(PyNotImplementedError::new_err(format!(
"libver with mode '{mode}': clawhdf5's in-place editor keeps the format \
versions the file already uses"
)));
}
if is_url(path) {
if mode != "r" {
return Err(PyValueError::new_err(format!(
@@ -113,6 +138,7 @@ impl PyFile {
// Absolute now: the file is written at close, possibly
// after the working directory changed.
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
libver,
root_datasets: Vec::new(),
root_attrs: Arc::new(Mutex::new(Vec::new())),
groups: Vec::new(),
@@ -460,10 +486,71 @@ fn parse_compression(
}
}
/// One of h5py's `libver` names as a bound; `high` says which end it is.
/// `Ok(None)` is 'earliest' as the low bound: the pre-1.8 format, which
/// clawhdf5 cannot write.
fn libver_name(name: &str, high: bool) -> PyResult<Option<LibVer>> {
Ok(Some(match name {
"earliest" if high => {
return Err(PyValueError::new_err(
"libver high bound 'earliest' (the pre-1.8 format) cannot be written by \
clawhdf5; the oldest format it writes is 'v108'",
));
}
"earliest" => return Ok(None),
"v108" => LibVer::V18,
"v110" => LibVer::V110,
"v112" => LibVer::V112,
"v114" => LibVer::V114,
"v200" => LibVer::V200,
"latest" => LibVer::Latest,
other => {
return Err(PyValueError::new_err(format!(
"unknown libver '{other}'; expected 'earliest', 'v108', 'v110', 'v112', \
'v114', 'v200' or 'latest'"
)));
}
}))
}
/// h5py's `libver=`: a name (the low bound, high bound 'latest') or a
/// `(low, high)` pair.
fn parse_libver(py: Python<'_>, v: &Bound<'_, PyAny>) -> PyResult<(LibVer, LibVer)> {
let (low, high): (String, String) = match v.extract::<String>() {
Ok(name) => (name, "latest".into()),
Err(_) => v.extract().map_err(|_| {
PyValueError::new_err("libver must be a string or a (low, high) tuple of strings")
})?,
};
let Some(high) = libver_name(&high, true)? else {
unreachable!("a high bound is never None")
};
let low = match libver_name(&low, false)? {
Some(low) => low,
None => {
PyModule::import(py, "warnings")?.getattr("warn")?.call1((
"libver 'earliest': clawhdf5 cannot write the pre-1.8 format; the file \
is written in the HDF5 1.8 format ('v108') instead",
py.get_type::<pyo3::exceptions::PyUserWarning>(),
))?;
LibVer::V18
}
};
if low > high {
return Err(PyValueError::new_err(format!(
"libver low bound {low} is newer than the high bound {high}"
)));
}
Ok((low, high))
}
/// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| {
let mut builder = clawhdf5_rs::FileBuilder::new();
if let Some((low, high)) = state.libver {
builder.libver_bounds(low, high);
}
// Root attributes
let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner());
@@ -520,6 +607,7 @@ mod tests {
let state = WriteState {
path: path.clone(),
libver: None,
root_datasets: vec![DatasetSpec {
name: "data".into(),
data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]),
+143
View File
@@ -0,0 +1,143 @@
"""`clawhdf5.File(path, 'w', libver=...)`: h5py's library version bounds.
A file written with libver='v108' must open in HDF5 1.8. Its h5dump is
found through CLAWHDF5_H5DUMP18 or at ~/.cache/hdf5-1.8.23/bin/h5dump
(scripts/build-hdf5-1.8.sh builds it); without it that check is skipped.
"""
import os
import subprocess
import numpy as np
import pytest
import clawhdf5
def superblock_version(path):
with open(path, "rb") as f:
head = f.read(9)
assert head[:8] == b"\x89HDF\r\n\x1a\n"
return head[8]
def h5dump18():
path = os.environ.get("CLAWHDF5_H5DUMP18") or os.path.expanduser(
"~/.cache/hdf5-1.8.23/bin/h5dump"
)
try:
out = subprocess.run([path, "--version"], capture_output=True, text=True)
except OSError:
return None
return path if "1.8." in out.stdout else None
def write_sample(path, libver):
with clawhdf5.File(path, "w", libver=libver) as f:
f.create_dataset("x", data=np.arange(10, dtype=np.float64))
f.create_dataset(
"chunked",
data=np.arange(100, dtype=np.int32),
chunks=(30,),
compression="gzip",
)
f.create_dataset("z", data=np.array([1 + 2j, -3j], dtype=np.complex128))
g = f.create_group("g")
g.create_dataset("y", data=np.ones(3, dtype=np.float32))
f.attrs["version"] = 1
@pytest.mark.parametrize(
"libver, sb",
[
(None, 3),
("v108", 2),
(("v108", "v108"), 2),
(("v108", "latest"), 2),
("v110", 3),
("v112", 3),
("v114", 3),
("v200", 3),
("latest", 3),
(("v110", "v200"), 3),
],
)
def test_libver_bounds(h5py, tmp_path, libver, sb):
path = str(tmp_path / "libver.h5")
write_sample(path, libver)
# v108 low bound: HDF5 1.8's version-2 superblock; 1.10 and later: 3.
assert superblock_version(path) == sb
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
np.testing.assert_array_equal(f["z"][:], [1 + 2j, -3j])
np.testing.assert_array_equal(f["g/y"][:], np.ones(3))
assert f.attrs["version"] == 1
with clawhdf5.File(path, "r") as f:
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
def test_libver_v108_opens_in_hdf5_1_8(tmp_path):
h5dump = h5dump18()
if h5dump is None:
pytest.skip("no HDF5 1.8 h5dump (CLAWHDF5_H5DUMP18)")
path = str(tmp_path / "v108.h5")
write_sample(path, "v108")
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode == 0, out.stderr
assert "h5dump error" not in out.stderr
assert 'DATASET "y"' in out.stdout
# The chunked, deflated dataset (a version-1 B-tree) reads in full.
out = subprocess.run(
[h5dump, "-w", "0", "-d", "/chunked", path], capture_output=True, text=True
)
assert out.returncode == 0, out.stderr
assert "(0): " + ", ".join(str(k) for k in range(100)) + "\n" in out.stdout
# The default (1.10 format) does not open in 1.8: the bound matters.
path = str(tmp_path / "default.h5")
write_sample(path, None)
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode != 0
def test_libver_earliest_writes_v108_with_a_warning(h5py, tmp_path):
path = str(tmp_path / "earliest.h5")
with pytest.warns(UserWarning, match="pre-1.8"):
write_sample(path, "earliest")
assert superblock_version(path) == 2
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
@pytest.mark.parametrize(
"libver, match",
[
(("v108", "earliest"), "high bound 'earliest'"),
("v109", "unknown libver 'v109'"),
(("latest", "v108"), "newer than the high bound"),
(("v108",), "tuple"),
(3, "tuple"),
],
)
def test_libver_errors(tmp_path, libver, match):
with pytest.raises(ValueError, match=match):
clawhdf5.File(str(tmp_path / "bad.h5"), "w", libver=libver)
def test_libver_v108_high_bound_writes_everything(tmp_path):
# Native complex numbers need HDF5 2.0; the Python writer stores complex
# as h5py's {r, i} compound, which 1.8 reads, so the 1.8 high bound is
# fine for everything clawhdf5.File writes.
path = str(tmp_path / "v18only.h5")
write_sample(path, ("v108", "v108"))
assert superblock_version(path) == 2
def test_libver_not_for_editing(tmp_path):
path = str(tmp_path / "e.h5")
write_sample(path, None)
with pytest.raises(NotImplementedError, match="libver"):
clawhdf5.File(path, "r+", libver="latest")
# Reading ignores it, as the bounds only affect what is written.
with clawhdf5.File(path, "r", libver="v108") as f:
assert f["x"].shape == (10,)
+22 -1
View File
@@ -603,7 +603,22 @@ impl Diff {
let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) {
(Some(x), Some(y), ..) => x.abs_diff(y).to_string(),
(_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64),
// Each part's absolute difference, as h5diff 2.x prints
// it (`1+0i`).
_ => match (&va, &vb) {
(
Value::Complex { re: ar, im: ai, .. },
Value::Complex { re: br, im: bi, .. },
) => match (number(ar), number(ai), number(br), number(bi)) {
(Some(ar), Some(ai), Some(br), Some(bi)) => format!(
"{}+{}i",
value::fmt_float((ar - br).abs(), 64),
value::fmt_float((ai - bi).abs(), 64)
),
_ => String::new(),
},
_ => String::new(),
},
};
rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}"));
}
@@ -713,6 +728,9 @@ impl Diff {
(Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => {
p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v))
}
(Value::Complex { re: pr, im: pi, .. }, Value::Complex { re: qr, im: qi, .. }) => {
self.equal(a, b, pr, qr) && self.equal(a, b, pi, qi)
}
(Value::Ref(None), Value::Ref(None)) => true,
(Value::Ref(Some(p)), Value::Ref(Some(q))) => {
// Addresses mean nothing across files: compare the paths the
@@ -846,7 +864,10 @@ fn types_comparable(a: &Datatype, b: &Datatype) -> bool {
(
Datatype::Enumeration { base_type: x, .. },
Datatype::Enumeration { base_type: y, .. },
) => types_comparable(x, y),
)
| (Datatype::Complex { base_type: x, .. }, Datatype::Complex { base_type: y, .. }) => {
types_comparable(x, y)
}
_ => true,
}
}
+34 -6
View File
@@ -135,9 +135,16 @@ pub fn short(dt: &Datatype) -> String {
return f.short.into();
}
match dt {
Datatype::Complex { size, base_type } => {
short(&Datatype::complex_as_compound(*size, base_type))
}
// numpy's names (`complex64` is two `float32`) for IEEE parts,
// `complex<part>` otherwise.
Datatype::Complex { size, base_type } => match base_type.as_ref() {
Datatype::FloatingPoint { byte_order, .. } if is_ieee(base_type) => format!(
"complex{}{}",
u64::from(*size) * 8,
if be(byte_order) { "-be" } else { "" }
),
_ => format!("complex<{}>", short(base_type)),
},
Datatype::FixedPoint {
size,
signed,
@@ -216,8 +223,9 @@ pub fn long(dt: &Datatype) -> String {
return f.long.into();
}
match dt {
Datatype::Complex { size, base_type } => {
long(&Datatype::complex_as_compound(*size, base_type))
// As h5ls 2.x prints a complex type it has no native name for.
Datatype::Complex { base_type, .. } => {
format!("complex number of\n {}", long(base_type))
}
Datatype::FixedPoint {
size,
@@ -361,6 +369,19 @@ fn atomic_ddl(dt: &Datatype) -> Option<String> {
)),
// As h5dump 2.x names them (checked against h5dump 2.2.0).
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
// h5dump 2.x's predefined complex names: IEEE binary16/32/64 parts.
Datatype::Complex { base_type, .. } => match base_type.as_ref() {
Datatype::FloatingPoint {
size, byte_order, ..
} if is_ieee(base_type) && !matches!(byte_order, DatatypeByteOrder::Vax) => {
Some(format!(
"H5T_COMPLEX_IEEE_F{}{}",
u64::from(*size) * 8,
order_suffix(byte_order)
))
}
_ => None,
},
Datatype::BitField {
size, byte_order, ..
} => Some(format!(
@@ -418,6 +439,10 @@ pub fn ddl(dt: &Datatype, ind: usize) -> String {
Datatype::VariableLength { base_type, .. } => {
format!("H5T_VLEN {{ {} }}", ddl(base_type, ind))
}
// Parts h5dump has no predefined complex name for.
Datatype::Complex { base_type, .. } => {
format!("H5T_COMPLEX {{ {} }}", ddl(base_type, ind))
}
Datatype::Opaque { size, tag } => {
let tag = String::from_utf8_lossy(tag);
format!(
@@ -498,6 +523,9 @@ fn string_ddl(
/// hdf5-json type object.
pub fn json(dt: &Datatype) -> J {
match dt {
// hdf5-json (h5json 2.0.0) has no complex class: a native complex
// is written as the `{r, i}` compound h5py uses for complex numbers,
// its values as `[re, im]` pairs.
Datatype::Complex { size, base_type } => {
json(&Datatype::complex_as_compound(*size, base_type))
}
@@ -598,7 +626,7 @@ fn pad_json(p: &StringPadding) -> &'static str {
/// Class name used to decide whether two datatypes can be compared.
pub fn class(dt: &Datatype) -> &'static str {
match dt {
Datatype::Complex { .. } => "compound",
Datatype::Complex { .. } => "complex",
Datatype::FixedPoint { .. } => "integer",
Datatype::FloatingPoint { .. } => "float",
Datatype::Time { .. } => "time",
+31 -5
View File
@@ -27,6 +27,15 @@ pub enum Value {
/// An enum member (name, when the value matches one) and its value.
Enum(Option<String>, i128),
Compound(Vec<(String, Value)>),
/// A native complex number (HDF5 2.0 class 11): real and imaginary
/// part. `sign_free` is set when h5dump joins the parts with a bare `+`
/// (parts it has no native C type for: binary16, non-IEEE), giving
/// `1+-2i`; otherwise the imaginary part carries its sign (`1-2i`).
Complex {
re: Box<Value>,
im: Box<Value>,
sign_free: bool,
},
Array(Vec<Value>),
/// A variable-length sequence.
Seq(Vec<Value>),
@@ -183,11 +192,18 @@ impl<'a> Decoder<'a> {
return Value::Error("short element".into());
};
match dt {
Datatype::Complex { size, base_type } => self.decode(
&Datatype::complex_as_compound(*size, base_type),
b,
depth + 1,
),
Datatype::Complex { base_type, .. } => {
let part = base_type.type_size() as usize;
let (re, im) = b.split_at(part.min(b.len()));
Value::Complex {
re: Box::new(self.decode(base_type, re, depth + 1)),
im: Box::new(self.decode(base_type, im, depth + 1)),
// h5dump prints `float`/`double` complex (whatever the
// byte order) with `%g%+gi`, anything else as
// `<re>+<im>i`.
sign_free: !(dtype::is_ieee(base_type) && matches!(part, 4 | 8)),
}
}
Datatype::FixedPoint { .. } => match decode_int(dt, b) {
Some(v) => Value::Int(v),
None => Value::Bytes(b.to_vec()),
@@ -354,6 +370,15 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
.collect::<Vec<_>>()
.join(", ")
),
Value::Complex { re, im, sign_free } => {
let (re, im) = (text(re, h5paths), text(im, h5paths));
let plus = if *sign_free || !im.starts_with('-') {
"+"
} else {
""
};
format!("{re}{plus}{im}i")
}
Value::Array(vs) => format!(
"[ {} ]",
vs.iter()
@@ -408,6 +433,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
Value::Complex { re, im, .. } => J::Array(vec![to_json(re, h5paths), to_json(im, h5paths)]),
Value::Array(vs) | Value::Seq(vs) => {
J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect())
}
@@ -0,0 +1,202 @@
//! `h5rs dump` and `h5rs ls` of HDF5 2.0 native complex numbers (datatype
//! class 11), against the output of h5dump 2.2.0 and h5ls 2.2.0 (the Debian
//! h5dump CI installs, 1.14.x, predates the type). The h5dump output is
//! stored next to the fixture; see
//! `crates/clawhdf5/tests/fixtures/gen_native_complex.py`.
use std::path::PathBuf;
use std::process::Command;
fn fixture(name: &str) -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name)
}
fn h5rs(args: &[&str]) -> String {
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(args)
.output()
.unwrap();
assert!(out.status.success(), "h5rs {args:?}: {out:?}");
String::from_utf8(out.stdout).unwrap()
}
/// Split a dump into its lines, with the data lines (`(i): ...`) replaced
/// by their index, and the data values in order.
fn split(ddl: &str) -> (Vec<String>, Vec<String>) {
let (mut frame, mut data) = (Vec::new(), Vec::new());
for line in ddl.lines() {
match line.split_once("):") {
Some((idx, body)) if idx.trim_start().starts_with('(') => {
frame.push(format!("{idx}):"));
// Values are split at `, `; array elements come bracketed.
data.extend(
body.split(", ")
.map(|s| s.trim().trim_matches(|c| c == '[' || c == ']' || c == ','))
.map(str::trim)
.filter(|s| !s.is_empty() && *s != "{")
.map(String::from),
);
}
_ => {
let t = line.trim().trim_end_matches(',');
if t.ends_with('i') || t.parse::<f64>().is_ok() {
// A compound member's value on its own line.
data.push(t.to_string());
} else {
frame.push(line.to_string());
}
}
}
}
(frame, data)
}
/// `a+bi`, `a-bi` or `a+-bi` as its two parts (a real number as one).
fn parts(v: &str) -> Vec<f64> {
let Some(z) = v.strip_suffix('i') else {
return vec![v.parse().unwrap()];
};
let b = z.as_bytes();
let at = (1..b.len())
.find(|&k| (b[k] == b'+' || b[k] == b'-') && !matches!(b[k - 1], b'e' | b'E' | b'+'))
.unwrap_or_else(|| panic!("not a complex value: {v}"));
let (re, im) = z.split_at(at);
let im = im.strip_prefix('+').unwrap_or(im);
vec![re.parse().unwrap(), im.parse().unwrap()]
}
#[test]
fn dump_matches_h5dump_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
let ours = h5rs(&["dump", path.to_str().unwrap()]);
let reference = std::fs::read_to_string(fixture("native_complex_hdf5_2.ddl")).unwrap();
let (mut our_frame, our_data) = split(&ours);
let (mut ref_frame, ref_data) = split(&reference);
// The first line names the file as it was given.
our_frame.remove(0);
ref_frame.remove(0);
// Everything but the values is h5dump's byte for byte: the types print
// as H5T_COMPLEX_IEEE_F64LE, H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }, ...
assert_eq!(our_frame, ref_frame);
assert_eq!(our_data.len(), ref_data.len(), "{our_data:?}\n{ref_data:?}");
assert_eq!(our_data.len(), 36);
// Values: h5dump prints `%g` (6 digits), inf/nan, and the binary16
// parts it has no C type for as `<re>+<im>i` (`1+-2i`); h5rs prints
// the shortest round-trip string and Inf/NaN, joined the same way.
for (i, (o, r)) in our_data.iter().zip(&ref_data).enumerate() {
assert_eq!(o.contains("+-"), r.contains("+-"), "[{i}] {o} vs {r}");
let (op, rp) = (parts(o), parts(r));
assert_eq!(op.len(), rp.len(), "[{i}] {o} vs {r}");
for (x, y) in op.iter().zip(&rp) {
let same = if y.is_nan() || y.is_infinite() {
x.is_nan() == y.is_nan() && (y.is_nan() || x == y)
} else {
*x as f32 == *y as f32 || (x - y).abs() <= 5e-6 * y.abs()
};
assert!(same, "[{i}] h5rs {o}, h5dump {r}");
assert_eq!(
x.is_sign_negative(),
y.is_sign_negative(),
"[{i}] {o} vs {r}"
);
}
}
}
#[test]
fn ls_describes_complex_like_h5ls_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
// h5ls 2.2.0 -v, `Type:` of the types it has no native C name for (it
// prints `native double _Complex` for the others, as it prints `native
// double` where h5rs prints `IEEE 64-bit little-endian float`).
for (name, long) in [
(
"c128be",
"complex number of\n IEEE 64-bit big-endian float",
),
(
"c128",
"complex number of\n IEEE 64-bit little-endian float",
),
(
"c32",
"complex number of\n IEEE 16-bit little-endian float",
),
(
"array",
"[2] complex number of\n IEEE 32-bit little-endian float",
),
] {
let target = format!("{}/{name}", path.display());
let verbose = h5rs(&["ls", "-v", &target]);
assert!(
verbose.contains(&format!(" Type: {long}\n")),
"{name}:\n{verbose}"
);
}
let listing = h5rs(&["ls", path.to_str().unwrap()]);
for (name, short) in [
("c64", "complex64"),
("c128", "complex128"),
("c128be", "complex128-be"),
("c32", "complex32"),
("scalar", "complex128"),
("array", "array[2]<complex64>"),
("compound", "compound{z: complex128, k: int64}"),
] {
assert!(
listing
.lines()
.any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))),
"{name}:\n{listing}"
);
}
}
#[test]
fn json_dump_uses_the_r_i_compound() {
// hdf5-json (h5json 2.0.0) has no complex class: the type is written as
// the `{r, i}` compound h5py uses, the values as `[re, im]` pairs.
let path = fixture("native_complex_hdf5_2.h5");
let j: serde_json::Value =
serde_json::from_str(&h5rs(&["dump", "--json", path.to_str().unwrap()])).unwrap();
let ds = j["datasets"]
.as_object()
.unwrap()
.values()
.find(|d| d["alias"][0] == "/scalar")
.unwrap();
assert_eq!(ds["type"]["class"], "H5T_COMPOUND");
assert_eq!(ds["type"]["fields"][0]["name"], "r");
assert_eq!(ds["type"]["fields"][1]["type"]["base"], "H5T_IEEE_F64LE");
assert_eq!(ds["value"], serde_json::json!([2.5, -0.5]));
}
#[test]
fn diff_compares_complex_values() {
let path = fixture("native_complex_hdf5_2.h5");
let p = path.to_str().unwrap();
// Same values, different byte order: no differences.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", p, p, "/c128", "/c128be"])
.output()
.unwrap();
assert!(out.status.success(), "{out:?}");
// One element differs in its imaginary part: one difference, exit
// status 1, as h5diff 2.2.0 reports it.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", "-r", p, p, "/c128", "/c128x"])
.output()
.unwrap();
assert_eq!(out.status.code(), Some(1), "{out:?}");
let text = String::from_utf8_lossy(&out.stdout);
assert!(text.contains("1 difference(s) found"), "{text}");
// h5diff 2.2.0: `[ 1 1 ] 0+1i 1+1i 1+0i`.
let row = text.lines().find(|l| l.starts_with("[ 1 1 ]")).unwrap();
assert_eq!(
row.split_whitespace().collect::<Vec<_>>(),
["[", "1", "1", "]", "0+1i", "1+1i", "1+0i"]
);
}
+85 -7
View File
@@ -312,11 +312,14 @@ impl Reader {
let base = array_base(dt);
let is_array = !std::ptr::eq(base, dt);
Ok(match base {
// Read as the part type: an array or complex element is its
// parts stored one after another (a complex number's real, then
// imaginary part).
Datatype::FloatingPoint { size, .. } if *size <= 4 => {
Data::F32(data_read::read_as_f32(raw, dt).map_err(err)?)
Data::F32(data_read::read_as_f32(raw, base).map_err(err)?)
}
Datatype::FloatingPoint { .. } => {
Data::F64(data_read::read_as_f64(raw, dt).map_err(err)?)
Data::F64(data_read::read_as_f64(raw, base).map_err(err)?)
}
Datatype::FixedPoint { size, signed, .. } => {
let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err);
@@ -393,10 +396,13 @@ fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T
.collect()
}
/// Innermost element type of (possibly nested) array datatypes.
/// Innermost element type of (possibly nested) array datatypes; for a
/// native complex type, its part type (each element holds two).
fn array_base(dt: &Datatype) -> &Datatype {
match dt {
Datatype::Array { base_type, .. } => array_base(base_type),
Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => {
array_base(base_type)
}
_ => dt,
}
}
@@ -411,6 +417,12 @@ fn element_shape(dt: &Datatype) -> Vec<u64> {
dims.extend(element_shape(base_type));
dims
}
// `[re, im]`: a complex element reads as its two parts.
Datatype::Complex { base_type, .. } => {
let mut dims = vec![2];
dims.extend(element_shape(base_type));
dims
}
_ => Vec::new(),
}
}
@@ -487,9 +499,7 @@ pub fn describe(dt: &Datatype) -> String {
}
}
match dt {
Datatype::Complex { size, base_type } => {
describe(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::Complex { base_type, .. } => format!("complex<{}>", describe(base_type)),
Datatype::FixedPoint {
size,
signed,
@@ -681,6 +691,74 @@ mod tests {
assert!(e.contains("not supported"), "{e}");
}
#[test]
fn native_complex_reads_as_re_im_pairs() {
let z = [[1.0f64, -2.0], [0.5, 0.0], [-0.0, 3.25], [f64::MAX, 1e-300]];
let mut b = FileBuilder::new();
b.create_dataset("z")
.with_native_complex_f64_data(&z)
.with_shape(&[2, 2]);
b.create_dataset("z32")
.with_native_complex_f32_data(&[[1.5f32, -2.5]]);
let r = Reader::open(b.finish().unwrap()).unwrap();
let i = r.info("z").unwrap();
assert_eq!(
(i.shape, i.dtype, i.element_shape),
(vec![2, 2], "complex<f64>".to_string(), vec![2])
);
let all = r.read("z", None).unwrap();
assert_eq!(all.shape, vec![2, 2, 2]);
assert_eq!(all.data, Data::F64(z.iter().flatten().copied().collect()));
let slab = Hyperslab {
start: vec![1, 0],
count: vec![1, 2],
stride: None,
block: None,
};
let part = r.read("z", Some(&slab)).unwrap();
assert_eq!(part.shape, vec![1, 2, 2]);
assert_eq!(part.data, Data::F64(vec![-0.0, 3.25, f64::MAX, 1e-300]));
let z32 = r.read("z32", None).unwrap();
assert_eq!(
(z32.shape, z32.data),
(vec![1, 2], Data::F32(vec![1.5, -2.5]))
);
}
#[test]
fn native_complex_written_by_libhdf5_2() {
// Written by h5py 3.16 / libhdf5 2.0.0 (gen_native_complex.py).
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5/tests/fixtures/native_complex_hdf5_2.h5"
);
let r = Reader::open(std::fs::read(path).unwrap()).unwrap();
let be = r.read("c128be", None).unwrap();
assert_eq!(be.shape, vec![2, 3, 2]);
assert_eq!(be.data, r.read("c128", None).unwrap().data);
assert_eq!(r.info("c128be").unwrap().dtype, "complex<f64 (big-endian)>");
// binary16 parts widen to f32.
assert_eq!(
r.read("c32", None).unwrap().data,
Data::F32(vec![1.0, -2.0, 0.5, 65504.0, 0.0, -0.0])
);
// An array of complex: array dimensions, then the two parts.
let i = r.info("array").unwrap();
assert_eq!(
(i.dtype.as_str(), i.element_shape),
("array[2]<complex<f32>>", vec![2, 2])
);
assert_eq!(
r.read("array", None).unwrap().data,
Data::F32(vec![1.0, 2.0, 3.0, 4.0, -1.0, -1.0, 0.0, 0.0])
);
let s = r.read("scalar", None).unwrap();
assert_eq!((s.shape, s.data), (vec![2], Data::F64(vec![2.5, -0.5])));
// A compound holding a complex member is still refused.
let e = r.read("compound", None).unwrap_err();
assert!(e.contains("compound{z: complex<f64>, k: i64}"), "{e}");
}
#[test]
fn garbage_is_an_error() {
assert!(Reader::open(vec![0u8; 64]).is_err());
+4 -1
View File
@@ -1342,7 +1342,10 @@ impl<'f> Dataset<'f> {
) -> Result<(Vec<T>, Vec<T>), Error> {
let dt = self.datatype()?;
let is_complex = match &dt {
// Class 11 parses to this same `{r, i}` compound.
Datatype::Complex { size, base_type } => {
matches!(base_type.as_ref(), Datatype::FloatingPoint { .. })
&& *size == 2 * base_type.type_size()
}
Datatype::Compound { size, members } => {
matches!(members.as_slice(), [r, i]
if r.name == "r" && i.name == "i"
+11
View File
@@ -41,6 +41,13 @@ pub enum DType {
Array(Box<DType>, Vec<u32>),
/// Variable-length UTF-8 string stored via a global heap reference.
VariableLengthString,
/// HDF5 2.0 native complex number (datatype class 11) over the given
/// floating-point part type: `Complex(F32)` is numpy `complex64`,
/// `Complex(F64)` `complex128`, and a half-precision part is
/// `Complex(Other("float16"))`. h5py's default complex encoding, the
/// compound `{r, i}`, stays a [`DType::Compound`]; `read_complex_f64`
/// and `read_complex_f32` read both.
Complex(Box<DType>),
/// Catch-all for HDF5 datatypes that do not map to a specific variant.
Other(String),
}
@@ -60,6 +67,7 @@ impl fmt::Display for DType {
DType::U64 => write!(f, "u64"),
DType::String => write!(f, "string"),
DType::VariableLengthString => write!(f, "vlen_string"),
DType::Complex(base) => write!(f, "complex<{base}>"),
DType::Compound(fields) => {
write!(f, "compound{{")?;
for (i, (name, dt)) in fields.iter().enumerate() {
@@ -148,6 +156,9 @@ pub(crate) fn classify_datatype(dt: &clawhdf5_format::datatype::Datatype) -> DTy
base_type,
dimensions,
} => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()),
Datatype::Complex { base_type, .. } => {
DType::Complex(Box::new(classify_datatype(base_type)))
}
_ => DType::Other(format!("{dt:?}")),
}
}
+75
View File
@@ -0,0 +1,75 @@
"""Generate native_complex_hdf5_2.h5: HDF5 2.0 native complex (class 11).
Datasets (every one class 11, or a type holding it):
c64 H5T_COMPLEX_IEEE_F32LE, 4 values incl. -0, inf, nan, subnormal
c128 H5T_COMPLEX_IEEE_F64LE, 2 x 3
c128be H5T_COMPLEX_IEEE_F64BE, the same values
c128x H5T_COMPLEX_IEEE_F64LE, c128 with element (1,1) changed
c32 H5T_COMPLEX_IEEE_F16LE, 3 values
scalar H5T_COMPLEX_IEEE_F64LE, scalar dataspace
compound {z: H5T_COMPLEX_IEEE_F64LE @0, k: H5T_STD_I64LE @16}
array H5T_ARRAY [2] of H5T_COMPLEX_IEEE_F32LE, 2 elements
plus the root attribute `zattr` (H5T_COMPLEX_IEEE_F64LE, 2 values).
Written with h5py 3.16 (libhdf5 2.0.0) through its low-level API, the file
type as the memory type (no conversion). Re-run only if the fixture ever
needs regenerating (generated 2026-09-29):
python gen_native_complex.py native_complex_hdf5_2.h5
The h5dump / h5ls 2.2.0 reference output the tests compare with was made
with the tools of the libhdf5 2.2.0 build described in gen_mx_floats.py:
h5dump native_complex_hdf5_2.h5 > native_complex_hdf5_2.ddl
h5ls -r -v native_complex_hdf5_2.h5 > native_complex_hdf5_2.ls
"""
import sys
import h5py
import numpy as np
from h5py import h5a, h5d, h5s, h5t
out = sys.argv[1]
f = h5py.File(out, "w", libver="latest")
def ds(name, tid, arr, shape=None):
shape = arr.shape if shape is None else shape
sp = h5s.create_simple(shape) if shape != () else h5s.create(h5s.SCALAR)
d = h5d.create(f.id, name.encode(), tid, sp)
d.write(h5s.ALL, h5s.ALL, np.ascontiguousarray(arr), mtype=tid)
nan, inf = float("nan"), float("inf")
c64 = np.array(
[1.5 - 2j, complex(0.0, -0.0), complex(inf, nan), complex(-inf, 1e-40)],
dtype=np.complex64,
)
ds("c64", h5t.COMPLEX_IEEE_F32LE, c64)
c128 = np.array(
[[1 + 2j, -3.5 + 4e300j, 0.1 + 0.2j], [5e-324 - 0j, 1j, -1]], dtype=np.complex128
)
ds("c128", h5t.COMPLEX_IEEE_F64LE, c128)
ds("c128be", h5t.COMPLEX_IEEE_F64BE, c128.astype(">c16"))
c128x = c128.copy()
c128x[1, 1] = 1 + 1j
ds("c128x", h5t.COMPLEX_IEEE_F64LE, c128x)
half = np.array([1.0, -2.0, 0.5, 65504.0, 0.0, -0.0], dtype=np.float16)
ds("c32", h5t.COMPLEX_IEEE_F16LE, half, shape=(3,))
ds("scalar", h5t.COMPLEX_IEEE_F64LE, np.array(2.5 - 0.5j), shape=())
ct = h5t.create(h5t.COMPOUND, 24)
ct.insert(b"z", 0, h5t.COMPLEX_IEEE_F64LE)
ct.insert(b"k", 16, h5t.STD_I64LE)
cd = np.array([(1 + 1j, 7), (-2.5 + 0j, -8)], dtype=[("z", "<c16"), ("k", "<i8")])
ds("compound", ct, cd)
at = h5t.array_create(h5t.COMPLEX_IEEE_F32LE, (2,))
ad = np.array([[1 + 2j, 3 + 4j], [-1 - 1j, 0j]], dtype=np.complex64)
ds("array", at, ad, shape=(2,))
a = h5a.create(f.id, b"zattr", h5t.COMPLEX_IEEE_F64LE, h5s.create_simple((2,)))
a.write(np.array([0.5 - 1.5j, 2j]), mtype=h5t.COMPLEX_IEEE_F64LE)
f.close()
@@ -0,0 +1,80 @@
HDF5 "native_complex_hdf5_2.h5" {
GROUP "/" {
ATTRIBUTE "zattr" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): 0.5-1.5i, 0+2i
}
}
DATASET "array" {
DATATYPE H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): [ 1+2i, 3+4i ], [ -1-1i, 0+0i ]
}
}
DATASET "c128" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128be" {
DATATYPE H5T_COMPLEX_IEEE_F64BE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128x" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 1+1i, -1+0i
}
}
DATASET "c32" {
DATATYPE H5T_COMPLEX_IEEE_F16LE
DATASPACE SIMPLE { ( 3 ) / ( 3 ) }
DATA {
(0): 1+-2i, 0.5+65504i, 0+-0i
}
}
DATASET "c64" {
DATATYPE H5T_COMPLEX_IEEE_F32LE
DATASPACE SIMPLE { ( 4 ) / ( 4 ) }
DATA {
(0): 1.5-2i, 0-0i, inf+nani, -inf+9.99995e-41i
}
}
DATASET "compound" {
DATATYPE H5T_COMPOUND {
H5T_COMPLEX_IEEE_F64LE "z";
H5T_STD_I64LE "k";
}
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): {
1+1i,
7
},
(1): {
-2.5+0i,
-8
}
}
}
DATASET "scalar" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SCALAR
DATA {
(0): 2.5-0.5i
}
}
}
}
Binary file not shown.
+19 -5
View File
@@ -1064,12 +1064,26 @@ fn complex_datasets_and_attributes_round_trip() {
let ds = file.dataset(name).unwrap();
assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}");
assert_eq!(ds.shape().unwrap(), vec![3]);
// Both encodings read as the same {r, i} compound.
assert_eq!(
ds.dtype().unwrap(),
// h5py's encoding is a compound, the native one a complex type.
let want = if name == "native64" {
DType::Complex(Box::new(DType::F32))
} else {
DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)])
);
};
assert_eq!(ds.dtype().unwrap(), want, "{name}");
}
assert_eq!(
file.dataset("native128").unwrap().dtype().unwrap(),
DType::Complex(Box::new(DType::F64))
);
assert_eq!(
file.dataset("native128").unwrap().raw_datatype().unwrap(),
clawhdf5::make_native_complex_f64_type()
);
assert_eq!(
DType::Complex(Box::new(DType::F64)).to_string(),
"complex<f64>"
);
for name in ["compound128", "native128"] {
same(
file.dataset(name).unwrap().read_complex_f64().unwrap(),
@@ -1094,7 +1108,7 @@ fn complex_datasets_and_attributes_round_trip() {
data,
}) => {
assert!(shape.is_empty());
assert_eq!(datatype.type_size(), 16);
assert_eq!(datatype, clawhdf5::make_native_complex_f64_type());
let fields =
clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap();
let part = |k: usize| {
+36 -11
View File
@@ -315,15 +315,12 @@ wrong data.
decoded, and external references are an error (object references
decode). A multi-dimensional numeric attribute is returned as a flat
array (its shape is not reported; `AttrValue::Raw` carries the shape).
HDF5 2.0's native complex type (class 11) is read as the equivalent
`{r, i}` compound: values are right (`Dataset::read_complex_f64`, and
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs
ls`/`dump` and the browser reader show a compound where h5dump 2.x
prints `H5T_COMPLEX_IEEE_F64LE`, so `h5rs dump` of such a file does not
match h5dump 2.x (h5dump 1.14 cannot read it at all). Writing class 11
(added 2026-09-28) is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`); the Python bindings write complex
arrays as h5py's compound only.
HDF5 2.0's native complex type (class 11) is read as its own type since
2026-09-29 ([fixed](#hdf5-20-native-complex-numbers-read-as-a-r-i-compound));
writing it is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`), and the Python bindings write complex
arrays as h5py's compound only. `h5rs dump --json` writes it as the
`{r, i}` compound: hdf5-json (h5json 2.0.0) has no complex class.
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
that: libhdf5 fails only the first metadata read of an image it cannot
load and then reads the file's own (possibly stale) metadata, where we
@@ -533,8 +530,10 @@ browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27.
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included; not in Firefox or Safari.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type (HDF5 2.0 native complex datasets
read, as `[re, im]` pairs); attributes of those types come back
as `value: null` with their `dtype`. (VL strings read, with h5py's
values, through the same `VlResolver` as `File` and `h5rs`.)
- No Zstd or SZIP (both link C): such datasets fail with
@@ -669,6 +668,32 @@ netCDF4-python-written files with netCDF-C itself, and
`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in
[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)).
## HDF5 2.0 native complex numbers read as a `{r, i}` compound
**Status:** fixed 2026-09-29 (branch `feat/complex-first-class`; PR not
yet opened). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0)
and `main` until then. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
A class-11 datatype (`H5T_COMPLEX_IEEE_F64LE`, ...) parsed into the
equivalent compound `{r, i}`. Values were right (`read_complex_f64`,
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs ls`/`dump`
and the browser reader showed a compound, so `h5rs dump` did not match
h5dump 2.x, the browser refused the dataset, and a parsed type written
back out became a compound. `Datatype::parse` now returns
`Datatype::Complex` (in compounds, arrays and VL types too),
`Dataset::dtype()` is `DType::Complex(..)`, `h5rs` prints what h5dump,
h5ls and h5diff 2.2.0 print (checked on tank, 2026-09-29, against a file
written by h5py 3.16 / libhdf5 2.0.0:
`cargo test -p clawhdf5-tools --test native_complex_dump`), and the
browser reads `[re, im]` pairs.
**What users must do:** code that matched a native complex type as a
`Datatype::Compound` (or a `DType::Compound` of `r`/`i`) must match
`Datatype::Complex` / `DType::Complex`, or take the compound view with
`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a
`Complex` arm. Files are unchanged; nothing needs rewriting.
## NetCDF-4: variables' dimensions are guessed from sizes
+6 -2
View File
@@ -123,8 +123,12 @@ else throws an `Error` naming the type.
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in
`docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with
- HDF5 2.0 native complex datasets read as `[re, im]` pairs: `dtype`
`complex<f64>`, `elementShape` ending in `2`, the parts interleaved in
the typed array.
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque and variable-length-sequence datasets are refused with
an error. Attributes of those types are listed with
`value: null` and their `dtype`.
- No Zstd or SZIP filters (they link C): such a dataset fails with
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
+6
View File
@@ -93,6 +93,12 @@ expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
# 3-D: leading dimension held at 0, window over the last two.
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
# HDF5 2.0 native complex (written when h5py's libhdf5 is 2.0 or later):
# each cell is the [re, im] pair.
if "$PY" -c 'import sys, h5py; sys.exit(not getattr(h5py.get_config(), "has_native_complex", False))'; then
expect $H5 "/native_c128" '<dd>complex&lt;f64&gt;</dd>' '<dd>(3, 2)</dd>' '<td>[1, -1.5]</td>' \
'<td>[5, -7.5]</td>'
fi
# Unsupported type: an error, not values.
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
+22 -3
View File
@@ -80,6 +80,15 @@ with h5py.File(h5, "w") as f:
**hdf5plugin.Zstd())
comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")])
f.create_dataset("table", data=comp)
if getattr(h5py.get_config(), "has_native_complex", False):
# HDF5 2.0 native complex (class 11): read as [re, im] pairs.
from h5py import h5d, h5s, h5t
for name, t, dt in [(b"native_c128", h5t.COMPLEX_IEEE_F64LE, "<c16"),
(b"native_c64_be", h5t.COMPLEX_IEEE_F32BE, ">c8")]:
z = (np.arange(6) - 1.5j * np.arange(6)).reshape(3, 2).astype(dt)
d = h5d.create(f.id, name, t, h5s.create_simple(z.shape))
d.write(h5s.ALL, h5s.ALL, z, mtype=t)
g = f.create_group("sensors")
g.attrs["location"] = "lab"
g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4"))
@@ -105,6 +114,8 @@ def kind(dt):
return "strings"
if dt.subdtype is not None:
return kind(dt.subdtype[0])
if dt.kind == "c":
return "f64" if dt.itemsize == 16 else "f32"
if dt.kind == "f":
return "f64" if dt.itemsize == 8 else "f32"
if dt.kind in "iu":
@@ -120,20 +131,25 @@ def flat(a, k):
return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()]
if k.startswith(("i", "u")):
return [str(int(x)) for x in a.ravel()]
if a.dtype.kind == "c":
# [re, im] pairs, in order.
a = np.stack([a.real, a.imag], axis=-1)
return [float(x) for x in a.ravel()]
def entry(ds, slab=None):
k = kind(ds.dtype)
data = ds[()]
e = {"kind": k, "shape": list(np.shape(data)), "values": flat(data, k)}
# A complex element reads as its two parts: one more dimension.
pair = [2] if ds.dtype.kind == "c" else []
e = {"kind": k, "shape": list(np.shape(data)) + pair, "values": flat(data, k)}
if slab:
start, count, stride = slab
idx = tuple(slice(s, s + (c - 1) * st + 1, st)
for s, c, st in zip(start, count, stride))
part = ds[idx]
e["slab"] = {"start": start, "count": count, "stride": stride,
"shape": list(part.shape), "values": flat(part, k)}
"shape": list(part.shape) + pair, "values": flat(part, k)}
return e
@@ -176,7 +192,10 @@ def describe(path):
"datasets": sorted(sets)}
for n, o in members.items():
walk(key.rstrip("/") + "/" + n, o)
elif obj.dtype.names:
elif obj.dtype.names or (
obj.dtype.kind == "c"
and obj.id.get_type().get_class() != getattr(h5py.h5t, "COMPLEX", None)):
# h5py's own complex numbers are a compound {r, i}: refused.
expected["errors"][key] = "compound"
else:
expected["datasets"][key] = entry(obj, slab_for(obj))
+2 -2
View File
@@ -399,7 +399,7 @@ async function remoteTests() {
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": `bytes 0-511/${statSync(join(fixDir, "fixture.h5")).size}` })),
});
await f.read("/grid");
}, /the server sent 0-511/, "wrong range");
@@ -550,7 +550,7 @@ async function floodTests() {
const range = new Headers(init.headers).get("Range");
if (range === "bytes=0-511") return fetch(url, init);
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/${statSync(join(fixDir, "fixture.h5")).size}` } });
},
}).catch((e) => e);
// The open itself may need a second range: flooded either way.