Chunk dimensions of 2^32 or more; copy-free unfiltered 4 GiB writes #32

Open
osobh wants to merge 7 commits from feat/huge-chunk-dims into main
22 changed files with 988 additions and 81 deletions
Showing only changes of commit e4ba09946f - Show all commits
+63
View File
@@ -56,6 +56,69 @@
and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4). and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4).
Affected v2.1.0 to v2.7.0. Affected v2.1.0 to v2.7.0.
### HDF5 2.0 native complex is its own type on read; Python `libver=` (2026-09-29)
- **`Datatype::parse` returns `Datatype::Complex { size, base_type }`** for
a class-11 message, also as a compound member, array base or
variable-length base. It used to return the equivalent `{r, i}`
compound (`docs/known-issues.md`, "HDF5 2.0 native complex numbers read
as a `{r, i}` compound"). **Breaking** for code that matched the
compound view of a native complex type: `Dataset::raw_datatype()` and
attribute datatypes are now `Datatype::Complex`;
`Datatype::complex_as_compound(size, base)` still gives the `{r, i}` view,
and `data_read::read_compound_fields` accepts a `Complex` directly.
Re-serializing a parsed type (e.g. copying it to another file) now writes
class 11 again instead of a compound.
- **Facade:** `DType::Complex(Box<DType>)` (new variant, **breaking** for
exhaustive `match`es over `DType`): `Complex(F32)` is numpy `complex64`,
`Complex(F64)` `complex128`, a binary16 part `Complex(Other("float16"))`
(the same `Other` name small floats already use); `Display` prints
`complex<f64>`. h5py's own complex encoding, the compound `{r, i}`, is
still `DType::Compound`. `read_complex_f64`/`read_complex_f32` read both.
- **`h5rs`:** `dump` prints native complex as h5dump 2.2.0 does:
`H5T_COMPLEX_IEEE_F{16,32,64}{LE,BE}` (else `H5T_COMPLEX { <base> }`), in
arrays and compounds too, and values as `1.5-2i` (`%g%+gi`; for binary16
parts, which h5dump has no C type for, `1+-2i` as it prints them). `ls`
shows `complex64`, `complex128-be`, `complex32` (and `complex<part>` for
non-IEEE parts); `ls -v` shows h5ls 2.2.0's `complex number of` /
`IEEE 64-bit big-endian float`. `diff` compares complex values part by
part and prints the difference as h5diff 2.2.0 does (`1+0i`); a native
complex and an `{r, i}` compound are no longer comparable (different
classes, as in h5diff). `dump --json` keeps the `{r, i}` compound and
`[re, im]` values: hdf5-json (h5json 2.0.0) has no complex class.
- **Browser (`clawhdf5-wasm`):** native complex datasets are readable, as
`[re, im]` pairs: `info().elementShape` ends in `2`, `read()` returns the
parts interleaved in a `Float32Array`/`Float64Array` (binary16 parts
widened to f32), and `dtype` is `complex<f64>` (`complex<f32
(big-endian)>`, `array[2]<complex<f32>>`). The viewer shows each element
as `[re, im]` with no change to the page. h5py's `{r, i}` compound is
still refused, as every compound is.
- **Python:** `clawhdf5.File(path, 'w', libver=...)` takes h5py's values:
`'v108'`, `'v110'`, `'v112'`, `'v114'`, `'v200'`, `'latest'` (the low
bound; high `'latest'`, as h5py) or a `(low, high)` tuple, mapped to
`FileBuilder::libver_bounds`. `'earliest'` as the low bound writes the
1.8 format with a `UserWarning` (clawhdf5 cannot write the pre-1.8
format); as the high bound it is a `ValueError`, as are unknown names and
a low bound above the high one. It is ignored for `'r'` and refused
(`NotImplementedError`) for `'r+'`/`'a'`. Native complex already read as
numpy `complex64`/`complex128` and still does
(`test_read_native_complex_from_h5py`).
- Tests (tank, 2026-09-29): `crates/clawhdf5-tools/tests/native_complex_dump.rs`
compares `h5rs dump` with h5dump 2.2.0's output stored next to a fixture
written by h5py 3.16 / libhdf5 2.0.0
(`crates/clawhdf5/tests/fixtures/gen_native_complex.py`:
`F32LE`/`F64LE`/`F64BE`/`F16LE`, scalar, compound member, array, root
attribute), line for line outside the values and value for value within
them, and `ls`/`diff` with h5ls/h5diff 2.2.0;
`integration_tests::complex_datasets_and_attributes_round_trip`
(`DType::Complex`, `raw_datatype()`); `clawhdf5-wasm` unit tests on the
fixture, `h5py_interop` and `examples/wasm-viewer/test` (Node:
`test.mjs`, 274 + 1324 checks; the page in headless Chromium for
`/native_c128`, `/native_c64_be`, `/pairs`) with two native complex
datasets added to `make_fixture.py` (when h5py's libhdf5 is 2.0+);
`crates/clawhdf5-py/tests/test_libver.py` (every bound; `'v108'` output
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
hard-codes the fixture's length.
### NetCDF-4: variables' dimensions come from the file (2026-09-28) ### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of - `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
+2 -2
View File
@@ -115,12 +115,12 @@ Limits and open issues, with dates, are in
|---|---|---|---| |---|---|---|---|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) | | **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | | **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 | | **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more | | **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | | **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) | | **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | | **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) | | **Bindings** | Python (read, `'w'` for numeric and complex arrays with h5py's `libver=`, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`; HDF5 2.0 native complex as `[re, im]` pairs); no Zstd/SZIP/pcodec, no compound (h5py's `{r, i}` complex included), reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`, Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`,
`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure `blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure
+35 -29
View File
@@ -145,13 +145,16 @@ pub enum Datatype {
/// imaginary, in rectangular form. `size` is twice the base size and the /// imaginary, in rectangular form. `size` is twice the base size and the
/// base is an IEEE float. /// base is an IEEE float.
/// ///
/// This variant exists for **writing** (see /// [`Datatype::parse`] returns this variant for every class-11 message,
/// also inside compounds, arrays and variable-length types. Readers that
/// want the member view (`{r, i}`, as h5py writes complex numbers by
/// default) use [`Datatype::complex_as_compound`]; `read_compound_fields`
/// accepts this variant directly.
///
/// Writing it is opt-in (see
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and /// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
/// newer can read class 11, so it is opt-in and h5py's compound `{r, i}` /// newer can read class 11, so h5py's compound `{r, i}` stays the
/// stays the default complex encoding. [`Datatype::parse`] still /// default complex encoding.
/// surfaces a class-11 message as that equivalent `{r, i}` compound, so
/// every compound reader handles both encodings; parsing what this
/// variant serializes therefore yields a `Compound`, not a `Complex`.
Complex { size: u32, base_type: Box<Datatype> }, Complex { size: u32, base_type: Box<Datatype> },
} }
@@ -857,9 +860,7 @@ impl Datatype {
// Complex number (HDF5 2.0, datatype version 5). The properties // Complex number (HDF5 2.0, datatype version 5). The properties
// are a single base floating-point datatype message; an element // are a single base floating-point datatype message; an element
// is two consecutive base-type values (real, imaginary). There // is two consecutive base-type values (real, imaginary). There
// is no member list. Surface it as the equivalent two-member // is no member list.
// compound `{r, i}` — the same shape h5py writes for numpy
// complex dtypes — so downstream compound readers work as-is.
if version != 5 { if version != 5 {
return Err(FormatError::InvalidDatatypeVersion { return Err(FormatError::InvalidDatatypeVersion {
class: class_id, class: class_id,
@@ -875,7 +876,13 @@ impl Datatype {
actual: size as usize, actual: size as usize,
}); });
} }
Ok((Self::complex_as_compound(size, &base_type), pos)) Ok((
Datatype::Complex {
size,
base_type: Box::new(base_type),
},
pos,
))
} }
_ => Err(FormatError::InvalidDatatypeClass(class_id)), _ => Err(FormatError::InvalidDatatypeClass(class_id)),
} }
@@ -1174,9 +1181,8 @@ impl Datatype {
/// The `{r, i}` compound equivalent to a native complex type of `size` /// The `{r, i}` compound equivalent to a native complex type of `size`
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the /// bytes over `base_type`: `r` at offset 0, `i` right after it — the
/// shape h5py writes for numpy complex dtypes, and what [`Self::parse`] /// shape h5py writes for numpy complex dtypes. Readers that want a
/// returns for a class-11 message. Readers that meet a /// [`Datatype::Complex`] as members handle it through this view.
/// [`Datatype::Complex`] handle it through this view.
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype { pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
let base_size = base_type.type_size(); let base_size = base_type.type_size();
Datatype::Compound { Datatype::Compound {
@@ -1849,21 +1855,24 @@ mod tests {
fn test_complex_v5_from_hdf5_2_0() { fn test_complex_v5_from_hdf5_2_0() {
let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap(); let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap();
assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len()); assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len());
match dt { match &dt {
Datatype::Compound { size, members } => { Datatype::Complex { size, base_type } => {
assert_eq!(size, 16); assert_eq!(*size, 16);
assert_eq!(members.len(), 2);
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
for m in &members {
assert!(matches!( assert!(matches!(
m.datatype, base_type.as_ref(),
Datatype::FloatingPoint { size: 8, .. } Datatype::FloatingPoint { size: 8, .. }
)); ));
} }
other => panic!("expected Complex, got {other:?}"),
} }
other => panic!("expected Compound, got {other:?}"), // The member view is h5py's `{r, i}` compound.
} let Datatype::Compound { members, .. } =
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
else {
unreachable!()
};
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
} }
#[test] #[test]
@@ -1887,7 +1896,7 @@ mod tests {
assert_eq!(members.len(), 2); assert_eq!(members.len(), 2);
assert!(matches!( assert!(matches!(
&members[0].datatype, &members[0].datatype,
Datatype::Compound { size: 16, members } if members.len() == 2 Datatype::Complex { size: 16, .. }
)); ));
assert_eq!( assert_eq!(
(members[1].name.as_str(), members[1].byte_offset), (members[1].name.as_str(), members[1].byte_offset),
@@ -1918,12 +1927,9 @@ mod tests {
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0); assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
assert_eq!(dt.type_size(), 16); assert_eq!(dt.type_size(), 16);
dt.check_encodable().unwrap(); dt.check_encodable().unwrap();
// Parsing surfaces class 11 as the equivalent `{r, i}` compound. // Parsing returns the native complex type unchanged.
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap(); let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
assert_eq!( assert_eq!(parsed, dt);
parsed,
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
);
let f32c = make_native_complex_f32_type().serialize(); let f32c = make_native_complex_f32_type().serialize();
assert_eq!( assert_eq!(
+16 -1
View File
@@ -104,7 +104,22 @@ writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and
`complex128` arrays; the file is written on `close()`. Complex arrays are `complex128` arrays; the file is written on `close()`. Complex arrays are
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the (not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
Rust API writes that on request). Rust API writes that on request). Both forms read back as numpy
`complex64`/`complex128`.
`libver=` sets the library version bounds as in h5py: `'v108'`,
`'v110'`, `'v112'`, `'v114'`, `'v200'` or `'latest'` (the low bound; the
high bound is then `'latest'`), or a `(low, high)` tuple. With `'v108'`
the file is written in the HDF5 1.8 format (version-2 superblock,
version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the
HDF5 1.10 format. `'earliest'` as the low bound writes the 1.8 format too,
with a `UserWarning`: clawhdf5 cannot write the pre-1.8 format. The
argument is ignored when reading and refused for `'r+'`/`'a'`.
```python
with clawhdf5.File("old.h5", "w", libver="v108") as f: # HDF5 1.8 reads it
f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip")
```
## Editing a file in place ## Editing a file in place
+90 -2
View File
@@ -13,10 +13,13 @@ use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group}; use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle; use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err}; use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
use clawhdf5_rs::LibVer;
/// Internal state for write mode. /// Internal state for write mode.
struct WriteState { struct WriteState {
path: PathBuf, path: PathBuf,
/// `libver=` as (low, high); `None` keeps the writer's default.
libver: Option<(LibVer, LibVer)>,
root_datasets: Vec<DatasetSpec>, root_datasets: Vec<DatasetSpec>,
root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>, root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
groups: Vec<Arc<Mutex<WriteGroupState>>>, groups: Vec<Arc<Mutex<WriteGroupState>>>,
@@ -80,10 +83,32 @@ impl PyFile {
/// `az://`; which schemes work depends on how the wheel was built) /// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`) /// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write /// mode: 'r' for read (default), 'w' for write
/// libver: library version bounds for mode 'w', as h5py's: one of
/// 'earliest', 'v108', 'v110', 'v112', 'v114', 'v200', 'latest' (the
/// low bound; the high bound is then 'latest') or a (low, high)
/// tuple of them. The low bound is the oldest HDF5 release whose
/// format the file uses ('v108': HDF5 1.8 can read it); the high
/// bound the newest whose features it may use. clawhdf5 cannot write
/// the pre-1.8 format, so a low bound of 'earliest' writes the 1.8
/// format (with a warning) and a high bound of 'earliest' is an
/// error. Default (None): the HDF5 1.10 format clawhdf5 has always
/// written. Ignored for reading; not supported with 'r+' / 'a'.
#[new] #[new]
#[pyo3(signature = (path, mode="r"))] #[pyo3(signature = (path, mode="r", libver=None))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> { fn new(
py: Python<'_>,
path: &str,
mode: &str,
libver: Option<&Bound<'_, PyAny>>,
) -> PyResult<Self> {
let filename = path.to_string(); let filename = path.to_string();
let libver = libver.map(|v| parse_libver(py, v)).transpose()?;
if libver.is_some() && matches!(mode, "r+" | "a") {
return Err(PyNotImplementedError::new_err(format!(
"libver with mode '{mode}': clawhdf5's in-place editor keeps the format \
versions the file already uses"
)));
}
if is_url(path) { if is_url(path) {
if mode != "r" { if mode != "r" {
return Err(PyValueError::new_err(format!( return Err(PyValueError::new_err(format!(
@@ -113,6 +138,7 @@ impl PyFile {
// Absolute now: the file is written at close, possibly // Absolute now: the file is written at close, possibly
// after the working directory changed. // after the working directory changed.
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)), path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
libver,
root_datasets: Vec::new(), root_datasets: Vec::new(),
root_attrs: Arc::new(Mutex::new(Vec::new())), root_attrs: Arc::new(Mutex::new(Vec::new())),
groups: Vec::new(), groups: Vec::new(),
@@ -460,10 +486,71 @@ fn parse_compression(
} }
} }
/// One of h5py's `libver` names as a bound; `high` says which end it is.
/// `Ok(None)` is 'earliest' as the low bound: the pre-1.8 format, which
/// clawhdf5 cannot write.
fn libver_name(name: &str, high: bool) -> PyResult<Option<LibVer>> {
Ok(Some(match name {
"earliest" if high => {
return Err(PyValueError::new_err(
"libver high bound 'earliest' (the pre-1.8 format) cannot be written by \
clawhdf5; the oldest format it writes is 'v108'",
));
}
"earliest" => return Ok(None),
"v108" => LibVer::V18,
"v110" => LibVer::V110,
"v112" => LibVer::V112,
"v114" => LibVer::V114,
"v200" => LibVer::V200,
"latest" => LibVer::Latest,
other => {
return Err(PyValueError::new_err(format!(
"unknown libver '{other}'; expected 'earliest', 'v108', 'v110', 'v112', \
'v114', 'v200' or 'latest'"
)));
}
}))
}
/// h5py's `libver=`: a name (the low bound, high bound 'latest') or a
/// `(low, high)` pair.
fn parse_libver(py: Python<'_>, v: &Bound<'_, PyAny>) -> PyResult<(LibVer, LibVer)> {
let (low, high): (String, String) = match v.extract::<String>() {
Ok(name) => (name, "latest".into()),
Err(_) => v.extract().map_err(|_| {
PyValueError::new_err("libver must be a string or a (low, high) tuple of strings")
})?,
};
let Some(high) = libver_name(&high, true)? else {
unreachable!("a high bound is never None")
};
let low = match libver_name(&low, false)? {
Some(low) => low,
None => {
PyModule::import(py, "warnings")?.getattr("warn")?.call1((
"libver 'earliest': clawhdf5 cannot write the pre-1.8 format; the file \
is written in the HDF5 1.8 format ('v108') instead",
py.get_type::<pyo3::exceptions::PyUserWarning>(),
))?;
LibVer::V18
}
};
if low > high {
return Err(PyValueError::new_err(format!(
"libver low bound {low} is newer than the high bound {high}"
)));
}
Ok((low, high))
}
/// Build and write the HDF5 file from accumulated write state. /// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> { fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| { crate::no_panic(|| {
let mut builder = clawhdf5_rs::FileBuilder::new(); let mut builder = clawhdf5_rs::FileBuilder::new();
if let Some((low, high)) = state.libver {
builder.libver_bounds(low, high);
}
// Root attributes // Root attributes
let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner()); let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner());
@@ -520,6 +607,7 @@ mod tests {
let state = WriteState { let state = WriteState {
path: path.clone(), path: path.clone(),
libver: None,
root_datasets: vec![DatasetSpec { root_datasets: vec![DatasetSpec {
name: "data".into(), name: "data".into(),
data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]), data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]),
+143
View File
@@ -0,0 +1,143 @@
"""`clawhdf5.File(path, 'w', libver=...)`: h5py's library version bounds.
A file written with libver='v108' must open in HDF5 1.8. Its h5dump is
found through CLAWHDF5_H5DUMP18 or at ~/.cache/hdf5-1.8.23/bin/h5dump
(scripts/build-hdf5-1.8.sh builds it); without it that check is skipped.
"""
import os
import subprocess
import numpy as np
import pytest
import clawhdf5
def superblock_version(path):
with open(path, "rb") as f:
head = f.read(9)
assert head[:8] == b"\x89HDF\r\n\x1a\n"
return head[8]
def h5dump18():
path = os.environ.get("CLAWHDF5_H5DUMP18") or os.path.expanduser(
"~/.cache/hdf5-1.8.23/bin/h5dump"
)
try:
out = subprocess.run([path, "--version"], capture_output=True, text=True)
except OSError:
return None
return path if "1.8." in out.stdout else None
def write_sample(path, libver):
with clawhdf5.File(path, "w", libver=libver) as f:
f.create_dataset("x", data=np.arange(10, dtype=np.float64))
f.create_dataset(
"chunked",
data=np.arange(100, dtype=np.int32),
chunks=(30,),
compression="gzip",
)
f.create_dataset("z", data=np.array([1 + 2j, -3j], dtype=np.complex128))
g = f.create_group("g")
g.create_dataset("y", data=np.ones(3, dtype=np.float32))
f.attrs["version"] = 1
@pytest.mark.parametrize(
"libver, sb",
[
(None, 3),
("v108", 2),
(("v108", "v108"), 2),
(("v108", "latest"), 2),
("v110", 3),
("v112", 3),
("v114", 3),
("v200", 3),
("latest", 3),
(("v110", "v200"), 3),
],
)
def test_libver_bounds(h5py, tmp_path, libver, sb):
path = str(tmp_path / "libver.h5")
write_sample(path, libver)
# v108 low bound: HDF5 1.8's version-2 superblock; 1.10 and later: 3.
assert superblock_version(path) == sb
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
np.testing.assert_array_equal(f["z"][:], [1 + 2j, -3j])
np.testing.assert_array_equal(f["g/y"][:], np.ones(3))
assert f.attrs["version"] == 1
with clawhdf5.File(path, "r") as f:
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
def test_libver_v108_opens_in_hdf5_1_8(tmp_path):
h5dump = h5dump18()
if h5dump is None:
pytest.skip("no HDF5 1.8 h5dump (CLAWHDF5_H5DUMP18)")
path = str(tmp_path / "v108.h5")
write_sample(path, "v108")
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode == 0, out.stderr
assert "h5dump error" not in out.stderr
assert 'DATASET "y"' in out.stdout
# The chunked, deflated dataset (a version-1 B-tree) reads in full.
out = subprocess.run(
[h5dump, "-w", "0", "-d", "/chunked", path], capture_output=True, text=True
)
assert out.returncode == 0, out.stderr
assert "(0): " + ", ".join(str(k) for k in range(100)) + "\n" in out.stdout
# The default (1.10 format) does not open in 1.8: the bound matters.
path = str(tmp_path / "default.h5")
write_sample(path, None)
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode != 0
def test_libver_earliest_writes_v108_with_a_warning(h5py, tmp_path):
path = str(tmp_path / "earliest.h5")
with pytest.warns(UserWarning, match="pre-1.8"):
write_sample(path, "earliest")
assert superblock_version(path) == 2
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
@pytest.mark.parametrize(
"libver, match",
[
(("v108", "earliest"), "high bound 'earliest'"),
("v109", "unknown libver 'v109'"),
(("latest", "v108"), "newer than the high bound"),
(("v108",), "tuple"),
(3, "tuple"),
],
)
def test_libver_errors(tmp_path, libver, match):
with pytest.raises(ValueError, match=match):
clawhdf5.File(str(tmp_path / "bad.h5"), "w", libver=libver)
def test_libver_v108_high_bound_writes_everything(tmp_path):
# Native complex numbers need HDF5 2.0; the Python writer stores complex
# as h5py's {r, i} compound, which 1.8 reads, so the 1.8 high bound is
# fine for everything clawhdf5.File writes.
path = str(tmp_path / "v18only.h5")
write_sample(path, ("v108", "v108"))
assert superblock_version(path) == 2
def test_libver_not_for_editing(tmp_path):
path = str(tmp_path / "e.h5")
write_sample(path, None)
with pytest.raises(NotImplementedError, match="libver"):
clawhdf5.File(path, "r+", libver="latest")
# Reading ignores it, as the bounds only affect what is written.
with clawhdf5.File(path, "r", libver="v108") as f:
assert f["x"].shape == (10,)
+22 -1
View File
@@ -603,7 +603,22 @@ impl Diff {
let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) { let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) {
(Some(x), Some(y), ..) => x.abs_diff(y).to_string(), (Some(x), Some(y), ..) => x.abs_diff(y).to_string(),
(_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64), (_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64),
// Each part's absolute difference, as h5diff 2.x prints
// it (`1+0i`).
_ => match (&va, &vb) {
(
Value::Complex { re: ar, im: ai, .. },
Value::Complex { re: br, im: bi, .. },
) => match (number(ar), number(ai), number(br), number(bi)) {
(Some(ar), Some(ai), Some(br), Some(bi)) => format!(
"{}+{}i",
value::fmt_float((ar - br).abs(), 64),
value::fmt_float((ai - bi).abs(), 64)
),
_ => String::new(), _ => String::new(),
},
_ => String::new(),
},
}; };
rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}")); rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}"));
} }
@@ -713,6 +728,9 @@ impl Diff {
(Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => { (Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => {
p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v)) p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v))
} }
(Value::Complex { re: pr, im: pi, .. }, Value::Complex { re: qr, im: qi, .. }) => {
self.equal(a, b, pr, qr) && self.equal(a, b, pi, qi)
}
(Value::Ref(None), Value::Ref(None)) => true, (Value::Ref(None), Value::Ref(None)) => true,
(Value::Ref(Some(p)), Value::Ref(Some(q))) => { (Value::Ref(Some(p)), Value::Ref(Some(q))) => {
// Addresses mean nothing across files: compare the paths the // Addresses mean nothing across files: compare the paths the
@@ -846,7 +864,10 @@ fn types_comparable(a: &Datatype, b: &Datatype) -> bool {
( (
Datatype::Enumeration { base_type: x, .. }, Datatype::Enumeration { base_type: x, .. },
Datatype::Enumeration { base_type: y, .. }, Datatype::Enumeration { base_type: y, .. },
) => types_comparable(x, y), )
| (Datatype::Complex { base_type: x, .. }, Datatype::Complex { base_type: y, .. }) => {
types_comparable(x, y)
}
_ => true, _ => true,
} }
} }
+34 -6
View File
@@ -135,9 +135,16 @@ pub fn short(dt: &Datatype) -> String {
return f.short.into(); return f.short.into();
} }
match dt { match dt {
Datatype::Complex { size, base_type } => { // numpy's names (`complex64` is two `float32`) for IEEE parts,
short(&Datatype::complex_as_compound(*size, base_type)) // `complex<part>` otherwise.
} Datatype::Complex { size, base_type } => match base_type.as_ref() {
Datatype::FloatingPoint { byte_order, .. } if is_ieee(base_type) => format!(
"complex{}{}",
u64::from(*size) * 8,
if be(byte_order) { "-be" } else { "" }
),
_ => format!("complex<{}>", short(base_type)),
},
Datatype::FixedPoint { Datatype::FixedPoint {
size, size,
signed, signed,
@@ -216,8 +223,9 @@ pub fn long(dt: &Datatype) -> String {
return f.long.into(); return f.long.into();
} }
match dt { match dt {
Datatype::Complex { size, base_type } => { // As h5ls 2.x prints a complex type it has no native name for.
long(&Datatype::complex_as_compound(*size, base_type)) Datatype::Complex { base_type, .. } => {
format!("complex number of\n {}", long(base_type))
} }
Datatype::FixedPoint { Datatype::FixedPoint {
size, size,
@@ -361,6 +369,19 @@ fn atomic_ddl(dt: &Datatype) -> Option<String> {
)), )),
// As h5dump 2.x names them (checked against h5dump 2.2.0). // As h5dump 2.x names them (checked against h5dump 2.2.0).
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()), Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
// h5dump 2.x's predefined complex names: IEEE binary16/32/64 parts.
Datatype::Complex { base_type, .. } => match base_type.as_ref() {
Datatype::FloatingPoint {
size, byte_order, ..
} if is_ieee(base_type) && !matches!(byte_order, DatatypeByteOrder::Vax) => {
Some(format!(
"H5T_COMPLEX_IEEE_F{}{}",
u64::from(*size) * 8,
order_suffix(byte_order)
))
}
_ => None,
},
Datatype::BitField { Datatype::BitField {
size, byte_order, .. size, byte_order, ..
} => Some(format!( } => Some(format!(
@@ -418,6 +439,10 @@ pub fn ddl(dt: &Datatype, ind: usize) -> String {
Datatype::VariableLength { base_type, .. } => { Datatype::VariableLength { base_type, .. } => {
format!("H5T_VLEN {{ {} }}", ddl(base_type, ind)) format!("H5T_VLEN {{ {} }}", ddl(base_type, ind))
} }
// Parts h5dump has no predefined complex name for.
Datatype::Complex { base_type, .. } => {
format!("H5T_COMPLEX {{ {} }}", ddl(base_type, ind))
}
Datatype::Opaque { size, tag } => { Datatype::Opaque { size, tag } => {
let tag = String::from_utf8_lossy(tag); let tag = String::from_utf8_lossy(tag);
format!( format!(
@@ -498,6 +523,9 @@ fn string_ddl(
/// hdf5-json type object. /// hdf5-json type object.
pub fn json(dt: &Datatype) -> J { pub fn json(dt: &Datatype) -> J {
match dt { match dt {
// hdf5-json (h5json 2.0.0) has no complex class: a native complex
// is written as the `{r, i}` compound h5py uses for complex numbers,
// its values as `[re, im]` pairs.
Datatype::Complex { size, base_type } => { Datatype::Complex { size, base_type } => {
json(&Datatype::complex_as_compound(*size, base_type)) json(&Datatype::complex_as_compound(*size, base_type))
} }
@@ -598,7 +626,7 @@ fn pad_json(p: &StringPadding) -> &'static str {
/// Class name used to decide whether two datatypes can be compared. /// Class name used to decide whether two datatypes can be compared.
pub fn class(dt: &Datatype) -> &'static str { pub fn class(dt: &Datatype) -> &'static str {
match dt { match dt {
Datatype::Complex { .. } => "compound", Datatype::Complex { .. } => "complex",
Datatype::FixedPoint { .. } => "integer", Datatype::FixedPoint { .. } => "integer",
Datatype::FloatingPoint { .. } => "float", Datatype::FloatingPoint { .. } => "float",
Datatype::Time { .. } => "time", Datatype::Time { .. } => "time",
+31 -5
View File
@@ -27,6 +27,15 @@ pub enum Value {
/// An enum member (name, when the value matches one) and its value. /// An enum member (name, when the value matches one) and its value.
Enum(Option<String>, i128), Enum(Option<String>, i128),
Compound(Vec<(String, Value)>), Compound(Vec<(String, Value)>),
/// A native complex number (HDF5 2.0 class 11): real and imaginary
/// part. `sign_free` is set when h5dump joins the parts with a bare `+`
/// (parts it has no native C type for: binary16, non-IEEE), giving
/// `1+-2i`; otherwise the imaginary part carries its sign (`1-2i`).
Complex {
re: Box<Value>,
im: Box<Value>,
sign_free: bool,
},
Array(Vec<Value>), Array(Vec<Value>),
/// A variable-length sequence. /// A variable-length sequence.
Seq(Vec<Value>), Seq(Vec<Value>),
@@ -183,11 +192,18 @@ impl<'a> Decoder<'a> {
return Value::Error("short element".into()); return Value::Error("short element".into());
}; };
match dt { match dt {
Datatype::Complex { size, base_type } => self.decode( Datatype::Complex { base_type, .. } => {
&Datatype::complex_as_compound(*size, base_type), let part = base_type.type_size() as usize;
b, let (re, im) = b.split_at(part.min(b.len()));
depth + 1, Value::Complex {
), re: Box::new(self.decode(base_type, re, depth + 1)),
im: Box::new(self.decode(base_type, im, depth + 1)),
// h5dump prints `float`/`double` complex (whatever the
// byte order) with `%g%+gi`, anything else as
// `<re>+<im>i`.
sign_free: !(dtype::is_ieee(base_type) && matches!(part, 4 | 8)),
}
}
Datatype::FixedPoint { .. } => match decode_int(dt, b) { Datatype::FixedPoint { .. } => match decode_int(dt, b) {
Some(v) => Value::Int(v), Some(v) => Value::Int(v),
None => Value::Bytes(b.to_vec()), None => Value::Bytes(b.to_vec()),
@@ -354,6 +370,15 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
.collect::<Vec<_>>() .collect::<Vec<_>>()
.join(", ") .join(", ")
), ),
Value::Complex { re, im, sign_free } => {
let (re, im) = (text(re, h5paths), text(im, h5paths));
let plus = if *sign_free || !im.starts_with('-') {
"+"
} else {
""
};
format!("{re}{plus}{im}i")
}
Value::Array(vs) => format!( Value::Array(vs) => format!(
"[ {} ]", "[ {} ]",
vs.iter() vs.iter()
@@ -408,6 +433,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)), Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths), Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()), Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
Value::Complex { re, im, .. } => J::Array(vec![to_json(re, h5paths), to_json(im, h5paths)]),
Value::Array(vs) | Value::Seq(vs) => { Value::Array(vs) | Value::Seq(vs) => {
J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect()) J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect())
} }
@@ -0,0 +1,202 @@
//! `h5rs dump` and `h5rs ls` of HDF5 2.0 native complex numbers (datatype
//! class 11), against the output of h5dump 2.2.0 and h5ls 2.2.0 (the Debian
//! h5dump CI installs, 1.14.x, predates the type). The h5dump output is
//! stored next to the fixture; see
//! `crates/clawhdf5/tests/fixtures/gen_native_complex.py`.
use std::path::PathBuf;
use std::process::Command;
fn fixture(name: &str) -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name)
}
fn h5rs(args: &[&str]) -> String {
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(args)
.output()
.unwrap();
assert!(out.status.success(), "h5rs {args:?}: {out:?}");
String::from_utf8(out.stdout).unwrap()
}
/// Split a dump into its lines, with the data lines (`(i): ...`) replaced
/// by their index, and the data values in order.
fn split(ddl: &str) -> (Vec<String>, Vec<String>) {
let (mut frame, mut data) = (Vec::new(), Vec::new());
for line in ddl.lines() {
match line.split_once("):") {
Some((idx, body)) if idx.trim_start().starts_with('(') => {
frame.push(format!("{idx}):"));
// Values are split at `, `; array elements come bracketed.
data.extend(
body.split(", ")
.map(|s| s.trim().trim_matches(|c| c == '[' || c == ']' || c == ','))
.map(str::trim)
.filter(|s| !s.is_empty() && *s != "{")
.map(String::from),
);
}
_ => {
let t = line.trim().trim_end_matches(',');
if t.ends_with('i') || t.parse::<f64>().is_ok() {
// A compound member's value on its own line.
data.push(t.to_string());
} else {
frame.push(line.to_string());
}
}
}
}
(frame, data)
}
/// `a+bi`, `a-bi` or `a+-bi` as its two parts (a real number as one).
fn parts(v: &str) -> Vec<f64> {
let Some(z) = v.strip_suffix('i') else {
return vec![v.parse().unwrap()];
};
let b = z.as_bytes();
let at = (1..b.len())
.find(|&k| (b[k] == b'+' || b[k] == b'-') && !matches!(b[k - 1], b'e' | b'E' | b'+'))
.unwrap_or_else(|| panic!("not a complex value: {v}"));
let (re, im) = z.split_at(at);
let im = im.strip_prefix('+').unwrap_or(im);
vec![re.parse().unwrap(), im.parse().unwrap()]
}
#[test]
fn dump_matches_h5dump_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
let ours = h5rs(&["dump", path.to_str().unwrap()]);
let reference = std::fs::read_to_string(fixture("native_complex_hdf5_2.ddl")).unwrap();
let (mut our_frame, our_data) = split(&ours);
let (mut ref_frame, ref_data) = split(&reference);
// The first line names the file as it was given.
our_frame.remove(0);
ref_frame.remove(0);
// Everything but the values is h5dump's byte for byte: the types print
// as H5T_COMPLEX_IEEE_F64LE, H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }, ...
assert_eq!(our_frame, ref_frame);
assert_eq!(our_data.len(), ref_data.len(), "{our_data:?}\n{ref_data:?}");
assert_eq!(our_data.len(), 36);
// Values: h5dump prints `%g` (6 digits), inf/nan, and the binary16
// parts it has no C type for as `<re>+<im>i` (`1+-2i`); h5rs prints
// the shortest round-trip string and Inf/NaN, joined the same way.
for (i, (o, r)) in our_data.iter().zip(&ref_data).enumerate() {
assert_eq!(o.contains("+-"), r.contains("+-"), "[{i}] {o} vs {r}");
let (op, rp) = (parts(o), parts(r));
assert_eq!(op.len(), rp.len(), "[{i}] {o} vs {r}");
for (x, y) in op.iter().zip(&rp) {
let same = if y.is_nan() || y.is_infinite() {
x.is_nan() == y.is_nan() && (y.is_nan() || x == y)
} else {
*x as f32 == *y as f32 || (x - y).abs() <= 5e-6 * y.abs()
};
assert!(same, "[{i}] h5rs {o}, h5dump {r}");
assert_eq!(
x.is_sign_negative(),
y.is_sign_negative(),
"[{i}] {o} vs {r}"
);
}
}
}
#[test]
fn ls_describes_complex_like_h5ls_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
// h5ls 2.2.0 -v, `Type:` of the types it has no native C name for (it
// prints `native double _Complex` for the others, as it prints `native
// double` where h5rs prints `IEEE 64-bit little-endian float`).
for (name, long) in [
(
"c128be",
"complex number of\n IEEE 64-bit big-endian float",
),
(
"c128",
"complex number of\n IEEE 64-bit little-endian float",
),
(
"c32",
"complex number of\n IEEE 16-bit little-endian float",
),
(
"array",
"[2] complex number of\n IEEE 32-bit little-endian float",
),
] {
let target = format!("{}/{name}", path.display());
let verbose = h5rs(&["ls", "-v", &target]);
assert!(
verbose.contains(&format!(" Type: {long}\n")),
"{name}:\n{verbose}"
);
}
let listing = h5rs(&["ls", path.to_str().unwrap()]);
for (name, short) in [
("c64", "complex64"),
("c128", "complex128"),
("c128be", "complex128-be"),
("c32", "complex32"),
("scalar", "complex128"),
("array", "array[2]<complex64>"),
("compound", "compound{z: complex128, k: int64}"),
] {
assert!(
listing
.lines()
.any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))),
"{name}:\n{listing}"
);
}
}
#[test]
fn json_dump_uses_the_r_i_compound() {
// hdf5-json (h5json 2.0.0) has no complex class: the type is written as
// the `{r, i}` compound h5py uses, the values as `[re, im]` pairs.
let path = fixture("native_complex_hdf5_2.h5");
let j: serde_json::Value =
serde_json::from_str(&h5rs(&["dump", "--json", path.to_str().unwrap()])).unwrap();
let ds = j["datasets"]
.as_object()
.unwrap()
.values()
.find(|d| d["alias"][0] == "/scalar")
.unwrap();
assert_eq!(ds["type"]["class"], "H5T_COMPOUND");
assert_eq!(ds["type"]["fields"][0]["name"], "r");
assert_eq!(ds["type"]["fields"][1]["type"]["base"], "H5T_IEEE_F64LE");
assert_eq!(ds["value"], serde_json::json!([2.5, -0.5]));
}
#[test]
fn diff_compares_complex_values() {
let path = fixture("native_complex_hdf5_2.h5");
let p = path.to_str().unwrap();
// Same values, different byte order: no differences.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", p, p, "/c128", "/c128be"])
.output()
.unwrap();
assert!(out.status.success(), "{out:?}");
// One element differs in its imaginary part: one difference, exit
// status 1, as h5diff 2.2.0 reports it.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", "-r", p, p, "/c128", "/c128x"])
.output()
.unwrap();
assert_eq!(out.status.code(), Some(1), "{out:?}");
let text = String::from_utf8_lossy(&out.stdout);
assert!(text.contains("1 difference(s) found"), "{text}");
// h5diff 2.2.0: `[ 1 1 ] 0+1i 1+1i 1+0i`.
let row = text.lines().find(|l| l.starts_with("[ 1 1 ]")).unwrap();
assert_eq!(
row.split_whitespace().collect::<Vec<_>>(),
["[", "1", "1", "]", "0+1i", "1+1i", "1+0i"]
);
}
+85 -7
View File
@@ -312,11 +312,14 @@ impl Reader {
let base = array_base(dt); let base = array_base(dt);
let is_array = !std::ptr::eq(base, dt); let is_array = !std::ptr::eq(base, dt);
Ok(match base { Ok(match base {
// Read as the part type: an array or complex element is its
// parts stored one after another (a complex number's real, then
// imaginary part).
Datatype::FloatingPoint { size, .. } if *size <= 4 => { Datatype::FloatingPoint { size, .. } if *size <= 4 => {
Data::F32(data_read::read_as_f32(raw, dt).map_err(err)?) Data::F32(data_read::read_as_f32(raw, base).map_err(err)?)
} }
Datatype::FloatingPoint { .. } => { Datatype::FloatingPoint { .. } => {
Data::F64(data_read::read_as_f64(raw, dt).map_err(err)?) Data::F64(data_read::read_as_f64(raw, base).map_err(err)?)
} }
Datatype::FixedPoint { size, signed, .. } => { Datatype::FixedPoint { size, signed, .. } => {
let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err); let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err);
@@ -393,10 +396,13 @@ fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T
.collect() .collect()
} }
/// Innermost element type of (possibly nested) array datatypes. /// Innermost element type of (possibly nested) array datatypes; for a
/// native complex type, its part type (each element holds two).
fn array_base(dt: &Datatype) -> &Datatype { fn array_base(dt: &Datatype) -> &Datatype {
match dt { match dt {
Datatype::Array { base_type, .. } => array_base(base_type), Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => {
array_base(base_type)
}
_ => dt, _ => dt,
} }
} }
@@ -411,6 +417,12 @@ fn element_shape(dt: &Datatype) -> Vec<u64> {
dims.extend(element_shape(base_type)); dims.extend(element_shape(base_type));
dims dims
} }
// `[re, im]`: a complex element reads as its two parts.
Datatype::Complex { base_type, .. } => {
let mut dims = vec![2];
dims.extend(element_shape(base_type));
dims
}
_ => Vec::new(), _ => Vec::new(),
} }
} }
@@ -487,9 +499,7 @@ pub fn describe(dt: &Datatype) -> String {
} }
} }
match dt { match dt {
Datatype::Complex { size, base_type } => { Datatype::Complex { base_type, .. } => format!("complex<{}>", describe(base_type)),
describe(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::FixedPoint { Datatype::FixedPoint {
size, size,
signed, signed,
@@ -681,6 +691,74 @@ mod tests {
assert!(e.contains("not supported"), "{e}"); assert!(e.contains("not supported"), "{e}");
} }
#[test]
fn native_complex_reads_as_re_im_pairs() {
let z = [[1.0f64, -2.0], [0.5, 0.0], [-0.0, 3.25], [f64::MAX, 1e-300]];
let mut b = FileBuilder::new();
b.create_dataset("z")
.with_native_complex_f64_data(&z)
.with_shape(&[2, 2]);
b.create_dataset("z32")
.with_native_complex_f32_data(&[[1.5f32, -2.5]]);
let r = Reader::open(b.finish().unwrap()).unwrap();
let i = r.info("z").unwrap();
assert_eq!(
(i.shape, i.dtype, i.element_shape),
(vec![2, 2], "complex<f64>".to_string(), vec![2])
);
let all = r.read("z", None).unwrap();
assert_eq!(all.shape, vec![2, 2, 2]);
assert_eq!(all.data, Data::F64(z.iter().flatten().copied().collect()));
let slab = Hyperslab {
start: vec![1, 0],
count: vec![1, 2],
stride: None,
block: None,
};
let part = r.read("z", Some(&slab)).unwrap();
assert_eq!(part.shape, vec![1, 2, 2]);
assert_eq!(part.data, Data::F64(vec![-0.0, 3.25, f64::MAX, 1e-300]));
let z32 = r.read("z32", None).unwrap();
assert_eq!(
(z32.shape, z32.data),
(vec![1, 2], Data::F32(vec![1.5, -2.5]))
);
}
#[test]
fn native_complex_written_by_libhdf5_2() {
// Written by h5py 3.16 / libhdf5 2.0.0 (gen_native_complex.py).
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5/tests/fixtures/native_complex_hdf5_2.h5"
);
let r = Reader::open(std::fs::read(path).unwrap()).unwrap();
let be = r.read("c128be", None).unwrap();
assert_eq!(be.shape, vec![2, 3, 2]);
assert_eq!(be.data, r.read("c128", None).unwrap().data);
assert_eq!(r.info("c128be").unwrap().dtype, "complex<f64 (big-endian)>");
// binary16 parts widen to f32.
assert_eq!(
r.read("c32", None).unwrap().data,
Data::F32(vec![1.0, -2.0, 0.5, 65504.0, 0.0, -0.0])
);
// An array of complex: array dimensions, then the two parts.
let i = r.info("array").unwrap();
assert_eq!(
(i.dtype.as_str(), i.element_shape),
("array[2]<complex<f32>>", vec![2, 2])
);
assert_eq!(
r.read("array", None).unwrap().data,
Data::F32(vec![1.0, 2.0, 3.0, 4.0, -1.0, -1.0, 0.0, 0.0])
);
let s = r.read("scalar", None).unwrap();
assert_eq!((s.shape, s.data), (vec![2], Data::F64(vec![2.5, -0.5])));
// A compound holding a complex member is still refused.
let e = r.read("compound", None).unwrap_err();
assert!(e.contains("compound{z: complex<f64>, k: i64}"), "{e}");
}
#[test] #[test]
fn garbage_is_an_error() { fn garbage_is_an_error() {
assert!(Reader::open(vec![0u8; 64]).is_err()); assert!(Reader::open(vec![0u8; 64]).is_err());
+4 -1
View File
@@ -1342,7 +1342,10 @@ impl<'f> Dataset<'f> {
) -> Result<(Vec<T>, Vec<T>), Error> { ) -> Result<(Vec<T>, Vec<T>), Error> {
let dt = self.datatype()?; let dt = self.datatype()?;
let is_complex = match &dt { let is_complex = match &dt {
// Class 11 parses to this same `{r, i}` compound. Datatype::Complex { size, base_type } => {
matches!(base_type.as_ref(), Datatype::FloatingPoint { .. })
&& *size == 2 * base_type.type_size()
}
Datatype::Compound { size, members } => { Datatype::Compound { size, members } => {
matches!(members.as_slice(), [r, i] matches!(members.as_slice(), [r, i]
if r.name == "r" && i.name == "i" if r.name == "r" && i.name == "i"
+11
View File
@@ -41,6 +41,13 @@ pub enum DType {
Array(Box<DType>, Vec<u32>), Array(Box<DType>, Vec<u32>),
/// Variable-length UTF-8 string stored via a global heap reference. /// Variable-length UTF-8 string stored via a global heap reference.
VariableLengthString, VariableLengthString,
/// HDF5 2.0 native complex number (datatype class 11) over the given
/// floating-point part type: `Complex(F32)` is numpy `complex64`,
/// `Complex(F64)` `complex128`, and a half-precision part is
/// `Complex(Other("float16"))`. h5py's default complex encoding, the
/// compound `{r, i}`, stays a [`DType::Compound`]; `read_complex_f64`
/// and `read_complex_f32` read both.
Complex(Box<DType>),
/// Catch-all for HDF5 datatypes that do not map to a specific variant. /// Catch-all for HDF5 datatypes that do not map to a specific variant.
Other(String), Other(String),
} }
@@ -60,6 +67,7 @@ impl fmt::Display for DType {
DType::U64 => write!(f, "u64"), DType::U64 => write!(f, "u64"),
DType::String => write!(f, "string"), DType::String => write!(f, "string"),
DType::VariableLengthString => write!(f, "vlen_string"), DType::VariableLengthString => write!(f, "vlen_string"),
DType::Complex(base) => write!(f, "complex<{base}>"),
DType::Compound(fields) => { DType::Compound(fields) => {
write!(f, "compound{{")?; write!(f, "compound{{")?;
for (i, (name, dt)) in fields.iter().enumerate() { for (i, (name, dt)) in fields.iter().enumerate() {
@@ -148,6 +156,9 @@ pub(crate) fn classify_datatype(dt: &clawhdf5_format::datatype::Datatype) -> DTy
base_type, base_type,
dimensions, dimensions,
} => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()), } => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()),
Datatype::Complex { base_type, .. } => {
DType::Complex(Box::new(classify_datatype(base_type)))
}
_ => DType::Other(format!("{dt:?}")), _ => DType::Other(format!("{dt:?}")),
} }
} }
+75
View File
@@ -0,0 +1,75 @@
"""Generate native_complex_hdf5_2.h5: HDF5 2.0 native complex (class 11).
Datasets (every one class 11, or a type holding it):
c64 H5T_COMPLEX_IEEE_F32LE, 4 values incl. -0, inf, nan, subnormal
c128 H5T_COMPLEX_IEEE_F64LE, 2 x 3
c128be H5T_COMPLEX_IEEE_F64BE, the same values
c128x H5T_COMPLEX_IEEE_F64LE, c128 with element (1,1) changed
c32 H5T_COMPLEX_IEEE_F16LE, 3 values
scalar H5T_COMPLEX_IEEE_F64LE, scalar dataspace
compound {z: H5T_COMPLEX_IEEE_F64LE @0, k: H5T_STD_I64LE @16}
array H5T_ARRAY [2] of H5T_COMPLEX_IEEE_F32LE, 2 elements
plus the root attribute `zattr` (H5T_COMPLEX_IEEE_F64LE, 2 values).
Written with h5py 3.16 (libhdf5 2.0.0) through its low-level API, the file
type as the memory type (no conversion). Re-run only if the fixture ever
needs regenerating (generated 2026-09-29):
python gen_native_complex.py native_complex_hdf5_2.h5
The h5dump / h5ls 2.2.0 reference output the tests compare with was made
with the tools of the libhdf5 2.2.0 build described in gen_mx_floats.py:
h5dump native_complex_hdf5_2.h5 > native_complex_hdf5_2.ddl
h5ls -r -v native_complex_hdf5_2.h5 > native_complex_hdf5_2.ls
"""
import sys
import h5py
import numpy as np
from h5py import h5a, h5d, h5s, h5t
out = sys.argv[1]
f = h5py.File(out, "w", libver="latest")
def ds(name, tid, arr, shape=None):
shape = arr.shape if shape is None else shape
sp = h5s.create_simple(shape) if shape != () else h5s.create(h5s.SCALAR)
d = h5d.create(f.id, name.encode(), tid, sp)
d.write(h5s.ALL, h5s.ALL, np.ascontiguousarray(arr), mtype=tid)
nan, inf = float("nan"), float("inf")
c64 = np.array(
[1.5 - 2j, complex(0.0, -0.0), complex(inf, nan), complex(-inf, 1e-40)],
dtype=np.complex64,
)
ds("c64", h5t.COMPLEX_IEEE_F32LE, c64)
c128 = np.array(
[[1 + 2j, -3.5 + 4e300j, 0.1 + 0.2j], [5e-324 - 0j, 1j, -1]], dtype=np.complex128
)
ds("c128", h5t.COMPLEX_IEEE_F64LE, c128)
ds("c128be", h5t.COMPLEX_IEEE_F64BE, c128.astype(">c16"))
c128x = c128.copy()
c128x[1, 1] = 1 + 1j
ds("c128x", h5t.COMPLEX_IEEE_F64LE, c128x)
half = np.array([1.0, -2.0, 0.5, 65504.0, 0.0, -0.0], dtype=np.float16)
ds("c32", h5t.COMPLEX_IEEE_F16LE, half, shape=(3,))
ds("scalar", h5t.COMPLEX_IEEE_F64LE, np.array(2.5 - 0.5j), shape=())
ct = h5t.create(h5t.COMPOUND, 24)
ct.insert(b"z", 0, h5t.COMPLEX_IEEE_F64LE)
ct.insert(b"k", 16, h5t.STD_I64LE)
cd = np.array([(1 + 1j, 7), (-2.5 + 0j, -8)], dtype=[("z", "<c16"), ("k", "<i8")])
ds("compound", ct, cd)
at = h5t.array_create(h5t.COMPLEX_IEEE_F32LE, (2,))
ad = np.array([[1 + 2j, 3 + 4j], [-1 - 1j, 0j]], dtype=np.complex64)
ds("array", at, ad, shape=(2,))
a = h5a.create(f.id, b"zattr", h5t.COMPLEX_IEEE_F64LE, h5s.create_simple((2,)))
a.write(np.array([0.5 - 1.5j, 2j]), mtype=h5t.COMPLEX_IEEE_F64LE)
f.close()
@@ -0,0 +1,80 @@
HDF5 "native_complex_hdf5_2.h5" {
GROUP "/" {
ATTRIBUTE "zattr" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): 0.5-1.5i, 0+2i
}
}
DATASET "array" {
DATATYPE H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): [ 1+2i, 3+4i ], [ -1-1i, 0+0i ]
}
}
DATASET "c128" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128be" {
DATATYPE H5T_COMPLEX_IEEE_F64BE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128x" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 1+1i, -1+0i
}
}
DATASET "c32" {
DATATYPE H5T_COMPLEX_IEEE_F16LE
DATASPACE SIMPLE { ( 3 ) / ( 3 ) }
DATA {
(0): 1+-2i, 0.5+65504i, 0+-0i
}
}
DATASET "c64" {
DATATYPE H5T_COMPLEX_IEEE_F32LE
DATASPACE SIMPLE { ( 4 ) / ( 4 ) }
DATA {
(0): 1.5-2i, 0-0i, inf+nani, -inf+9.99995e-41i
}
}
DATASET "compound" {
DATATYPE H5T_COMPOUND {
H5T_COMPLEX_IEEE_F64LE "z";
H5T_STD_I64LE "k";
}
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): {
1+1i,
7
},
(1): {
-2.5+0i,
-8
}
}
}
DATASET "scalar" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SCALAR
DATA {
(0): 2.5-0.5i
}
}
}
}
Binary file not shown.
+19 -5
View File
@@ -1064,12 +1064,26 @@ fn complex_datasets_and_attributes_round_trip() {
let ds = file.dataset(name).unwrap(); let ds = file.dataset(name).unwrap();
assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}"); assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}");
assert_eq!(ds.shape().unwrap(), vec![3]); assert_eq!(ds.shape().unwrap(), vec![3]);
// Both encodings read as the same {r, i} compound. // h5py's encoding is a compound, the native one a complex type.
assert_eq!( let want = if name == "native64" {
ds.dtype().unwrap(), DType::Complex(Box::new(DType::F32))
} else {
DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)]) DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)])
); };
assert_eq!(ds.dtype().unwrap(), want, "{name}");
} }
assert_eq!(
file.dataset("native128").unwrap().dtype().unwrap(),
DType::Complex(Box::new(DType::F64))
);
assert_eq!(
file.dataset("native128").unwrap().raw_datatype().unwrap(),
clawhdf5::make_native_complex_f64_type()
);
assert_eq!(
DType::Complex(Box::new(DType::F64)).to_string(),
"complex<f64>"
);
for name in ["compound128", "native128"] { for name in ["compound128", "native128"] {
same( same(
file.dataset(name).unwrap().read_complex_f64().unwrap(), file.dataset(name).unwrap().read_complex_f64().unwrap(),
@@ -1094,7 +1108,7 @@ fn complex_datasets_and_attributes_round_trip() {
data, data,
}) => { }) => {
assert!(shape.is_empty()); assert!(shape.is_empty());
assert_eq!(datatype.type_size(), 16); assert_eq!(datatype, clawhdf5::make_native_complex_f64_type());
let fields = let fields =
clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap(); clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap();
let part = |k: usize| { let part = |k: usize| {
+36 -11
View File
@@ -315,15 +315,12 @@ wrong data.
decoded, and external references are an error (object references decoded, and external references are an error (object references
decode). A multi-dimensional numeric attribute is returned as a flat decode). A multi-dimensional numeric attribute is returned as a flat
array (its shape is not reported; `AttrValue::Raw` carries the shape). array (its shape is not reported; `AttrValue::Raw` carries the shape).
HDF5 2.0's native complex type (class 11) is read as the equivalent HDF5 2.0's native complex type (class 11) is read as its own type since
`{r, i}` compound: values are right (`Dataset::read_complex_f64`, and 2026-09-29 ([fixed](#hdf5-20-native-complex-numbers-read-as-a-r-i-compound));
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs writing it is opt-in (`with_native_complex_f64_data`,
ls`/`dump` and the browser reader show a compound where h5dump 2.x `make_native_complex_f64_type`), and the Python bindings write complex
prints `H5T_COMPLEX_IEEE_F64LE`, so `h5rs dump` of such a file does not arrays as h5py's compound only. `h5rs dump --json` writes it as the
match h5dump 2.x (h5dump 1.14 cannot read it at all). Writing class 11 `{r, i}` compound: hdf5-json (h5json 2.0.0) has no complex class.
(added 2026-09-28) is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`); the Python bindings write complex
arrays as h5py's compound only.
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in - **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
that: libhdf5 fails only the first metadata read of an image it cannot that: libhdf5 fails only the first metadata read of an image it cannot
load and then reads the file's own (possibly stale) metadata, where we load and then reads the file's own (possibly stale) metadata, where we
@@ -533,8 +530,10 @@ browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27.
again retries it; what was fetched stays cached). again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against - Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included; not in Firefox or Safari. a local server, cross-origin included; not in Firefox or Safari.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are - Compound (h5py's complex numbers, a compound `{r, i}`, included),
refused with an error naming the type; attributes of those types come back reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type (HDF5 2.0 native complex datasets
read, as `[re, im]` pairs); attributes of those types come back
as `value: null` with their `dtype`. (VL strings read, with h5py's as `value: null` with their `dtype`. (VL strings read, with h5py's
values, through the same `VlResolver` as `File` and `h5rs`.) values, through the same `VlResolver` as `File` and `h5rs`.)
- No Zstd or SZIP (both link C): such datasets fail with - No Zstd or SZIP (both link C): such datasets fail with
@@ -669,6 +668,32 @@ netCDF4-python-written files with netCDF-C itself, and
`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in `corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in
[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)). [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)).
## HDF5 2.0 native complex numbers read as a `{r, i}` compound
**Status:** fixed 2026-09-29 (branch `feat/complex-first-class`; PR not
yet opened). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0)
and `main` until then. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
A class-11 datatype (`H5T_COMPLEX_IEEE_F64LE`, ...) parsed into the
equivalent compound `{r, i}`. Values were right (`read_complex_f64`,
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs ls`/`dump`
and the browser reader showed a compound, so `h5rs dump` did not match
h5dump 2.x, the browser refused the dataset, and a parsed type written
back out became a compound. `Datatype::parse` now returns
`Datatype::Complex` (in compounds, arrays and VL types too),
`Dataset::dtype()` is `DType::Complex(..)`, `h5rs` prints what h5dump,
h5ls and h5diff 2.2.0 print (checked on tank, 2026-09-29, against a file
written by h5py 3.16 / libhdf5 2.0.0:
`cargo test -p clawhdf5-tools --test native_complex_dump`), and the
browser reads `[re, im]` pairs.
**What users must do:** code that matched a native complex type as a
`Datatype::Compound` (or a `DType::Compound` of `r`/`i`) must match
`Datatype::Complex` / `DType::Complex`, or take the compound view with
`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a
`Complex` arm. Files are unchanged; nothing needs rewriting.
## NetCDF-4: variables' dimensions are guessed from sizes ## NetCDF-4: variables' dimensions are guessed from sizes
+6 -2
View File
@@ -123,8 +123,12 @@ else throws an `Error` naming the type.
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32). `readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in Every response body is cut off past the length asked for. More in
`docs/known-issues.md`. `docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are - HDF5 2.0 native complex datasets read as `[re, im]` pairs: `dtype`
refused with an error. Attributes of those types are listed with `complex<f64>`, `elementShape` ending in `2`, the parts interleaved in
the typed array.
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque and variable-length-sequence datasets are refused with
an error. Attributes of those types are listed with
`value: null` and their `dtype`. `value: null` and their `dtype`.
- No Zstd or SZIP filters (they link C): such a dataset fails with - No Zstd or SZIP filters (they link C): such a dataset fails with
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and `unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
+6
View File
@@ -93,6 +93,12 @@ expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>' expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
# 3-D: leading dimension held at 0, window over the last two. # 3-D: leading dimension held at 0, window over the last two.
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0' expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
# HDF5 2.0 native complex (written when h5py's libhdf5 is 2.0 or later):
# each cell is the [re, im] pair.
if "$PY" -c 'import sys, h5py; sys.exit(not getattr(h5py.get_config(), "has_native_complex", False))'; then
expect $H5 "/native_c128" '<dd>complex&lt;f64&gt;</dd>' '<dd>(3, 2)</dd>' '<td>[1, -1.5]</td>' \
'<td>[5, -7.5]</td>'
fi
# Unsupported type: an error, not values. # Unsupported type: an error, not values.
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported' expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that # Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
+22 -3
View File
@@ -80,6 +80,15 @@ with h5py.File(h5, "w") as f:
**hdf5plugin.Zstd()) **hdf5plugin.Zstd())
comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")]) comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")])
f.create_dataset("table", data=comp) f.create_dataset("table", data=comp)
if getattr(h5py.get_config(), "has_native_complex", False):
# HDF5 2.0 native complex (class 11): read as [re, im] pairs.
from h5py import h5d, h5s, h5t
for name, t, dt in [(b"native_c128", h5t.COMPLEX_IEEE_F64LE, "<c16"),
(b"native_c64_be", h5t.COMPLEX_IEEE_F32BE, ">c8")]:
z = (np.arange(6) - 1.5j * np.arange(6)).reshape(3, 2).astype(dt)
d = h5d.create(f.id, name, t, h5s.create_simple(z.shape))
d.write(h5s.ALL, h5s.ALL, z, mtype=t)
g = f.create_group("sensors") g = f.create_group("sensors")
g.attrs["location"] = "lab" g.attrs["location"] = "lab"
g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4")) g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4"))
@@ -105,6 +114,8 @@ def kind(dt):
return "strings" return "strings"
if dt.subdtype is not None: if dt.subdtype is not None:
return kind(dt.subdtype[0]) return kind(dt.subdtype[0])
if dt.kind == "c":
return "f64" if dt.itemsize == 16 else "f32"
if dt.kind == "f": if dt.kind == "f":
return "f64" if dt.itemsize == 8 else "f32" return "f64" if dt.itemsize == 8 else "f32"
if dt.kind in "iu": if dt.kind in "iu":
@@ -120,20 +131,25 @@ def flat(a, k):
return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()] return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()]
if k.startswith(("i", "u")): if k.startswith(("i", "u")):
return [str(int(x)) for x in a.ravel()] return [str(int(x)) for x in a.ravel()]
if a.dtype.kind == "c":
# [re, im] pairs, in order.
a = np.stack([a.real, a.imag], axis=-1)
return [float(x) for x in a.ravel()] return [float(x) for x in a.ravel()]
def entry(ds, slab=None): def entry(ds, slab=None):
k = kind(ds.dtype) k = kind(ds.dtype)
data = ds[()] data = ds[()]
e = {"kind": k, "shape": list(np.shape(data)), "values": flat(data, k)} # A complex element reads as its two parts: one more dimension.
pair = [2] if ds.dtype.kind == "c" else []
e = {"kind": k, "shape": list(np.shape(data)) + pair, "values": flat(data, k)}
if slab: if slab:
start, count, stride = slab start, count, stride = slab
idx = tuple(slice(s, s + (c - 1) * st + 1, st) idx = tuple(slice(s, s + (c - 1) * st + 1, st)
for s, c, st in zip(start, count, stride)) for s, c, st in zip(start, count, stride))
part = ds[idx] part = ds[idx]
e["slab"] = {"start": start, "count": count, "stride": stride, e["slab"] = {"start": start, "count": count, "stride": stride,
"shape": list(part.shape), "values": flat(part, k)} "shape": list(part.shape) + pair, "values": flat(part, k)}
return e return e
@@ -176,7 +192,10 @@ def describe(path):
"datasets": sorted(sets)} "datasets": sorted(sets)}
for n, o in members.items(): for n, o in members.items():
walk(key.rstrip("/") + "/" + n, o) walk(key.rstrip("/") + "/" + n, o)
elif obj.dtype.names: elif obj.dtype.names or (
obj.dtype.kind == "c"
and obj.id.get_type().get_class() != getattr(h5py.h5t, "COMPLEX", None)):
# h5py's own complex numbers are a compound {r, i}: refused.
expected["errors"][key] = "compound" expected["errors"][key] = "compound"
else: else:
expected["datasets"][key] = entry(obj, slab_for(obj)) expected["datasets"][key] = entry(obj, slab_for(obj))
+2 -2
View File
@@ -399,7 +399,7 @@ async function remoteTests() {
await fails(async () => { await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, { const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512, blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })), fetch: tamper(async (r) => withHeaders(r, { "Content-Range": `bytes 0-511/${statSync(join(fixDir, "fixture.h5")).size}` })),
}); });
await f.read("/grid"); await f.read("/grid");
}, /the server sent 0-511/, "wrong range"); }, /the server sent 0-511/, "wrong range");
@@ -550,7 +550,7 @@ async function floodTests() {
const range = new Headers(init.headers).get("Range"); const range = new Headers(init.headers).get("Range");
if (range === "bytes=0-511") return fetch(url, init); if (range === "bytes=0-511") return fetch(url, init);
const m = /^bytes=(\d+)-(\d+)$/.exec(range); const m = /^bytes=(\d+)-(\d+)$/.exec(range);
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } }); return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/${statSync(join(fixDir, "fixture.h5")).size}` } });
}, },
}).catch((e) => e); }).catch((e) => e);
// The open itself may need a second range: flooded either way. // The open itself may need a second range: flooded either way.