Chunk dimensions of 2^32 or more; copy-free unfiltered 4 GiB writes #32

Open
osobh wants to merge 7 commits from feat/huge-chunk-dims into main
55 changed files with 4599 additions and 880 deletions
+36
View File
@@ -601,6 +601,42 @@ about 150 per dataset here) against a Fixed Array's few blocks. Writing costs th
same; files grow by about 36 bytes per chunk. The default therefore stays
the 1.10 format; the 1.8 format is opt-in.
### Streaming `FileBuilder::write` (2026-09-29, tank)
`FileBuilder::write` now streams the file through a buffered writer
instead of building it in memory first (`feat/huge-chunk-dims`, the change
that removes copies of 4 GiB unfiltered chunks). Measured 2026-09-29 on
tank (AMD Ryzen 7 7800X3D), idle (1-minute load average 1.66 to 1.83 at the
start of each run): `main` at `4260af4` against the stacked branch at
`549e442`, alternating, three full runs each; median (min-max) of
criterion's estimate, ms.
> **Run:** `cargo bench -p clawhdf5-bench --bench h5bench_write -- --noplot --warm-up-time 1 --measurement-time 3`
The first candidate used a 1 MiB write buffer and was 1.35x to 1.83x slower
on every 512 x 512 (1 MiB) chunked case when the full suite ran (e.g.
`write_2d_chunked/512x512` 0.454 -> 0.754 ms): the buffer is above glibc's
128 KiB mmap threshold, and after the smaller cases it was mapped afresh on
each write (1 629 368 minor page faults over the `write_2d_chunked` group
against 12 888 on `main`). Run alone, the case showed no difference. With a
64 KiB buffer (`549e442`):
| benchmark | main | candidate | ratio |
|---|---:|---:|---:|
| write_1d_contiguous/100000 | 0.216 (0.215-0.218) | 0.206 (0.204-0.207) | 0.96x |
| write_2d_chunked/32x32 | 0.062 (0.061-0.063) | 0.061 (0.061-0.064) | 0.99x |
| write_2d_chunked/128x128 | 0.102 (0.100-0.103) | 0.101 (0.099-0.102) | 0.98x |
| write_2d_chunked/512x512 | 0.460 (0.457-0.463) | 0.466 (0.463-0.477) | 1.01x |
| write_2d_chunked_zstd/deflate-6/512x512 | 0.457 (0.450-0.465) | 0.465 (0.458-0.477) | 1.02x |
| write_2d_chunked_zstd/zstd-3/512x512 | 0.376 (0.369-0.381) | 0.387 (0.379-0.390) | 1.03x |
| write_2d_chunked_pcodec/pcodec/512x512 | 0.906 (0.895-0.908) | 0.901 (0.898-0.914) | 0.99x |
| write_multi_dataset/64 | 0.136 (0.135-0.138) | 0.136 (0.135-0.136) | 1.00x |
| write_with_attrs/64 | 0.063 (0.062-0.063) | 0.063 (0.062-0.063) | 1.00x |
The other 18 cases are within 2%. Minor page faults per full run: 59 602 to
64 486 on both sides. The write path is as fast as before; what changes is
peak memory for large unfiltered chunks (see CHANGELOG, 2026-09-29).
### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank)
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load
+176
View File
@@ -2,6 +2,182 @@
## Unreleased
### NetCDF-4: phony dimensions, skipped types and order as in netCDF-C (2026-09-29)
- `clawhdf5-netcdf4` now reads a file's metadata as netCDF-C 4.9.3 does
(`libhdf5/hdf5open.c`), for the whole file on first use
(`src/model.rs`): links in creation order when the group tracks it,
else in name order; a group's datasets before its subgroups; dimension
ids file-wide (`_Netcdf4Dimid`, else the next free id); then variables'
dimensions, subgroups first: `_Netcdf4Coordinates` ids (looked up
file-wide), else the scales `DIMENSION_LIST` attaches (when the first
axis has one), else phony dimensions.
- **Files without dimension scales** get netCDF-C's phony dimensions,
`phony_dim_<id>` (`create_phony_dims`): per axis, the first dimension of
the variable's group of the same length and unlimitedness not used by an
earlier axis of the variable (real dimensions included), else a new one;
a length of 0 is always unlimited. They used to be one dimension per 1-D
dataset, named after it, and `dim_<size>` for other axes.
- **Datasets netCDF-C skips are not variables:** references, bit fields,
time and array types, and compounds, enums and VLENs whose members or
base type are not netCDF atomic types (4- and 8-byte floats only) or a
type read before — replaying netCDF-C's file-wide type list, which also
keeps a type it failed to read, so the second dataset of a compound with
a reference member is a variable, as in netCDF-C.
- `NcType` gains `Enum`, `Compound`, `VLen` and `Opaque` and is now
`#[non_exhaustive]` (breaking for exhaustive matches);
`Variable::nc_type` is the type netCDF-C gives the variable: a 1-byte
fixed-length string is `Char`, a longer one `String` (both were
`String`); user-defined types have their class (they were `Char`).
`dtype_to_nctype` maps compounds and enums to their classes.
- Groups (`group_names`), variables (`variables`, `variable_names`) and
dimensions (`dimensions`, by id) come in netCDF-C's order; a zero-length
dimension scale is unlimited; an unlimited dimension's length is the
longest extent along it of the variables in its group and below
(`nc4_find_dim_len`; was: of the variables its `REFERENCE_LIST` names).
`NetCDF4File::group` and `NetCDF4Group::group` take a path (`"a/b"`).
- Deliberate differences: floats of other than 4 or 8 bytes are
`Float`/`Double` (netCDF-C 4.9.3 on libhdf5 1.14.6 labels them
`NC_STRING`); an axis netCDF-C leaves without a dimension (where it
reads uninitialised memory) gets one by the phony rule.
- New `clawhdf5_format::group_v2::links_in_creation_order_in`: a group's
link names in creation order, or `None` when it does not track it.
- Tests compare with netCDF-C itself (`tests/netcdf_c_view.py` calls the
libnetcdf netCDF4-python bundles through ctypes, since netCDF4-python
hides variables of types it does not support): `interop_tests` cases of
h5py files without dimension scales (sharing, unlimited and zero-length
axes, subgroups, a real dimension taken by length), of every HDF5 type
class, of creation and name order in compact and dense groups, and a
netCDF4-python file; and the gated `tests/corpus_vs_netcdf_c.rs`
(`CLAWHDF5_NETCDF_CORPUS=<dir>`): over the conformance corpus 420 of the
429 files netCDF-C 4.9.3 opens match in groups, dimensions, variables,
types, shapes and numeric values (up to 5000 elements); `main` at
`4260af4` matched 68. The other 9 (external links, values the HDF5
reader refuses) are explained in `tests/corpus_known_differences.txt`
and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4).
Affected v2.1.0 to v2.7.0.
### HDF5 2.0 native complex is its own type on read; Python `libver=` (2026-09-29)
- **`Datatype::parse` returns `Datatype::Complex { size, base_type }`** for
a class-11 message, also as a compound member, array base or
variable-length base. It used to return the equivalent `{r, i}`
compound (`docs/known-issues.md`, "HDF5 2.0 native complex numbers read
as a `{r, i}` compound"). **Breaking** for code that matched the
compound view of a native complex type: `Dataset::raw_datatype()` and
attribute datatypes are now `Datatype::Complex`;
`Datatype::complex_as_compound(size, base)` still gives the `{r, i}` view,
and `data_read::read_compound_fields` accepts a `Complex` directly.
Re-serializing a parsed type (e.g. copying it to another file) now writes
class 11 again instead of a compound.
- **Facade:** `DType::Complex(Box<DType>)` (new variant, **breaking** for
exhaustive `match`es over `DType`): `Complex(F32)` is numpy `complex64`,
`Complex(F64)` `complex128`, a binary16 part `Complex(Other("float16"))`
(the same `Other` name small floats already use); `Display` prints
`complex<f64>`. h5py's own complex encoding, the compound `{r, i}`, is
still `DType::Compound`. `read_complex_f64`/`read_complex_f32` read both.
- **`h5rs`:** `dump` prints native complex as h5dump 2.2.0 does:
`H5T_COMPLEX_IEEE_F{16,32,64}{LE,BE}` (else `H5T_COMPLEX { <base> }`), in
arrays and compounds too, and values as `1.5-2i` (`%g%+gi`; for binary16
parts, which h5dump has no C type for, `1+-2i` as it prints them). `ls`
shows `complex64`, `complex128-be`, `complex32` (and `complex<part>` for
non-IEEE parts); `ls -v` shows h5ls 2.2.0's `complex number of` /
`IEEE 64-bit big-endian float`. `diff` compares complex values part by
part and prints the difference as h5diff 2.2.0 does (`1+0i`); a native
complex and an `{r, i}` compound are no longer comparable (different
classes, as in h5diff). `dump --json` keeps the `{r, i}` compound and
`[re, im]` values: hdf5-json (h5json 2.0.0) has no complex class.
- **Browser (`clawhdf5-wasm`):** native complex datasets are readable, as
`[re, im]` pairs: `info().elementShape` ends in `2`, `read()` returns the
parts interleaved in a `Float32Array`/`Float64Array` (binary16 parts
widened to f32), and `dtype` is `complex<f64>` (`complex<f32
(big-endian)>`, `array[2]<complex<f32>>`). The viewer shows each element
as `[re, im]` with no change to the page. h5py's `{r, i}` compound is
still refused, as every compound is.
- **Python:** `clawhdf5.File(path, 'w', libver=...)` takes h5py's values:
`'v108'`, `'v110'`, `'v112'`, `'v114'`, `'v200'`, `'latest'` (the low
bound; high `'latest'`, as h5py) or a `(low, high)` tuple, mapped to
`FileBuilder::libver_bounds`. `'earliest'` as the low bound writes the
1.8 format with a `UserWarning` (clawhdf5 cannot write the pre-1.8
format); as the high bound it is a `ValueError`, as are unknown names and
a low bound above the high one. It is ignored for `'r'` and refused
(`NotImplementedError`) for `'r+'`/`'a'`. Native complex already read as
numpy `complex64`/`complex128` and still does
(`test_read_native_complex_from_h5py`).
- Tests (tank, 2026-09-29): `crates/clawhdf5-tools/tests/native_complex_dump.rs`
compares `h5rs dump` with h5dump 2.2.0's output stored next to a fixture
written by h5py 3.16 / libhdf5 2.0.0
(`crates/clawhdf5/tests/fixtures/gen_native_complex.py`:
`F32LE`/`F64LE`/`F64BE`/`F16LE`, scalar, compound member, array, root
attribute), line for line outside the values and value for value within
them, and `ls`/`diff` with h5ls/h5diff 2.2.0;
`integration_tests::complex_datasets_and_attributes_round_trip`
(`DType::Complex`, `raw_datatype()`); `clawhdf5-wasm` unit tests on the
fixture, `h5py_interop` and `examples/wasm-viewer/test` (Node:
`test.mjs`, 274 + 1324 checks; the page in headless Chromium for
`/native_c128`, `/native_c64_be`, `/pairs`) with two native complex
datasets added to `make_fixture.py` (when h5py's libhdf5 is 2.0+);
`crates/clawhdf5-py/tests/test_libver.py` (every bound; `'v108'` output
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
hard-codes the fixture's length.
### Chunk dimensions of 2^32 or more; 4 GiB chunks written without copies (2026-09-29)
- **Chunk dimensions of 2^32 or more** (libhdf5 2.x writes them with
layout message version 5, up to 8 bytes per dimension; a 4 GiB chunk of
1-byte elements needs one) are read and written. They were refused
(`InvalidChunkDimensions`). **Public API changes:**
`DataLayout::Chunked::chunk_dimensions` is `Vec<u64>` (was `Vec<u32>`),
and the `chunk_dimensions` argument of `chunked_read::{collect_chunk_info_checked,
collect_chunk_info_checked_in, generate_implicit_chunks_in_grid}`,
`fixed_array::{read_fixed_array_chunks, read_fixed_array_chunks_in}`,
`extensible_array::{read_extensible_array_chunks,
read_extensible_array_chunks_in}`, and the `chunk_dims` argument of
`chunked_write::serialize_v4_single_chunk_pub`, are `&[u64]` (were
`&[u32]`). A chunk whose size overflows 64 bits is refused when a dataset is read
(`InvalidChunkDimensions`), and the writer refuses a chunk dimension of
0 up front. A version-3 layout (4-byte dimensions) is never written for
such a chunk: over 4 GiB, it always takes version 5, as in libhdf5.
- **Writing without copies:** `FileWriter` no longer copies chunks into a
buffer per pass and then into the file's buffer. A chunk that is a
contiguous run of the dataset's data (a dataset stored as one chunk of
its shape, blocks of whole rows) is borrowed when unfiltered and
compressed straight from it when filtered, and contiguous datasets are
not copied. New `FileWriter::finish_with(put)` hands the file out piece
by piece; `FileBuilder::write` uses it (through a 1 MiB `BufWriter`,
still to a temporary file renamed into place), so a file is never
assembled in memory. New `DatasetBuilder::with_u8_data_owned(Vec<u8>)`
takes the data without the copy `with_u8_data` makes. Peak resident
memory of one unfiltered 1 GiB chunk of `u8`
(`cargo run --release -p clawhdf5 --example write_one_chunk --
1073741824 OUT owned|slice`, tank, 2026-09-29): 5.0 GiB before (the
writer of `4260af4`), 1.0 GiB after with `owned`; 6.0 and 2.0 GiB with
`slice`; 4 GiB + 7 bytes `owned`, 4.00 GiB. Same bytes written. Write
speed was not re-measured (the machine was not idle): the A/B of
`FileBuilder::write` and `finish` on small and large files is still to
be run.
- `chunked_write::extract_chunk` (behind `split_into_chunks`) no longer
panics when the data is shorter than the shape; the missing part is
zeros, as it was meant to be.
- Tests: `crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5` (31 KB,
libhdf5 2.0.0 via h5py 3.16, `gen_huge_chunks.py dims`: `u8` chunks of
2^32 + 7 deflated twice, Single Chunk and Extensible Array), and in
`huge_chunks_interop`: `huge_chunk_dims_list` (always), and with
`CLAWHDF5_HUGE_CHUNKS=1` `huge_chunk_dims_read`,
`huge_chunk_dims_unfiltered_read` (a sparse file h5py writes, Single
Chunk and Fixed Array, mmap and positioned reads),
`writer_huge_chunk_dims_round_trip` (deflated, read back by clawhdf5,
h5py 3.16 and — layouts — h5dump 2.2.0),
`writer_unfiltered_huge_chunk_round_trip` (an unfiltered 4 GiB + 7 byte
chunk: clawhdf5, h5py, and h5dump 2.2.0 printing values past 2^32) and
`writer_huge_chunks_lz4_zstd_round_trip` (with the `lz4`/`zstd`
features; h5py through hdf5plugin 7.1.0). The whole opt-in suite
(`cargo test --release -p clawhdf5 --features lz4,zstd --test
huge_chunks_interop -- --test-threads=1`, h5py and h5dump 2.2.0
included) passed with a peak of 10.04 GiB resident (tank, 2026-09-29,
commit `0065c6b`). Unit tests: 5-byte dimension encoding, contiguous
chunk detection, `finish_with` against `finish`. The wasm package test
lists the new fixture and gets `FormatError::Overflow` reading it.
`docs/known-issues.md`: two fixed entries; the open "Chunks of 4 GiB or
more: limits" keeps the refused filters, the editor and memory.
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
+6 -5
View File
@@ -102,7 +102,8 @@ writes 85.2 µs vs 877 µs (10.3x); 64 group creates 130 µs vs 1.37 ms
(10.6x); a 512×512 `f32` chunked deflate-6 write 1.44 ms vs 65.0 ms
(re-measured 2026-09-23 with the pure-Rust deflate: 1.46 ms vs 51.4 ms,
35x); a 100K `f32` sequential write is a tie. The writer (`FileBuilder`)
assembles a file in memory and writes it once, which is part of that
lays out the whole file before writing it (when these were measured, in
one buffer in memory), which is part of that
difference; read the caveats in [BENCHMARKS.md](BENCHMARKS.md#caveats)
before quoting these.
@@ -115,12 +116,12 @@ Limits and open issues, with dates, are in
|---|---|---|---|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more and chunk dimensions of 2^32 or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error) |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
| **Bindings** | Python (read, `'w'` for numeric and complex arrays with h5py's `libver=`, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`; HDF5 2.0 native complex as `[re, im]` pairs); no Zstd/SZIP/pcodec, no compound (h5py's `{r, i}` complex included), reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`,
`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure
@@ -386,7 +387,7 @@ filter).
| `clawhdf5-filters` | Deflate backends (zlib-rs default, zlib-ng, Apple Compression) |
| `clawhdf5-io` | I/O helpers: mmap, async, an HSDS client, `mpi-io` (not collective I/O) |
| `clawhdf5-remote` | HTTP(S) and object-store files through a block cache |
| `clawhdf5-netcdf4` | NetCDF-4 dimensions, variables, CF attributes |
| `clawhdf5-netcdf4` | NetCDF-4 dimensions, variables, CF attributes, as netCDF-C reports them (phony dimensions included) |
| `clawhdf5-derive` | Derive macros for HDF5-serialisable structs |
| `clawhdf5-tools` | `h5rs`: `ls`, `dump`, `stat`, `diff`, `check` |
| **Bindings** | |
+42 -30
View File
@@ -555,7 +555,7 @@ pub(crate) fn checked_byte_len(elements: u64, elem_size: usize) -> Result<usize,
/// HDF5 2.0 writes them with layout version 5. A zero chunk dimension used
/// to read as all fill values, and a huge one to hang the reader.
pub(crate) fn chunk_geometry(
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
layout_version: u8,
dataspace: &Dataspace,
elem_size: usize,
@@ -576,15 +576,26 @@ pub(crate) fn chunk_geometry(
"chunk size must be > 0, dim = {d}"
)));
}
let bytes = spatial
.iter()
.fold(elem_size as u128, |acc, &c| acc * u128::from(c));
// Saturating: 33 dimensions of up to 64 bits each can overflow u128.
let bytes = spatial.iter().fold(elem_size as u128, |acc, &c| {
acc.saturating_mul(u128::from(c))
});
// So the indexes can compute a chunk's size in 64 bits.
if bytes > u128::from(u64::MAX) {
return Err(FormatError::InvalidChunkDimensions(format!(
"chunk size overflows 64 bits (chunk {spatial:?} of {elem_size}-byte elements)"
)));
}
if layout_version < 4 && bytes > u128::from(u32::MAX) {
return Err(FormatError::InvalidChunkDimensions(format!(
"chunk size must be < 4GB with v1 b-tree index (chunk {spatial:?} of {elem_size}-byte elements)"
)));
}
Ok((rank, spatial.iter().map(|&c| c as usize).collect()))
let spatial = spatial
.iter()
.map(|&c| to_usize(c))
.collect::<Result<Vec<usize>, _>>()?;
Ok((rank, spatial))
}
/// The size of one element of `dt` as stored in the file: a
@@ -625,7 +636,7 @@ pub fn check_chunk_element_size(
return Ok(());
};
let expected = stored_element_size(datatype, offset_size);
if u64::from(stored) != expected {
if stored != expected {
return Err(FormatError::InvalidChunkDimensions(format!(
"stored datatype size in chunk layout does not match datatype description \
(layout {stored} bytes, datatype {expected})"
@@ -759,7 +770,7 @@ pub fn collect_chunk_info_in<S: Storage + ?Sized>(
pub fn collect_chunk_info_checked(
file_data: &[u8],
btree_address: u64,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
@@ -776,7 +787,7 @@ pub fn collect_chunk_info_checked(
pub fn collect_chunk_info_checked_in<S: Storage + ?Sized>(
file_data: &S,
btree_address: u64,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
offset_size: u8,
length_size: u8,
) -> Result<Vec<ChunkInfo>, FormatError> {
@@ -808,7 +819,7 @@ pub fn collect_chunk_info_checked_in<S: Storage + ?Sized>(
.iter_mut()
.zip(chunk.offsets.iter().zip(chunk_dimensions))
{
*w = o / u64::from(d);
*w = o / d;
}
wanted[ndims - 1] = 0;
if let Some(i) = root.find(&wanted) {
@@ -849,9 +860,9 @@ impl ChunkNode {
}
}
fn scale_keys(&mut self, dims: &[u32]) {
fn scale_keys(&mut self, dims: &[u64]) {
for (k, &d) in self.keys.iter_mut().zip(dims.iter().cycle()) {
*k /= u64::from(d);
*k /= d;
}
if let ChunkChildren::Nodes(nodes) = &mut self.children {
for n in nodes {
@@ -919,9 +930,9 @@ fn btree_cmp3(lt: &[u64], scaled: &[u64], rt: &[u64]) -> core::cmp::Ordering {
/// Check one v1 B-tree chunk key's offsets (see
/// [`collect_chunk_info_checked`]).
fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u32]) -> Result<(), FormatError> {
fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u64]) -> Result<(), FormatError> {
for (&offset, &dim) in offsets.iter().zip(chunk_dimensions) {
if dim == 0 || offset % u64::from(dim) != 0 {
if dim == 0 || offset % dim != 0 {
return Err(FormatError::ChunkedReadError(format!(
"bad coordinate offset {offsets:?} for chunk dimensions {chunk_dimensions:?}"
)));
@@ -937,7 +948,7 @@ fn read_key_offsets(
file_data: &[u8],
pos: usize,
ndims: usize,
chunk_dimensions: Option<&[u32]>,
chunk_dimensions: Option<&[u64]>,
out: &mut Vec<u64>,
) -> Result<(), FormatError> {
let start = out.len();
@@ -966,7 +977,7 @@ fn parse_chunk_node<S: Storage + ?Sized>(
file_data: &S,
btree_address: u64,
ndims: usize,
chunk_dimensions: Option<&[u32]>,
chunk_dimensions: Option<&[u64]>,
offset_size: u8,
depth: usize,
stored: &mut Vec<ChunkInfo>,
@@ -1075,7 +1086,7 @@ fn parse_chunk_node<S: Storage + ?Sized>(
pub fn generate_implicit_chunks(
base_address: u64,
dataset_dims: &[u64],
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
) -> Vec<ChunkInfo> {
generate_implicit_chunks_in_grid(
@@ -1097,17 +1108,18 @@ pub fn generate_implicit_chunks_in_grid(
base_address: u64,
dataset_dims: &[u64],
max_dims: &[u64],
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
) -> Vec<ChunkInfo> {
let rank = chunk_dimensions.len();
let chunk_byte_size: u64 =
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
let chunk_byte_size: u64 = chunk_dimensions
.iter()
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
let mut num_chunks_per_dim = Vec::with_capacity(rank);
let mut grid_per_dim = Vec::with_capacity(rank);
for d in 0..rank {
let ch = chunk_dimensions[d] as u64;
let ch = chunk_dimensions[d];
let n = dataset_dims[d].div_ceil(ch);
num_chunks_per_dim.push(n);
grid_per_dim.push(max_dims.get(d).map_or(n, |m| m.div_ceil(ch)).max(n));
@@ -1125,7 +1137,7 @@ pub fn generate_implicit_chunks_in_grid(
let nchunks = num_chunks_per_dim[d];
let chunk_idx = remaining % nchunks;
remaining /= nchunks;
offsets[d] = chunk_idx * chunk_dimensions[d] as u64;
offsets[d] = chunk_idx * chunk_dimensions[d];
grid_idx = grid_idx.saturating_add(chunk_idx.saturating_mul(down));
down = down.saturating_mul(grid_per_dim[d]);
}
@@ -1350,7 +1362,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
}
(4, Some(2)) => {
// Implicit index — use spatial chunk dims only
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
generate_implicit_chunks_in_grid(
addr,
&dataspace.dimensions,
@@ -1364,7 +1376,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
}
(4, Some(3)) => {
// Fixed Array — use spatial chunk dims only
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
let header = FixedArrayHeader::parse_in(
file_data,
checked_addr(addr)?,
@@ -1384,7 +1396,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
}
(4, Some(4)) => {
// Extensible Array — use spatial chunk dims only
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
let header = ExtensibleArrayHeader::parse_in(
file_data,
checked_addr(addr)?,
@@ -2435,12 +2447,12 @@ mod tests {
/// A leaf whose final key is the one libhdf5 writes past the last chunk:
/// each coordinate of the last chunk plus its chunk dimension (the
/// element size last).
fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u32], offset_size: u8) -> Vec<u8> {
fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u64], offset_size: u8) -> Vec<u8> {
let last = &chunks.last().expect("a chunk").offsets;
let end: Vec<u64> = dims
.iter()
.enumerate()
.map(|(d, &c)| last.get(d).copied().unwrap_or(0) + u64::from(c))
.map(|(d, &c)| last.get(d).copied().unwrap_or(0) + c)
.collect();
build_chunk_btree_leaf_to(chunks, &end, offset_size)
}
@@ -2771,7 +2783,7 @@ mod tests {
}
// Build B-tree at offset 0x100
let dims = [chunk_size_elems as u32, elem_size as u32];
let dims = [chunk_size_elems as u64, elem_size as u64];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
@@ -3145,7 +3157,7 @@ mod tests {
data_offset += compressed.len() + 16; // some padding
}
let dims = [chunk_elems as u32, elem_size as u32];
let dims = [chunk_elems as u64, elem_size as u64];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
@@ -3228,7 +3240,7 @@ mod tests {
}
}
let dims = [chunk_dims[0] as u32, chunk_dims[1] as u32, elem_size as u32];
let dims = [chunk_dims[0] as u64, chunk_dims[1] as u64, elem_size as u64];
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
let btree_addr = 0x100usize;
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
@@ -3408,7 +3420,7 @@ mod tests {
}
let layout = DataLayout::Chunked {
chunk_dimensions: vec![chunk_elems as u32, elem_size as u32],
chunk_dimensions: vec![chunk_elems as u64, elem_size as u64],
btree_address: Some(data_addr as u64),
version: 4,
chunk_index_type: Some(1),
+350 -101
View File
@@ -5,12 +5,14 @@ extern crate alloc;
use crate::addr::saturating_usize;
#[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec};
use alloc::{borrow::Cow, format, vec, vec::Vec};
#[cfg(feature = "std")]
use std::borrow::Cow;
use crate::btree_v1_write;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::checksum::jenkins_lookup3;
use crate::chunk_cache::{CACHE_LINE_SIZE, align_to_cache_line};
use crate::chunk_cache::CACHE_LINE_SIZE;
use crate::chunk_grid::ChunkGrid;
use crate::ea_writer;
use crate::error::FormatError;
@@ -464,7 +466,9 @@ fn extract_chunk(
let dst: usize = (0..rank).map(|d| idx[d] * chunk_strides[d]).sum::<usize>() * element_size;
// Whole elements only, as far as `raw_data` reaches.
let n = row.min(raw_data.len().saturating_sub(src)) / element_size * element_size;
if n > 0 {
chunk[dst..dst + n].copy_from_slice(&raw_data[src..src + n]);
}
// Next row: advance every dimension but the last.
let mut d = rank - 1;
loop {
@@ -481,6 +485,66 @@ fn extract_chunk(
}
}
/// The bytes of the `linear_idx`-th chunk (see [`extract_chunk`]),
/// borrowed from `raw_data` when the chunk is one contiguous run of it: a
/// chunk inside the dataset (not past its edge) whose dimensions after the
/// first one above 1 equal the dataset's. A dataset stored as one chunk of
/// its own shape is such a run, so its chunk is never copied. Other chunks
/// are copied out, as [`extract_chunk`] does.
fn chunk_bytes_of<'a>(
raw_data: &'a [u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
linear_idx: u64,
) -> Cow<'a, [u8]> {
if shape.is_empty() {
return Cow::Borrowed(raw_data);
}
if let Some(range) = contiguous_chunk(shape, chunk_dims, element_size, linear_idx)
&& let (Ok(start), Ok(end)) = (usize::try_from(range.start), usize::try_from(range.end))
&& end <= raw_data.len()
{
return Cow::Borrowed(&raw_data[start..end]);
}
Cow::Owned(extract_chunk(raw_data, shape, chunk_dims, element_size, linear_idx).1)
}
/// The byte range of the `linear_idx`-th chunk in the row-major dataset,
/// when the chunk is a contiguous run of it (see [`chunk_bytes_of`]).
fn contiguous_chunk(
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
linear_idx: u64,
) -> Option<core::ops::Range<u64>> {
let rank = shape.len();
// Past the first chunk dimension above 1, every chunk dimension must
// span the dataset's.
let first_wide = chunk_dims.iter().position(|&c| c > 1).unwrap_or(rank);
if (first_wide + 1..rank).any(|d| chunk_dims[d] != shape[d]) {
return None;
}
let mut remaining = linear_idx;
let mut start = 0u64;
let mut stride = element_size as u64;
for d in (0..rank).rev() {
let n = shape[d].div_ceil(chunk_dims[d]);
let offset = (remaining % n).checked_mul(chunk_dims[d])?;
remaining /= n;
// A chunk past the dataset's edge is padded: not a run of it.
if offset.checked_add(chunk_dims[d])? > shape[d] {
return None;
}
start = start.checked_add(offset.checked_mul(stride)?)?;
stride = stride.checked_mul(shape[d])?;
}
let len = chunk_dims
.iter()
.try_fold(element_size as u64, |acc, &c| acc.checked_mul(c))?;
Some(start..start.checked_add(len)?)
}
/// Parallel compression threshold: use rayon when chunk count exceeds this.
///
/// Lowered to 2 to enable parallel compression for typical 4-chunk workloads
@@ -508,25 +572,21 @@ const PARALLEL_COMPRESS_MAX_CHUNK_BYTES: u64 = 64 << 20;
/// at most [`PARALLEL_COMPRESS_MAX_CHUNK_BYTES`], compression runs across
/// rayon threads; otherwise it is sequential. Output order matches chunk
/// order, so per-chunk bytes are identical to the sequential path.
fn compress_all_chunks(
raw_data: &[u8],
fn compress_all_chunks<'a>(
raw_data: &'a [u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
chunk_bytes: u64,
pipeline: &Option<FilterPipeline>,
) -> Result<Vec<(u64, Vec<u8>, u32)>, FormatError> {
let one = |i: u64| -> Result<(u64, Vec<u8>, u32), FormatError> {
let (_, raw) = if shape.is_empty() {
(Vec::new(), raw_data.to_vec())
} else {
extract_chunk(raw_data, shape, chunk_dims, element_size, i)
};
) -> Result<Vec<PreparedChunk<'a>>, FormatError> {
let one = |i: u64| -> Result<PreparedChunk<'a>, FormatError> {
let raw = chunk_bytes_of(raw_data, shape, chunk_dims, element_size, i);
let raw_size = raw.len() as u64;
match pipeline {
Some(pl) => {
let (stored, mask) = compress_chunk_masked(&raw, pl, element_size as u32)?;
Ok((raw_size, stored, mask))
Ok((raw_size, Cow::Owned(stored), mask))
}
None => Ok((raw_size, raw, 0)),
}
@@ -557,7 +617,7 @@ fn compress_all_chunks(
/// layout/pipeline messages. `base_address` is where the blob will be placed in the file.
/// Serialize a v4 single chunk layout message (public for OH size estimation).
pub fn serialize_v4_single_chunk_pub(
chunk_dims: &[u32],
chunk_dims: &[u64],
chunk_address: u64,
filtered_size: Option<u64>,
filter_mask: Option<u32>,
@@ -578,7 +638,7 @@ pub fn serialize_v4_single_chunk_pub(
/// Serialize a v4 (or, for a chunk of 4 GiB or more, v5) single chunk
/// layout message.
fn serialize_v4_single_chunk(
chunk_dims: &[u32],
chunk_dims: &[u64],
chunk_address: u64,
filtered_size: Option<u64>,
filter_mask: Option<u32>,
@@ -622,7 +682,7 @@ fn serialize_v4_single_chunk(
/// Serialize a v4 Fixed Array layout message.
fn serialize_v4_fixed_array(
chunk_dims: &[u32],
chunk_dims: &[u64],
fixed_array_address: u64,
offset_size: u8,
element_size: u32,
@@ -653,7 +713,8 @@ fn serialize_v4_fixed_array(
/// dimensions, then the element size). Each takes the fewest bytes that hold
/// the largest, as libhdf5 computes it (`H5D__chunk_set_sizes`:
/// `(log2(dim) + 8) / 8`); HDF5 2.0.0 refuses any other width.
pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u32], element_size: u32) {
pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u64], element_size: u32) {
let element_size = u64::from(element_size);
let max_dim = chunk_dims
.iter()
.copied()
@@ -661,14 +722,14 @@ pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u32], element_
.max()
.unwrap_or(1)
.max(1);
let width = (32 - max_dim.leading_zeros()).div_ceil(8) as usize;
let width = (64 - max_dim.leading_zeros()).div_ceil(8) as usize;
buf.push(width as u8);
for &d in chunk_dims.iter().chain(core::iter::once(&element_size)) {
buf.extend_from_slice(&d.to_le_bytes()[..width]);
}
}
fn layout_v4_chunked_prefix(chunk_dims: &[u32], element_size: u32, version: u8) -> Vec<u8> {
fn layout_v4_chunked_prefix(chunk_dims: &[u64], element_size: u32, version: u8) -> Vec<u8> {
let mut buf = Vec::new();
buf.push(version);
buf.push(2); // class = chunked
@@ -884,7 +945,66 @@ pub fn precompress_chunks(
element_size: usize,
options: &ChunkOptions,
) -> Result<PrecompressedChunks, FormatError> {
let (_, chunk_bytes) = checked_chunk_dims(chunk_dims, element_size)?;
let p = prepare_chunks(raw_data, shape, chunk_dims, element_size, options)?;
Ok(PrecompressedChunks {
chunks: p
.chunks
.into_iter()
.map(|(raw, stored, mask)| (raw, stored.into_owned(), mask))
.collect(),
has_filters: p.has_filters,
element_size: p.element_size,
shape: p.shape,
chunk_dims: p.chunk_dims,
pipeline_message: p.pipeline_message,
})
}
/// One chunk of [`PreparedChunks`]: its raw size, its stored bytes and its
/// filter mask.
pub(crate) type PreparedChunk<'a> = (u64, Cow<'a, [u8]>, u32);
/// [`PrecompressedChunks`] whose stored bytes may borrow the dataset's data:
/// an unfiltered chunk that is a contiguous run of it (a dataset stored as
/// one chunk, say) is not copied. What the file writer lays out.
pub(crate) struct PreparedChunks<'a> {
pub(crate) chunks: Vec<PreparedChunk<'a>>,
pub(crate) has_filters: bool,
pub(crate) element_size: usize,
pub(crate) shape: Vec<u64>,
pub(crate) chunk_dims: Vec<u64>,
pub(crate) pipeline_message: Option<Vec<u8>>,
}
impl PrecompressedChunks {
/// A view of these chunks as [`PreparedChunks`] (their bytes borrowed).
fn prepared(&self) -> PreparedChunks<'_> {
PreparedChunks {
chunks: self
.chunks
.iter()
.map(|(raw, stored, mask)| (*raw, Cow::Borrowed(&stored[..]), *mask))
.collect(),
has_filters: self.has_filters,
element_size: self.element_size,
shape: self.shape.clone(),
chunk_dims: self.chunk_dims.clone(),
pipeline_message: self.pipeline_message.clone(),
}
}
}
/// [`precompress_chunks`], borrowing what it can (see [`PreparedChunks`]).
/// A filtered chunk that is a contiguous run of `raw_data` is compressed
/// straight from it, without a copy of the raw chunk.
pub(crate) fn prepare_chunks<'a>(
raw_data: &'a [u8],
shape: &[u64],
chunk_dims: &[u64],
element_size: usize,
options: &ChunkOptions,
) -> Result<PreparedChunks<'a>, FormatError> {
let chunk_bytes = checked_chunk_dims(chunk_dims, element_size)?;
if chunk_bytes > MAX_V4_CHUNK_BYTES {
check_huge_chunk_filters(options, chunk_bytes)?;
}
@@ -902,7 +1022,7 @@ pub fn precompress_chunks(
&pipeline,
)?;
Ok(PrecompressedChunks {
Ok(PreparedChunks {
chunks,
has_filters,
element_size,
@@ -912,26 +1032,87 @@ pub fn precompress_chunks(
})
}
/// The chunk dimensions as the layout message stores them (each below
/// 2^32), and one chunk's size in bytes. A chunk dimension of 2^32 or more
/// (which HDF5 2.0 can store, in wider fields) is refused, as it is when
/// read: clawhdf5 holds chunk dimensions as `u32`. So is a chunk whose size
/// overflows 64 bits, or that this platform cannot hold in memory (a chunk
/// of 4 GiB or more on a 32-bit target).
fn checked_chunk_dims(
chunk_dims: &[u64],
element_size: usize,
) -> Result<(Vec<u32>, u64), FormatError> {
let dims = chunk_dims
.iter()
.map(|&d| {
u32::try_from(d).map_err(|_| {
FormatError::InvalidChunkDimensions(format!(
"chunk dimension {d} is 2^32 or more, which clawhdf5 does not support"
))
})
})
.collect::<Result<Vec<u32>, _>>()?;
/// A piece of a chunked dataset's data blob as laid out by
/// [`place_chunked_data`]: zero padding, a chunk's stored bytes (by index
/// into [`PreparedChunks::chunks`]), or index structures.
#[derive(Debug)]
pub(crate) enum Piece {
Zeros(usize),
Chunk(usize),
Bytes(Vec<u8>),
}
/// A chunked dataset laid out at an address, without its chunk bytes
/// copied: the pieces of its data blob, in order, and its messages.
pub(crate) struct ChunkedPlacement {
pub(crate) pieces: Vec<Piece>,
/// Length of the data blob in bytes.
pub(crate) len: u64,
pub(crate) layout_message: Vec<u8>,
pub(crate) pipeline_message: Option<Vec<u8>>,
}
impl ChunkedPlacement {
/// The data blob as one buffer (the chunks copied in).
fn into_result(self, pre: &PreparedChunks<'_>) -> ChunkedDataResult {
let mut data_bytes = Vec::with_capacity(saturating_usize(self.len));
for piece in &self.pieces {
match piece {
Piece::Zeros(n) => data_bytes.resize(data_bytes.len() + n, 0),
Piece::Chunk(i) => data_bytes.extend_from_slice(&pre.chunks[*i].1),
Piece::Bytes(b) => data_bytes.extend_from_slice(b),
}
}
ChunkedDataResult {
data_bytes,
layout_message: self.layout_message,
pipeline_message: self.pipeline_message,
}
}
}
/// Builds the pieces of a data blob, tracking its length.
#[derive(Default)]
struct Pieces {
pieces: Vec<Piece>,
len: u64,
}
impl Pieces {
/// Pad to the next cache line.
fn align(&mut self) {
let aligned = align_chunk_offset(self.len);
if aligned > self.len {
self.pieces
.push(Piece::Zeros(saturating_usize(aligned - self.len)));
self.len = aligned;
}
}
fn chunk(&mut self, i: usize, len: usize) {
self.pieces.push(Piece::Chunk(i));
self.len += len as u64;
}
fn bytes(&mut self, b: Vec<u8>) {
self.len += b.len() as u64;
self.pieces.push(Piece::Bytes(b));
}
}
/// One chunk's size in bytes. A chunk dimension of 0 is refused, and so is
/// a chunk whose size overflows 64 bits, or that this platform cannot hold
/// in memory (a chunk of 4 GiB or more on a 32-bit target). Chunk
/// dimensions of 2^32 or more are allowed, as in libhdf5 2.x: such a chunk
/// is over 4 GiB, so it takes layout message version 5, whose dimension
/// fields are up to 8 bytes wide (a version-3 layout, with 4-byte fields,
/// is never written for it).
fn checked_chunk_dims(chunk_dims: &[u64], element_size: usize) -> Result<u64, FormatError> {
if chunk_dims.contains(&0) {
return Err(FormatError::InvalidChunkDimensions(format!(
"chunk dimensions {chunk_dims:?} include 0"
)));
}
let bytes = chunk_dims
.iter()
.try_fold(element_size as u64, |acc, &d| acc.checked_mul(d))
@@ -941,7 +1122,7 @@ fn checked_chunk_dims(
"a chunk of {chunk_dims:?} x {element_size} bytes exceeds this platform's address space"
))
})?;
Ok((dims, bytes))
Ok(bytes)
}
/// The filters clawhdf5 can apply to a chunk of more than `u32::MAX` bytes:
@@ -1010,7 +1191,22 @@ pub fn build_chunked_data_from_precompressed_libver(
low: LibVer,
high: LibVer,
) -> Result<ChunkedDataResult, FormatError> {
let (_, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?;
let pre = pre.prepared();
let placement = place_chunked_data(&pre, base_address, maxshape, low, high)?;
Ok(placement.into_result(&pre))
}
/// [`build_chunked_data_from_precompressed_libver`] without copying the
/// chunks into one buffer: the data blob as [`Piece`]s referring to
/// `pre`'s chunks. The file writer writes the pieces out one by one.
pub(crate) fn place_chunked_data(
pre: &PreparedChunks<'_>,
base_address: u64,
maxshape: Option<&[u64]>,
low: LibVer,
high: LibVer,
) -> Result<ChunkedPlacement, FormatError> {
let chunk_bytes = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?;
if chunk_bytes > MAX_V4_CHUNK_BYTES {
if high < LibVer::V200 {
return Err(FormatError::LibverBound {
@@ -1020,7 +1216,7 @@ pub fn build_chunked_data_from_precompressed_libver(
});
}
} else if low < LibVer::V110 {
return build_btree_v1_chunked_data(pre, base_address, maxshape);
return place_btree_v1_chunked_data(pre, base_address, maxshape);
}
let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?;
let offset_size: u8 = 8;
@@ -1028,17 +1224,14 @@ pub fn build_chunked_data_from_precompressed_libver(
let num_chunks = pre.chunks.len();
let element_size = pre.element_size;
let mut data_buf = Vec::new();
let mut blob = Pieces::default();
let mut written_chunks = Vec::with_capacity(num_chunks);
for (raw_size, compressed, filter_mask) in &pre.chunks {
let aligned_offset = align_to_cache_line(data_buf.len());
if aligned_offset > data_buf.len() {
data_buf.resize(aligned_offset, 0u8);
}
let address = base_address + data_buf.len() as u64;
for (i, (raw_size, compressed, filter_mask)) in pre.chunks.iter().enumerate() {
blob.align();
let address = base_address + blob.len;
let compressed_size = compressed.len() as u64;
data_buf.extend_from_slice(compressed);
blob.chunk(i, compressed.len());
written_chunks.push(WrittenChunk {
address,
compressed_size,
@@ -1047,28 +1240,22 @@ pub fn build_chunked_data_from_precompressed_libver(
});
}
let (chunk_dims_u32, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, element_size)?;
let version = layout_version_for(chunk_bytes);
let aligned_idx = align_to_cache_line(data_buf.len());
if aligned_idx > data_buf.len() {
data_buf.resize(aligned_idx, 0u8);
}
blob.align();
let layout_message = match &index {
ChunkIndexPlan::ExtensibleArray(grid) => {
let ea_address = base_address + data_buf.len() as u64;
let ea_address = base_address + blob.len;
let slots = index_slots(grid, &pre.shape, &pre.chunk_dims, &written_chunks, None)?;
let ea_bytes = ea_writer::build_extensible_array_at(
blob.bytes(ea_writer::build_extensible_array_at(
&slots,
offset_size,
length_size,
pre.has_filters,
ea_address,
);
data_buf.extend_from_slice(&ea_bytes);
));
ea_writer::serialize_v4_extensible_array(
&chunk_dims_u32,
&pre.chunk_dims,
ea_address,
offset_size,
element_size as u32,
@@ -1084,7 +1271,7 @@ pub fn build_chunked_data_from_precompressed_libver(
};
let filter_mask = pre.has_filters.then_some(written_chunks[0].filter_mask);
serialize_v4_single_chunk(
&chunk_dims_u32,
&pre.chunk_dims,
chunk_addr,
filtered_size,
filter_mask,
@@ -1094,7 +1281,7 @@ pub fn build_chunked_data_from_precompressed_libver(
)
}
ChunkIndexPlan::FixedArray(grid, nslots) => {
let fa_address = base_address + data_buf.len() as u64;
let fa_address = base_address + blob.len;
let slots = index_slots(
grid,
&pre.shape,
@@ -1102,16 +1289,15 @@ pub fn build_chunked_data_from_precompressed_libver(
&written_chunks,
Some(*nslots),
)?;
let fa_bytes = build_fixed_array_at(
blob.bytes(build_fixed_array_at(
&slots,
offset_size,
length_size,
pre.has_filters,
fa_address,
);
data_buf.extend_from_slice(&fa_bytes);
));
serialize_v4_fixed_array(
&chunk_dims_u32,
&pre.chunk_dims,
fa_address,
offset_size,
element_size as u32,
@@ -1120,7 +1306,7 @@ pub fn build_chunked_data_from_precompressed_libver(
)
}
ChunkIndexPlan::BTreeV2 => {
let bt_address = base_address + data_buf.len() as u64;
let bt_address = base_address + blob.len;
let records: Vec<(Vec<u64>, &WrittenChunk)> = written_chunks
.iter()
.enumerate()
@@ -1134,9 +1320,9 @@ pub fn build_chunked_data_from_precompressed_libver(
pre.has_filters,
bt_address,
)?;
data_buf.extend_from_slice(&bt_bytes);
blob.bytes(bt_bytes);
serialize_v4_btree_v2(
&chunk_dims_u32,
&pre.chunk_dims,
bt_address,
offset_size,
element_size as u32,
@@ -1146,8 +1332,9 @@ pub fn build_chunked_data_from_precompressed_libver(
}
};
Ok(ChunkedDataResult {
data_bytes: data_buf,
Ok(ChunkedPlacement {
pieces: blob.pieces,
len: blob.len,
layout_message,
pipeline_message: pre.pipeline_message.clone(),
})
@@ -1156,11 +1343,11 @@ pub fn build_chunked_data_from_precompressed_libver(
/// Lay out precompressed chunks at `base_address` followed by a version-1
/// B-tree chunk index, with a version-3 layout message: what libhdf5 writes
/// for a chunked dataset under a low bound of 1.8.
fn build_btree_v1_chunked_data(
pre: &PrecompressedChunks,
fn place_btree_v1_chunked_data(
pre: &PreparedChunks<'_>,
base_address: u64,
maxshape: Option<&[u64]>,
) -> Result<ChunkedDataResult, FormatError> {
) -> Result<ChunkedPlacement, FormatError> {
if let Some(ms) = maxshape {
let bad = |what: &str| FormatError::ChunkedReadError(format!("maxshape: {what}"));
if ms.len() != pre.shape.len() {
@@ -1171,20 +1358,17 @@ fn build_btree_v1_chunked_data(
}
}
let offset_size: u8 = 8;
let mut data_buf = Vec::new();
let mut blob = Pieces::default();
let mut entries = Vec::with_capacity(pre.chunks.len());
for (i, (_raw_size, stored, filter_mask)) in pre.chunks.iter().enumerate() {
let aligned_offset = align_to_cache_line(data_buf.len());
if aligned_offset > data_buf.len() {
data_buf.resize(aligned_offset, 0u8);
}
blob.align();
entries.push(btree_v1_write::ChunkEntry {
scaled: scaled_coords(&pre.shape, &pre.chunk_dims, i),
nbytes: stored.len() as u64,
filter_mask: *filter_mask,
address: base_address + data_buf.len() as u64,
address: base_address + blob.len,
});
data_buf.extend_from_slice(stored);
blob.chunk(i, stored.len());
}
let element_size = u32::try_from(pre.element_size)
.map_err(|_| FormatError::Overflow("element size".into()))?;
@@ -1193,11 +1377,8 @@ fn build_btree_v1_chunked_data(
let btree_address = if entries.is_empty() {
u64::MAX
} else {
let aligned_idx = align_to_cache_line(data_buf.len());
if aligned_idx > data_buf.len() {
data_buf.resize(aligned_idx, 0u8);
}
let addr = base_address + data_buf.len() as u64;
blob.align();
let addr = base_address + blob.len;
let tree = btree_v1_write::build_chunk_btree_v1_at(
&entries,
&pre.chunk_dims,
@@ -1205,13 +1386,14 @@ fn build_btree_v1_chunked_data(
addr,
offset_size,
)?;
data_buf.extend_from_slice(&tree);
blob.bytes(tree);
addr
};
let layout_message =
serialize_v3_chunked(&pre.chunk_dims, btree_address, offset_size, element_size)?;
Ok(ChunkedDataResult {
data_bytes: data_buf,
Ok(ChunkedPlacement {
pieces: blob.pieces,
len: blob.len,
layout_message,
pipeline_message: pre.pipeline_message.clone(),
})
@@ -1424,7 +1606,7 @@ fn build_btree_v2_chunk_index_at(
/// Serialize a v4 layout message for a version-2 B-tree chunk index.
fn serialize_v4_btree_v2(
chunk_dims: &[u32],
chunk_dims: &[u64],
btree_address: u64,
offset_size: u8,
element_size: u32,
@@ -2184,7 +2366,7 @@ mod tests {
/// filtered index element stores the chunk's size in 8 bytes.
#[test]
fn huge_chunk_layout_messages_and_index_elements() {
let dims = [HUGE_DIM as u32];
let dims = [HUGE_DIM];
let parsed = |msg: &[u8]| {
assert_eq!(msg[0], 5, "layout message version");
// Class chunked, then (after the flags) 2 dimensions of 4 bytes.
@@ -2197,7 +2379,7 @@ mod tests {
chunk_index_type,
..
} => {
assert_eq!(chunk_dimensions, vec![HUGE_DIM as u32, 8]);
assert_eq!(chunk_dimensions, vec![HUGE_DIM, 8]);
chunk_index_type.unwrap()
}
other => panic!("{other:?}"),
@@ -2245,22 +2427,89 @@ mod tests {
assert_eq!(u16::from_le_bytes([bt[10], bt[11]]), 28);
}
/// A chunk dimension of 2^32 or more (a 1-byte element, 2^32 + 7 per
/// chunk): the dimensions take 5 bytes each, as libhdf5 2.x writes
/// them (`(log2(dim) + 8) / 8`), and parse back whole.
#[test]
fn chunk_dimensions_of_2_pow_32_or_more_take_5_bytes() {
const DIM: u64 = (1 << 32) + 7;
let msg = serialize_v4_single_chunk(&[DIM], 0x800, None, None, 8, 1, 5);
assert_eq!(&msg[..5], &[5, 2, 0, 2, 5]);
assert_eq!(&msg[5..10], &[7, 0, 0, 0, 1]);
assert_eq!(&msg[10..15], &[1, 0, 0, 0, 0]);
assert_eq!(msg[15], 1, "Single Chunk index");
match DataLayout::parse(&msg, 8, 8).unwrap() {
DataLayout::Chunked {
chunk_dimensions, ..
} => assert_eq!(chunk_dimensions, [DIM, 1]),
other => panic!("{other:?}"),
}
// The largest dimension takes all 8 bytes.
let mut buf = Vec::new();
push_v4_chunk_dims(&mut buf, &[u64::MAX, 3], 4);
assert_eq!(buf[0], 8);
assert_eq!(buf.len(), 1 + 3 * 8);
// A version-3 layout has 4-byte dimensions: refused, not truncated.
assert!(serialize_v3_chunked(&[DIM], 0x800, 8, 1).is_err());
}
/// Which chunks are contiguous runs of the dataset's data (borrowed,
/// not copied, when unfiltered).
#[test]
fn contiguous_chunks_are_borrowed() {
// Rows of a 2-D dataset: chunk (2, 5) of shape (5, 5).
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 0), Some(0..40));
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 1), Some(40..80));
// The last row chunk runs past the edge: padded, so copied.
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 2), None);
// Columns split: not contiguous.
assert_eq!(contiguous_chunk(&[4, 6], &[2, 3], 1, 0), None);
// Leading dimensions of 1, then a partial row: contiguous.
assert_eq!(contiguous_chunk(&[2, 3, 8], &[1, 1, 4], 2, 5), Some(40..48));
// The whole dataset as one chunk.
assert_eq!(contiguous_chunk(&[3, 4], &[3, 4], 8, 0), Some(0..96));
let raw: Vec<u8> = (0..100).collect();
let whole = chunk_bytes_of(&raw, &[10, 10], &[10, 10], 1, 0);
assert!(matches!(whole, Cow::Borrowed(b) if b.as_ptr() == raw.as_ptr()));
let rows = chunk_bytes_of(&raw, &[10, 10], &[3, 10], 1, 1);
assert!(matches!(rows, Cow::Borrowed(b) if b == &raw[30..60]));
// Data shorter than the shape is copied (and zero-padded) as before.
let short = chunk_bytes_of(&raw[..50], &[10, 10], &[10, 10], 1, 0);
assert!(matches!(&short, Cow::Owned(v) if v.len() == 100));
// Borrowed or copied, the bytes are what extract_chunk gives.
for (shape, chunk) in [
(&[10u64, 10][..], &[3u64, 10][..]),
(&[10, 10], &[3, 4]),
(&[4, 5, 5], &[1, 5, 5]),
(&[4, 5, 5], &[1, 2, 5]),
] {
for i in 0..chunk_count(shape, chunk) {
let got = chunk_bytes_of(&raw, shape, chunk, 1, i);
assert_eq!(&got[..], &extract_chunk(&raw, shape, chunk, 1, i).1[..]);
}
}
}
#[test]
fn huge_chunks_refused_where_unsupported() {
// A chunk dimension of 2^32 or more.
// A chunk dimension of 0.
assert!(matches!(
checked_chunk_dims(&[1 << 32], 1),
Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32")
checked_chunk_dims(&[4, 0], 1),
Err(FormatError::InvalidChunkDimensions(_))
));
// A chunk dimension of 2^32 or more is fine (on a 64-bit target).
#[cfg(target_pointer_width = "64")]
assert_eq!(
checked_chunk_dims(&[(1 << 32) + 5], 1).unwrap(),
(1 << 32) + 5
);
// A chunk size that overflows u64.
assert!(matches!(
checked_chunk_dims(&[u32::MAX.into(), u32::MAX.into(), 2], 8),
Err(FormatError::Overflow(_))
));
assert_eq!(
checked_chunk_dims(&[HUGE_DIM], 8).unwrap(),
(vec![HUGE_DIM as u32], HUGE_BYTES)
);
assert_eq!(checked_chunk_dims(&[HUGE_DIM], 8).unwrap(), HUGE_BYTES);
// Filters that cannot take a chunk that large.
for plugin in [
PluginFilter::Lzf,
+19 -19
View File
@@ -34,7 +34,7 @@ const MAX_LAYOUT_NDIMS: usize = 33;
/// (`H5O__layout_decode`): at most [`MAX_LAYOUT_NDIMS`], no dimension 0, and
/// before version 4 at least one dataspace dimension plus the element size.
/// A zero chunk dimension used to read the dataset as all fill values.
fn check_chunk_dims(dims: Vec<u32>, layout_version: u8) -> Result<Vec<u32>, FormatError> {
fn check_chunk_dims(dims: Vec<u64>, layout_version: u8) -> Result<Vec<u64>, FormatError> {
if dims.len() > MAX_LAYOUT_NDIMS {
return Err(FormatError::InvalidChunkDimensions(
"dimensionality is too large".into(),
@@ -72,7 +72,7 @@ pub enum DataLayout {
/// Chunked: data stored in chunks via a B-tree.
Chunked {
/// Chunk dimension sizes.
chunk_dimensions: Vec<u32>,
chunk_dimensions: Vec<u64>,
/// B-tree address, or `None` if undefined.
btree_address: Option<u64>,
/// Layout version (3 or 4). Version 1/2 messages (HDF5 1.4/1.6-era)
@@ -404,11 +404,11 @@ impl DataLayout {
_ => return Err(FormatError::InvalidLayoutClass(layout_class)),
};
ensure_len(data, p, dimensionality * 4)?;
let dims: Vec<u32> = data[p..p + dimensionality * 4]
let dims: Vec<u64> = data[p..p + dimensionality * 4]
.as_chunks::<4>()
.0
.iter()
.map(|c| u32::from_le_bytes(*c))
.map(|c| u64::from(u32::from_le_bytes(*c)))
.collect();
p += dimensionality * 4;
match layout_class {
@@ -424,7 +424,7 @@ impl DataLayout {
1 => {
let size = dims
.iter()
.try_fold(1u64, |acc, &d| acc.checked_mul(d as u64))
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
.ok_or_else(|| {
FormatError::Overflow(format!("contiguous layout size {dims:?}"))
})?;
@@ -489,7 +489,12 @@ impl DataLayout {
ensure_len(data, p, dimensionality * 4)?;
let mut chunk_dimensions = Vec::with_capacity(dimensionality);
for _ in 0..dimensionality {
let dim = u32::from_le_bytes([data[p], data[p + 1], data[p + 2], data[p + 3]]);
let dim = u64::from(u32::from_le_bytes([
data[p],
data[p + 1],
data[p + 2],
data[p + 3],
]));
chunk_dimensions.push(dim);
p += 4;
}
@@ -565,14 +570,8 @@ impl DataLayout {
.iter()
.rev()
.fold(0u64, |acc, &b| (acc << 8) | u64::from(b));
// Chunk dimensions are held as u32; HDF5 2.0 can write
// larger ones (layout version 5), which are refused
// rather than truncated.
let val = u32::try_from(val).map_err(|_| {
FormatError::InvalidChunkDimensions(format!(
"chunk dimension {val} is larger than 2^32 - 1, which is not supported"
))
})?;
// HDF5 2.0 writes dimensions of 2^32 or more (layout
// version 5, chunks over 4 GiB) in 5 to 8 bytes.
chunk_dimensions.push(val);
p += dim_size_encoded_length;
}
@@ -858,7 +857,7 @@ mod tests {
.unwrap_or_else(|e| panic!("width {width}: {e:?}"));
assert!(
matches!(&layout, DataLayout::Chunked { chunk_dimensions, .. }
if chunk_dimensions.iter().map(|&d| u64::from(d)).eq(dims)),
if chunk_dimensions[..] == dims),
"width {width}: {layout:?}"
);
}
@@ -871,13 +870,14 @@ mod tests {
)
);
}
// A dimension past u32 cannot be represented and is refused, not
// truncated.
// HDF5 2.0 writes dimensions of 2^32 or more (up to 8 bytes each).
for (width, dim) in [(5u8, (1u64 << 32) + 3), (8, u64::MAX)] {
assert!(matches!(
DataLayout::parse(&v4_chunked_msg(5, &[1 << 32, 8]), 8, 8),
Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32")
DataLayout::parse(&v4_chunked_msg(width, &[dim, 1]), 8, 8),
Ok(DataLayout::Chunked { chunk_dimensions, .. }) if chunk_dimensions == [dim, 1]
));
}
}
#[test]
fn v1v2_rejects_bad_class_dimensionality_and_truncation() {
+35 -29
View File
@@ -145,13 +145,16 @@ pub enum Datatype {
/// imaginary, in rectangular form. `size` is twice the base size and the
/// base is an IEEE float.
///
/// This variant exists for **writing** (see
/// [`Datatype::parse`] returns this variant for every class-11 message,
/// also inside compounds, arrays and variable-length types. Readers that
/// want the member view (`{r, i}`, as h5py writes complex numbers by
/// default) use [`Datatype::complex_as_compound`]; `read_compound_fields`
/// accepts this variant directly.
///
/// Writing it is opt-in (see
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
/// newer can read class 11, so it is opt-in and h5py's compound `{r, i}`
/// stays the default complex encoding. [`Datatype::parse`] still
/// surfaces a class-11 message as that equivalent `{r, i}` compound, so
/// every compound reader handles both encodings; parsing what this
/// variant serializes therefore yields a `Compound`, not a `Complex`.
/// newer can read class 11, so h5py's compound `{r, i}` stays the
/// default complex encoding.
Complex { size: u32, base_type: Box<Datatype> },
}
@@ -857,9 +860,7 @@ impl Datatype {
// Complex number (HDF5 2.0, datatype version 5). The properties
// are a single base floating-point datatype message; an element
// is two consecutive base-type values (real, imaginary). There
// is no member list. Surface it as the equivalent two-member
// compound `{r, i}` — the same shape h5py writes for numpy
// complex dtypes — so downstream compound readers work as-is.
// is no member list.
if version != 5 {
return Err(FormatError::InvalidDatatypeVersion {
class: class_id,
@@ -875,7 +876,13 @@ impl Datatype {
actual: size as usize,
});
}
Ok((Self::complex_as_compound(size, &base_type), pos))
Ok((
Datatype::Complex {
size,
base_type: Box::new(base_type),
},
pos,
))
}
_ => Err(FormatError::InvalidDatatypeClass(class_id)),
}
@@ -1174,9 +1181,8 @@ impl Datatype {
/// The `{r, i}` compound equivalent to a native complex type of `size`
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the
/// shape h5py writes for numpy complex dtypes, and what [`Self::parse`]
/// returns for a class-11 message. Readers that meet a
/// [`Datatype::Complex`] handle it through this view.
/// shape h5py writes for numpy complex dtypes. Readers that want a
/// [`Datatype::Complex`] as members handle it through this view.
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
let base_size = base_type.type_size();
Datatype::Compound {
@@ -1849,21 +1855,24 @@ mod tests {
fn test_complex_v5_from_hdf5_2_0() {
let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap();
assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len());
match dt {
Datatype::Compound { size, members } => {
assert_eq!(size, 16);
assert_eq!(members.len(), 2);
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
for m in &members {
match &dt {
Datatype::Complex { size, base_type } => {
assert_eq!(*size, 16);
assert!(matches!(
m.datatype,
base_type.as_ref(),
Datatype::FloatingPoint { size: 8, .. }
));
}
other => panic!("expected Complex, got {other:?}"),
}
other => panic!("expected Compound, got {other:?}"),
}
// The member view is h5py's `{r, i}` compound.
let Datatype::Compound { members, .. } =
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
else {
unreachable!()
};
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
}
#[test]
@@ -1887,7 +1896,7 @@ mod tests {
assert_eq!(members.len(), 2);
assert!(matches!(
&members[0].datatype,
Datatype::Compound { size: 16, members } if members.len() == 2
Datatype::Complex { size: 16, .. }
));
assert_eq!(
(members[1].name.as_str(), members[1].byte_offset),
@@ -1918,12 +1927,9 @@ mod tests {
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
assert_eq!(dt.type_size(), 16);
dt.check_encodable().unwrap();
// Parsing surfaces class 11 as the equivalent `{r, i}` compound.
// Parsing returns the native complex type unchanged.
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
assert_eq!(
parsed,
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
);
assert_eq!(parsed, dt);
let f32c = make_native_complex_f32_type().serialize();
assert_eq!(
+1 -1
View File
@@ -14,7 +14,7 @@ use crate::chunked_write::{
/// Serialize a v4 Extensible Array layout message.
pub(crate) fn serialize_v4_extensible_array(
chunk_dims: &[u32],
chunk_dims: &[u64],
ea_address: u64,
offset_size: u8,
element_size: u32,
@@ -454,7 +454,7 @@ pub fn read_extensible_array_chunks(
header: &ExtensibleArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
length_size: u8,
@@ -480,7 +480,7 @@ pub fn read_extensible_array_chunks_in<S: Storage + ?Sized>(
header: &ExtensibleArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
_length_size: u8,
@@ -489,12 +489,12 @@ pub fn read_extensible_array_chunks_in<S: Storage + ?Sized>(
// Linear indexes follow the maximum dimensions, with the unlimited
// dimension swizzled to the slowest position (see `chunk_grid`).
let dims_u64: Vec<u64> = chunk_dimensions.iter().map(|&d| d as u64).collect();
let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, &dims_u64)?;
let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, chunk_dimensions)?;
let grid = &grid;
let chunk_byte_size: u64 =
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
let chunk_byte_size: u64 = chunk_dimensions
.iter()
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
// Parse index block (EAIB): signature(4) + version(1) + client_id(1)
// + header address(offset_size), then the inline elements, then the
@@ -930,7 +930,7 @@ mod tests {
let header = ExtensibleArrayHeader::parse(&file_data, aehd_offset, os, ls).unwrap();
let ds_dims = vec![40u64]; // 2 chunks × 20 elements
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_extensible_array_chunks(
&file_data,
&header,
@@ -1057,7 +1057,7 @@ mod tests {
let file_data = build_inline_plus_data_blocks();
let header = ExtensibleArrayHeader::parse(&file_data, 0x100, os, ls).unwrap();
let ds_dims = vec![40u64];
let chunk_dims = vec![10u32];
let chunk_dims = vec![10u64];
let chunks = read_extensible_array_chunks(
&file_data,
&header,
+231 -56
View File
@@ -3,15 +3,14 @@
//! Produces valid HDF5 files with v3 superblock, v2 object headers,
//! link messages, contiguous datasets, inline and dense attributes.
use crate::addr::saturating_usize;
use crate::addr::{saturating_usize, to_usize};
#[cfg(not(feature = "std"))]
use alloc::{format, string::String, vec, vec::Vec};
use crate::attribute::AttributeMessage;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::chunked_write::{
ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed_libver,
precompress_chunks,
ChunkOptions, Piece, PreparedChunks, place_chunked_data, prepare_chunks,
};
use crate::data_layout::VdsMapping;
use crate::dataspace::{Dataspace, DataspaceType};
@@ -1612,7 +1611,30 @@ impl FileWriter {
self
}
/// The file, in memory. Its datasets' data is copied into the buffer
/// once (a dataset stored as one chunk of its own shape, or unfiltered
/// chunks that are contiguous runs of its data, are not copied before);
/// to write a file without holding it all, use [`Self::finish_with`].
pub fn finish(self) -> Result<Vec<u8>, FormatError> {
let mut buf = Vec::new();
self.finish_into(&mut buf)?;
Ok(buf)
}
/// The file, handed to `put` in order, a piece at a time, without
/// assembling it in memory: an unfiltered chunk or contiguous dataset
/// is passed straight from the data the dataset was given. Writing a
/// dataset stored as one unfiltered chunk this way needs little memory
/// beyond the dataset's own data. `put`'s first error stops the write
/// and is returned.
pub fn finish_with<E: From<FormatError>>(
self,
put: impl FnMut(&[u8]) -> Result<(), E>,
) -> Result<(), E> {
self.finish_into(&mut FnSink(put))
}
fn finish_into<S: Sink>(self, sink: &mut S) -> Result<(), S::Error> {
let page_size = self.page_size;
if let Some(ps) = page_size
&& !(MIN_FILE_SPACE_PAGE_SIZE..=MAX_FILE_SPACE_PAGE_SIZE).contains(&ps)
@@ -1620,7 +1642,8 @@ impl FileWriter {
return Err(FormatError::SerializationError(format!(
"file space page size {ps} is outside libhdf5's \
{MIN_FILE_SPACE_PAGE_SIZE}..={MAX_FILE_SPACE_PAGE_SIZE} bytes"
)));
))
.into());
}
let (low, high) = (self.low, self.high);
@@ -1774,15 +1797,16 @@ impl FileWriter {
})
.collect::<Result<_, _>>()?;
struct DataBlob {
struct DataBlob<'a> {
data: Vec<u8>,
oh_bytes: Vec<u8>,
/// Cached compressed chunks for chunked datasets; reused in Pass 2
/// to avoid re-compressing the same data.
precompressed: Option<PrecompressedChunks>,
/// to avoid re-compressing the same data. Unfiltered chunks
/// borrow the dataset's data where they can.
precompressed: Option<PreparedChunks<'a>>,
}
let mut dummy_blobs: Vec<DataBlob> = Vec::new();
let mut dummy_blobs: Vec<DataBlob<'_>> = Vec::new();
let mut dummy_cursor = 0u64;
for (i, d) in all_ds.iter().enumerate() {
let dense_blob = ds_dense[i]
@@ -1820,21 +1844,16 @@ impl FileWriter {
.resolve_chunk_dims_for(&d.ds.dimensions, elem_size);
// Compress once in Pass 1; cache the result so Pass 2 can skip
// re-compression and just rebuild the index with real addresses.
let pre = precompress_chunks(
let pre = prepare_chunks(
&d.raw,
&d.ds.dimensions,
&chunk_dims,
elem_size,
&d.chunk_options,
)?;
let result = build_chunked_data_from_precompressed_libver(
&pre,
dummy_cursor,
d.maxshape.as_deref(),
low,
high,
)?;
dummy_cursor += result.data_bytes.len() as u64;
let result =
place_chunked_data(&pre, dummy_cursor, d.maxshape.as_deref(), low, high)?;
dummy_cursor += result.len;
let oh = build_chunked_dataset_oh(
&d.dt,
&d.ds,
@@ -1849,7 +1868,7 @@ impl FileWriter {
d.refcount,
)?;
dummy_blobs.push(DataBlob {
data: result.data_bytes,
data: Vec::new(),
oh_bytes: oh,
precompressed: Some(pre),
});
@@ -1955,7 +1974,7 @@ impl FileWriter {
})
.collect::<Result<_, FormatError>>()?;
let mut ds_blobs2: Vec<DataBlob> = Vec::new();
let mut ds_blobs2: Vec<OutBlob<'_>> = Vec::new();
let global_align_threshold = self.alignment_threshold;
let global_align_bytes = self.alignment_bytes;
for (i, d) in all_ds.iter().enumerate() {
@@ -1977,26 +1996,21 @@ impl FileWriter {
&d.fill_message,
d.refcount,
)?;
ds_blobs2.push(DataBlob {
data: gcol_bytes.clone(),
ds_blobs2.push(OutBlob {
data: vec![Out::Slice(gcol_bytes)],
oh_bytes: oh,
precompressed: None,
});
} else if is_chunked[i] {
let base_address = cursor2 as u64;
// Reuse precompressed chunks from Pass 1 — avoids re-compressing
// the same data a second time.
let result = build_chunked_data_from_precompressed_libver(
dummy_blobs[i]
let pre = dummy_blobs[i]
.precompressed
.as_ref()
.expect("chunked dataset missing precompressed cache"),
base_address,
d.maxshape.as_deref(),
low,
high,
)?;
cursor2 += result.data_bytes.len();
.expect("chunked dataset missing precompressed cache");
let result =
place_chunked_data(pre, base_address, d.maxshape.as_deref(), low, high)?;
cursor2 += to_usize(result.len)?;
let oh = build_chunked_dataset_oh(
&d.dt,
&d.ds,
@@ -2010,11 +2024,16 @@ impl FileWriter {
&d.fill_message,
d.refcount,
)?;
ds_blobs2.push(DataBlob {
data: result.data_bytes,
oh_bytes: oh,
precompressed: None,
});
let data = result
.pieces
.into_iter()
.map(|p| match p {
Piece::Zeros(n) => Out::Zeros(n),
Piece::Chunk(c) => Out::Slice(&pre.chunks[c].1),
Piece::Bytes(b) => Out::Bytes(b),
})
.collect();
ds_blobs2.push(OutBlob { data, oh_bytes: oh });
} else if is_compact[i] {
// Compact: data is inline in the object header, no external blob
let oh = build_compact_dataset_oh(
@@ -2030,10 +2049,9 @@ impl FileWriter {
d.refcount,
layout_version,
)?;
ds_blobs2.push(DataBlob {
ds_blobs2.push(OutBlob {
data: vec![],
oh_bytes: oh,
precompressed: None,
});
} else {
// Determine alignment: per-dataset overrides global
@@ -2060,13 +2078,10 @@ impl FileWriter {
d.refcount,
layout_version,
)?;
let mut data = vec![0u8; padding];
data.extend_from_slice(&d.raw);
cursor2 += d.raw.len();
ds_blobs2.push(DataBlob {
data,
ds_blobs2.push(OutBlob {
data: vec![Out::Zeros(padding), Out::Slice(&d.raw)],
oh_bytes: oh,
precompressed: None,
});
}
}
@@ -2080,7 +2095,8 @@ impl FileWriter {
cursor2 = cursor2.next_multiple_of(ps as usize);
}
let eof_addr2 = cursor2 as u64;
let mut buf = Vec::with_capacity(cursor2);
sink.reserve(cursor2);
let mut out = Counted { sink, written: 0 };
let sb = Superblock {
version: superblock_version,
@@ -2103,9 +2119,9 @@ impl FileWriter {
checksum: None,
page_size: None,
};
buf.extend_from_slice(&sb.serialize());
out.put(&sb.serialize())?;
if let Some(ref ext) = sb_ext {
buf.extend_from_slice(ext);
out.put(ext)?;
}
// Group OHs + dense blobs (link blob, then attr blob, matching pass 2)
@@ -2132,32 +2148,120 @@ impl FileWriter {
g.refcount,
)?;
debug_assert_eq!(oh.len(), group_oh_sizes[gi]);
debug_assert_eq!(buf.len() as u64, group_addrs2[gi]);
buf.extend_from_slice(&oh);
debug_assert_eq!(out.written as u64, group_addrs2[gi]);
out.put(&oh)?;
if let Some(ref b) = link_blob {
buf.extend_from_slice(&b.blob);
out.put(&b.blob)?;
}
if let Some(ref blob) = group_dense_blobs[gi] {
buf.extend_from_slice(&blob.blob);
out.put(&blob.blob)?;
}
}
// Dataset OHs + dense blobs
for (i, blob) in ds_blobs2.iter().enumerate() {
buf.extend_from_slice(&blob.oh_bytes);
out.put(&blob.oh_bytes)?;
if let Some(ref dense) = ds_dense_blobs[i] {
buf.extend_from_slice(&dense.blob);
out.put(&dense.blob)?;
}
}
// Data
for blob in &ds_blobs2 {
buf.extend_from_slice(&blob.data);
for piece in &blob.data {
match piece {
Out::Bytes(b) => out.put(b)?,
Out::Slice(s) => out.put(s)?,
Out::Zeros(n) => out.zeros(*n)?,
}
}
}
debug_assert_eq!(buf.len(), data_end);
buf.resize(cursor2, 0);
Ok(buf)
debug_assert_eq!(out.written, data_end);
out.zeros(cursor2 - data_end)?;
Ok(())
}
}
/// A piece of a dataset's data as [`FileWriter`] writes it out.
enum Out<'a> {
Bytes(Vec<u8>),
/// Borrowed from the dataset's data or its prepared chunks.
Slice(&'a [u8]),
Zeros(usize),
}
/// A dataset's object header and the pieces of its data.
struct OutBlob<'a> {
oh_bytes: Vec<u8>,
data: Vec<Out<'a>>,
}
/// Where [`FileWriter`] puts a file's bytes, in order.
trait Sink {
type Error: From<FormatError>;
/// The file will be `_len` bytes long.
fn reserve(&mut self, _len: usize) {}
fn put(&mut self, bytes: &[u8]) -> Result<(), Self::Error>;
fn zeros(&mut self, n: usize) -> Result<(), Self::Error> {
const ZEROS: [u8; 4096] = [0; 4096];
let mut left = n;
while left > 0 {
let k = left.min(ZEROS.len());
self.put(&ZEROS[..k])?;
left -= k;
}
Ok(())
}
}
impl Sink for Vec<u8> {
type Error = FormatError;
fn reserve(&mut self, len: usize) {
self.reserve_exact(len.saturating_sub(self.len()));
}
fn put(&mut self, bytes: &[u8]) -> Result<(), FormatError> {
self.extend_from_slice(bytes);
Ok(())
}
fn zeros(&mut self, n: usize) -> Result<(), FormatError> {
self.resize(self.len() + n, 0);
Ok(())
}
}
/// A [`Sink`] calling a function ([`FileWriter::finish_with`]).
struct FnSink<F>(F);
impl<E: From<FormatError>, F: FnMut(&[u8]) -> Result<(), E>> Sink for FnSink<F> {
type Error = E;
fn put(&mut self, bytes: &[u8]) -> Result<(), E> {
(self.0)(bytes)
}
}
/// A [`Sink`] and the number of bytes put into it.
struct Counted<'s, S> {
sink: &'s mut S,
written: usize,
}
impl<S: Sink> Counted<'_, S> {
fn put(&mut self, bytes: &[u8]) -> Result<(), S::Error> {
self.written += bytes.len();
self.sink.put(bytes)
}
fn zeros(&mut self, n: usize) -> Result<(), S::Error> {
self.written += n;
self.sink.zeros(n)
}
}
@@ -3015,4 +3119,75 @@ mod tests {
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 3);
assert_eq!(layout_of(&bytes, "x")[0], 3);
}
/// `finish_with` hands out exactly the bytes `finish` returns, for every
/// storage (contiguous, compact, chunked filtered and not, chunks
/// borrowed and copied, a version-1 B-tree, a paged file), and stops at
/// the sink's first error.
#[test]
fn finish_with_matches_finish() {
let build = |low: LibVer, paged: bool| {
let mut fw = FileWriter::new();
fw.libver_bounds(low, LibVer::Latest);
if paged {
fw.with_page_size(4096);
}
let v: Vec<f64> = (0..300).map(f64::from).collect();
fw.create_dataset("contig").with_f64_data(&v);
fw.create_dataset("compact")
.with_f64_data(&[1.0, 2.0])
.compact();
fw.create_dataset("one_chunk")
.with_f64_data(&v)
.with_shape(&[30, 10])
.with_chunks(&[30, 10]);
fw.create_dataset("rows")
.with_f64_data(&v)
.with_shape(&[30, 10])
.with_chunks(&[7, 10]);
fw.create_dataset("tiles")
.with_f64_data(&v)
.with_shape(&[30, 10])
.with_chunks(&[7, 3]);
fw.create_dataset("deflated")
.with_f64_data(&v)
.with_chunks(&[64])
.with_shuffle()
.with_deflate(4);
fw.create_dataset("grows")
.with_u8_data_owned((0..=255).collect())
.with_chunks(&[100])
.with_maxshape(&[u64::MAX]);
fw
};
for (low, paged) in [
(LibVer::V18, false),
(LibVer::V110, false),
(LibVer::Latest, true),
] {
let bytes = build(low, paged).finish().unwrap();
let mut streamed = Vec::new();
build(low, paged)
.finish_with(|b: &[u8]| {
streamed.extend_from_slice(b);
Ok::<(), FormatError>(())
})
.unwrap();
assert_eq!(streamed, bytes, "{low:?} paged {paged}");
}
let mut calls = 0;
let err = build(LibVer::V110, false)
.finish_with(|_| {
calls += 1;
if calls == 3 {
Err(FormatError::SerializationError("full".into()))
} else {
Ok(())
}
})
.unwrap_err();
assert_eq!(err, FormatError::SerializationError("full".into()));
assert_eq!(calls, 3);
}
}
+9 -9
View File
@@ -155,7 +155,7 @@ pub fn read_fixed_array_chunks(
header: &FixedArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
length_size: u8,
@@ -180,7 +180,7 @@ pub fn read_fixed_array_chunks_in<S: Storage + ?Sized>(
header: &FixedArrayHeader,
dataset_dims: &[u64],
max_dims: Option<&[u64]>,
chunk_dimensions: &[u32],
chunk_dimensions: &[u64],
element_size: u32,
offset_size: u8,
_length_size: u8,
@@ -227,11 +227,11 @@ pub fn read_fixed_array_chunks_in<S: Storage + ?Sized>(
// The index is laid out over the chunk grid of the *maximum* dimensions
// (row-major), so a dataset smaller than its maxshape has gaps.
let dims_u64: Vec<u64> = chunk_dimensions.iter().map(|&d| d as u64).collect();
let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, &dims_u64)?;
let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, chunk_dimensions)?;
let chunk_byte_size: u64 =
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
let chunk_byte_size: u64 = chunk_dimensions
.iter()
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
let mut chunks = Vec::new();
// `rel` is relative to the data block, whose bytes are in `w`.
@@ -670,7 +670,7 @@ mod tests {
let header =
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
let ds_dims = vec![100u64];
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_fixed_array_chunks(
&file_data,
&header,
@@ -747,7 +747,7 @@ mod tests {
let header =
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
let ds_dims = vec![60u64];
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_fixed_array_chunks(
&file_data,
&header,
@@ -848,7 +848,7 @@ mod tests {
assert_eq!(header.num_elements, 11);
let ds_dims = vec![11u64 * 20];
let chunk_dims = vec![20u32];
let chunk_dims = vec![20u64];
let chunks = read_fixed_array_chunks(
&file_data,
&header,
+43
View File
@@ -731,6 +731,49 @@ fn group_children<S: Storage + ?Sized>(
Ok(entries)
}
/// The names of the links of the group at `group_address` in creation
/// order — the order libhdf5 iterates them in with `H5_INDEX_CRT_ORDER`
/// (`H5Literate`) — when the group tracks the creation order of its links;
/// `None` when it does not (a version-1 group never does), in which case
/// libhdf5 can only iterate by name (`H5_INDEX_NAME`: byte order of the
/// names). Hard, soft and external links are listed (user-defined ones,
/// which cannot be followed, are not); links without a creation order
/// value, which a tracking group should not have, come last in the order
/// they are stored.
pub fn links_in_creation_order_in<S: Storage + ?Sized>(
file_data: &S,
superblock: &Superblock,
group_address: u64,
) -> Result<Option<Vec<String>>, FormatError> {
let os = superblock.offset_size;
let ls = superblock.length_size;
let header = ObjectHeader::parse_in(file_data, checked_addr(group_address)?, os, ls)?;
if !is_v2_group(&header) {
return Ok(None);
}
let link_info = find_link_info(&header, os)?;
if link_info.max_creation_order.is_none() {
return Ok(None);
}
let mut links: Vec<(u64, String)> = Vec::new();
let mut visit = |link: LinkMessage| {
links.push((link.creation_order.unwrap_or(u64::MAX), link.name));
};
if let Some(fh_addr) = link_info.fractal_heap_address {
for_each_dense_link(file_data, &link_info, fh_addr, os, ls, false, visit)?;
} else {
for msg in &header.messages {
if msg.msg_type == MessageType::Link
&& let Some(link) = parse_link(&msg.data, os)?
{
visit(link);
}
}
}
links.sort_by_key(|&(order, _)| order);
Ok(Some(links.into_iter().map(|(_, name)| name).collect()))
}
/// Soft links followed while resolving one path. Guards against link cycles
/// (`a -> b -> a`), which are legal to create.
const MAX_SOFT_LINK_DEPTH: u8 = 16;
+7 -1
View File
@@ -663,11 +663,17 @@ impl DatasetBuilder {
}
pub fn with_u8_data(&mut self, data: &[u8]) -> &mut Self {
self.with_u8_data_owned(data.to_vec())
}
/// [`Self::with_u8_data`] without the copy: the builder takes the
/// vector. For large data this halves the memory a write needs.
pub fn with_u8_data_owned(&mut self, data: Vec<u8>) -> &mut Self {
self.datatype = Some(make_u8_type());
self.data = Some(data.to_vec());
if self.shape.is_none() {
self.shape = Some(vec![data.len() as u64]);
}
self.data = Some(data);
self
}
+50 -15
View File
@@ -36,23 +36,46 @@ let values: Vec<f64> = temp.read_f64()?;
| Item | What |
|---|---|
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable` (a name or a path into a subgroup), `global_attrs`, `group` (a name or a path), `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `variable_names`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `stored_shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable |
| `NcType` | the NetCDF type of a variable: an atomic type (`Byte` ... `UInt64`, `Float`, `Double`, `Char`, `String`) or the class of a user-defined one (`Enum`, `Compound`, `VLen`, `Opaque`); `#[non_exhaustive]` |
Variables and dimensions follow netCDF-C:
Groups, dimensions and variables are what netCDF-C 4.9.3 reports
(`libhdf5/hdf5open.c`), names, order and types included. The first call
that needs them reads the metadata of the whole file, as `nc_open` does:
- A variable's dimensions are the ones the file names: the ids in its
`_Netcdf4Coordinates` attribute, else the dimension scales its
`DIMENSION_LIST` references, found in its group or a parent group. Only
an axis the file names no dimension for (an HDF5 file not written by a
netCDF library) gets the first dimension of the group of the same size,
else an anonymous `dim_<size>`.
- Dimension scales that are only dimensions are not variables; a dataset
- Order: a group's links in creation order when it tracks it (netCDF-4
files do), else in name order (h5py's default); groups and variables in
that order, dimensions by id.
- A dimension scale defines a dimension (its `_Netcdf4Dimid`, else the
next free id; unlimited when its first axis is or its length is 0);
scales that are only dimensions are not variables, and a dataset
`_nc4_non_coord_<name>` is the variable `<name>`.
- A variable's dimensions are the ones the file names: the ids in its
`_Netcdf4Coordinates` attribute (any group), else — when its first axis
has one — the dimension scales its `DIMENSION_LIST` attaches, found in
its group or a parent group.
- Otherwise (an HDF5 file not written by a netCDF library) its axes get
netCDF-C's **phony dimensions** `phony_dim_<n>`: per axis, the first
dimension of the variable's group with the same length and
unlimitedness that an earlier axis of the variable does not use, else a
new one. The numbers run file-wide, a group's subgroups before its own
variables.
- Datasets of types netCDF-C cannot represent are not variables:
references, bit fields, time and array types, and compounds, enums and
VLENs over members or base types that are not netCDF atomic types (4-
and 8-byte floats only) or types netCDF-C has read before in the file.
Enum, compound, VLEN and opaque datasets are variables of those classes
(netCDF4-python itself leaves out opaque ones).
- Two deliberate differences: floats of other than 4 or 8 bytes (half,
bfloat16, 4/6/8-bit floats, `long double`) are `Float`/`Double` and
read as numbers, where netCDF-C 4.9.3 on libhdf5 1.14.6 labels them
`NC_STRING`; and an axis netCDF-C leaves without a dimension (an id or
scale it cannot find, or no scale on an axis after a first one that has
one, where it reads uninitialised memory) gets one by the phony rule.
- A variable along an unlimited dimension has the dimension's length:
`shape` is that length, and the reads return that many values, the
records the variable has not written as its fill value (`_FillValue`,
@@ -60,11 +83,23 @@ Variables and dimensions follow netCDF-C:
`stored_shape` is the HDF5 dataset's extent.
No cargo features. Tests compare against files written by netCDF4-python,
h5py dimension scales, h5netcdf and xarray, variable by variable with what
netCDF4-python reads (`tests/interop_tests.rs`; the CI job requires them
with `CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is
not installed). What the HDF5 reader underneath cannot read is listed in
[`docs/known-issues.md`](../../docs/known-issues.md).
h5py (with and without dimension scales, every HDF5 type class), h5netcdf
and xarray, with what netCDF4-python reads and with what netCDF-C itself
reports (`tests/interop_tests.rs`; `tests/netcdf_c_view.py` calls the
libnetcdf netCDF4-python bundles; the CI job requires them with
`CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is not
installed). `tests/corpus_vs_netcdf_c.rs` compares every file of a corpus
netCDF-C opens, when `CLAWHDF5_NETCDF_CORPUS` names one: over the
conformance corpus, 420 of 429 files match (tank, 2026-09-29), and the
other 9 are explained in `tests/corpus_known_differences.txt`:
```sh
CLAWHDF5_NETCDF_CORPUS=conformance/.cache/corpus CLAWHDF5_PYTHON=$PWD/.venv/bin/python \
cargo test -p clawhdf5-netcdf4 --test corpus_vs_netcdf_c -- --nocapture
```
The differences from netCDF-C, and what the HDF5 reader underneath cannot
read, are listed in [`docs/known-issues.md`](../../docs/known-issues.md).
## License
+7 -204
View File
@@ -1,16 +1,15 @@
//! NetCDF-4 dimension representation.
//!
//! Dimensions in NetCDF-4 are stored as HDF5 datasets with the CLASS=DIMENSION_SCALE
//! attribute and a `_Netcdf4Dimid` attribute. Unlimited dimensions are detected via
//! the HDF5 dataspace max_dimensions (u64::MAX indicates unlimited); their length
//! is the largest extent of the variables attached to them.
//! attribute and a `_Netcdf4Dimid` attribute; variables name theirs in
//! `_Netcdf4Coordinates` and `DIMENSION_LIST`. A file without them gets
//! netCDF-C's phony dimensions. How they are put together is in
//! `crate::model`.
use std::collections::HashMap;
use clawhdf5::AttrValue;
use crate::error::Error;
/// A NetCDF-4 dimension.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Dimension {
@@ -22,202 +21,10 @@ pub struct Dimension {
pub is_unlimited: bool,
}
/// A dimension scale of one group: the dataset that defines a dimension.
#[derive(Debug, Clone)]
pub(crate) struct Scale {
/// Object header address of the scale's dataset (what a variable's
/// `DIMENSION_LIST` references).
pub address: u64,
/// Its `_Netcdf4Dimid` (what a variable's `_Netcdf4Coordinates` lists).
pub dimid: Option<i64>,
/// Index of its dimension in [`GroupDims::dims`].
pub dim: usize,
}
/// The dimensions a group defines, with the scales that define them.
#[derive(Debug, Clone, Default)]
pub(crate) struct GroupDims {
/// The group's dimensions, in `_Netcdf4Dimid` order (then discovery order).
pub dims: Vec<Dimension>,
/// The dimension scales behind `dims`; empty when the group has no
/// dimension scale and `dims` were inferred from 1-D datasets.
pub scales: Vec<Scale>,
}
impl GroupDims {
/// The dimension defined by the scale at `address`.
pub fn by_address(&self, address: u64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.address == address)
.map(|s| &self.dims[s.dim])
}
/// The dimension whose scale has `_Netcdf4Dimid` `id`.
pub fn by_dimid(&self, id: i64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.dimid == Some(id))
.map(|s| &self.dims[s.dim])
}
}
/// The dimensions of an HDF5 group (root or subgroup).
///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// dimension's size is the dataset's first (and typically only) shape extent.
/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`. A group with no dimension
/// scale at all (not written by a netCDF library) gets one dimension per
/// 1-D dataset instead.
pub(crate) fn group_dims(
file: &clawhdf5::File,
group: &clawhdf5::Group<'_>,
) -> Result<GroupDims, Error> {
let addresses: HashMap<String, u64> = group.entries()?.into_iter().collect();
let dataset_names = group.datasets()?;
// (dimid, dimension, scale address), in discovery order.
let mut found: Vec<(Option<i64>, Dimension, u64)> = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let attrs = ds.attrs()?;
if !is_dimension_scale(&attrs) {
continue;
}
let Some(&address) = addresses.get(ds_name) else {
continue;
};
let shape = ds.shape()?;
let is_unlimited = is_unlimited(&ds);
let size = if is_unlimited {
unlimited_len(file, &attrs, &shape)
} else {
shape.first().copied().unwrap_or(0)
};
let dim = Dimension {
name: ds_name.clone(),
size,
is_unlimited,
};
found.push((get_dimid(&attrs), dim, address));
}
if found.is_empty() {
// Fallback: infer dimensions from dataset shapes and names.
// In NetCDF-4, coordinate variables are datasets whose name matches
// a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
let mut dims = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
if shape.len() == 1 {
dims.push(Dimension {
name: ds_name.clone(),
size: shape[0],
is_unlimited: is_unlimited(&ds),
});
}
}
return Ok(GroupDims {
dims,
scales: Vec::new(),
});
}
// By dimid; scales without one keep their discovery order after those
// with one (the sort is stable).
found.sort_by_key(|(id, ..)| (id.is_none(), id.unwrap_or(0)));
let mut out = GroupDims::default();
for (i, (dimid, dim, address)) in found.into_iter().enumerate() {
out.dims.push(dim);
out.scales.push(Scale {
address,
dimid,
dim: i,
});
}
Ok(out)
}
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that
/// is only a dimension, not also a (coordinate) variable.
const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF variable";
/// The current length of an unlimited dimension, as netCDF-C reports it
/// (`NC4_inq_dim` → `nc4_find_dim_len`): the largest current extent, along
/// the dimension, of the variables that use it, in any group; 0 when none
/// has been written. netCDF-C does not extend a dimension scale that is not
/// also a variable, so such a scale's own extent (0) is not counted; a
/// coordinate variable's is. The variables are the scale's attachments,
/// listed with the axis they use in its `REFERENCE_LIST` attribute (the
/// mirror of each variable's `DIMENSION_LIST`). Attachments that cannot be
/// read are skipped; without a readable `REFERENCE_LIST` the length is the
/// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0);
let is_variable = !is_pure_dimension(attrs);
let Some(refs) = reference_list(file, attrs) else {
return own;
};
refs.into_iter()
.filter_map(|(address, axis)| {
let shape = file.dataset_at(address).ok()?.shape().ok()?;
shape.get(usize::try_from(axis).ok()?).copied()
})
.chain(is_variable.then_some(own))
.max()
.unwrap_or(0)
}
/// The `(dataset address, axis)` pairs of a dimension scale's
/// `REFERENCE_LIST` attribute (HDF5 dimension scales: a compound of an
/// object reference `dataset` and an integer `dimension`), or `None` when it
/// is missing or not in that form.
fn reference_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<(u64, u64)>> {
use clawhdf5_format::data_read::{read_compound_field, read_object_references};
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("REFERENCE_LIST") else {
return None;
};
let dataset = read_compound_field(data, datatype, "dataset").ok()?;
let addresses = read_object_references(
&dataset.raw_data,
&dataset.datatype,
file.superblock().offset_size,
)
.ok()?;
let dimension = read_compound_field(data, datatype, "dimension").ok()?;
let Datatype::FixedPoint {
size, byte_order, ..
} = dimension.datatype
else {
return None;
};
let size = usize::try_from(size).ok().filter(|s| (1..=8).contains(s))?;
let axes = dimension.raw_data.chunks_exact(size).map(|b| {
let mut v = [0u8; 8];
match byte_order {
DatatypeByteOrder::BigEndian => {
v[8 - size..].copy_from_slice(b);
u64::from_be_bytes(v)
}
_ => {
v[..size].copy_from_slice(b);
u64::from_le_bytes(v)
}
}
});
if axes.len() != addresses.len() {
return None;
}
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
}
/// The dimension scale attached to each axis of a variable, from its
/// `DIMENSION_LIST` attribute (HDF5 dimension scales: one variable-length
/// sequence of object references per axis) — the address of the scale
@@ -281,13 +88,9 @@ pub(crate) fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
pub(crate) fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
match attrs.get("_Netcdf4Dimid") {
Some(AttrValue::I64(id)) => Some(*id),
Some(AttrValue::U64(id)) => Some(*id as i64),
Some(AttrValue::U64(id)) => i64::try_from(*id).ok(),
Some(AttrValue::I64Array(ids)) if ids.len() == 1 => Some(ids[0]),
Some(AttrValue::U64Array(ids)) if ids.len() == 1 => i64::try_from(ids[0]).ok(),
_ => None,
}
}
/// Whether a dataset's first axis is unlimited (`max_dimensions[0] ==
/// u64::MAX` in its dataspace).
fn is_unlimited(ds: &clawhdf5::Dataset<'_>) -> bool {
matches!(ds.max_dimensions(), Ok(Some(max_dims)) if max_dims.first() == Some(&u64::MAX))
}
+35 -31
View File
@@ -7,36 +7,36 @@ use std::collections::HashMap;
use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension};
use crate::dimension::Dimension;
use crate::error::Error;
use crate::scope::{self, Scope};
use crate::model::Model;
use crate::variable::Variable;
/// A NetCDF-4 group corresponding to an HDF5 group.
pub struct NetCDF4Group<'f> {
/// Group name.
name: String,
/// Path of the group from the root (`/`-separated).
path: String,
/// Underlying HDF5 file.
file: &'f clawhdf5::File,
/// Underlying HDF5 group.
hdf5_group: clawhdf5::Group<'f>,
/// The file's groups, dimensions and variables.
model: &'f Model,
/// This group's index in `model`.
index: usize,
}
impl<'f> NetCDF4Group<'f> {
/// Create a new NetCDF4Group from an HDF5 group.
/// The group `index` of `model`.
pub(crate) fn new(
name: String,
path: String,
file: &'f clawhdf5::File,
hdf5_group: clawhdf5::Group<'f>,
model: &'f Model,
index: usize,
) -> Self {
Self {
name,
path,
file,
hdf5_group,
model,
index,
}
}
@@ -45,52 +45,56 @@ impl<'f> NetCDF4Group<'f> {
&self.name
}
/// List dimensions defined in this group (not those of its parent
/// groups, which its variables can also use).
/// List dimensions defined in this group, in dimension-id order (not
/// those of its parent groups, which its variables can also use).
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
Ok(dimension::group_dims(self.file, &self.hdf5_group)?.dims)
Ok(self.model.dimensions(self.index))
}
/// List variables in this group: its datasets, except the dimension
/// scales that are only dimensions. Their dimensions can be defined in
/// this group or a parent group.
/// List variables in this group, in netCDF-C's order: its datasets,
/// except the dimension scales that are only dimensions and datasets
/// of types netCDF-C cannot represent. Their dimensions can be defined
/// in this group or a parent group.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
Scope::new(self.file, &self.path)?.variables()
self.model.variables(self.file, self.index)
}
/// Get a specific variable by name.
/// Get a specific variable by name (or by path, `"sub/var"`).
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
scope::variable_at(self.file, &self.path, name)
self.model.variable(self.file, self.index, name)
}
/// Read all attributes of this group.
pub fn attrs(&self) -> Result<HashMap<String, AttrValue>, Error> {
Ok(self.hdf5_group.attrs()?)
Ok(self
.file
.group_at(self.model.group_address(self.index))
.attrs()?)
}
/// List subgroup names.
/// List subgroup names, in netCDF-C's order.
pub fn group_names(&self) -> Result<Vec<String>, Error> {
Ok(self.hdf5_group.groups()?)
Ok(self.model.group_names(self.index))
}
/// Get a subgroup by name.
/// Get a subgroup by name (or by path, `"a/b"`).
pub fn group(&self, name: &str) -> Result<NetCDF4Group<'f>, Error> {
let hdf5_group = self
.hdf5_group
.group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?;
let index = self
.model
.find_group(self.index, name)
.ok_or_else(|| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(
name.to_string(),
format!("{}/{name}", self.path),
self.file,
hdf5_group,
self.model,
index,
))
}
/// The names of this group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
Scope::new(self.file, &self.path)?.variable_names()
Ok(self.model.variable_names(self.index))
}
}
+48 -23
View File
@@ -27,7 +27,7 @@ pub mod cf;
pub mod dimension;
pub mod error;
pub mod group;
mod scope;
mod model;
pub mod types;
pub mod variable;
@@ -40,26 +40,49 @@ pub use types::NcType;
pub use variable::Variable;
use std::collections::HashMap;
use std::sync::OnceLock;
use model::Model;
/// A NetCDF-4 file reader.
///
/// Wraps a clawhdf5 File and provides NetCDF-4 semantics: dimensions,
/// variables with CF attributes, groups, and type mapping.
///
/// The groups, dimensions and variables are what netCDF-C reports for the
/// file (see the crate README): the first call that needs them reads the
/// metadata of the whole file, as `nc_open` does, and keeps it; the values
/// are read when asked for.
pub struct NetCDF4File {
hdf5: clawhdf5::File,
model: OnceLock<Model>,
}
impl NetCDF4File {
/// Open a NetCDF-4 file from a filesystem path.
pub fn open<P: AsRef<std::path::Path>>(path: P) -> Result<Self, Error> {
let hdf5 = clawhdf5::File::open(path)?;
Ok(Self { hdf5 })
Ok(Self::from_hdf5(clawhdf5::File::open(path)?))
}
/// Open a NetCDF-4 file from in-memory bytes.
pub fn from_bytes(data: Vec<u8>) -> Result<Self, Error> {
let hdf5 = clawhdf5::File::from_bytes(data)?;
Ok(Self { hdf5 })
Ok(Self::from_hdf5(clawhdf5::File::from_bytes(data)?))
}
fn from_hdf5(hdf5: clawhdf5::File) -> Self {
Self {
hdf5,
model: OnceLock::new(),
}
}
/// The file's groups, dimensions and variables, read on first use.
fn model(&self) -> Result<&Model, Error> {
if let Some(model) = self.model.get() {
return Ok(model);
}
let model = Model::build(&self.hdf5)?;
Ok(self.model.get_or_init(|| model))
}
/// Get the _NCProperties root attribute, if present.
@@ -74,27 +97,29 @@ impl NetCDF4File {
}
}
/// List dimensions defined in the root group.
/// List dimensions defined in the root group, in dimension-id order.
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
Ok(dimension::group_dims(&self.hdf5, &self.hdf5.root())?.dims)
Ok(self.model()?.dimensions(0))
}
/// List all variables in the root group: its datasets, except the
/// dimension scales that are only dimensions (netCDF-C does not list
/// them either).
/// List the variables of the root group, in netCDF-C's order: its
/// datasets, except the dimension scales that are only dimensions and
/// datasets of types netCDF-C cannot represent (see
/// [`NcType`]).
pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> {
scope::Scope::new(&self.hdf5, "/")?.variables()
self.model()?.variables(&self.hdf5, 0)
}
/// The names of the root group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
scope::Scope::new(&self.hdf5, "/")?.variable_names()
Ok(self.model()?.variable_names(0))
}
/// Get a specific variable by name from the root group.
/// Get a specific variable by name from the root group; the name may be
/// a path into a subgroup (`"sub/var"`).
pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> {
scope::variable_at(&self.hdf5, "", name)
self.model()?.variable(&self.hdf5, 0, name)
}
/// Read all global (root group) attributes.
@@ -102,22 +127,22 @@ impl NetCDF4File {
Ok(self.hdf5.root().attrs()?)
}
/// List subgroup names in the root group.
/// List subgroup names in the root group, in netCDF-C's order.
pub fn group_names(&self) -> Result<Vec<String>, Error> {
Ok(self.hdf5.root().groups()?)
Ok(self.model()?.group_names(0))
}
/// Get a subgroup by name.
/// Get a subgroup by name (or by path, `"a/b"`).
pub fn group(&self, name: &str) -> Result<NetCDF4Group<'_>, Error> {
let hdf5_group = self
.hdf5
.group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?;
let model = self.model()?;
let index = model
.find_group(0, name)
.ok_or_else(|| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(
name.to_string(),
name.to_string(),
&self.hdf5,
hdf5_group,
model,
index,
))
}
+853
View File
@@ -0,0 +1,853 @@
//! A file's groups, dimensions and variables, as netCDF-C reads them.
//!
//! netCDF-C (`libhdf5/hdf5open.c`, 4.9.3) reads a file in two passes, and
//! the names, order and sharing of the dimensions depend on both, so this
//! module replays them for the whole file at once:
//!
//! 1. `rec_read_metadata`: each group's links in creation order when the
//! group tracks it, else in name order; a group's datasets and named
//! datatypes before its subgroups, which follow in the same order. A
//! dimension scale (`CLASS` `DIMENSION_SCALE`) defines a dimension with
//! the id in its `_Netcdf4Dimid`, else the next free id (ids are
//! file-wide), its first extent as length, unlimited when that axis is
//! (or the length is 0); it is a variable too unless its `NAME` says it
//! is only a dimension. Every other dataset is a variable, unless its
//! type is one netCDF-C cannot represent ([`VarTypes::nc_type`]); a
//! dataset `_nc4_non_coord_<name>` is the variable `<name>`.
//! 2. `rec_match_dimscales`, subgroups first, then the group's variables in
//! order: a variable gets the dimensions whose ids its
//! `_Netcdf4Coordinates` lists (looked up file-wide), else — when its
//! `DIMENSION_LIST` attaches a scale to its first axis — the scales that
//! list attaches, looked up in its group and then each parent, else
//! "phony" dimensions (`create_phony_dims`): per axis, the first
//! dimension of the variable's group of that length and unlimitedness
//! not already used by an earlier axis of the variable, else a new one
//! called `phony_dim_<id>`.
//!
//! An unlimited dimension's length is the largest extent, along it, of the
//! variables that use it in its group and the groups below
//! (`nc4_find_dim_len`).
//!
//! Where netCDF-C 4.9.3 leaves an axis without a dimension (an id or a scale
//! it cannot find, or an axis without a scale of a variable whose first
//! axis has one — there it reads uninitialised memory), netCDF4-python
//! cannot open the file; this crate gives such an axis a dimension by the
//! phony rule instead.
use std::collections::HashMap;
use clawhdf5::AttrValue;
use clawhdf5_format::datatype::{Datatype, DatatypeByteOrder};
use clawhdf5_format::object_header::{ObjectClass, ObjectHeader};
use crate::dimension::{self, Dimension};
use crate::error::Error;
use crate::types::NcType;
use crate::variable::Variable;
/// The prefix netCDF-C gives the dataset of a variable that has a
/// dimension's name but is not that dimension's coordinate variable (the
/// dimension's scale holds the name).
const NON_COORD_PREFIX: &str = "_nc4_non_coord_";
/// Groups read at most, a guard against files whose groups are hard-linked
/// into each other many times over (netCDF-C reads each link as its own
/// group).
const MAX_GROUPS: usize = 100_000;
/// A file as netCDF-C sees it.
#[derive(Debug)]
pub(crate) struct Model {
/// Every group; the root is the first.
groups: Vec<Group>,
/// Every dimension, in the order they were created.
dims: Vec<Dim>,
}
#[derive(Debug)]
struct Group {
/// Object header address of the HDF5 group.
address: u64,
/// Its subgroups, `(name, index in Model::groups)`, in netCDF-C's order.
children: Vec<(String, usize)>,
parent: Option<usize>,
/// The dimensions it defines (indexes in `Model::dims`), in creation
/// order.
dims: Vec<usize>,
/// Its variables, in netCDF-C's order.
vars: Vec<Var>,
}
#[derive(Debug)]
struct Dim {
id: i64,
name: String,
/// The length it was created with (for an unlimited dimension, replaced
/// by its current length once every variable has its dimensions).
len: u64,
unlimited: bool,
/// Object header address of its dimension scale; `None` for a phony
/// dimension.
scale: Option<u64>,
}
#[derive(Debug)]
struct Var {
/// The netCDF name.
name: String,
/// Whether its dataset is `_nc4_non_coord_<name>`.
non_coord: bool,
address: u64,
nc_type: NcType,
extent: Vec<u64>,
/// Whether each axis is unlimited in the dataspace.
unlimited: Vec<bool>,
/// How netCDF-C finds its dimensions.
source: DimSource,
/// Its dimensions (indexes in `Model::dims`), one per axis, once found.
dims: Vec<usize>,
}
#[derive(Debug)]
enum DimSource {
/// A one-dimensional coordinate variable: its own scale's dimension.
OwnScale(usize),
/// The ids in `_Netcdf4Coordinates`; for a multi-dimensional coordinate
/// variable, also its own dimension (for the first axis, should an id
/// not be found).
Coordinates(Vec<i64>, Option<usize>),
/// The scale `DIMENSION_LIST` attaches to each axis (the first has one).
Scales(Vec<Option<u64>>, Option<usize>),
/// None: phony dimensions.
Phony(Option<usize>),
}
impl Model {
/// Read the metadata of `file` as netCDF-C does.
pub fn build(file: &clawhdf5::File) -> Result<Self, Error> {
let mut builder = Builder {
file,
model: Model {
groups: Vec::new(),
dims: Vec::new(),
},
next_id: 0,
types: VarTypes::default(),
};
let root = file.superblock().root_group_address;
builder.model.groups.push(Group {
address: root,
children: Vec::new(),
parent: None,
dims: Vec::new(),
vars: Vec::new(),
});
builder.read_group(0, &mut vec![root])?;
builder.match_dims(0);
builder.unlimited_lengths();
Ok(builder.model)
}
/// The group at `path` (`/`-separated names), from the group `from`.
pub fn find_group(&self, from: usize, path: &str) -> Option<usize> {
path.split('/')
.filter(|p| !p.is_empty())
.try_fold(from, |g, name| {
self.groups[g]
.children
.iter()
.find(|(n, _)| n == name)
.map(|&(_, i)| i)
})
}
/// The object header address of group `g`.
pub fn group_address(&self, g: usize) -> u64 {
self.groups[g].address
}
/// The names of group `g`'s subgroups, in netCDF-C's order.
pub fn group_names(&self, g: usize) -> Vec<String> {
self.groups[g]
.children
.iter()
.map(|(n, _)| n.clone())
.collect()
}
/// The dimensions group `g` defines, in id order (as `nc_inq_dimids`).
pub fn dimensions(&self, g: usize) -> Vec<Dimension> {
let mut dims: Vec<&Dim> = self.groups[g].dims.iter().map(|&d| &self.dims[d]).collect();
dims.sort_by_key(|d| d.id);
dims.into_iter().map(Dim::dimension).collect()
}
/// The names of group `g`'s variables, in netCDF-C's order.
pub fn variable_names(&self, g: usize) -> Vec<String> {
self.groups[g].vars.iter().map(|v| v.name.clone()).collect()
}
/// Group `g`'s variables, in netCDF-C's order.
pub fn variables<'f>(
&self,
file: &'f clawhdf5::File,
g: usize,
) -> Result<Vec<Variable<'f>>, Error> {
self.groups[g]
.vars
.iter()
.map(|v| self.open(file, v))
.collect()
}
/// The variable `name` of group `g`; `name` may be a path (`"sub/var"`)
/// relative to it. Of two variables of one name (datasets `<name>` and
/// `_nc4_non_coord_<name>`), the second.
pub fn variable<'f>(
&self,
file: &'f clawhdf5::File,
g: usize,
name: &str,
) -> Result<Variable<'f>, Error> {
let not_found = || Error::VariableNotFound(name.to_string());
let trimmed = name.trim_start_matches('/');
let (g, leaf) = match trimmed.rsplit_once('/') {
Some((dir, leaf)) => (self.find_group(g, dir).ok_or_else(not_found)?, leaf),
None => (g, trimmed),
};
let vars = &self.groups[g].vars;
let var = vars
.iter()
.find(|v| v.name == leaf && v.non_coord)
.or_else(|| vars.iter().find(|v| v.name == leaf))
.ok_or_else(not_found)?;
self.open(file, var)
}
fn open<'f>(&self, file: &'f clawhdf5::File, var: &Var) -> Result<Variable<'f>, Error> {
let ds = file.dataset_at(var.address)?;
let attrs = ds.attrs().unwrap_or_default();
let dims = var.dims.iter().map(|&d| self.dims[d].dimension()).collect();
Ok(Variable::new(
var.name.clone(),
ds,
dims,
attrs,
var.nc_type,
))
}
}
impl Dim {
fn dimension(&self) -> Dimension {
Dimension {
name: self.name.clone(),
size: self.len,
is_unlimited: self.unlimited,
}
}
}
struct Builder<'f> {
file: &'f clawhdf5::File,
model: Model,
/// netCDF-C's `next_dimid`.
next_id: i64,
types: VarTypes,
}
impl Builder<'_> {
/// Pass 1 for group `g` and, after its own links, its subgroups.
/// `ancestors` holds the addresses of the groups from the root to `g`,
/// so that a group linked into itself is not read forever.
fn read_group(&mut self, g: usize, ancestors: &mut Vec<u64>) -> Result<(), Error> {
let address = self.model.groups[g].address;
let sb = self.file.superblock();
let mut subgroups = Vec::new();
for (name, child) in ordered_entries(self.file, address)? {
let Ok(header) =
ObjectHeader::parse_in(self.file.storage(), child, sb.offset_size, sb.length_size)
else {
continue;
};
match header.object_class() {
Some(ObjectClass::Dataset) => self.read_dataset(g, name, child),
Some(ObjectClass::NamedDatatype) => {
if let Some(dt) = header_datatype(&header) {
self.types.named(&dt);
}
}
_ if is_group(&header) => subgroups.push((name, child)),
_ => {}
}
}
for (name, child) in subgroups {
if ancestors.contains(&child) || self.model.groups.len() >= MAX_GROUPS {
continue;
}
let index = self.model.groups.len();
self.model.groups.push(Group {
address: child,
children: Vec::new(),
parent: Some(g),
dims: Vec::new(),
vars: Vec::new(),
});
self.model.groups[g].children.push((name, index));
ancestors.push(child);
let read = self.read_group(index, ancestors);
ancestors.pop();
read?;
}
Ok(())
}
/// `read_dataset`: a dimension for a dimension scale, a variable for
/// the rest. A dataset that cannot be opened is skipped.
fn read_dataset(&mut self, g: usize, name: String, address: u64) {
let Ok(ds) = self.file.dataset_at(address) else {
return;
};
let attrs = ds.attrs().unwrap_or_default();
let Ok(extent) = ds.shape() else {
return;
};
let max = ds.max_dimensions().ok().flatten();
let unlimited: Vec<bool> = (0..extent.len())
.map(|i| matches!(&max, Some(m) if m.get(i) == Some(&u64::MAX)))
.collect();
let mut own = None;
if dimension::is_dimension_scale(&attrs) && !extent.is_empty() {
// read_scale
let id = match dimension::get_dimid(&attrs) {
Some(id) => {
if id >= self.next_id {
self.next_id = id.saturating_add(1);
}
id
}
None => self.take_id(),
};
let len = extent[0];
own = Some(self.add_dim(
g,
Dim {
id,
name: name.clone(),
len,
unlimited: unlimited[0] || len == 0,
scale: Some(address),
},
));
if dimension::is_pure_dimension(&attrs) {
return;
}
}
// read_var: a type netCDF-C cannot represent drops the variable
// (the dimension of a scale stays).
let Some(nc_type) = ds
.raw_datatype()
.ok()
.and_then(|dt| self.types.nc_type(&dt))
else {
return;
};
let rank = extent.len();
let coordinates = coordinates(&attrs).filter(|ids| ids.len() == rank && rank > 0);
let source = match (own, coordinates) {
(Some(dim), _) if rank == 1 => DimSource::OwnScale(dim),
(_, Some(ids)) => DimSource::Coordinates(ids, own),
_ => match dimension::dimension_list(self.file, &attrs) {
Some(scales)
if scales.len() == rank && scales.first().is_some_and(Option::is_some) =>
{
DimSource::Scales(scales, own)
}
_ => DimSource::Phony(own),
},
};
let (name, non_coord) = match name.strip_prefix(NON_COORD_PREFIX) {
Some(rest) if !rest.is_empty() => (rest.to_string(), true),
_ => (name, false),
};
self.model.groups[g].vars.push(Var {
name,
non_coord,
address,
nc_type,
extent,
unlimited,
source,
dims: Vec::new(),
});
}
fn take_id(&mut self) -> i64 {
let id = self.next_id;
self.next_id = self.next_id.saturating_add(1);
id
}
fn add_dim(&mut self, g: usize, dim: Dim) -> usize {
let index = self.model.dims.len();
self.model.dims.push(dim);
self.model.groups[g].dims.push(index);
index
}
/// Pass 2 (`rec_match_dimscales`): subgroups first, then the group's
/// variables in order.
fn match_dims(&mut self, g: usize) {
let children: Vec<usize> = self.model.groups[g]
.children
.iter()
.map(|&(_, c)| c)
.collect();
for child in children {
self.match_dims(child);
}
for v in 0..self.model.groups[g].vars.len() {
let var = &self.model.groups[g].vars[v];
let rank = var.extent.len();
let mut found: Vec<Option<usize>> = vec![None; rank];
match &var.source {
DimSource::OwnScale(dim) => found[0] = Some(*dim),
DimSource::Coordinates(ids, own) => {
for (slot, id) in found.iter_mut().zip(ids) {
*slot = self.dim_by_id(*id);
}
if found[0].is_none() {
found[0] = *own;
}
}
DimSource::Scales(scales, own) => {
for (slot, scale) in found.iter_mut().zip(scales) {
*slot = scale.and_then(|s| self.dim_by_scale(g, s));
}
if found[0].is_none() {
found[0] = *own;
}
}
DimSource::Phony(own) => {
if rank > 0 {
found[0] = *own;
}
}
}
let mut dims: Vec<usize> = Vec::with_capacity(rank);
for (axis, found) in found.into_iter().enumerate() {
let dim = match found {
Some(dim) => dim,
None => {
let var = &self.model.groups[g].vars[v];
let (len, unlimited) = (var.extent[axis], var.unlimited[axis]);
self.phony_dim(g, len, unlimited, &dims)
}
};
dims.push(dim);
}
self.model.groups[g].vars[v].dims = dims;
}
}
/// The dimension with id `id`: the last one created with it
/// (`nc4_find_dim` looks ids up in a file-wide table, where a later
/// dimension of the same id replaces an earlier one).
fn dim_by_id(&self, id: i64) -> Option<usize> {
self.model.dims.iter().rposition(|d| d.id == id)
}
/// The dimension of the scale at `address`, in group `g` or the nearest
/// parent that has it.
fn dim_by_scale(&self, g: usize, address: u64) -> Option<usize> {
let mut group = Some(g);
while let Some(i) = group {
let found = self.model.groups[i]
.dims
.iter()
.copied()
.find(|&d| self.model.dims[d].scale == Some(address));
if found.is_some() {
return found;
}
group = self.model.groups[i].parent;
}
None
}
/// `create_phony_dims` for one axis: the first dimension of group `g`
/// of length `len` and unlimitedness `unlimited` that no earlier axis
/// of the variable uses (`taken`), else a new `phony_dim_<id>`.
fn phony_dim(&mut self, g: usize, len: u64, unlimited: bool, taken: &[usize]) -> usize {
let dims = &self.model.dims;
let existing = self.model.groups[g].dims.iter().copied().find(|&d| {
let dim = &dims[d];
dim.len == len
&& dim.unlimited == unlimited
&& !taken.iter().any(|&t| dims[t].id == dim.id)
});
if let Some(dim) = existing {
return dim;
}
let id = self.take_id();
self.add_dim(
g,
Dim {
id,
name: format!("phony_dim_{id}"),
len,
// `nc4_dim_list_add`: a length of 0 is NC_UNLIMITED.
unlimited: unlimited || len == 0,
scale: None,
},
)
}
/// Each unlimited dimension's length: the largest extent along it of
/// the variables using it in its group and the groups below.
fn unlimited_lengths(&mut self) {
let mut owner = vec![0usize; self.model.dims.len()];
for (g, group) in self.model.groups.iter().enumerate() {
for &d in &group.dims {
owner[d] = g;
}
}
let mut lens = vec![0u64; self.model.dims.len()];
for (g, group) in self.model.groups.iter().enumerate() {
for var in &group.vars {
for (&d, &e) in var.dims.iter().zip(&var.extent) {
if self.model.dims[d].unlimited && self.is_within(g, owner[d]) {
lens[d] = lens[d].max(e);
}
}
}
}
for (dim, len) in self.model.dims.iter_mut().zip(lens) {
if dim.unlimited {
dim.len = len;
}
}
}
/// Whether group `g` is `ancestor` or below it.
fn is_within(&self, g: usize, ancestor: usize) -> bool {
let mut group = Some(g);
while let Some(i) = group {
if i == ancestor {
return true;
}
group = self.model.groups[i].parent;
}
false
}
}
/// A group's entries in netCDF-C's order: creation order when the group
/// tracks it, else the byte order of the names.
fn ordered_entries(file: &clawhdf5::File, address: u64) -> Result<Vec<(String, u64)>, Error> {
let mut entries = file.group_at(address).entries()?;
let order = clawhdf5_format::group_v2::links_in_creation_order_in(
file.storage(),
file.superblock(),
address,
)
.ok()
.flatten();
match order {
Some(names) => {
let position: HashMap<&str, usize> = names
.iter()
.enumerate()
.map(|(i, n)| (n.as_str(), i))
.collect();
entries.sort_by_key(|(n, _)| position.get(n.as_str()).copied().unwrap_or(usize::MAX));
}
None => entries.sort_by(|a, b| a.0.as_bytes().cmp(b.0.as_bytes())),
}
Ok(entries)
}
fn is_group(header: &ObjectHeader) -> bool {
use clawhdf5_format::message_type::MessageType;
header.messages.iter().any(|m| {
matches!(
m.msg_type,
MessageType::LinkInfo | MessageType::Link | MessageType::SymbolTable
)
})
}
/// The datatype a named datatype's header holds.
fn header_datatype(header: &ObjectHeader) -> Option<Datatype> {
use clawhdf5_format::message_type::MessageType;
let msg = header
.messages
.iter()
.find(|m| m.msg_type == MessageType::Datatype)?;
Datatype::parse_in_header(&msg.data, header.version)
.ok()
.map(|(dt, _)| dt)
}
/// A variable's `_Netcdf4Coordinates`: the id of the dimension of each
/// axis.
fn coordinates(attrs: &HashMap<String, AttrValue>) -> Option<Vec<i64>> {
match attrs.get("_Netcdf4Coordinates")? {
AttrValue::I64Array(ids) => Some(ids.clone()),
AttrValue::I64(id) => Some(vec![*id]),
AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(),
AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]),
_ => None,
}
}
/// The user-defined types netCDF-C has read so far, and the netCDF type of
/// a dataset.
///
/// netCDF-C keeps every type `read_type` meets in a file-wide list and
/// finds one again by `H5Tequal` on the native types — also a type it
/// failed to read: `read_type` adds the type (with its class, for a
/// compound, enum, variable-length or opaque type) before it looks at the
/// members or base type, and does not remove it when one of those fails.
/// So a dataset of a type netCDF-C skipped once becomes a variable the
/// second time (a compound with a reference member, say), and a compound
/// member of a type it skipped (a bit field) is accepted. This replays
/// that, except that a dataset whose skipped type has no class (a bit
/// field, time, array or complex type) stays hidden: netCDF-C lists the
/// second such dataset with an invalid type (class 0).
#[derive(Debug, Default)]
pub(crate) struct VarTypes {
/// Every type read, with its class when it has one netCDF knows.
known: Vec<(Datatype, Option<NcType>)>,
}
impl VarTypes {
/// The netCDF type netCDF-C gives a dataset of type `dt`
/// (`get_type_info2`), or `None` when it skips the dataset
/// (`NC_EBADTYPID`): a reference, bit field, time or array type, a
/// compound with a member, or an enum or variable-length type with a
/// base type, that is not a netCDF atomic type or a type read before
/// (see the type docs).
///
/// Floats other than 4 and 8 bytes (half floats, bfloat16, the 4-, 6-
/// and 8-bit floats, `long double`) are `Float` (up to 4 bytes) and
/// `Double` here; netCDF-C 4.9.3 on libhdf5 1.14.6 labels them
/// `NC_STRING` (their native type matches none of its own).
pub fn nc_type(&mut self, dt: &Datatype) -> Option<NcType> {
match dt {
Datatype::FixedPoint { size, signed, .. } => int_type(*size, *signed),
Datatype::FloatingPoint { size, .. } => Some(if *size <= 4 {
NcType::Float
} else {
NcType::Double
}),
Datatype::String { size, .. } => Some(if *size > 1 {
NcType::String
} else {
NcType::Char
}),
Datatype::VariableLength {
is_string: true, ..
} => Some(NcType::String),
_ => match self.find(dt) {
Some(class) => class,
None => self.read_type(dt),
},
}
}
/// A named datatype of the file (`read_type` on a committed type).
pub fn named(&mut self, dt: &Datatype) {
if self.find(dt).is_none() {
self.read_type(dt);
}
}
/// The type read before that is `dt`, by its class.
fn find(&self, dt: &Datatype) -> Option<Option<NcType>> {
self.known
.iter()
.find(|(k, _)| native_eq(k, dt))
.map(|&(_, class)| class)
}
/// `read_type`: remember `dt` (not a reference type), then its class if
/// netCDF-C can represent its members or base type.
fn read_type(&mut self, dt: &Datatype) -> Option<NcType> {
let class = match dt {
Datatype::Reference { .. } => return None,
Datatype::Compound { .. } => Some(NcType::Compound),
Datatype::VariableLength { .. } => Some(NcType::VLen),
Datatype::Opaque { .. } => Some(NcType::Opaque),
Datatype::Enumeration { .. } => Some(NcType::Enum),
_ => None,
};
self.known.push((dt.clone(), class));
let parts_ok = match dt {
Datatype::Compound { members, .. } => members.iter().all(|m| match &m.datatype {
Datatype::Array { base_type, .. } => self.is_member_type(base_type),
other => self.is_member_type(other),
}),
Datatype::VariableLength { base_type, .. }
| Datatype::Enumeration { base_type, .. } => self.is_member_type(base_type),
_ => true,
};
class.filter(|_| parts_ok)
}
/// `get_netcdf_type`: whether netCDF-C takes `dt` as the type of a
/// compound member or the base of an enum or variable-length type — an
/// atomic type (4- and 8-byte floats only) or a type read before.
fn is_member_type(&self, dt: &Datatype) -> bool {
match dt {
Datatype::FixedPoint { size, signed, .. } => int_type(*size, *signed).is_some(),
Datatype::FloatingPoint { size: 4 | 8, .. }
| Datatype::String { .. }
| Datatype::VariableLength {
is_string: true, ..
} => true,
_ => self.find(dt).is_some(),
}
}
}
/// The netCDF integer type of an HDF5 integer: the native integer libhdf5
/// converts it to (`H5Tget_native_type`: the smallest at least as wide).
fn int_type(size: u32, signed: bool) -> Option<NcType> {
Some(match (size, signed) {
(1, true) => NcType::Byte,
(1, false) => NcType::UByte,
(2, true) => NcType::Short,
(2, false) => NcType::UShort,
(3..=4, true) => NcType::Int,
(3..=4, false) => NcType::UInt,
(5..=8, true) => NcType::Int64,
(5..=8, false) => NcType::UInt64,
_ => return None,
})
}
/// Whether two datatypes have the same native type (`H5Tequal` after
/// `H5Tget_native_type`): byte order, padding and compound member offsets
/// do not count.
fn native_eq(a: &Datatype, b: &Datatype) -> bool {
use Datatype as D;
match (a, b) {
(
D::FixedPoint {
size: s1,
signed: g1,
..
},
D::FixedPoint {
size: s2,
signed: g2,
..
},
) => g1 == g2 && int_type(*s1, *g1) == int_type(*s2, *g2),
(D::FloatingPoint { size: s1, .. }, D::FloatingPoint { size: s2, .. }) => s1 == s2,
(
D::String {
size: s1,
padding: p1,
charset: c1,
},
D::String {
size: s2,
padding: p2,
charset: c2,
},
) => s1 == s2 && p1 == p2 && c1 == c2,
(
D::VariableLength {
is_string: i1,
base_type: b1,
charset: c1,
..
},
D::VariableLength {
is_string: i2,
base_type: b2,
charset: c2,
..
},
) => i1 == i2 && if *i1 { c1 == c2 } else { native_eq(b1, b2) },
(D::Compound { members: m1, .. }, D::Compound { members: m2, .. }) => {
m1.len() == m2.len()
&& m1
.iter()
.zip(m2)
.all(|(x, y)| x.name == y.name && native_eq(&x.datatype, &y.datatype))
}
(
D::Enumeration {
base_type: b1,
members: m1,
..
},
D::Enumeration {
base_type: b2,
members: m2,
..
},
) => {
native_eq(b1, b2)
&& m1.len() == m2.len()
&& m1.iter().zip(m2).all(|(x, y)| {
x.name == y.name && enum_value(&x.value, b1) == enum_value(&y.value, b2)
})
}
(D::Opaque { size: s1, tag: t1 }, D::Opaque { size: s2, tag: t2 }) => s1 == s2 && t1 == t2,
(
D::Array {
base_type: b1,
dimensions: d1,
},
D::Array {
base_type: b2,
dimensions: d2,
},
) => d1 == d2 && native_eq(b1, b2),
(
D::Reference {
size: s1,
ref_type: r1,
},
D::Reference {
size: s2,
ref_type: r2,
},
) => s1 == s2 && r1 == r2,
(D::BitField { size: s1, .. }, D::BitField { size: s2, .. })
| (D::Time { size: s1, .. }, D::Time { size: s2, .. }) => s1 == s2,
(
D::Complex {
size: s1,
base_type: b1,
},
D::Complex {
size: s2,
base_type: b2,
},
) => s1 == s2 && native_eq(b1, b2),
_ => false,
}
}
/// An enum member's value, in the order of its base type's bytes.
fn enum_value(value: &[u8], base: &Datatype) -> Vec<u8> {
let big = matches!(
base,
Datatype::FixedPoint {
byte_order: DatatypeByteOrder::BigEndian,
..
}
);
let mut v = value.to_vec();
if big {
v.reverse();
}
v
}
-231
View File
@@ -1,231 +0,0 @@
//! A group's variables and the dimensions they are defined on.
//!
//! netCDF-C (`libhdf5/hdf5open.c`) gives a variable its dimensions from the
//! file, never by size: the dimension ids in its `_Netcdf4Coordinates`
//! attribute (each dimension scale's `_Netcdf4Dimid`), else the dimension
//! scales its `DIMENSION_LIST` attribute references, looked up in the
//! variable's group and then each parent group up to the root. Only an axis
//! with neither (a file not written by a netCDF library) gets a dimension
//! by size. Dimension scales that are only dimensions are not variables, and
//! a variable stored as `_nc4_non_coord_<name>` (a variable sharing a
//! dimension's name without being its coordinate variable) is `<name>`.
use std::collections::{HashMap, HashSet};
use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension, GroupDims};
use crate::error::Error;
use crate::variable::Variable;
/// The prefix netCDF-C gives the dataset of a variable that has a
/// dimension's name but is not that dimension's coordinate variable (the
/// dimension's scale holds the name).
const NON_COORD_PREFIX: &str = "_nc4_non_coord_";
/// A group, with the dimensions visible from it.
pub(crate) struct Scope<'f> {
file: &'f clawhdf5::File,
group: clawhdf5::Group<'f>,
/// This group's dimensions, then its parent's, and so on to the root's.
levels: Vec<GroupDims>,
}
impl<'f> Scope<'f> {
/// The group at `path` (`/`-separated from the root; `""` or `"/"` is
/// the root).
pub fn new(file: &'f clawhdf5::File, path: &str) -> Result<Self, Error> {
let parts: Vec<&str> = path.split('/').filter(|p| !p.is_empty()).collect();
let mut levels = Vec::with_capacity(parts.len() + 1);
for n in (0..=parts.len()).rev() {
let group = file.group(&parts[..n].join("/"))?;
levels.push(dimension::group_dims(file, &group)?);
}
let group = file.group(&parts.join("/"))?;
Ok(Self {
file,
group,
levels,
})
}
/// The group's datasets, as `(dataset name, object header address)` in
/// listing order.
fn datasets(&self) -> Result<Vec<(String, u64)>, Error> {
let datasets: HashSet<String> = self.group.datasets()?.into_iter().collect();
Ok(self
.group
.entries()?
.into_iter()
.filter(|(name, _)| datasets.contains(name))
.collect())
}
/// The group's variables: every dataset but the dimension scales that
/// are only dimensions.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let mut variables = Vec::new();
for (ds_name, address) in self.datasets()? {
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
continue;
}
variables.push(self.variable_from(nc_name(&ds_name), address, ds, attrs)?);
}
Ok(variables)
}
/// The names of the group's variables.
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
let mut names = Vec::new();
for (ds_name, address) in self.datasets()? {
let attrs = self.file.dataset_at(address)?.attrs()?;
if !dimension::is_pure_dimension(&attrs) {
names.push(nc_name(&ds_name));
}
}
Ok(names)
}
/// The variable called `name`: the dataset `_nc4_non_coord_<name>` if
/// there is one, else the dataset `<name>` unless it is only a
/// dimension.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let not_found = || Error::VariableNotFound(name.to_string());
let datasets = self.datasets()?;
let prefixed = format!("{NON_COORD_PREFIX}{name}");
let address = datasets
.iter()
.find(|(n, _)| *n == prefixed)
.or_else(|| datasets.iter().find(|(n, _)| n == name))
.map(|&(_, address)| address)
.ok_or_else(not_found)?;
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
return Err(not_found());
}
self.variable_from(nc_name(name), address, ds, attrs)
}
fn variable_from(
&self,
name: String,
address: u64,
ds: clawhdf5::Dataset<'f>,
attrs: HashMap<String, AttrValue>,
) -> Result<Variable<'f>, Error> {
let shape = ds.shape()?;
let dims = self.variable_dims(address, &attrs, &shape);
Ok(Variable::new(name, ds, dims, attrs))
}
/// The first dimension, searching this group and then its ancestors,
/// that `find` picks.
fn find<'a>(
&'a self,
find: impl Fn(&'a GroupDims) -> Option<&'a Dimension>,
) -> Option<Dimension> {
self.levels.iter().find_map(find).cloned()
}
/// The dimensions of the dataset at `address`, one per axis of `shape`,
/// as netCDF-C resolves them (see the module docs).
fn variable_dims(
&self,
address: u64,
attrs: &HashMap<String, AttrValue>,
shape: &[u64],
) -> Vec<Dimension> {
let rank = shape.len();
let mut dims: Vec<Option<Dimension>> = vec![None; rank];
if rank == 0 {
return Vec::new();
}
// A coordinate variable is the scale of its (first) dimension.
dims[0] = self.levels[0].by_address(address).cloned();
if let Some(ids) = coordinates(attrs).filter(|ids| ids.len() == rank) {
for (slot, id) in dims.iter_mut().zip(ids) {
if slot.is_none() {
*slot = self.find(|level| level.by_dimid(id));
}
}
}
if dims.iter().any(Option::is_none)
&& let Some(scales) =
dimension::dimension_list(self.file, attrs).filter(|s| s.len() == rank)
{
for (slot, scale) in dims.iter_mut().zip(scales) {
if slot.is_none()
&& let Some(scale) = scale
{
*slot = self.find(|level| level.by_address(scale));
}
}
}
// Neither: the first dimension of this group of the same size not
// already taken by another such axis, else an anonymous one.
let own = &self.levels[0].dims;
let mut used = vec![false; own.len()];
dims.into_iter()
.zip(shape)
.map(|(dim, &size)| {
dim.unwrap_or_else(|| {
match own
.iter()
.enumerate()
.find(|&(i, d)| !used[i] && d.size == size)
{
Some((i, d)) => {
used[i] = true;
d.clone()
}
None => Dimension {
name: format!("dim_{size}"),
size,
is_unlimited: false,
},
}
})
})
.collect()
}
}
/// The variable `name` of the group at `group_path`; `name` may itself be
/// a path (`"sub/var"`), relative to that group.
pub(crate) fn variable_at<'f>(
file: &'f clawhdf5::File,
group_path: &str,
name: &str,
) -> Result<Variable<'f>, Error> {
match name.trim_start_matches('/').rsplit_once('/') {
Some((dir, leaf)) => Scope::new(file, &format!("{group_path}/{dir}"))
.map_err(|_| Error::VariableNotFound(name.to_string()))?
.variable(leaf),
None => Scope::new(file, group_path)?.variable(name.trim_start_matches('/')),
}
}
/// The netCDF name of the dataset `ds_name`.
fn nc_name(ds_name: &str) -> String {
ds_name
.strip_prefix(NON_COORD_PREFIX)
.unwrap_or(ds_name)
.to_string()
}
/// A variable's `_Netcdf4Coordinates`: the `_Netcdf4Dimid` of the dimension
/// of each axis.
fn coordinates(attrs: &HashMap<String, AttrValue>) -> Option<Vec<i64>> {
match attrs.get("_Netcdf4Coordinates")? {
AttrValue::I64Array(ids) => Some(ids.clone()),
AttrValue::I64(id) => Some(vec![*id]),
AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(),
AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]),
_ => None,
}
}
+28 -3
View File
@@ -4,8 +4,16 @@
use clawhdf5::DType;
/// NetCDF-4 data types corresponding to the standard NetCDF type system.
/// NetCDF-4 data types corresponding to the standard NetCDF type system:
/// the atomic types, and the class of a user-defined type.
///
/// A dataset whose type netCDF-C cannot represent — a reference, bit
/// field, time or array type; a compound with such a member, a
/// half-precision float member, or a member of a user-defined type not
/// read before it; an enum or variable-length type over such a base — is
/// not a variable, as in netCDF-C.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
#[non_exhaustive]
pub enum NcType {
/// NC_BYTE: signed 8-bit integer
Byte,
@@ -29,8 +37,18 @@ pub enum NcType {
Double,
/// NC_STRING: variable-length string
String,
/// NC_CHAR: fixed-length string / character data
/// NC_CHAR: a fixed-length string of one byte (longer ones are
/// `String`, as netCDF-C reads them)
Char,
/// NC_ENUM: an enumeration (user-defined type)
Enum,
/// NC_COMPOUND: a compound (user-defined type)
Compound,
/// NC_VLEN: a variable-length sequence (user-defined type)
VLen,
/// NC_OPAQUE: an opaque type (user-defined type; netCDF4-python skips
/// such variables)
Opaque,
}
impl std::fmt::Display for NcType {
@@ -48,11 +66,16 @@ impl std::fmt::Display for NcType {
NcType::Double => write!(f, "NC_DOUBLE"),
NcType::String => write!(f, "NC_STRING"),
NcType::Char => write!(f, "NC_CHAR"),
NcType::Enum => write!(f, "NC_ENUM"),
NcType::Compound => write!(f, "NC_COMPOUND"),
NcType::VLen => write!(f, "NC_VLEN"),
NcType::Opaque => write!(f, "NC_OPAQUE"),
}
}
}
/// Map a clawhdf5 DType to a NetCDF type.
/// Map a clawhdf5 DType to a NetCDF type. A variable's own type, as
/// netCDF-C reads it, is [`Variable::nc_type`](crate::Variable::nc_type).
pub fn dtype_to_nctype(dtype: &DType) -> NcType {
match dtype {
DType::I8 => NcType::Byte,
@@ -66,6 +89,8 @@ pub fn dtype_to_nctype(dtype: &DType) -> NcType {
DType::F32 => NcType::Float,
DType::F64 => NcType::Double,
DType::String | DType::VariableLengthString => NcType::String,
DType::Compound(_) => NcType::Compound,
DType::Enum(_) => NcType::Enum,
_ => NcType::Char, // fallback for other types
}
}
+18 -9
View File
@@ -16,7 +16,7 @@ use clawhdf5::AttrValue;
use crate::cf::{self, CfAttributes, FillValue};
use crate::dimension::Dimension;
use crate::error::Error;
use crate::types::{NcType, dtype_to_nctype};
use crate::types::NcType;
/// A NetCDF-4 variable backed by an HDF5 dataset.
pub struct Variable<'f> {
@@ -28,6 +28,8 @@ pub struct Variable<'f> {
dims: Vec<Dimension>,
/// The dataset's attributes.
attrs: HashMap<String, AttrValue>,
/// Its netCDF type.
nc_type: NcType,
}
impl<'f> Variable<'f> {
@@ -37,12 +39,14 @@ impl<'f> Variable<'f> {
dataset: clawhdf5::Dataset<'f>,
dims: Vec<Dimension>,
attrs: HashMap<String, AttrValue>,
nc_type: NcType,
) -> Self {
Self {
name,
dataset,
dims,
attrs,
nc_type,
}
}
@@ -51,11 +55,12 @@ impl<'f> Variable<'f> {
&self.name
}
/// The dimensions of this variable, one per axis: the ones the file
/// gives it (`_Netcdf4Coordinates`, else `DIMENSION_LIST`), found in its
/// group or a parent group. An axis the file gives no dimension (a file
/// not written by a netCDF library) gets the first dimension of the
/// variable's group of the same size, else an anonymous `dim_<size>`.
/// The dimensions of this variable, one per axis, as netCDF-C gives
/// them: the ones the file names (`_Netcdf4Coordinates`, else the
/// scales `DIMENSION_LIST` attaches), found in its group or a parent
/// group; for a dataset without them (a file not written by a netCDF
/// library), netCDF-C's phony dimensions `phony_dim_<n>`, shared by
/// the variables of a group by length.
pub fn dimensions(&self) -> &[Dimension] {
&self.dims
}
@@ -75,10 +80,11 @@ impl<'f> Variable<'f> {
Ok(self.dataset.shape()?)
}
/// The NetCDF data type of this variable.
/// The NetCDF data type of this variable, as netCDF-C reads it (a
/// fixed-length string longer than one byte is `String`, one byte
/// long `Char`).
pub fn nc_type(&self) -> Result<NcType, Error> {
let dtype = self.dataset.dtype()?;
Ok(dtype_to_nctype(&dtype))
Ok(self.nc_type)
}
/// Read all attributes as a HashMap.
@@ -297,6 +303,9 @@ fn default_fill(nc_type: NcType) -> FillValue {
NcType::Double => FillValue::Float(9.969_209_968_386_869e36),
NcType::String => FillValue::String(String::new()),
NcType::Char => FillValue::Int(0),
// User-defined types have no default fill value in netCDF-C; their
// values are not read through these methods.
_ => FillValue::Int(0),
}
}
+281
View File
@@ -0,0 +1,281 @@
//! clawhdf5-netcdf4 compared with netCDF-C itself: `tests/netcdf_c_view.py`
//! prints what netCDF-C (the libnetcdf netCDF4-python bundles) reports for
//! a file, and [`differences`] lists how clawhdf5-netcdf4's reading differs.
#![allow(dead_code)]
use std::collections::{BTreeMap, BTreeSet};
use std::path::{Path, PathBuf};
use std::process::Command;
use std::sync::atomic::{AtomicUsize, Ordering};
use clawhdf5_netcdf4::{NcType, NetCDF4File, NetCDF4Group, Variable};
/// The Python with netCDF4-python (`CLAWHDF5_PYTHON`, else `python3`).
pub fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
/// netCDF-C's view of one file.
#[derive(Default)]
pub struct View {
/// `G`, `D` and `V` lines, in order.
pub meta: Vec<String>,
/// `(group, variable)` → values, from the `X` lines.
pub values: BTreeMap<(String, String), Vec<f64>>,
}
/// netCDF-C's view of each file (`None`: netCDF-C cannot open it), from
/// `tests/netcdf_c_view.py`.
pub fn netcdf_c_views(files: &[PathBuf]) -> BTreeMap<PathBuf, Option<View>> {
let dir = tempfile::tempdir().unwrap();
let list = dir.path().join("files.txt");
let mut paths = String::new();
for f in files {
paths.push_str(&f.to_string_lossy());
paths.push('\n');
}
std::fs::write(&list, paths).unwrap();
let script = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/netcdf_c_view.py");
let out = Command::new(python())
.arg(&script)
.arg("--list")
.arg(&list)
.output()
.expect("failed to run python");
assert!(
out.status.success(),
"netcdf_c_view.py failed: {}",
String::from_utf8_lossy(&out.stderr)
);
parse_views(&String::from_utf8_lossy(&out.stdout))
}
/// The views in the output of `tests/netcdf_c_view.py`.
pub fn parse_views(text: &str) -> BTreeMap<PathBuf, Option<View>> {
let mut views = BTreeMap::new();
let mut current: Option<(PathBuf, Option<View>)> = None;
for line in text.lines() {
let fields: Vec<&str> = line.split('\t').collect();
match fields[0] {
"FILE" => current = Some((PathBuf::from(fields[1]), Some(View::default()))),
"END" => {
let (path, view) = current.take().expect("END without FILE");
views.insert(path, view);
}
"ERROR" => current.as_mut().expect("ERROR without FILE").1 = None,
"X" => {
let view = current.as_mut().and_then(|c| c.1.as_mut()).unwrap();
let vals = if fields[3] == "-" {
Vec::new()
} else {
fields[3]
.split(' ')
.map(|v| v.parse().expect("value"))
.collect()
};
view.values
.insert((fields[1].to_string(), fields[2].to_string()), vals);
}
_ => {
let view = current.as_mut().and_then(|c| c.1.as_mut()).unwrap();
view.meta.push(line.to_string());
}
}
}
views
}
/// Variables whose type is labelled as netCDF-C labels it rather than as
/// clawhdf5-netcdf4 does (see [`nc_type_label`]).
pub static RELABELLED: AtomicUsize = AtomicUsize::new(0);
/// Values not compared because this build of the HDF5 reader lacks a
/// filter (a cargo feature).
pub static FILTER_SKIPPED: AtomicUsize = AtomicUsize::new(0);
/// The same `G`/`D`/`V` lines from clawhdf5-netcdf4.
pub fn our_meta(file: &NetCDF4File) -> Result<Vec<String>, String> {
let e = |e: clawhdf5_netcdf4::Error| e.to_string();
let mut out = Vec::new();
out.push("G\t/".to_string());
push_group(
&mut out,
file.hdf5_file(),
"/",
file.dimensions().map_err(e)?,
file.variables().map_err(e)?,
)?;
for name in file.group_names().map_err(e)? {
walk(
&mut out,
file.hdf5_file(),
&format!("/{name}"),
&file.group(&name).map_err(e)?,
)?;
}
Ok(out)
}
fn walk(
out: &mut Vec<String>,
hdf5: &clawhdf5::File,
path: &str,
group: &NetCDF4Group<'_>,
) -> Result<(), String> {
let e = |e: clawhdf5_netcdf4::Error| e.to_string();
out.push(format!("G\t{path}"));
push_group(
out,
hdf5,
path,
group.dimensions().map_err(e)?,
group.variables().map_err(e)?,
)?;
for name in group.group_names().map_err(e)? {
walk(
out,
hdf5,
&format!("{path}/{name}"),
&group.group(&name).map_err(e)?,
)?;
}
Ok(())
}
/// The type of variable `name` of the group at `path` as the `V` line
/// shows it: its `NcType`, except for the one deliberate difference from
/// netCDF-C — a float of other than 4 or 8 bytes (half, bfloat16, the 4-,
/// 6- and 8-bit floats, `long double`), `NC_FLOAT`/`NC_DOUBLE` here, is
/// labelled `NC_STRING` by netCDF-C 4.9.3 on libhdf5 1.14.6; such a
/// variable is shown as netCDF-C shows it, and counted.
fn nc_type_label(hdf5: &clawhdf5::File, path: &str, var: &Variable<'_>) -> Result<String, String> {
let nc_type = var.nc_type().map_err(|e| e.to_string())?;
if matches!(nc_type, NcType::Float | NcType::Double) {
let dataset = format!("{}/{}", path.trim_end_matches('/'), var.name());
if let Ok(ds) = hdf5.dataset(&dataset)
&& let Ok(clawhdf5_format::datatype::Datatype::FloatingPoint { size, .. }) =
ds.raw_datatype()
&& size != 4
&& size != 8
{
RELABELLED.fetch_add(1, Ordering::Relaxed);
return Ok("NC_STRING".to_string());
}
}
Ok(nc_type.to_string())
}
fn push_group(
out: &mut Vec<String>,
hdf5: &clawhdf5::File,
path: &str,
dims: Vec<clawhdf5_netcdf4::Dimension>,
vars: Vec<Variable<'_>>,
) -> Result<(), String> {
for d in dims {
out.push(format!(
"D\t{path}\t{}\t{}\t{}",
d.name,
d.size,
u8::from(d.is_unlimited)
));
}
for v in vars {
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
let shape: Vec<String> = v
.shape()
.map_err(|e| e.to_string())?
.iter()
.map(u64::to_string)
.collect();
let or_dash = |s: String| if s.is_empty() { "-".to_string() } else { s };
out.push(format!(
"V\t{path}\t{}\t{}\t{}\t{}",
v.name(),
nc_type_label(hdf5, path, &v)?,
or_dash(dims.join(",")),
or_dash(shape.join(","))
));
}
Ok(())
}
/// The values of variable `name` of group `group`.
fn our_values(file: &NetCDF4File, group: &str, name: &str) -> Result<Vec<f64>, String> {
let var = if group == "/" {
file.variable(name)
} else {
file.variable(&format!("{}/{name}", group.trim_start_matches('/')))
}
.map_err(|e| e.to_string())?;
match var.nc_type().map_err(|e| e.to_string())? {
NcType::String | NcType::Char => Err("not numeric".into()),
_ => var.read_raw_f64().map_err(|e| e.to_string()),
}
}
/// How clawhdf5-netcdf4's reading of `path` differs from `want`; empty
/// when it does not.
pub fn differences(path: &Path, want: &View) -> Vec<String> {
let file = match NetCDF4File::open(path) {
Ok(f) => f,
Err(e) => return vec![format!("netCDF-C opens it, clawhdf5-netcdf4 does not: {e}")],
};
let meta = match std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| our_meta(&file))) {
Ok(Ok(meta)) => meta,
Ok(Err(e)) => return vec![format!("error: {e}")],
Err(_) => return vec!["panic".to_string()],
};
let mut diffs = Vec::new();
if meta != want.meta {
let ours: BTreeSet<&String> = meta.iter().collect();
let theirs: BTreeSet<&String> = want.meta.iter().collect();
for line in want.meta.iter().filter(|l| !ours.contains(l)) {
diffs.push(format!("netCDF-C: {line}"));
}
for line in meta.iter().filter(|l| !theirs.contains(l)) {
diffs.push(format!("ours: {line}"));
}
if diffs.is_empty() {
diffs.push("same lines, different order".to_string());
}
}
for ((group, name), want) in &want.values {
let got = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
our_values(&file, group, name)
}))
.unwrap_or_else(|_| Err("panic".to_string()));
match got {
Ok(got) => {
let same = got.len() == want.len()
&& got
.iter()
.zip(want)
.all(|(a, b)| a.to_bits() == b.to_bits() || (a.is_nan() && b.is_nan()));
if !same {
diffs.push(format!("values of {group} {name} differ"));
}
}
Err(e) if e.contains("this build lacks the") => {
FILTER_SKIPPED.fetch_add(1, Ordering::Relaxed);
}
Err(e) => diffs.push(format!("values of {group} {name}: {e}")),
}
}
diffs
}
/// clawhdf5-netcdf4 reads the file at `path` as netCDF-C does: the same
/// groups, dimensions, variables and values (see [`differences`]).
pub fn assert_matches_netcdf_c(path: &Path) {
let views = netcdf_c_views(&[path.to_path_buf()]);
let Some(Some(want)) = views.get(path) else {
panic!("netCDF-C cannot open {}", path.display());
};
let diffs = differences(path, want);
assert!(
diffs.is_empty(),
"{} differs from netCDF-C:\n{}",
path.display(),
diffs.join("\n")
);
}
@@ -0,0 +1,26 @@
# Files of the conformance corpus (conformance/.cache/corpus) that netCDF-C
# opens and clawhdf5-netcdf4 reads differently, with the reason.
# <path relative to the corpus root><TAB><reason>
# Read by tests/corpus_vs_netcdf_c.rs; see docs/known-issues.md
# ("NetCDF-4: differences from netCDF-C").
#
# External links: libhdf5 follows them into the other file (present in the
# corpus); clawhdf5 does not follow external links (known-issues: "External
# links and external raw data are not followed"), so the linked groups and
# datasets are missing and, in files without dimension scales, the phony
# dimension numbers after them shift.
hdf5/test/testfiles/be_extlink1.h5 external link not followed
hdf5/test/testfiles/le_extlink1.h5 external link not followed
hdf5/tools/test/testfiles/h5diff_ext2softlink_src.h5 external link not followed
hdf5/tools/test/testfiles/h5diff_grp_recurse_ext2-1.h5 external link not followed
hdf5/tools/test/testfiles/h5diff_grp_recurse_ext2-2.h5 external link not followed
#
# Values the HDF5 reader refuses and libhdf5 1.14.6 (the netCDF4-python
# wheel's) returns; metadata matches. The first three are the scale-offset
# and N-Bit ref-bug / our-error files of CONFORMANCE.md (libhdf5 reads past
# the stored data); bad_nbit_decompress.h5 is not in the conformance run and
# is not investigated yet.
cve_hdf5/cvefiles/cve-2025-2308.h5 values: scale-offset chunk shorter than its values (libhdf5 over-read)
cve_hdf5/cvefiles/cve-2025-44904.h5 values: unfiltered chunks shorter than a chunk (libhdf5 over-read)
hdf5/test/testfiles/bad_nbit_parms_walk.h5 values: N-Bit parameters too short (conformance our-error/ref-bug)
hdf5/test/testfiles/bad_nbit_decompress.h5 values: N-Bit chunk refused ("element count exceeds chunk size"); not investigated
@@ -0,0 +1,173 @@
//! Every file of a corpus that netCDF-C opens, read by clawhdf5-netcdf4 and
//! compared with what netCDF-C reports: groups (order), dimensions (names,
//! lengths, unlimited, order), variables (names, order, types, dimensions,
//! shapes) and the values of numeric variables of at most 5000 elements.
//!
//! Gated: set `CLAWHDF5_NETCDF_CORPUS` to a directory (relative to the
//! workspace root, or absolute) — the conformance corpus
//! (`conformance/.cache/corpus`) or any tree of HDF5/netCDF-4 files — and
//! `CLAWHDF5_PYTHON` to a Python with netCDF4-python (whose bundled
//! libnetcdf `tests/netcdf_c_view.py` calls). Without the variable the test
//! does nothing. netCDF-C reads each file in its own process (4 at a time,
//! `CLAWHDF5_NETCDF_JOBS`), under a 60 s timeout and a 4 GiB address-space
//! limit: 40 s to 3 minutes for the conformance corpus on tank.
//! `CLAWHDF5_NETCDF_VIEW` may name the saved output of an earlier
//! `netcdf_c_view.py --list <file of paths>` run over the same (absolute)
//! paths, to skip that.
//!
//! A file whose differences are explained is listed, with the reason, in
//! `tests/corpus_known_differences.txt` (and in `docs/known-issues.md`); a
//! listed file that matches fails the test too, so the list stays true.
//!
//! ```sh
//! CLAWHDF5_NETCDF_CORPUS=conformance/.cache/corpus CLAWHDF5_PYTHON=$PWD/.venv/bin/python \
//! cargo test -p clawhdf5-netcdf4 --test corpus_vs_netcdf_c -- --nocapture
//! ```
mod common;
use std::collections::BTreeMap;
use std::path::{Path, PathBuf};
use std::sync::atomic::Ordering;
use common::{FILTER_SKIPPED, RELABELLED, differences, netcdf_c_views, parse_views};
/// The file extensions `conformance/list_files.py` sweeps.
const EXTS: [&str; 7] = ["h5", "hdf5", "he5", "nc", "nc4", "hdf", "h5f"];
/// The files of the corpus: those with an HDF5/netCDF-4 extension, except
/// netCDF classic files (magic `CDF`), following directory symlinks, in
/// byte order of their paths.
fn corpus_files(root: &Path) -> Vec<PathBuf> {
fn walk(dir: &Path, out: &mut Vec<PathBuf>, depth: usize) {
let Ok(entries) = std::fs::read_dir(dir) else {
return;
};
for entry in entries.flatten() {
let path = entry.path();
if path.file_name().is_some_and(|n| n == ".git") {
continue;
}
let Ok(meta) = std::fs::metadata(&path) else {
continue;
};
if meta.is_dir() && depth < 32 {
walk(&path, out, depth + 1);
} else if meta.is_file()
&& path
.extension()
.and_then(|e| e.to_str())
.is_some_and(|e| EXTS.contains(&e.to_ascii_lowercase().as_str()))
{
let mut magic = [0u8; 3];
let classic = std::fs::File::open(&path)
.and_then(|mut f| std::io::Read::read_exact(&mut f, &mut magic))
.is_ok()
&& &magic == b"CDF";
if !classic {
out.push(path);
}
}
}
}
let mut out = Vec::new();
walk(root, &mut out, 0);
out.sort_by(|a, b| {
a.as_os_str()
.as_encoded_bytes()
.cmp(b.as_os_str().as_encoded_bytes())
});
out
}
/// The files `tests/corpus_known_differences.txt` explains: path relative
/// to the corpus root → reason.
fn known_differences() -> BTreeMap<String, String> {
let list = Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/corpus_known_differences.txt");
std::fs::read_to_string(list)
.expect("tests/corpus_known_differences.txt")
.lines()
.filter(|l| !l.trim().is_empty() && !l.starts_with('#'))
.map(|l| {
let (path, reason) = l.split_once('\t').unwrap_or((l, ""));
(path.trim().to_string(), reason.trim().to_string())
})
.collect()
}
#[test]
fn corpus_matches_netcdf_c() {
let Ok(root) = std::env::var("CLAWHDF5_NETCDF_CORPUS") else {
eprintln!("SKIP: set CLAWHDF5_NETCDF_CORPUS to a corpus directory");
return;
};
// A relative path is from the workspace root (cargo runs the test in
// the crate's directory).
let root = Path::new(env!("CARGO_MANIFEST_DIR"))
.join("../..")
.join(root);
let root = std::fs::canonicalize(&root).unwrap_or(root);
let files = corpus_files(&root);
assert!(!files.is_empty(), "no files under {}", root.display());
// `CLAWHDF5_NETCDF_VIEW`: the output of an earlier run of
// `tests/netcdf_c_view.py` over the same files, which takes a few
// minutes over the conformance corpus.
let views = match std::env::var("CLAWHDF5_NETCDF_VIEW") {
Ok(saved) => parse_views(&std::fs::read_to_string(saved).expect("CLAWHDF5_NETCDF_VIEW")),
Err(_) => netcdf_c_views(&files),
};
let known = known_differences();
let (mut opened, mut matched) = (0, 0);
let mut explained = Vec::new();
let mut unexplained = Vec::new();
let mut stale = Vec::new();
for path in &files {
let Some(Some(want)) = views.get(path) else {
continue;
};
opened += 1;
let rel = path
.strip_prefix(&root)
.unwrap_or(path)
.to_string_lossy()
.into_owned();
let diffs = differences(path, want);
match (diffs.is_empty(), known.get(&rel)) {
(true, None) => matched += 1,
(true, Some(_)) => stale.push(rel),
(false, Some(reason)) => explained.push((rel, reason.clone(), diffs)),
(false, None) => unexplained.push((rel, diffs)),
}
}
eprintln!(
"{} files; netCDF-C opens {opened}; {matched} match; {} differ as explained; \
{} differ unexplained; {} listed but match; {} variables relabelled \
(floats of other than 4 or 8 bytes); {} variables' values not compared \
(filter not in this build)",
files.len(),
explained.len(),
unexplained.len(),
stale.len(),
RELABELLED.load(Ordering::Relaxed),
FILTER_SKIPPED.load(Ordering::Relaxed),
);
for (rel, reason, diffs) in &explained {
eprintln!("explained: {rel} ({reason}): {} differences", diffs.len());
}
for (rel, diffs) in &unexplained {
eprintln!("DIFFERS: {rel}");
for d in diffs.iter().take(20) {
eprintln!(" {d}");
}
}
assert!(
unexplained.is_empty(),
"{} files differ from netCDF-C unexplained",
unexplained.len()
);
assert!(
stale.is_empty(),
"listed in corpus_known_differences.txt but match: {stale:?}"
);
}
@@ -2,6 +2,8 @@
//!
//! Tests are skipped if python3 or netCDF4/xarray Python packages are not available.
mod common;
use std::process::Command;
use clawhdf5_netcdf4::{AttrValue, NcType, NetCDF4File};
@@ -842,3 +844,267 @@ ds.to_netcdf({path:?}, engine={engine:?}, unlimited_dims=["time"])
assert_same_view(&path);
}
}
// ===========================================================================
// Compared with netCDF-C itself (tests/netcdf_c_view.py): groups,
// dimensions, variables, types and values, in netCDF-C's order
// ===========================================================================
/// The dimensions of the variable `name` (a path) of `file`.
fn dim_names(file: &NetCDF4File, name: &str) -> Vec<String> {
file.variable(name)
.unwrap()
.dimensions()
.iter()
.map(|d| d.name.clone())
.collect()
}
/// An h5py file without dimension scales: netCDF-C's phony dimensions,
/// numbered file-wide (subgroups before their parent's variables, each
/// group's datasets in name order, as h5py does not track creation order),
/// shared by length within a group — but not between two axes of one
/// variable, nor between a fixed and an unlimited axis — with a length of 0
/// always unlimited and never shared with a fixed axis, and the real
/// dimension of a scale taken by length too.
#[test]
fn phony_dimensions_match_netcdf_c() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("phony.h5");
run_python(&format!(
r#"
import h5py
import numpy as np
with h5py.File({path:?}, "w") as f:
f["zz"] = np.arange(12.0).reshape(2, 2, 3)
f["aa"] = np.arange(6, dtype="i4").reshape(3, 2)
f.create_dataset("un", data=np.ones((2, 3), "f4"), maxshape=(None, 3))
f.create_dataset("un2", data=np.ones(2, "f4"), maxshape=(None,))
f["zero"] = np.zeros((0,))
f.create_dataset("zero_un", (0,), "f4", maxshape=(None,))
f["zero2"] = np.zeros((0, 2))
f["s"] = 1.5
g = f.create_group("g")
g["x"] = np.arange(5, dtype="i2")
g["y"] = np.arange(2, dtype="u1")
g.create_group("h")["q"] = np.arange(7.0)
f.create_group("b")["w"] = np.arange(18, dtype="i8").reshape(2, 9)
f["sc"] = np.arange(6.0)
f["sc"].make_scale("sc")
f["second_only"] = np.zeros((2, 6))
f["second_only"].dims[1].attach_scale(f["sc"])
"#,
path = path.display().to_string()
));
common::assert_matches_netcdf_c(&path);
let file = NetCDF4File::open(&path).unwrap();
// `sc` is dimension 0; /b gets 1 and 2, /g/h 3, /g 4 and 5, / from 6.
assert_eq!(dim_names(&file, "b/w"), ["phony_dim_1", "phony_dim_2"]);
assert_eq!(dim_names(&file, "g/h/q"), ["phony_dim_3"]);
let h = file.group("g/h").unwrap();
assert_eq!(h.variable_names().unwrap(), ["q"]);
assert_eq!(h.dimensions().unwrap()[0].name, "phony_dim_3");
assert_eq!(dim_names(&file, "aa"), ["phony_dim_6", "phony_dim_7"]);
assert_eq!(
dim_names(&file, "zz"),
["phony_dim_7", "phony_dim_11", "phony_dim_6"]
);
assert_eq!(dim_names(&file, "second_only"), ["phony_dim_7", "sc"]);
assert_eq!(dim_names(&file, "un"), ["phony_dim_8", "phony_dim_6"]);
assert_eq!(dim_names(&file, "un2"), ["phony_dim_8"]);
assert_eq!(dim_names(&file, "zero"), ["phony_dim_9"]);
assert_eq!(dim_names(&file, "zero_un"), ["phony_dim_9"]);
assert_eq!(dim_names(&file, "zero2"), ["phony_dim_10", "phony_dim_7"]);
let root: Vec<(String, u64, bool)> = file
.dimensions()
.unwrap()
.into_iter()
.map(|d| (d.name, d.size, d.is_unlimited))
.collect();
assert_eq!(root[0], ("sc".to_string(), 6, false));
assert!(root.contains(&("phony_dim_8".to_string(), 2, true)));
assert!(root.contains(&("phony_dim_9".to_string(), 0, true)));
assert!(root.contains(&("phony_dim_10".to_string(), 0, true)));
assert_eq!(
file.variable_names().unwrap(),
[
"aa",
"s",
"sc",
"second_only",
"un",
"un2",
"zero",
"zero2",
"zero_un",
"zz"
]
);
}
/// Datasets of types netCDF-C cannot represent are not variables:
/// references, bit fields, array types, a compound with a reference member
/// or a half-float member, a compound nesting a compound not seen before;
/// enum, compound, variable-length and opaque types are variables of those
/// classes. netCDF-C also remembers a type it failed to read, so the second
/// dataset of a compound with a reference member is a variable, and a
/// nested compound is accepted once a dataset of the inner type was read.
#[test]
fn types_netcdf_c_skips_are_not_variables() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("types.h5");
run_python(&format!(
r#"
import h5py
import numpy as np
inner = np.dtype([("x", "i2"), ("y", "f4")])
with h5py.File({path:?}, "w") as f:
f["i4"] = np.arange(3, dtype="i4")
f["f2"] = np.arange(3, dtype="f2")
f["s1"] = np.array([b"a", b"b"], dtype="S1")
f["s5"] = np.array([b"abc", b"b"], dtype="S5")
f["vs"] = np.array(["x", "yy"], dtype=h5py.string_dtype())
f["bool"] = np.array([True, False])
f["cmp"] = np.array([(1, 2.0)], dtype=[("a", "i4"), ("b", "f8")])
f["en"] = np.array([0, 1], dtype=h5py.enum_dtype({{"A": 0, "B": 1}}, basetype="i1"))
d = f.create_dataset("vl", (2,), dtype=h5py.vlen_dtype("i4"))
d[0] = [1, 2]
d[1] = [3]
f["op"] = np.array([b"ab", b"cd"], dtype="V2")
f.create_dataset("ref", (1,), dtype=h5py.ref_dtype)[0] = f["i4"].ref
f["cref1"] = np.array([(1, f["i4"].ref)], dtype=[("a", "i4"), ("r", h5py.ref_dtype)])
f["cref2"] = np.array([(2, f["i4"].ref)], dtype=[("a", "i4"), ("r", h5py.ref_dtype)])
f["cmp_f2"] = np.zeros(2, dtype=[("a", "f2")])
f["in1"] = np.zeros(2, dtype=inner)
f["nested"] = np.zeros(2, dtype=[("a", "i4"), ("in", inner)])
f["a_nested"] = np.zeros(2, dtype=[("b", "i4"), ("in", np.dtype([("p", "i1")]))])
sid = h5py.h5s.create_simple((2,))
h5py.h5d.create(f.id, b"bitf", h5py.h5t.STD_B8LE.copy(), sid)
h5py.h5d.create(f.id, b"arr", h5py.h5t.array_create(h5py.h5t.NATIVE_INT32, (3,)), sid)
"#,
path = path.display().to_string()
));
common::assert_matches_netcdf_c(&path);
let file = NetCDF4File::open(&path).unwrap();
assert_eq!(
file.variable_names().unwrap(),
[
"bool", "cmp", "cref2", "en", "f2", "i4", "in1", "nested", "op", "s1", "s5", "vl", "vs"
]
);
let nc_type = |name: &str| file.variable(name).unwrap().nc_type().unwrap();
assert_eq!(nc_type("bool"), NcType::Enum);
assert_eq!(nc_type("cmp"), NcType::Compound);
assert_eq!(nc_type("vl"), NcType::VLen);
assert_eq!(nc_type("op"), NcType::Opaque);
assert_eq!(nc_type("s1"), NcType::Char);
assert_eq!(nc_type("s5"), NcType::String);
// netCDF-C 4.9.3 on libhdf5 1.14.6 says NC_STRING (deliberate
// difference, see the README).
assert_eq!(nc_type("f2"), NcType::Float);
assert_eq!(
file.variable("f2").unwrap().read_raw_f32().unwrap(),
[0.0, 1.0, 2.0]
);
assert!(matches!(
file.variable("ref"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
}
/// The order of groups and variables: creation order where the group
/// tracks it (every netCDF-4 file; also h5py with `track_order`), compact
/// or dense (more than 8 links), else name order (h5py by default).
#[test]
fn group_and_variable_order_match_netcdf_c() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let nc = dir.path().join("order.nc");
let h5 = dir.path().join("order.h5");
run_python(&format!(
r#"
import h5py
import netCDF4 as nc
import numpy as np
names = ["zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c"]
with nc.Dataset({nc:?}, "w") as f:
f.createDimension("x", 2)
for i, n in enumerate(names):
f.createVariable(n, "i4", ("x",))[:] = [i, i + 1]
for n in ["gz", "ga", "gm"]:
f.createGroup(n).createVariable("v", "f8", ("x",))[:] = [1, 2]
small = f.createGroup("small")
for n in ["q", "b", "p"]:
small.createVariable(n, "i2", ("x",))[:] = [3, 4]
with h5py.File({h5:?}, "w") as f:
for i, n in enumerate(names):
f[n] = np.arange(i + 1)
t = f.create_group("tracked", track_order=True)
for i, n in enumerate(names):
t[n] = np.arange(3, dtype="i1")
for n in ["gz", "ga"]:
f.create_group(n)["v"] = np.arange(4.0)
"#,
nc = nc.display().to_string(),
h5 = h5.display().to_string()
));
common::assert_matches_netcdf_c(&nc);
common::assert_matches_netcdf_c(&h5);
let file = NetCDF4File::open(&nc).unwrap();
assert_eq!(
file.variable_names().unwrap(),
[
"zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c"
]
);
assert_eq!(file.group_names().unwrap(), ["gz", "ga", "gm", "small"]);
let file = NetCDF4File::open(&h5).unwrap();
assert_eq!(file.group_names().unwrap(), ["ga", "gz", "tracked"]);
assert_eq!(
file.variable_names().unwrap(),
[
"a", "alpha", "beta", "c", "gamma", "k", "mid", "omega", "zeta", "zz"
]
);
assert_eq!(
file.group("tracked").unwrap().variable_names().unwrap(),
[
"zeta", "alpha", "mid", "beta", "omega", "gamma", "k", "a", "zz", "c"
]
);
}
/// The files of the earlier tests, compared with netCDF-C itself too:
/// dimension scales of h5py, netCDF4-python files with groups, unlimited
/// dimensions and non-coordinate variables named like a dimension.
#[test]
fn netcdf4_python_files_match_netcdf_c() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("mixed.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("time", None)
f.createDimension("p", 2)
f.createDimension("q", 2)
f.createVariable("a", "i4", ("time",))[0:2] = [1, 2]
f.createVariable("b", "f4", ("time", "q"))[0:5, :] = np.arange(10).reshape(5, 2)
f.createVariable("q", "f4", ("q",))[:] = [0, 1]
f.createVariable("p", "f4", ("q", "p"))[:] = np.array([[0, 1], [2, 3]])
g = f.createGroup("g")
g.createDimension("r", 3)
g.createVariable("w", "i4", ("r", "time"))[:, 0:1] = np.ones((3, 1))
g.createGroup("h").createVariable("z", "i8", ("p", "r"))[:] = np.arange(6).reshape(2, 3)
"#,
path = path.display().to_string()
));
common::assert_matches_netcdf_c(&path);
}
@@ -776,9 +776,12 @@ fn test_pure_dimensions_hidden_and_non_coord_names() {
.with_shape(&[2]);
let file = NetCDF4File::from_bytes(b.finish().unwrap()).unwrap();
// `x`, and the phony dimension netCDF-C gives the 3 values of
// `_nc4_non_coord_x` (no dimension scale attached).
let dims = file.dimensions().unwrap();
assert_eq!(dims.len(), 1);
assert_eq!(dims[0].name, "x");
let names: Vec<&str> = dims.iter().map(|d| d.name.as_str()).collect();
assert_eq!(names, ["x", "phony_dim_1"]);
assert_eq!(dims[1].size, 3);
let mut names = file.variable_names().unwrap();
names.sort();
assert_eq!(names, ["v", "x"]);
@@ -786,7 +789,8 @@ fn test_pure_dimensions_hidden_and_non_coord_names() {
assert_eq!(x.name(), "x");
assert_eq!(x.read_raw_f64().unwrap(), [1.0, 2.0, 3.0]);
assert!(!x.is_coordinate());
// No DIMENSION_LIST: `v` gets `x` by size, as before.
// No DIMENSION_LIST: `v` gets the group's first dimension of its length
// (netCDF-C's phony rule, which also takes real dimensions).
assert_eq!(file.variable("v").unwrap().dimensions()[0].name, "x");
}
@@ -0,0 +1,238 @@
#!/usr/bin/env python3
"""netcdf_c_view.py FILE...: print each file as netCDF-C sees it.
The metadata comes from netCDF-C itself (the libnetcdf that netCDF4-python
bundles, called through ctypes), not from netCDF4-python's objects, which
leave out variables of types netCDF-C supports but netCDF4-python does not
(opaque, compounds of vlen strings, ...). Values come from netCDF4-python.
Each file is read in its own process under a timeout, so a file that crashes
or hangs libnetcdf only costs that file.
Output, per file, fields separated by tabs:
FILE <path>
ERROR <message> netCDF-C cannot open it; nothing else
G <group> every group, pre-order, children in
netCDF-C's order ("/" is the root)
D <group> <name> <length> <0|1> its dimensions in dimension-id order
(1: unlimited)
V <group> <name> <type> <dims> <shape>
its variables in netCDF-C's order;
<dims> and <shape> comma-separated,
"-" when there are none; <type> as
clawhdf5_netcdf4::NcType prints it
X <group> <name> <values> the values of a numeric variable of
at most MAX_VALUES elements, as
space-separated Python float reprs
("-" when there are none)
END
Used by tests/corpus_vs_netcdf_c.rs (CLAWHDF5_NETCDF_CORPUS).
"""
import concurrent.futures
import ctypes
import glob
import os
import subprocess
import sys
MAX_VALUES = 5000
TIMEOUT = 60
MEMORY = 4 << 30 # address space of each file's process
ATOMIC = {
1: "NC_BYTE", 2: "NC_CHAR", 3: "NC_SHORT", 4: "NC_INT", 5: "NC_FLOAT",
6: "NC_DOUBLE", 7: "NC_UBYTE", 8: "NC_USHORT", 9: "NC_UINT",
10: "NC_INT64", 11: "NC_UINT64", 12: "NC_STRING",
}
USER_CLASS = {13: "NC_VLEN", 14: "NC_OPAQUE", 15: "NC_ENUM", 16: "NC_COMPOUND"}
NUMERIC = {"NC_BYTE", "NC_SHORT", "NC_INT", "NC_FLOAT", "NC_DOUBLE",
"NC_UBYTE", "NC_USHORT", "NC_UINT", "NC_INT64", "NC_UINT64"}
NAME = 257 # NC_MAX_NAME + 1
MAXDIMS = 1024
def libnetcdf():
"""The libnetcdf netCDF4-python is linked with (same process, same copy)."""
import netCDF4
site = os.path.dirname(os.path.dirname(netCDF4.__file__))
found = glob.glob(os.path.join(site, "netcdf4.libs", "libnetcdf*.so*")) + \
glob.glob(os.path.join(site, "netCDF4.libs", "libnetcdf*.so*"))
if found:
return ctypes.CDLL(found[0])
return ctypes.CDLL("libnetcdf.so")
def view(path, out):
nc = libnetcdf()
ncid = ctypes.c_int()
rc = nc.nc_open(path.encode(), 0, ctypes.byref(ncid)) # NC_NOWRITE
if rc != 0:
nc.nc_strerror.restype = ctypes.c_char_p
out.append("ERROR\t" + nc.nc_strerror(rc).decode(errors="replace"))
return
numeric = [] # (group path, variable name)
try:
walk(nc, ncid.value, "/", out, numeric)
finally:
nc.nc_close(ncid)
values(path, numeric, out)
def check(rc, what):
if rc != 0:
raise RuntimeError(f"{what} failed: {rc}")
def name_of(fn, *args):
buf = ctypes.create_string_buffer(NAME)
check(fn(*args, buf), fn.__name__)
return buf.value.decode(errors="replace")
def walk(nc, gid, path, out, numeric):
out.append(f"G\t{path}")
n = ctypes.c_int()
ids = (ctypes.c_int * MAXDIMS)()
check(nc.nc_inq_dimids(gid, ctypes.byref(n), ids, 0), "nc_inq_dimids")
for dimid in ids[: n.value]:
length = ctypes.c_size_t()
dname = ctypes.create_string_buffer(NAME)
check(nc.nc_inq_dim(gid, dimid, dname, ctypes.byref(length)), "nc_inq_dim")
out.append(f"D\t{path}\t{dname.value.decode(errors='replace')}\t{length.value}\t"
f"{int(is_unlimited(nc, gid, dimid))}")
nvars = ctypes.c_int()
varids = (ctypes.c_int * 65536)()
check(nc.nc_inq_varids(gid, ctypes.byref(nvars), varids), "nc_inq_varids")
for varid in varids[: nvars.value]:
vname = ctypes.create_string_buffer(NAME)
xtype = ctypes.c_int()
ndims = ctypes.c_int()
dimids = (ctypes.c_int * MAXDIMS)()
natts = ctypes.c_int()
check(nc.nc_inq_var(gid, varid, vname, ctypes.byref(xtype), ctypes.byref(ndims),
dimids, ctypes.byref(natts)), "nc_inq_var")
names, shape = [], []
for dimid in dimids[: ndims.value]:
dname = ctypes.create_string_buffer(NAME)
length = ctypes.c_size_t()
if nc.nc_inq_dim(gid, dimid, dname, ctypes.byref(length)) != 0:
names.append("?")
shape.append("?")
continue
names.append(dname.value.decode(errors="replace"))
shape.append(str(length.value))
tname = type_name(nc, gid, xtype.value)
vn = vname.value.decode(errors="replace")
out.append(f"V\t{path}\t{vn}\t{tname}\t{','.join(names) or '-'}\t{','.join(shape) or '-'}")
if tname in NUMERIC and "?" not in shape:
count = 1
for s in shape:
count *= int(s)
if count <= MAX_VALUES:
numeric.append((path, vn))
ngrps = ctypes.c_int()
grps = (ctypes.c_int * 65536)()
check(nc.nc_inq_grps(gid, ctypes.byref(ngrps), grps), "nc_inq_grps")
for child in grps[: ngrps.value]:
cname = ctypes.create_string_buffer(NAME)
check(nc.nc_inq_grpname(child, cname), "nc_inq_grpname")
cpath = path.rstrip("/") + "/" + cname.value.decode(errors="replace")
walk(nc, child, cpath, out, numeric)
def is_unlimited(nc, gid, dimid):
n = ctypes.c_int()
ids = (ctypes.c_int * MAXDIMS)()
# Unlimited dimensions visible from this group include its parents'.
if nc.nc_inq_unlimdims(gid, ctypes.byref(n), ids) != 0:
return False
return dimid in ids[: n.value]
def type_name(nc, gid, xtype):
if xtype in ATOMIC:
return ATOMIC[xtype]
size = ctypes.c_size_t()
base = ctypes.c_int()
nfields = ctypes.c_size_t()
klass = ctypes.c_int()
tname = ctypes.create_string_buffer(NAME)
if nc.nc_inq_user_type(gid, xtype, tname, ctypes.byref(size), ctypes.byref(base),
ctypes.byref(nfields), ctypes.byref(klass)) != 0:
return f"type{xtype}"
return USER_CLASS.get(klass.value, f"class{klass.value}")
def values(path, numeric, out):
if not numeric:
return
import numpy as np
import netCDF4
try:
ds = netCDF4.Dataset(path)
except Exception: # noqa: BLE001 - netCDF4-python refuses some files netCDF-C opens
return
with ds:
for gpath, name in numeric:
try:
group = ds if gpath == "/" else ds[gpath]
var = group.variables[name]
var.set_auto_maskandscale(False)
# A whole-variable read through netCDF-C 4.9.3 lays out a
# variable shorter than an unlimited dimension that is not its
# first wrongly (written values first); reads of one index of
# the leading axis are right.
if var.ndim >= 2:
data = np.stack([np.asarray(var[i]) for i in range(var.shape[0])]) \
if var.shape[0] else np.zeros(var.shape)
else:
data = np.asarray(var[...])
flat = np.asarray(data, dtype=np.float64).ravel()
except Exception: # noqa: BLE001
continue
vals = " ".join(repr(float(v)) for v in flat) or "-"
out.append(f"X\t{gpath}\t{name}\t{vals}")
def limit_memory():
import resource
resource.setrlimit(resource.RLIMIT_AS, (MEMORY, MEMORY))
def one(path):
"""Run view() for one file in a child process."""
try:
proc = subprocess.run([sys.executable, __file__, "--one", path],
capture_output=True, timeout=TIMEOUT, text=True,
preexec_fn=limit_memory)
except subprocess.TimeoutExpired:
return [f"FILE\t{path}", "ERROR\ttimeout", "END"]
lines = proc.stdout.splitlines()
if proc.returncode != 0 or not lines or lines[-1] != "END":
err = (proc.stderr.strip().splitlines() or [f"exit {proc.returncode}"])[-1]
return [f"FILE\t{path}", f"ERROR\tchild failed: {err}", "END"]
return lines
def main(argv):
if argv and argv[0] == "--one":
out = [f"FILE\t{argv[1]}"]
try:
view(argv[1], out)
except Exception as e: # noqa: BLE001
out = [f"FILE\t{argv[1]}", f"ERROR\t{e}"]
out.append("END")
sys.stdout.write("\n".join(out) + "\n")
return
if argv and argv[0] == "--list":
with open(argv[1]) as fh:
argv = [line.rstrip("\n") for line in fh if line.strip()]
jobs = int(os.environ.get("CLAWHDF5_NETCDF_JOBS", "4"))
with concurrent.futures.ThreadPoolExecutor(jobs) as pool:
for lines in pool.map(one, argv):
sys.stdout.write("\n".join(lines) + "\n")
if __name__ == "__main__":
main(sys.argv[1:])
+16 -1
View File
@@ -104,7 +104,22 @@ writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and
`complex128` arrays; the file is written on `close()`. Complex arrays are
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
Rust API writes that on request).
Rust API writes that on request). Both forms read back as numpy
`complex64`/`complex128`.
`libver=` sets the library version bounds as in h5py: `'v108'`,
`'v110'`, `'v112'`, `'v114'`, `'v200'` or `'latest'` (the low bound; the
high bound is then `'latest'`), or a `(low, high)` tuple. With `'v108'`
the file is written in the HDF5 1.8 format (version-2 superblock,
version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the
HDF5 1.10 format. `'earliest'` as the low bound writes the 1.8 format too,
with a `UserWarning`: clawhdf5 cannot write the pre-1.8 format. The
argument is ignored when reading and refused for `'r+'`/`'a'`.
```python
with clawhdf5.File("old.h5", "w", libver="v108") as f: # HDF5 1.8 reads it
f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip")
```
## Editing a file in place
+90 -2
View File
@@ -13,10 +13,13 @@ use crate::attrs::PyAttrs;
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
use crate::handle::Handle;
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
use clawhdf5_rs::LibVer;
/// Internal state for write mode.
struct WriteState {
path: PathBuf,
/// `libver=` as (low, high); `None` keeps the writer's default.
libver: Option<(LibVer, LibVer)>,
root_datasets: Vec<DatasetSpec>,
root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
groups: Vec<Arc<Mutex<WriteGroupState>>>,
@@ -80,10 +83,32 @@ impl PyFile {
/// `az://`; which schemes work depends on how the wheel was built)
/// to read the file remotely with default options (see `open_url`)
/// mode: 'r' for read (default), 'w' for write
/// libver: library version bounds for mode 'w', as h5py's: one of
/// 'earliest', 'v108', 'v110', 'v112', 'v114', 'v200', 'latest' (the
/// low bound; the high bound is then 'latest') or a (low, high)
/// tuple of them. The low bound is the oldest HDF5 release whose
/// format the file uses ('v108': HDF5 1.8 can read it); the high
/// bound the newest whose features it may use. clawhdf5 cannot write
/// the pre-1.8 format, so a low bound of 'earliest' writes the 1.8
/// format (with a warning) and a high bound of 'earliest' is an
/// error. Default (None): the HDF5 1.10 format clawhdf5 has always
/// written. Ignored for reading; not supported with 'r+' / 'a'.
#[new]
#[pyo3(signature = (path, mode="r"))]
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
#[pyo3(signature = (path, mode="r", libver=None))]
fn new(
py: Python<'_>,
path: &str,
mode: &str,
libver: Option<&Bound<'_, PyAny>>,
) -> PyResult<Self> {
let filename = path.to_string();
let libver = libver.map(|v| parse_libver(py, v)).transpose()?;
if libver.is_some() && matches!(mode, "r+" | "a") {
return Err(PyNotImplementedError::new_err(format!(
"libver with mode '{mode}': clawhdf5's in-place editor keeps the format \
versions the file already uses"
)));
}
if is_url(path) {
if mode != "r" {
return Err(PyValueError::new_err(format!(
@@ -113,6 +138,7 @@ impl PyFile {
// Absolute now: the file is written at close, possibly
// after the working directory changed.
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
libver,
root_datasets: Vec::new(),
root_attrs: Arc::new(Mutex::new(Vec::new())),
groups: Vec::new(),
@@ -460,10 +486,71 @@ fn parse_compression(
}
}
/// One of h5py's `libver` names as a bound; `high` says which end it is.
/// `Ok(None)` is 'earliest' as the low bound: the pre-1.8 format, which
/// clawhdf5 cannot write.
fn libver_name(name: &str, high: bool) -> PyResult<Option<LibVer>> {
Ok(Some(match name {
"earliest" if high => {
return Err(PyValueError::new_err(
"libver high bound 'earliest' (the pre-1.8 format) cannot be written by \
clawhdf5; the oldest format it writes is 'v108'",
));
}
"earliest" => return Ok(None),
"v108" => LibVer::V18,
"v110" => LibVer::V110,
"v112" => LibVer::V112,
"v114" => LibVer::V114,
"v200" => LibVer::V200,
"latest" => LibVer::Latest,
other => {
return Err(PyValueError::new_err(format!(
"unknown libver '{other}'; expected 'earliest', 'v108', 'v110', 'v112', \
'v114', 'v200' or 'latest'"
)));
}
}))
}
/// h5py's `libver=`: a name (the low bound, high bound 'latest') or a
/// `(low, high)` pair.
fn parse_libver(py: Python<'_>, v: &Bound<'_, PyAny>) -> PyResult<(LibVer, LibVer)> {
let (low, high): (String, String) = match v.extract::<String>() {
Ok(name) => (name, "latest".into()),
Err(_) => v.extract().map_err(|_| {
PyValueError::new_err("libver must be a string or a (low, high) tuple of strings")
})?,
};
let Some(high) = libver_name(&high, true)? else {
unreachable!("a high bound is never None")
};
let low = match libver_name(&low, false)? {
Some(low) => low,
None => {
PyModule::import(py, "warnings")?.getattr("warn")?.call1((
"libver 'earliest': clawhdf5 cannot write the pre-1.8 format; the file \
is written in the HDF5 1.8 format ('v108') instead",
py.get_type::<pyo3::exceptions::PyUserWarning>(),
))?;
LibVer::V18
}
};
if low > high {
return Err(PyValueError::new_err(format!(
"libver low bound {low} is newer than the high bound {high}"
)));
}
Ok((low, high))
}
/// Build and write the HDF5 file from accumulated write state.
fn finalize_write(state: WriteState) -> PyResult<()> {
crate::no_panic(|| {
let mut builder = clawhdf5_rs::FileBuilder::new();
if let Some((low, high)) = state.libver {
builder.libver_bounds(low, high);
}
// Root attributes
let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner());
@@ -520,6 +607,7 @@ mod tests {
let state = WriteState {
path: path.clone(),
libver: None,
root_datasets: vec![DatasetSpec {
name: "data".into(),
data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]),
+1 -6
View File
@@ -189,12 +189,7 @@ pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Optio
{
clawhdf5_format::data_layout::DataLayout::Chunked {
chunk_dimensions, ..
} if chunk_dimensions.len() >= rank => Some(
chunk_dimensions[..rank]
.iter()
.map(|&d| u64::from(d))
.collect(),
),
} if chunk_dimensions.len() >= rank => Some(chunk_dimensions[..rank].to_vec()),
_ => None,
}
}
+143
View File
@@ -0,0 +1,143 @@
"""`clawhdf5.File(path, 'w', libver=...)`: h5py's library version bounds.
A file written with libver='v108' must open in HDF5 1.8. Its h5dump is
found through CLAWHDF5_H5DUMP18 or at ~/.cache/hdf5-1.8.23/bin/h5dump
(scripts/build-hdf5-1.8.sh builds it); without it that check is skipped.
"""
import os
import subprocess
import numpy as np
import pytest
import clawhdf5
def superblock_version(path):
with open(path, "rb") as f:
head = f.read(9)
assert head[:8] == b"\x89HDF\r\n\x1a\n"
return head[8]
def h5dump18():
path = os.environ.get("CLAWHDF5_H5DUMP18") or os.path.expanduser(
"~/.cache/hdf5-1.8.23/bin/h5dump"
)
try:
out = subprocess.run([path, "--version"], capture_output=True, text=True)
except OSError:
return None
return path if "1.8." in out.stdout else None
def write_sample(path, libver):
with clawhdf5.File(path, "w", libver=libver) as f:
f.create_dataset("x", data=np.arange(10, dtype=np.float64))
f.create_dataset(
"chunked",
data=np.arange(100, dtype=np.int32),
chunks=(30,),
compression="gzip",
)
f.create_dataset("z", data=np.array([1 + 2j, -3j], dtype=np.complex128))
g = f.create_group("g")
g.create_dataset("y", data=np.ones(3, dtype=np.float32))
f.attrs["version"] = 1
@pytest.mark.parametrize(
"libver, sb",
[
(None, 3),
("v108", 2),
(("v108", "v108"), 2),
(("v108", "latest"), 2),
("v110", 3),
("v112", 3),
("v114", 3),
("v200", 3),
("latest", 3),
(("v110", "v200"), 3),
],
)
def test_libver_bounds(h5py, tmp_path, libver, sb):
path = str(tmp_path / "libver.h5")
write_sample(path, libver)
# v108 low bound: HDF5 1.8's version-2 superblock; 1.10 and later: 3.
assert superblock_version(path) == sb
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
np.testing.assert_array_equal(f["z"][:], [1 + 2j, -3j])
np.testing.assert_array_equal(f["g/y"][:], np.ones(3))
assert f.attrs["version"] == 1
with clawhdf5.File(path, "r") as f:
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
def test_libver_v108_opens_in_hdf5_1_8(tmp_path):
h5dump = h5dump18()
if h5dump is None:
pytest.skip("no HDF5 1.8 h5dump (CLAWHDF5_H5DUMP18)")
path = str(tmp_path / "v108.h5")
write_sample(path, "v108")
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode == 0, out.stderr
assert "h5dump error" not in out.stderr
assert 'DATASET "y"' in out.stdout
# The chunked, deflated dataset (a version-1 B-tree) reads in full.
out = subprocess.run(
[h5dump, "-w", "0", "-d", "/chunked", path], capture_output=True, text=True
)
assert out.returncode == 0, out.stderr
assert "(0): " + ", ".join(str(k) for k in range(100)) + "\n" in out.stdout
# The default (1.10 format) does not open in 1.8: the bound matters.
path = str(tmp_path / "default.h5")
write_sample(path, None)
out = subprocess.run([h5dump, path], capture_output=True, text=True)
assert out.returncode != 0
def test_libver_earliest_writes_v108_with_a_warning(h5py, tmp_path):
path = str(tmp_path / "earliest.h5")
with pytest.warns(UserWarning, match="pre-1.8"):
write_sample(path, "earliest")
assert superblock_version(path) == 2
with h5py.File(path, "r") as f:
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
@pytest.mark.parametrize(
"libver, match",
[
(("v108", "earliest"), "high bound 'earliest'"),
("v109", "unknown libver 'v109'"),
(("latest", "v108"), "newer than the high bound"),
(("v108",), "tuple"),
(3, "tuple"),
],
)
def test_libver_errors(tmp_path, libver, match):
with pytest.raises(ValueError, match=match):
clawhdf5.File(str(tmp_path / "bad.h5"), "w", libver=libver)
def test_libver_v108_high_bound_writes_everything(tmp_path):
# Native complex numbers need HDF5 2.0; the Python writer stores complex
# as h5py's {r, i} compound, which 1.8 reads, so the 1.8 high bound is
# fine for everything clawhdf5.File writes.
path = str(tmp_path / "v18only.h5")
write_sample(path, ("v108", "v108"))
assert superblock_version(path) == 2
def test_libver_not_for_editing(tmp_path):
path = str(tmp_path / "e.h5")
write_sample(path, None)
with pytest.raises(NotImplementedError, match="libver"):
clawhdf5.File(path, "r+", libver="latest")
# Reading ignores it, as the bounds only affect what is written.
with clawhdf5.File(path, "r", libver="v108") as f:
assert f["x"].shape == (10,)
+4 -4
View File
@@ -873,9 +873,9 @@ impl Checker<'_> {
self.counts.chunk_index_checksummed += 1;
}
let filtered = matches!(&info.filters, Ok(Some(p)) if !p.filters.is_empty());
let chunk_bytes = cdims.iter().try_fold(u64::from(dt.type_size()), |a, &d| {
a.checked_mul(u64::from(d))
});
let chunk_bytes = cdims
.iter()
.try_fold(u64::from(dt.type_size()), |a, &d| a.checked_mul(d));
let max = ds.max_dimensions.clone();
let mut seen: HashSet<Vec<u64>> = HashSet::with_capacity(chunks.len().min(1 << 20));
let mut reported = 0usize;
@@ -893,7 +893,7 @@ impl Checker<'_> {
));
} else {
for (i, (&o, &cd)) in c.offsets.iter().zip(cdims).enumerate() {
if o % u64::from(cd) != 0 {
if o % cd != 0 {
bad.push(format!(
"offset {o} in dimension {i} is not a multiple of the chunk size {cd}"
));
+22 -1
View File
@@ -603,7 +603,22 @@ impl Diff {
let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) {
(Some(x), Some(y), ..) => x.abs_diff(y).to_string(),
(_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64),
// Each part's absolute difference, as h5diff 2.x prints
// it (`1+0i`).
_ => match (&va, &vb) {
(
Value::Complex { re: ar, im: ai, .. },
Value::Complex { re: br, im: bi, .. },
) => match (number(ar), number(ai), number(br), number(bi)) {
(Some(ar), Some(ai), Some(br), Some(bi)) => format!(
"{}+{}i",
value::fmt_float((ar - br).abs(), 64),
value::fmt_float((ai - bi).abs(), 64)
),
_ => String::new(),
},
_ => String::new(),
},
};
rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}"));
}
@@ -713,6 +728,9 @@ impl Diff {
(Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => {
p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v))
}
(Value::Complex { re: pr, im: pi, .. }, Value::Complex { re: qr, im: qi, .. }) => {
self.equal(a, b, pr, qr) && self.equal(a, b, pi, qi)
}
(Value::Ref(None), Value::Ref(None)) => true,
(Value::Ref(Some(p)), Value::Ref(Some(q))) => {
// Addresses mean nothing across files: compare the paths the
@@ -846,7 +864,10 @@ fn types_comparable(a: &Datatype, b: &Datatype) -> bool {
(
Datatype::Enumeration { base_type: x, .. },
Datatype::Enumeration { base_type: y, .. },
) => types_comparable(x, y),
)
| (Datatype::Complex { base_type: x, .. }, Datatype::Complex { base_type: y, .. }) => {
types_comparable(x, y)
}
_ => true,
}
}
+34 -6
View File
@@ -135,9 +135,16 @@ pub fn short(dt: &Datatype) -> String {
return f.short.into();
}
match dt {
Datatype::Complex { size, base_type } => {
short(&Datatype::complex_as_compound(*size, base_type))
}
// numpy's names (`complex64` is two `float32`) for IEEE parts,
// `complex<part>` otherwise.
Datatype::Complex { size, base_type } => match base_type.as_ref() {
Datatype::FloatingPoint { byte_order, .. } if is_ieee(base_type) => format!(
"complex{}{}",
u64::from(*size) * 8,
if be(byte_order) { "-be" } else { "" }
),
_ => format!("complex<{}>", short(base_type)),
},
Datatype::FixedPoint {
size,
signed,
@@ -216,8 +223,9 @@ pub fn long(dt: &Datatype) -> String {
return f.long.into();
}
match dt {
Datatype::Complex { size, base_type } => {
long(&Datatype::complex_as_compound(*size, base_type))
// As h5ls 2.x prints a complex type it has no native name for.
Datatype::Complex { base_type, .. } => {
format!("complex number of\n {}", long(base_type))
}
Datatype::FixedPoint {
size,
@@ -361,6 +369,19 @@ fn atomic_ddl(dt: &Datatype) -> Option<String> {
)),
// As h5dump 2.x names them (checked against h5dump 2.2.0).
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
// h5dump 2.x's predefined complex names: IEEE binary16/32/64 parts.
Datatype::Complex { base_type, .. } => match base_type.as_ref() {
Datatype::FloatingPoint {
size, byte_order, ..
} if is_ieee(base_type) && !matches!(byte_order, DatatypeByteOrder::Vax) => {
Some(format!(
"H5T_COMPLEX_IEEE_F{}{}",
u64::from(*size) * 8,
order_suffix(byte_order)
))
}
_ => None,
},
Datatype::BitField {
size, byte_order, ..
} => Some(format!(
@@ -418,6 +439,10 @@ pub fn ddl(dt: &Datatype, ind: usize) -> String {
Datatype::VariableLength { base_type, .. } => {
format!("H5T_VLEN {{ {} }}", ddl(base_type, ind))
}
// Parts h5dump has no predefined complex name for.
Datatype::Complex { base_type, .. } => {
format!("H5T_COMPLEX {{ {} }}", ddl(base_type, ind))
}
Datatype::Opaque { size, tag } => {
let tag = String::from_utf8_lossy(tag);
format!(
@@ -498,6 +523,9 @@ fn string_ddl(
/// hdf5-json type object.
pub fn json(dt: &Datatype) -> J {
match dt {
// hdf5-json (h5json 2.0.0) has no complex class: a native complex
// is written as the `{r, i}` compound h5py uses for complex numbers,
// its values as `[re, im]` pairs.
Datatype::Complex { size, base_type } => {
json(&Datatype::complex_as_compound(*size, base_type))
}
@@ -598,7 +626,7 @@ fn pad_json(p: &StringPadding) -> &'static str {
/// Class name used to decide whether two datatypes can be compared.
pub fn class(dt: &Datatype) -> &'static str {
match dt {
Datatype::Complex { .. } => "compound",
Datatype::Complex { .. } => "complex",
Datatype::FixedPoint { .. } => "integer",
Datatype::FloatingPoint { .. } => "float",
Datatype::Time { .. } => "time",
+1 -1
View File
@@ -292,7 +292,7 @@ impl Ls<'_> {
.unwrap_or(0);
let bytes = dims
.iter()
.try_fold(esize, |a, &d| a.checked_mul(u64::from(d)))
.try_fold(esize, |a, &d| a.checked_mul(d))
.map(|b| b.to_string())
.unwrap_or_else(|| "?".into());
writeln!(
+31 -5
View File
@@ -27,6 +27,15 @@ pub enum Value {
/// An enum member (name, when the value matches one) and its value.
Enum(Option<String>, i128),
Compound(Vec<(String, Value)>),
/// A native complex number (HDF5 2.0 class 11): real and imaginary
/// part. `sign_free` is set when h5dump joins the parts with a bare `+`
/// (parts it has no native C type for: binary16, non-IEEE), giving
/// `1+-2i`; otherwise the imaginary part carries its sign (`1-2i`).
Complex {
re: Box<Value>,
im: Box<Value>,
sign_free: bool,
},
Array(Vec<Value>),
/// A variable-length sequence.
Seq(Vec<Value>),
@@ -183,11 +192,18 @@ impl<'a> Decoder<'a> {
return Value::Error("short element".into());
};
match dt {
Datatype::Complex { size, base_type } => self.decode(
&Datatype::complex_as_compound(*size, base_type),
b,
depth + 1,
),
Datatype::Complex { base_type, .. } => {
let part = base_type.type_size() as usize;
let (re, im) = b.split_at(part.min(b.len()));
Value::Complex {
re: Box::new(self.decode(base_type, re, depth + 1)),
im: Box::new(self.decode(base_type, im, depth + 1)),
// h5dump prints `float`/`double` complex (whatever the
// byte order) with `%g%+gi`, anything else as
// `<re>+<im>i`.
sign_free: !(dtype::is_ieee(base_type) && matches!(part, 4 | 8)),
}
}
Datatype::FixedPoint { .. } => match decode_int(dt, b) {
Some(v) => Value::Int(v),
None => Value::Bytes(b.to_vec()),
@@ -354,6 +370,15 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
.collect::<Vec<_>>()
.join(", ")
),
Value::Complex { re, im, sign_free } => {
let (re, im) = (text(re, h5paths), text(im, h5paths));
let plus = if *sign_free || !im.starts_with('-') {
"+"
} else {
""
};
format!("{re}{plus}{im}i")
}
Value::Array(vs) => format!(
"[ {} ]",
vs.iter()
@@ -408,6 +433,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
Value::Complex { re, im, .. } => J::Array(vec![to_json(re, h5paths), to_json(im, h5paths)]),
Value::Array(vs) | Value::Seq(vs) => {
J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect())
}
@@ -0,0 +1,202 @@
//! `h5rs dump` and `h5rs ls` of HDF5 2.0 native complex numbers (datatype
//! class 11), against the output of h5dump 2.2.0 and h5ls 2.2.0 (the Debian
//! h5dump CI installs, 1.14.x, predates the type). The h5dump output is
//! stored next to the fixture; see
//! `crates/clawhdf5/tests/fixtures/gen_native_complex.py`.
use std::path::PathBuf;
use std::process::Command;
fn fixture(name: &str) -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../clawhdf5/tests/fixtures")
.join(name)
}
fn h5rs(args: &[&str]) -> String {
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(args)
.output()
.unwrap();
assert!(out.status.success(), "h5rs {args:?}: {out:?}");
String::from_utf8(out.stdout).unwrap()
}
/// Split a dump into its lines, with the data lines (`(i): ...`) replaced
/// by their index, and the data values in order.
fn split(ddl: &str) -> (Vec<String>, Vec<String>) {
let (mut frame, mut data) = (Vec::new(), Vec::new());
for line in ddl.lines() {
match line.split_once("):") {
Some((idx, body)) if idx.trim_start().starts_with('(') => {
frame.push(format!("{idx}):"));
// Values are split at `, `; array elements come bracketed.
data.extend(
body.split(", ")
.map(|s| s.trim().trim_matches(|c| c == '[' || c == ']' || c == ','))
.map(str::trim)
.filter(|s| !s.is_empty() && *s != "{")
.map(String::from),
);
}
_ => {
let t = line.trim().trim_end_matches(',');
if t.ends_with('i') || t.parse::<f64>().is_ok() {
// A compound member's value on its own line.
data.push(t.to_string());
} else {
frame.push(line.to_string());
}
}
}
}
(frame, data)
}
/// `a+bi`, `a-bi` or `a+-bi` as its two parts (a real number as one).
fn parts(v: &str) -> Vec<f64> {
let Some(z) = v.strip_suffix('i') else {
return vec![v.parse().unwrap()];
};
let b = z.as_bytes();
let at = (1..b.len())
.find(|&k| (b[k] == b'+' || b[k] == b'-') && !matches!(b[k - 1], b'e' | b'E' | b'+'))
.unwrap_or_else(|| panic!("not a complex value: {v}"));
let (re, im) = z.split_at(at);
let im = im.strip_prefix('+').unwrap_or(im);
vec![re.parse().unwrap(), im.parse().unwrap()]
}
#[test]
fn dump_matches_h5dump_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
let ours = h5rs(&["dump", path.to_str().unwrap()]);
let reference = std::fs::read_to_string(fixture("native_complex_hdf5_2.ddl")).unwrap();
let (mut our_frame, our_data) = split(&ours);
let (mut ref_frame, ref_data) = split(&reference);
// The first line names the file as it was given.
our_frame.remove(0);
ref_frame.remove(0);
// Everything but the values is h5dump's byte for byte: the types print
// as H5T_COMPLEX_IEEE_F64LE, H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }, ...
assert_eq!(our_frame, ref_frame);
assert_eq!(our_data.len(), ref_data.len(), "{our_data:?}\n{ref_data:?}");
assert_eq!(our_data.len(), 36);
// Values: h5dump prints `%g` (6 digits), inf/nan, and the binary16
// parts it has no C type for as `<re>+<im>i` (`1+-2i`); h5rs prints
// the shortest round-trip string and Inf/NaN, joined the same way.
for (i, (o, r)) in our_data.iter().zip(&ref_data).enumerate() {
assert_eq!(o.contains("+-"), r.contains("+-"), "[{i}] {o} vs {r}");
let (op, rp) = (parts(o), parts(r));
assert_eq!(op.len(), rp.len(), "[{i}] {o} vs {r}");
for (x, y) in op.iter().zip(&rp) {
let same = if y.is_nan() || y.is_infinite() {
x.is_nan() == y.is_nan() && (y.is_nan() || x == y)
} else {
*x as f32 == *y as f32 || (x - y).abs() <= 5e-6 * y.abs()
};
assert!(same, "[{i}] h5rs {o}, h5dump {r}");
assert_eq!(
x.is_sign_negative(),
y.is_sign_negative(),
"[{i}] {o} vs {r}"
);
}
}
}
#[test]
fn ls_describes_complex_like_h5ls_2_2() {
let path = fixture("native_complex_hdf5_2.h5");
// h5ls 2.2.0 -v, `Type:` of the types it has no native C name for (it
// prints `native double _Complex` for the others, as it prints `native
// double` where h5rs prints `IEEE 64-bit little-endian float`).
for (name, long) in [
(
"c128be",
"complex number of\n IEEE 64-bit big-endian float",
),
(
"c128",
"complex number of\n IEEE 64-bit little-endian float",
),
(
"c32",
"complex number of\n IEEE 16-bit little-endian float",
),
(
"array",
"[2] complex number of\n IEEE 32-bit little-endian float",
),
] {
let target = format!("{}/{name}", path.display());
let verbose = h5rs(&["ls", "-v", &target]);
assert!(
verbose.contains(&format!(" Type: {long}\n")),
"{name}:\n{verbose}"
);
}
let listing = h5rs(&["ls", path.to_str().unwrap()]);
for (name, short) in [
("c64", "complex64"),
("c128", "complex128"),
("c128be", "complex128-be"),
("c32", "complex32"),
("scalar", "complex128"),
("array", "array[2]<complex64>"),
("compound", "compound{z: complex128, k: int64}"),
] {
assert!(
listing
.lines()
.any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))),
"{name}:\n{listing}"
);
}
}
#[test]
fn json_dump_uses_the_r_i_compound() {
// hdf5-json (h5json 2.0.0) has no complex class: the type is written as
// the `{r, i}` compound h5py uses, the values as `[re, im]` pairs.
let path = fixture("native_complex_hdf5_2.h5");
let j: serde_json::Value =
serde_json::from_str(&h5rs(&["dump", "--json", path.to_str().unwrap()])).unwrap();
let ds = j["datasets"]
.as_object()
.unwrap()
.values()
.find(|d| d["alias"][0] == "/scalar")
.unwrap();
assert_eq!(ds["type"]["class"], "H5T_COMPOUND");
assert_eq!(ds["type"]["fields"][0]["name"], "r");
assert_eq!(ds["type"]["fields"][1]["type"]["base"], "H5T_IEEE_F64LE");
assert_eq!(ds["value"], serde_json::json!([2.5, -0.5]));
}
#[test]
fn diff_compares_complex_values() {
let path = fixture("native_complex_hdf5_2.h5");
let p = path.to_str().unwrap();
// Same values, different byte order: no differences.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", p, p, "/c128", "/c128be"])
.output()
.unwrap();
assert!(out.status.success(), "{out:?}");
// One element differs in its imaginary part: one difference, exit
// status 1, as h5diff 2.2.0 reports it.
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["diff", "-r", p, p, "/c128", "/c128x"])
.output()
.unwrap();
assert_eq!(out.status.code(), Some(1), "{out:?}");
let text = String::from_utf8_lossy(&out.stdout);
assert!(text.contains("1 difference(s) found"), "{text}");
// h5diff 2.2.0: `[ 1 1 ] 0+1i 1+1i 1+0i`.
let row = text.lines().find(|l| l.starts_with("[ 1 1 ]")).unwrap();
assert_eq!(
row.split_whitespace().collect::<Vec<_>>(),
["[", "1", "1", "]", "0+1i", "1+1i", "1+0i"]
);
}
+85 -7
View File
@@ -312,11 +312,14 @@ impl Reader {
let base = array_base(dt);
let is_array = !std::ptr::eq(base, dt);
Ok(match base {
// Read as the part type: an array or complex element is its
// parts stored one after another (a complex number's real, then
// imaginary part).
Datatype::FloatingPoint { size, .. } if *size <= 4 => {
Data::F32(data_read::read_as_f32(raw, dt).map_err(err)?)
Data::F32(data_read::read_as_f32(raw, base).map_err(err)?)
}
Datatype::FloatingPoint { .. } => {
Data::F64(data_read::read_as_f64(raw, dt).map_err(err)?)
Data::F64(data_read::read_as_f64(raw, base).map_err(err)?)
}
Datatype::FixedPoint { size, signed, .. } => {
let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err);
@@ -393,10 +396,13 @@ fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T
.collect()
}
/// Innermost element type of (possibly nested) array datatypes.
/// Innermost element type of (possibly nested) array datatypes; for a
/// native complex type, its part type (each element holds two).
fn array_base(dt: &Datatype) -> &Datatype {
match dt {
Datatype::Array { base_type, .. } => array_base(base_type),
Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => {
array_base(base_type)
}
_ => dt,
}
}
@@ -411,6 +417,12 @@ fn element_shape(dt: &Datatype) -> Vec<u64> {
dims.extend(element_shape(base_type));
dims
}
// `[re, im]`: a complex element reads as its two parts.
Datatype::Complex { base_type, .. } => {
let mut dims = vec![2];
dims.extend(element_shape(base_type));
dims
}
_ => Vec::new(),
}
}
@@ -487,9 +499,7 @@ pub fn describe(dt: &Datatype) -> String {
}
}
match dt {
Datatype::Complex { size, base_type } => {
describe(&Datatype::complex_as_compound(*size, base_type))
}
Datatype::Complex { base_type, .. } => format!("complex<{}>", describe(base_type)),
Datatype::FixedPoint {
size,
signed,
@@ -681,6 +691,74 @@ mod tests {
assert!(e.contains("not supported"), "{e}");
}
#[test]
fn native_complex_reads_as_re_im_pairs() {
let z = [[1.0f64, -2.0], [0.5, 0.0], [-0.0, 3.25], [f64::MAX, 1e-300]];
let mut b = FileBuilder::new();
b.create_dataset("z")
.with_native_complex_f64_data(&z)
.with_shape(&[2, 2]);
b.create_dataset("z32")
.with_native_complex_f32_data(&[[1.5f32, -2.5]]);
let r = Reader::open(b.finish().unwrap()).unwrap();
let i = r.info("z").unwrap();
assert_eq!(
(i.shape, i.dtype, i.element_shape),
(vec![2, 2], "complex<f64>".to_string(), vec![2])
);
let all = r.read("z", None).unwrap();
assert_eq!(all.shape, vec![2, 2, 2]);
assert_eq!(all.data, Data::F64(z.iter().flatten().copied().collect()));
let slab = Hyperslab {
start: vec![1, 0],
count: vec![1, 2],
stride: None,
block: None,
};
let part = r.read("z", Some(&slab)).unwrap();
assert_eq!(part.shape, vec![1, 2, 2]);
assert_eq!(part.data, Data::F64(vec![-0.0, 3.25, f64::MAX, 1e-300]));
let z32 = r.read("z32", None).unwrap();
assert_eq!(
(z32.shape, z32.data),
(vec![1, 2], Data::F32(vec![1.5, -2.5]))
);
}
#[test]
fn native_complex_written_by_libhdf5_2() {
// Written by h5py 3.16 / libhdf5 2.0.0 (gen_native_complex.py).
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../clawhdf5/tests/fixtures/native_complex_hdf5_2.h5"
);
let r = Reader::open(std::fs::read(path).unwrap()).unwrap();
let be = r.read("c128be", None).unwrap();
assert_eq!(be.shape, vec![2, 3, 2]);
assert_eq!(be.data, r.read("c128", None).unwrap().data);
assert_eq!(r.info("c128be").unwrap().dtype, "complex<f64 (big-endian)>");
// binary16 parts widen to f32.
assert_eq!(
r.read("c32", None).unwrap().data,
Data::F32(vec![1.0, -2.0, 0.5, 65504.0, 0.0, -0.0])
);
// An array of complex: array dimensions, then the two parts.
let i = r.info("array").unwrap();
assert_eq!(
(i.dtype.as_str(), i.element_shape),
("array[2]<complex<f32>>", vec![2, 2])
);
assert_eq!(
r.read("array", None).unwrap().data,
Data::F32(vec![1.0, 2.0, 3.0, 4.0, -1.0, -1.0, 0.0, 0.0])
);
let s = r.read("scalar", None).unwrap();
assert_eq!((s.shape, s.data), (vec![2], Data::F64(vec![2.5, -0.5])));
// A compound holding a complex member is still refused.
let e = r.read("compound", None).unwrap_err();
assert!(e.contains("compound{z: complex<f64>, k: i64}"), "{e}");
}
#[test]
fn garbage_is_an_error() {
assert!(Reader::open(vec![0u8; 64]).is_err());
@@ -0,0 +1,46 @@
//! Write one dataset of `u8` stored as a single chunk, for measuring the
//! writer's peak memory (`/usr/bin/time -v`):
//!
//! ```text
//! cargo run --release -p clawhdf5 --example write_one_chunk -- BYTES OUT.h5 [slice|owned] [deflate|lz4|zstd]
//! ```
//!
//! `slice` passes the data by reference (`with_u8_data`, which copies it),
//! `owned` hands the builder the vector (`with_u8_data_owned`). The data is
//! `i % 251`, one byte per element; with no filter the chunk is stored raw.
fn main() {
let args: Vec<String> = std::env::args().collect();
let n: u64 = args[1].parse().expect("BYTES");
let out = &args[2];
let mode = args.get(3).map_or("slice", String::as_str);
let filter = args.get(4).map(String::as_str);
let data: Vec<u8> = (0..n).map(|i| (i % 251) as u8).collect();
let mut b = clawhdf5::FileBuilder::new();
let ds = b.create_dataset("x");
match mode {
"slice" => {
ds.with_u8_data(&data);
}
"owned" => {
ds.with_u8_data_owned(data);
}
m => panic!("mode {m}"),
}
ds.with_chunks(&[n]);
match filter {
None => {}
Some("deflate") => {
ds.with_deflate(1);
}
#[cfg(feature = "lz4")]
Some("lz4") => {
ds.with_lz4();
}
#[cfg(feature = "zstd")]
Some("zstd") => {
ds.with_zstd(1);
}
Some(f) => panic!("filter {f} (not built in?)"),
}
b.write(out).unwrap();
}
+1 -4
View File
@@ -408,10 +408,7 @@ impl<'t> ChunkedEdit<'t> {
"chunk rank differs from the dataspace".into(),
));
}
let cd: Vec<u64> = chunk_dimensions[..rank]
.iter()
.map(|&c| u64::from(c))
.collect();
let cd: Vec<u64> = chunk_dimensions[..rank].to_vec();
// Chunks of 4 GiB or more (HDF5 2.0 writes them with layout message
// version 5) are read, but not rewritten: the editor holds a chunk
// it rewrites in memory, compressed and not. Writing values, and a
+4 -1
View File
@@ -1342,7 +1342,10 @@ impl<'f> Dataset<'f> {
) -> Result<(Vec<T>, Vec<T>), Error> {
let dt = self.datatype()?;
let is_complex = match &dt {
// Class 11 parses to this same `{r, i}` compound.
Datatype::Complex { size, base_type } => {
matches!(base_type.as_ref(), Datatype::FloatingPoint { .. })
&& *size == 2 * base_type.type_size()
}
Datatype::Compound { size, members } => {
matches!(members.as_slice(), [r, i]
if r.name == "r" && i.name == "i"
+11
View File
@@ -41,6 +41,13 @@ pub enum DType {
Array(Box<DType>, Vec<u32>),
/// Variable-length UTF-8 string stored via a global heap reference.
VariableLengthString,
/// HDF5 2.0 native complex number (datatype class 11) over the given
/// floating-point part type: `Complex(F32)` is numpy `complex64`,
/// `Complex(F64)` `complex128`, and a half-precision part is
/// `Complex(Other("float16"))`. h5py's default complex encoding, the
/// compound `{r, i}`, stays a [`DType::Compound`]; `read_complex_f64`
/// and `read_complex_f32` read both.
Complex(Box<DType>),
/// Catch-all for HDF5 datatypes that do not map to a specific variant.
Other(String),
}
@@ -60,6 +67,7 @@ impl fmt::Display for DType {
DType::U64 => write!(f, "u64"),
DType::String => write!(f, "string"),
DType::VariableLengthString => write!(f, "vlen_string"),
DType::Complex(base) => write!(f, "complex<{base}>"),
DType::Compound(fields) => {
write!(f, "compound{{")?;
for (i, (name, dt)) in fields.iter().enumerate() {
@@ -148,6 +156,9 @@ pub(crate) fn classify_datatype(dt: &clawhdf5_format::datatype::Datatype) -> DTy
base_type,
dimensions,
} => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()),
Datatype::Complex { base_type, .. } => {
DType::Complex(Box::new(classify_datatype(base_type)))
}
_ => DType::Other(format!("{dt:?}")),
}
}
+31 -8
View File
@@ -137,10 +137,15 @@ impl FileBuilder {
Ok(self.writer.finish()?)
}
/// Serialize and write the file to the given path.
/// Serialize and write the file to the given path. The file is written
/// as it is serialized, not assembled in memory first: unfiltered
/// chunks and contiguous data go straight from the datasets' data.
pub fn write<P: AsRef<std::path::Path>>(self, path: P) -> Result<(), Error> {
let bytes = self.finish()?;
write_file_atomically(path.as_ref(), &bytes).map_err(Error::Io)
use std::io::Write;
write_file_atomically_with(path.as_ref(), |f| {
self.writer
.finish_with(|bytes| f.write_all(bytes).map_err(Error::Io))
})
}
}
@@ -259,8 +264,19 @@ pub fn create_datasets_parallel(specs: Vec<DatasetSpec>) -> Result<Vec<u8>, Erro
/// file or the complete new one — never a truncated mix. `std::fs::write`
/// truncates the destination first, so dying mid-write used to destroy the
/// existing file.
#[cfg(test)]
fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Result<()> {
use std::io::Write;
write_file_atomically_with(path, |f| f.write_all(bytes))
}
/// [`write_file_atomically`] with the contents written by `write` (into a
/// buffered writer over the temporary file).
fn write_file_atomically_with<E: From<std::io::Error>>(
path: &std::path::Path,
write: impl FnOnce(&mut std::io::BufWriter<std::fs::File>) -> Result<(), E>,
) -> Result<(), E> {
use std::io::Write;
// Same directory as the target, so the rename stays on one filesystem.
let mut tmp_name = path
@@ -272,11 +288,18 @@ fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Resul
tmp_name.push(format!(".tmp-{}", std::process::id()));
let tmp_path = path.with_file_name(tmp_name);
let result = (|| {
let mut f = std::fs::File::create(&tmp_path)?;
f.write_all(bytes)?;
f.sync_all()?;
std::fs::rename(&tmp_path, path)
let result = (|| -> Result<(), E> {
// 64 KiB, below glibc's 128 KiB mmap threshold: a 1 MiB buffer was
// mapped afresh by each write once the threshold had not risen, and
// its page faults made a 1 MiB chunked write 1.6x slower (criterion
// `write_2d_chunked/512x512` after the smaller cases, 2026-09-29).
// Writes larger than the buffer go straight to the file.
let mut f = std::io::BufWriter::with_capacity(64 << 10, std::fs::File::create(&tmp_path)?);
write(&mut f)?;
f.flush()?;
f.get_ref().sync_all()?;
std::fs::rename(&tmp_path, path)?;
Ok(())
})();
if result.is_err() {
let _ = std::fs::remove_file(&tmp_path);
+57 -2
View File
@@ -2,6 +2,12 @@
python gen_huge_chunks.py filtered OUT.h5 # huge_chunks_filtered.h5
python gen_huge_chunks.py unfiltered OUT.h5 # generated at test time
python gen_huge_chunks.py dims OUT.h5 # huge_chunk_dims.h5
python gen_huge_chunks.py dims-unfiltered OUT.h5 # generated at test time
`dims` and `dims-unfiltered` (see `make_dims`) hold `u1` datasets whose
chunk dimension is N8 = 2**32 + 7: a chunk dimension of 2**32 or more,
which a layout message of version 5 stores in 5 bytes.
libhdf5 2.x writes a chunk of more than 0xFFFFFFFF bytes with layout message
version 5 (`H5D__chunk_construct`: "chunk size > 4GB requires
@@ -38,7 +44,8 @@ allocated early (see `make`).
libhdf5 holds a whole filtered chunk in memory while it writes it, so the
`filtered` run needs about 4 GiB of memory and 100 s (tank).
Generated 2026-09-28 with h5py 3.16.0 (libhdf5 2.0.0).
huge_chunks_filtered.h5 was generated 2026-09-28 and huge_chunk_dims.h5
2026-09-29, both with h5py 3.16.0 (libhdf5 2.0.0).
"""
import sys
@@ -92,12 +99,60 @@ def make(fid, name, index, filtered):
dsid.close()
N8 = 2**32 + 7
def make_dims(fid, name, index, filtered):
"""A `u1` dataset whose chunks are N8 elements (a chunk dimension of
2**32 or more). `d[0:10] = 0..9` and `d[N8-10:N8] = 100..109`; where
there is a second chunk (`farray`, `earray`: shape N8 + 10),
`d[N8:N8+10] = 200..209`.
filtered (`dims`, committed): deflate level 9 twice, fill value 7;
`single` (Single Chunk) and `earray` (Extensible Array, maxshape None).
Writing it holds a 4 GiB chunk: 4.1 GiB peak resident memory and about
a minute (tank, 2026-09-29).
unfiltered (`dims-unfiltered`): fill time "never", sparse; `single`
(allocated early, as in `make`) and `farray` (Fixed Array).
"""
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
if index == "single":
shape, maxshape = (N8,), (N8,)
elif index == "farray":
shape, maxshape = (N8 + 10,), (N8 + 10,)
elif index == "earray":
shape, maxshape = (N8 + 10,), (UNLIM,)
dcpl.set_chunk((N8,))
if filtered:
dcpl.set_deflate(9)
dcpl.set_deflate(9)
dcpl.set_fill_value(np.array(7, dtype="u1"))
else:
dcpl.set_fill_time(h5py.h5d.FILL_TIME_NEVER)
if index == "single":
dcpl.set_alloc_time(h5py.h5d.ALLOC_TIME_EARLY)
space = h5py.h5s.create_simple(shape, maxshape)
dsid = h5py.h5d.create(fid, name.encode(), h5py.h5t.STD_U8LE, space, dcpl=dcpl)
ds = h5py.Dataset(dsid)
ds[0:10] = np.arange(10, dtype="u1")
ds[N8 - 10 : N8] = np.arange(100, 110, dtype="u1")
if shape[0] > N8:
ds[N8 : N8 + 10] = np.arange(200, 210, dtype="u1")
dsid.close()
def main():
mode, out = sys.argv[1], sys.argv[2]
filtered = {"filtered": True, "unfiltered": False}[mode]
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
fapl.set_libver_bounds(h5py.h5f.LIBVER_EARLIEST, h5py.h5f.LIBVER_V200)
fid = h5py.h5f.create(out.encode(), h5py.h5f.ACC_TRUNC, fapl=fapl)
if mode in ("dims", "dims-unfiltered"):
filtered = mode == "dims"
for index in ["single", "earray"] if filtered else ["single", "farray"]:
make_dims(fid, index, index, filtered)
fid.close()
return
filtered = {"filtered": True, "unfiltered": False}[mode]
indexes = ["single", "farray", "earray", "btree2"]
if not filtered:
indexes.insert(1, "implicit")
+75
View File
@@ -0,0 +1,75 @@
"""Generate native_complex_hdf5_2.h5: HDF5 2.0 native complex (class 11).
Datasets (every one class 11, or a type holding it):
c64 H5T_COMPLEX_IEEE_F32LE, 4 values incl. -0, inf, nan, subnormal
c128 H5T_COMPLEX_IEEE_F64LE, 2 x 3
c128be H5T_COMPLEX_IEEE_F64BE, the same values
c128x H5T_COMPLEX_IEEE_F64LE, c128 with element (1,1) changed
c32 H5T_COMPLEX_IEEE_F16LE, 3 values
scalar H5T_COMPLEX_IEEE_F64LE, scalar dataspace
compound {z: H5T_COMPLEX_IEEE_F64LE @0, k: H5T_STD_I64LE @16}
array H5T_ARRAY [2] of H5T_COMPLEX_IEEE_F32LE, 2 elements
plus the root attribute `zattr` (H5T_COMPLEX_IEEE_F64LE, 2 values).
Written with h5py 3.16 (libhdf5 2.0.0) through its low-level API, the file
type as the memory type (no conversion). Re-run only if the fixture ever
needs regenerating (generated 2026-09-29):
python gen_native_complex.py native_complex_hdf5_2.h5
The h5dump / h5ls 2.2.0 reference output the tests compare with was made
with the tools of the libhdf5 2.2.0 build described in gen_mx_floats.py:
h5dump native_complex_hdf5_2.h5 > native_complex_hdf5_2.ddl
h5ls -r -v native_complex_hdf5_2.h5 > native_complex_hdf5_2.ls
"""
import sys
import h5py
import numpy as np
from h5py import h5a, h5d, h5s, h5t
out = sys.argv[1]
f = h5py.File(out, "w", libver="latest")
def ds(name, tid, arr, shape=None):
shape = arr.shape if shape is None else shape
sp = h5s.create_simple(shape) if shape != () else h5s.create(h5s.SCALAR)
d = h5d.create(f.id, name.encode(), tid, sp)
d.write(h5s.ALL, h5s.ALL, np.ascontiguousarray(arr), mtype=tid)
nan, inf = float("nan"), float("inf")
c64 = np.array(
[1.5 - 2j, complex(0.0, -0.0), complex(inf, nan), complex(-inf, 1e-40)],
dtype=np.complex64,
)
ds("c64", h5t.COMPLEX_IEEE_F32LE, c64)
c128 = np.array(
[[1 + 2j, -3.5 + 4e300j, 0.1 + 0.2j], [5e-324 - 0j, 1j, -1]], dtype=np.complex128
)
ds("c128", h5t.COMPLEX_IEEE_F64LE, c128)
ds("c128be", h5t.COMPLEX_IEEE_F64BE, c128.astype(">c16"))
c128x = c128.copy()
c128x[1, 1] = 1 + 1j
ds("c128x", h5t.COMPLEX_IEEE_F64LE, c128x)
half = np.array([1.0, -2.0, 0.5, 65504.0, 0.0, -0.0], dtype=np.float16)
ds("c32", h5t.COMPLEX_IEEE_F16LE, half, shape=(3,))
ds("scalar", h5t.COMPLEX_IEEE_F64LE, np.array(2.5 - 0.5j), shape=())
ct = h5t.create(h5t.COMPOUND, 24)
ct.insert(b"z", 0, h5t.COMPLEX_IEEE_F64LE)
ct.insert(b"k", 16, h5t.STD_I64LE)
cd = np.array([(1 + 1j, 7), (-2.5 + 0j, -8)], dtype=[("z", "<c16"), ("k", "<i8")])
ds("compound", ct, cd)
at = h5t.array_create(h5t.COMPLEX_IEEE_F32LE, (2,))
ad = np.array([[1 + 2j, 3 + 4j], [-1 - 1j, 0j]], dtype=np.complex64)
ds("array", at, ad, shape=(2,))
a = h5a.create(f.id, b"zattr", h5t.COMPLEX_IEEE_F64LE, h5s.create_simple((2,)))
a.write(np.array([0.5 - 1.5j, 2j]), mtype=h5t.COMPLEX_IEEE_F64LE)
f.close()
Binary file not shown.
@@ -0,0 +1,80 @@
HDF5 "native_complex_hdf5_2.h5" {
GROUP "/" {
ATTRIBUTE "zattr" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): 0.5-1.5i, 0+2i
}
}
DATASET "array" {
DATATYPE H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): [ 1+2i, 3+4i ], [ -1-1i, 0+0i ]
}
}
DATASET "c128" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128be" {
DATATYPE H5T_COMPLEX_IEEE_F64BE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
}
}
DATASET "c128x" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
DATA {
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
(1,0): 4.94066e-324-0i, 1+1i, -1+0i
}
}
DATASET "c32" {
DATATYPE H5T_COMPLEX_IEEE_F16LE
DATASPACE SIMPLE { ( 3 ) / ( 3 ) }
DATA {
(0): 1+-2i, 0.5+65504i, 0+-0i
}
}
DATASET "c64" {
DATATYPE H5T_COMPLEX_IEEE_F32LE
DATASPACE SIMPLE { ( 4 ) / ( 4 ) }
DATA {
(0): 1.5-2i, 0-0i, inf+nani, -inf+9.99995e-41i
}
}
DATASET "compound" {
DATATYPE H5T_COMPOUND {
H5T_COMPLEX_IEEE_F64LE "z";
H5T_STD_I64LE "k";
}
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
DATA {
(0): {
1+1i,
7
},
(1): {
-2.5+0i,
-8
}
}
}
DATASET "scalar" {
DATATYPE H5T_COMPLEX_IEEE_F64LE
DATASPACE SCALAR
DATA {
(0): 2.5-0.5i
}
}
}
}
Binary file not shown.
+348 -1
View File
@@ -81,8 +81,15 @@ fn layout_of(
let lm = msg(MessageType::DataLayout);
let layout = DataLayout::parse(&lm.data, sb.offset_size, sb.length_size).unwrap();
let space = Dataspace::parse(&msg(MessageType::Dataspace).data, sb.length_size).unwrap();
// The element size is the layout's last dimension.
let elem = match &layout {
DataLayout::Chunked {
chunk_dimensions, ..
} => *chunk_dimensions.last().unwrap() as usize,
_ => 8,
};
let (chunks, _) =
list_chunks(data, &layout, &space, 8, sb.offset_size, sb.length_size).unwrap();
list_chunks(data, &layout, &space, elem, sb.offset_size, sb.length_size).unwrap();
(lm.data[0], layout, space, chunks)
}
@@ -516,3 +523,343 @@ print("ok")
}
}
}
// ---- Chunk dimensions of 2^32 or more ----
/// Elements per chunk of the `u8` datasets below: a chunk dimension of
/// 2^32 or more (4 GiB + 7 bytes), which a version-5 layout message stores
/// in 5 bytes.
const N8: u64 = (1 << 32) + 7;
fn dims_fixture() -> PathBuf {
Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures/huge_chunk_dims.h5")
}
/// `u8` values of `sel` in dataset `name` of `file`.
fn sel_u8(file: &File, name: &str, start: u64, count: u64) -> Vec<u8> {
let ds = file.dataset(name).unwrap();
let s = Selection::Hyperslab {
start: vec![start],
stride: vec![1],
count: vec![count],
block: vec![1],
};
ds.read_selection(&s).unwrap()
}
fn u8s(range: std::ops::Range<u8>) -> Vec<u8> {
range.collect()
}
/// `fixtures/huge_chunk_dims.h5` (libhdf5 2.0.0 through h5py 3.16,
/// `gen_huge_chunks.py dims`): `u8` datasets chunked by N8 = 2^32 + 7, in a
/// Single Chunk and an Extensible Array index. Their layouts parse to the
/// whole dimension (it was refused: chunk dimensions were `u32`) and their
/// chunks list at the right origins.
#[test]
fn huge_chunk_dims_list() {
let data = std::fs::read(dims_fixture()).unwrap();
for (name, index, shape, origins) in [
("single", 1u8, N8, &[0u64][..]),
("earray", 4, N8 + 10, &[0, N8]),
] {
let (version, layout, space, mut chunks) = layout_of(&data, name);
assert_eq!(version, 5, "{name}");
let DataLayout::Chunked {
chunk_dimensions,
chunk_index_type,
..
} = &layout
else {
panic!("{name}: {layout:?}");
};
assert_eq!(chunk_dimensions, &[N8, 1], "{name}");
assert_eq!(*chunk_index_type, Some(index), "{name}");
assert_eq!(space.dimensions, [shape], "{name}");
chunks.sort_by_key(|c| c.offsets[0]);
let got: Vec<u64> = chunks.iter().map(|c| c.offsets[0]).collect();
assert_eq!(got, origins, "{name}");
}
}
/// Decoding the fixture's chunks (4 GiB each, one at a time): written
/// values, the fill value (7) around them, and the second chunk.
#[test]
fn huge_chunk_dims_read() {
if !heavy() {
return;
}
let file = File::open(dims_fixture()).unwrap();
let mut first = u8s(0..10);
first.extend([7, 7]);
let mut last = vec![7, 7];
last.extend(u8s(100..110));
for name in ["single", "earray"] {
assert_eq!(sel_u8(&file, name, 0, 12), first, "{name}");
assert_eq!(sel_u8(&file, name, N8 - 12, 12), last, "{name}");
}
let mut across = u8s(108..110);
across.extend(u8s(200..210));
assert_eq!(sel_u8(&file, "earray", N8 - 2, 12), across);
}
/// Unfiltered chunks of N8 bytes that libhdf5 2.0.0 wrote (a sparse file,
/// `gen_huge_chunks.py dims-unfiltered`), in a Single Chunk and a Fixed
/// Array index: read through a memory map, and through positioned reads
/// that fetch only the rows selected.
#[test]
fn huge_chunk_dims_unfiltered_read() {
if !heavy() {
return;
}
let dir = scratch();
let Some(path) = generate(dir.path(), "dims-unfiltered") else {
return;
};
let check = |file: &File| {
let mut first = u8s(0..10);
first.extend([0, 0]);
let mut last = vec![0, 0];
last.extend(u8s(100..110));
for name in ["single", "farray"] {
assert_eq!(sel_u8(file, name, 0, 12), first, "{name}");
assert_eq!(sel_u8(file, name, N8 - 12, 12), last, "{name}");
}
let mut across = u8s(108..110);
across.extend(u8s(200..210));
assert_eq!(sel_u8(file, "farray", N8 - 2, 12), across);
};
check(&File::open(&path).unwrap());
let file = std::fs::File::open(&path).unwrap();
let len = file.metadata().unwrap().len();
let storage = std::sync::Arc::new(Counting {
file,
len,
read: 0.into(),
});
check(&File::open_storage(storage.clone()).unwrap());
let read = storage.read.load(std::sync::atomic::Ordering::Relaxed);
assert!(read < 1 << 20, "{read} bytes read");
}
/// Run `script` under h5py with `path` as `sys.argv[1]`; skipped without
/// h5py (a failure under `CLAWHDF5_REQUIRE_INTEROP=1`).
fn h5py_check(script: &str, path: &Path) {
if !python_available() {
assert!(!interop_required(), "h5py not available");
eprintln!("SKIP: python3 with h5py not available");
return;
}
let out = Command::new(python())
.arg("-c")
.arg(script)
.arg(path)
.output()
.unwrap();
assert!(
out.status.success(),
"h5py:\n{}",
String::from_utf8_lossy(&out.stderr)
);
}
/// `h5dump` 2.x's output for `args` on `path`, checked to succeed without
/// errors; `None` without `CLAWHDF5_H5DUMP2`.
fn h5dump2_text(args: &[&str], path: &Path) -> Option<String> {
let h5dump = h5dump2()?;
let out = Command::new(&h5dump).args(args).arg(path).output().unwrap();
let text = format!(
"{}{}",
String::from_utf8_lossy(&out.stdout),
String::from_utf8_lossy(&out.stderr)
);
assert!(out.status.success(), "h5dump {args:?}:\n{text}");
assert!(!text.contains("rror"), "h5dump {args:?}:\n{text}");
Some(text)
}
/// clawhdf5 writes deflated `u8` chunks of N8 elements (a chunk dimension
/// of 2^32 or more) in a Single Chunk and an Extensible Array index; it,
/// h5py (libhdf5 2.0.0) and h5dump 2.2.0 (layouts only: it cannot inflate
/// chunks over 4 GiB, see known-issues) read them.
#[test]
fn writer_huge_chunk_dims_round_trip() {
if !heavy() {
return;
}
let dir = scratch();
let path = dir.path().join("dims.h5");
let values = u8s(0..10);
let mut b = clawhdf5::FileBuilder::new();
for (name, max) in [("single", N8), ("earray", u64::MAX)] {
b.create_dataset(name)
.with_u8_data(&values)
.with_maxshape(&[max])
.with_chunks(&[N8])
.with_deflate(1);
}
b.write(&path).unwrap();
let data = std::fs::read(&path).unwrap();
// Deflate level 1 leaves about 40 MiB of each 4 GiB chunk of zeros.
assert!(data.len() < 256 << 20, "{} bytes", data.len());
for (name, index) in [("single", 1), ("earray", 4)] {
let (version, layout, _, chunks) = layout_of(&data, name);
assert_eq!(version, 5, "{name}");
assert!(
matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(t), .. }
if *t == index && chunk_dimensions == &[N8, 1]),
"{name}: {layout:?}"
);
assert_eq!(chunks.len(), 1, "{name}");
}
let file = File::open(&path).unwrap();
for name in ["single", "earray"] {
assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}");
}
drop(file);
h5py_check(
r#"
import sys, h5py
with h5py.File(sys.argv[1], "r") as f:
for name in ("single", "earray"):
d = f[name]
assert d.chunks == (2**32 + 7,), (name, d.chunks)
assert list(d[:]) == list(range(10)), (name, d[:])
"#,
&path,
);
for name in ["single", "earray"] {
if let Some(text) = h5dump2_text(&["-H", "-p", "-d", name], &path) {
assert!(text.contains("CHUNKED ( 4294967303 )"), "{name}:\n{text}");
}
}
}
/// An unfiltered chunk of N8 bytes (a chunk dimension of 2^32 or more)
/// written by clawhdf5 from data it was handed (`with_u8_data_owned`):
/// `FileBuilder::write` streams it to the file without copying it, so the
/// write needs about the data's own 4 GiB. Read back by clawhdf5, h5py
/// (a few elements) and h5dump 2.2.0.
#[test]
fn writer_unfiltered_huge_chunk_round_trip() {
if !heavy() {
return;
}
let dir = scratch();
let path = dir.path().join("raw.h5");
let value = |i: u64| (i % 251) as u8;
let mut b = clawhdf5::FileBuilder::new();
b.create_dataset("single")
.with_u8_data_owned((0..N8).map(value).collect())
.with_chunks(&[N8]);
b.write(&path).unwrap();
let file = File::open(&path).unwrap();
let data = file.as_bytes();
let (version, layout, space, chunks) = layout_of(data, "single");
assert_eq!(version, 5);
assert!(
matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(1), .. }
if chunk_dimensions == &[N8, 1]),
"{layout:?}"
);
assert_eq!(space.dimensions, [N8]);
assert_eq!(chunks.len(), 1);
assert_eq!(chunks[0].chunk_size, N8);
for start in [0, (1 << 32) - 3, N8 - 6] {
let want: Vec<u8> = (start..start + 6).map(value).collect();
assert_eq!(sel_u8(&file, "single", start, 6), want, "at {start}");
}
drop(file);
h5py_check(
r#"
import sys, h5py, numpy as np
n = 2**32 + 7
with h5py.File(sys.argv[1], "r") as f:
d = f["single"]
assert d.chunks == (n,) and d.shape == (n,) and d.compression is None
for start in (0, 2**32 - 3, n - 6):
want = [i % 251 for i in range(start, start + 6)]
assert list(d[start:start + 6]) == want, (start, d[start:start + 6])
"#,
&path,
);
let start = (1u64 << 32) - 3;
if let Some(text) = h5dump2_text(
&["-d", "single", "-s", &start.to_string(), "-c", "6"],
&path,
) {
let want: Vec<String> = (start..start + 6).map(|i| value(i).to_string()).collect();
assert!(text.contains(&want.join(", ")), "{text}");
}
}
/// LZ4 and Zstandard chunks of 4 GiB or more: written by clawhdf5, read
/// back by it and by h5py (through `hdf5plugin`'s LZ4 and Zstandard
/// filters when installed).
#[cfg(any(feature = "lz4", feature = "zstd"))]
#[test]
fn writer_huge_chunks_lz4_zstd_round_trip() {
if !heavy() {
return;
}
let dir = scratch();
let path = dir.path().join("codecs.h5");
let values = u8s(0..10);
let mut b = clawhdf5::FileBuilder::new();
let mut names = Vec::new();
#[cfg(feature = "lz4")]
{
b.create_dataset("lz4")
.with_u8_data(&values)
.with_maxshape(&[N8])
.with_chunks(&[N8])
.with_lz4();
names.push("lz4");
}
#[cfg(feature = "zstd")]
{
b.create_dataset("zstd")
.with_u8_data(&values)
.with_maxshape(&[N8])
.with_chunks(&[N8])
.with_zstd(3);
names.push("zstd");
}
b.write(&path).unwrap();
let data = std::fs::read(&path).unwrap();
assert!(data.len() < 256 << 20, "{} bytes", data.len());
for name in &names {
let (version, _, _, chunks) = layout_of(&data, name);
assert_eq!(version, 5, "{name}");
assert_eq!(chunks.len(), 1, "{name}");
}
drop(data);
let file = File::open(&path).unwrap();
for name in &names {
assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}");
}
drop(file);
h5py_check(
&format!(
r#"
import sys, h5py
try:
import hdf5plugin
except ImportError:
print("SKIP: no hdf5plugin")
sys.exit(0)
with h5py.File(sys.argv[1], "r") as f:
for name in {names:?}:
d = f[name]
assert d.chunks == (2**32 + 7,), (name, d.chunks)
assert list(d[0:10]) == list(range(10)), (name, d[0:10])
"#,
names = names
),
&path,
);
}
+19 -5
View File
@@ -1064,12 +1064,26 @@ fn complex_datasets_and_attributes_round_trip() {
let ds = file.dataset(name).unwrap();
assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}");
assert_eq!(ds.shape().unwrap(), vec![3]);
// Both encodings read as the same {r, i} compound.
assert_eq!(
ds.dtype().unwrap(),
// h5py's encoding is a compound, the native one a complex type.
let want = if name == "native64" {
DType::Complex(Box::new(DType::F32))
} else {
DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)])
);
};
assert_eq!(ds.dtype().unwrap(), want, "{name}");
}
assert_eq!(
file.dataset("native128").unwrap().dtype().unwrap(),
DType::Complex(Box::new(DType::F64))
);
assert_eq!(
file.dataset("native128").unwrap().raw_datatype().unwrap(),
clawhdf5::make_native_complex_f64_type()
);
assert_eq!(
DType::Complex(Box::new(DType::F64)).to_string(),
"complex<f64>"
);
for name in ["compound128", "native128"] {
same(
file.dataset(name).unwrap().read_complex_f64().unwrap(),
@@ -1094,7 +1108,7 @@ fn complex_datasets_and_attributes_round_trip() {
data,
}) => {
assert!(shape.is_empty());
assert_eq!(datatype.type_size(), 16);
assert_eq!(datatype, clawhdf5::make_native_complex_f64_type());
let fields =
clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap();
let part = |k: usize| {
+251 -32
View File
@@ -15,9 +15,12 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since |
|---|---|---|
| [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
@@ -33,6 +36,83 @@ Checked against `main` at `9b5803f` on 2026-09-28.
---
## NetCDF-4: differences from netCDF-C
**Status:** open (documented 2026-09-29): deliberate differences, and the
corpus files that still read differently. `clawhdf5-netcdf4` reports a
file's groups, dimensions and variables — names, order, types, shapes — as
netCDF-C 4.9.3 does (`libhdf5/hdf5open.c`; see the crate README), and
`crates/clawhdf5-netcdf4/tests/corpus_vs_netcdf_c.rs` compares it with
netCDF-C itself (the libnetcdf netCDF4-python bundles, called through
ctypes by `tests/netcdf_c_view.py`) over a corpus. Over the conformance
corpus (tank, 2026-09-29, netCDF4-python 1.7.4: netCDF-C 4.9.3 on libhdf5
1.14.6; `CLAWHDF5_NETCDF_CORPUS=<main checkout>/conformance/.cache/corpus
CLAWHDF5_PYTHON=$PWD/.venv/bin/python cargo test -p clawhdf5-netcdf4 --test
corpus_vs_netcdf_c -- --nocapture`): of 692 files, netCDF-C opens 429;
420 match in groups, dimensions (names, lengths, unlimited, order),
variables (names, order, types, dimensions, shapes) and the values of
every numeric variable of up to 5000 elements; the other 9 are listed
below and in `tests/corpus_known_differences.txt`, which the test holds to.
Deliberate:
- **Floats of other than 4 or 8 bytes** (half, bfloat16, the 4-, 6- and
8-bit floats, `long double`) are `NcType::Float` (up to 4 bytes) or
`Double`, and read as numbers. netCDF-C 4.9.3 on libhdf5 1.14.6 labels
them `NC_STRING`: their native type (libhdf5 1.14.6 has a native
`_Float16`) matches none of its own. 27 corpus variables; the comparison
shows them as netCDF-C does and counts them.
- **An axis netCDF-C leaves without a dimension** gets one by the phony
rule (the group's first dimension of its length, else a new
`phony_dim_<n>`): a `_Netcdf4Coordinates` id or a `DIMENSION_LIST` scale
it cannot find, and an axis without a scale of a variable whose first
axis has one — there netCDF-C 4.9.3 reads uninitialised memory
(`get_attached_info` mallocs the object ids and `dimscale_visitor` never
fills that axis's): in one run it gave a 2-long axis the group's first
phony dimension, 3 long; in another netCDF4-python could not open the
file. No corpus file has such an axis.
- **Types netCDF-C remembers after failing to read them.** netCDF-C adds a
type to its list before reading its members, and keeps it when that
fails, so the second dataset of such a type is a variable. For a
compound, enum or variable-length type it has a class and is shown here
too (`interop_tests::types_netcdf_c_skips_are_not_variables`); a bit
field, time, array or complex type has none, and netCDF-C lists the
second dataset with an invalid type (class 0): it stays hidden here.
- **Files netCDF-C refuses** are read as well as they can be: a
multi-dimensional dimension scale without `_Netcdf4Coordinates`, a
`_Netcdf4Coordinates` of the wrong length, a named datatype netCDF-C
cannot represent, an unreadable dataset or attribute (netCDF-C fails
`nc_open`; this crate skips it).
- **Cost:** the first call that needs the metadata reads that of the whole
file (every group, and every dataset's attributes), as `nc_open` does;
it was one group's per call. Not measured; worth measuring on a file
with many groups and variables before a release.
- A whole-variable read of a variable shorter than an unlimited dimension
that is not its first: see [the entry of #27](#netcdf-4-variables-dimensions-are-guessed-from-sizes).
Corpus files that differ (2026-09-29):
- **External links** (`hdf5/test/testfiles/be_extlink1.h5`,
`le_extlink1.h5`, `hdf5/tools/test/testfiles/h5diff_ext2softlink_src.h5`,
`h5diff_grp_recurse_ext2-1.h5`, `h5diff_grp_recurse_ext2-2.h5`): libhdf5
follows them into the other file; clawhdf5 does not
([External links](#external-links-and-external-raw-data-are-not-followed)),
so the linked objects are missing and the phony dimension numbers after
them shift.
- **Values the HDF5 reader refuses** and libhdf5 1.14.6 returns (metadata
matches): `cve_hdf5/cvefiles/cve-2025-2308.h5`, `cve-2025-44904.h5` and
`hdf5/test/testfiles/bad_nbit_parms_walk.h5`, the scale-offset and N-Bit
files of CONFORMANCE.md's ref-bug/our-error list (libhdf5 reads past the
stored data); and `hdf5/test/testfiles/bad_nbit_decompress.h5`, whose
`/Nbit_float_data_le` clawhdf5 refuses ("nbit: element count exceeds
chunk size", also `h5rs dump`) while libhdf5 1.14.6 and h5py 3.16
(libhdf5 2.0.0) return values. That file is not in the conformance
report and has not been investigated.
- Not compared: the values of 18 variables whose filter (SZIP, Blosc,
bzip2) this crate's default build of the HDF5 reader lacks. Blosc and
bzip2 are pure-Rust cargo features of `clawhdf5` (`blosc`, `bzip2`) that
a dependent can turn on; SZIP needs libaec.
## In-place modification (`FileEditor`) limits
**Status:** open (documented 2026-09-26, updated when version-2 B-tree
@@ -166,7 +246,10 @@ Values are correct in every case; this is cost only. Selections other than
## Chunks of 4 GiB or more: limits
**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more
**Status:** open (documented 2026-09-28; updated 2026-09-29, when
[chunk dimensions of 2^32 or more](#chunk-dimensions-of-232-or-more-were-refused)
and [unfiltered writes](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)
were fixed). HDF5 2.0 writes a chunk of more
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
also makes a filtered chunk index element store the chunk's size in "size
@@ -178,34 +261,40 @@ The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
What remains:
- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under
`H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`)
and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec<u32>`.
A 4 GiB chunk of 1-byte elements needs one.
- **Filters the writer refuses for such a chunk** (`FilterError`, before
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
filters record sizes or block lengths in 32 bits, or cannot take such a
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
written; only shuffle + deflate is tested end to end (LZ4's framing and
the 256 MiB limit it had are unit-tested; Zstd is not tested at this
size).
- **Unfiltered chunks this size are written but not tested end to end:**
`FileBuilder` assembles the whole file in memory, several copies of the
chunk (well over the 12 GiB the tests may use). Their layout messages
are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`).
written and tested end to end (`writer_huge_chunks_round_trip`,
`writer_huge_chunk_dims_round_trip`,
`writer_huge_chunks_lz4_zstd_round_trip`: read back by clawhdf5 and
h5py 3.16, LZ4 and Zstd through hdf5plugin 7.1.0).
- **`FileEditor`** refuses to rewrite such chunks (see
[its limits](#in-place-modification-fileeditor-limits)).
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
twice when it is shuffled (the inflated and the unshuffled copy, as
libhdf5's filters do too); writing one holds the chunk and its shuffled
copy. A selection decodes only the chunks it touches; one of an
unfiltered chunk reads only the rows it selects, through a memory map or
libhdf5's filters do too). Writing a filtered chunk holds its compressed
copy and, where it is not a contiguous run of the dataset's data (a
chunk past the dataset's edge is zero-padded), the raw chunk; shuffled,
the shuffled copy too. An unfiltered chunk is not copied when it is such
a run (a dataset stored as one chunk of its own shape): writing one
unfiltered chunk of 2^32 + 7 bytes with `with_u8_data_owned` and
`FileBuilder::write` peaked at 4.00 GiB resident, the data itself
(`write_one_chunk 4294967303 OUT owned`, tank, 2026-09-29, commit
`0065c6b`; see [the fix](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)).
`FileWriter::finish` (a `Vec`) holds the data and the file.
A selection decodes only the chunks it touches; one of an unfiltered
chunk reads only the rows it selects, through a memory map or
positioned reads (`File::open_storage`): every selection of
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop
-- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py
included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection
reads of the double-deflated fixture alone, 4.0 GiB.
memory of `cargo test --release -p clawhdf5 --features lz4,zstd --test
huge_chunks_interop -- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1`
(all twelve tests, h5py and h5dump included) was 10.04 GiB (tank,
2026-09-29, commit `0065c6b`), in `writer_huge_chunks_lz4_zstd_round_trip`
(the other tests, 8.2 GiB or less; the first run of the suite, six
tests, 8.06 GiB on 2026-09-28). Which step of that test holds 10 GiB
(clawhdf5's LZ4 or Zstd write or read, or hdf5plugin's) was not
measured separately.
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
reading or writing one is `FormatError::Overflow` ("exceeds the
addressable size" / "address space"); the file still opens and lists
@@ -240,15 +329,12 @@ wrong data.
decoded, and external references are an error (object references
decode). A multi-dimensional numeric attribute is returned as a flat
array (its shape is not reported; `AttrValue::Raw` carries the shape).
HDF5 2.0's native complex type (class 11) is read as the equivalent
`{r, i}` compound: values are right (`Dataset::read_complex_f64`, and
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs
ls`/`dump` and the browser reader show a compound where h5dump 2.x
prints `H5T_COMPLEX_IEEE_F64LE`, so `h5rs dump` of such a file does not
match h5dump 2.x (h5dump 1.14 cannot read it at all). Writing class 11
(added 2026-09-28) is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`); the Python bindings write complex
arrays as h5py's compound only.
HDF5 2.0's native complex type (class 11) is read as its own type since
2026-09-29 ([fixed](#hdf5-20-native-complex-numbers-read-as-a-r-i-compound));
writing it is opt-in (`with_native_complex_f64_data`,
`make_native_complex_f64_type`), and the Python bindings write complex
arrays as h5py's compound only. `h5rs dump --json` writes it as the
`{r, i}` compound: hdf5-json (h5json 2.0.0) has no complex class.
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
that: libhdf5 fails only the first metadata read of an image it cannot
load and then reads the file's own (possibly stale) metadata, where we
@@ -458,8 +544,10 @@ browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27.
again retries it; what was fetched stays cached).
- Tested under Node 22 and headless Chromium (Playwright's build) against
a local server, cross-origin included; not in Firefox or Safari.
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type; attributes of those types come back
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque, bitfield, time and VL-sequence datasets are
refused with an error naming the type (HDF5 2.0 native complex datasets
read, as `[re, im]` pairs); attributes of those types come back
as `value: null` with their `dtype`. (VL strings read, with h5py's
values, through the same `VlResolver` as `File` and `h5rs`.)
- No Zstd or SZIP (both link C): such datasets fail with
@@ -539,6 +627,136 @@ Newest first. "Before any release" means no tagged release (v2.7.0 and
earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date
given.
## NetCDF-4: files without dimension scales, types and order differed from netCDF-C
**Status:** fixed 2026-09-29 (#30). Affected
every release (v2.1.0 to v2.7.0) and `main` after #27. Wrong metadata only
(names, order and types of dimensions and variables); values were read
right. Users: the dimensions of a file without dimension scales are now
netCDF-C's `phony_dim_<n>`, so code that matched `dim_<size>` or took a
1-D dataset's name for a dimension must use the new names; datasets of
types netCDF-C skips (references, bit fields, array types, compounds of
those, ...) are no longer variables; lists come in netCDF-C's order.
`NcType` has four new variants (`Enum`, `Compound`, `VLen`, `Opaque`) and
is `#[non_exhaustive]`: an exhaustive `match` needs a wildcard arm.
Found 2026-09-28 by the one-off corpus comparison of #27 (52 of the 78
files netCDF4-python opened matched; netCDF4-python itself hides variables
of opaque and other types it does not support, so the comparison now asks
netCDF-C through its C API). Under that comparison `main` at `4260af4`
matched 68 of the 429 corpus files netCDF-C opens (tank, 2026-09-29, the
command above run on a copy of `4260af4` with the new test files). What
differed:
- **Files without dimension scales.** netCDF-C gives every axis a phony
dimension `phony_dim_<id>` (`create_phony_dims`), ids file-wide after the
real dimensions, subgroups numbered before their parent's variables,
shared by length and unlimitedness within a group but not between two
axes of one variable, a length of 0 always unlimited. The crate reported
one dimension per 1-D dataset, named after it (an h5py file with `x(5)`
and `v(5)` had dimensions `x` and `v`, and `v` on `x`), and other axes as
`dim_<size>`.
- **Datasets netCDF-C skips** (`NC_EBADTYPID`: references, bit fields,
time and array types, a compound with such a member or a half-float
member, an enum or VLEN over one) were variables, with `NcType::Char`;
so were enums, compounds, VLENs and opaque types, and 1-byte strings
were `String` where netCDF-C says `NC_CHAR` (longer ones `NC_STRING`).
- **Order.** Groups and variables came in the order of the HDF5 links as
stored (hash order in a dense group); netCDF-C uses creation order when
the group tracks it, else name order. Dimensions without
`_Netcdf4Dimid` came after the others instead of taking the next id in
reading order; a zero-length dimension scale was not unlimited.
- `_Netcdf4Coordinates` ids were looked up only in the variable's group
and its parents (netCDF-C: file-wide), and an unlimited dimension's
length came from its scale's `REFERENCE_LIST` in any group
(`nc4_find_dim_len`: the variables in its group and below).
The crate now reads the metadata of the whole file in netCDF-C's two
passes (`crates/clawhdf5-netcdf4/src/model.rs`), with a new
`clawhdf5_format::group_v2::links_in_creation_order_in` for the order.
Tests: `interop_tests::phony_dimensions_match_netcdf_c`,
`types_netcdf_c_skips_are_not_variables`,
`group_and_variable_order_match_netcdf_c` and
`netcdf4_python_files_match_netcdf_c` compare h5py- and
netCDF4-python-written files with netCDF-C itself, and
`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in
[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)).
## HDF5 2.0 native complex numbers read as a `{r, i}` compound
**Status:** fixed 2026-09-29 (#31). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0)
and `main` until then. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
A class-11 datatype (`H5T_COMPLEX_IEEE_F64LE`, ...) parsed into the
equivalent compound `{r, i}`. Values were right (`read_complex_f64`,
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs ls`/`dump`
and the browser reader showed a compound, so `h5rs dump` did not match
h5dump 2.x, the browser refused the dataset, and a parsed type written
back out became a compound. `Datatype::parse` now returns
`Datatype::Complex` (in compounds, arrays and VL types too),
`Dataset::dtype()` is `DType::Complex(..)`, `h5rs` prints what h5dump,
h5ls and h5diff 2.2.0 print (checked on tank, 2026-09-29, against a file
written by h5py 3.16 / libhdf5 2.0.0:
`cargo test -p clawhdf5-tools --test native_complex_dump`), and the
browser reads `[re, im]` pairs.
**What users must do:** code that matched a native complex type as a
`Datatype::Compound` (or a `DType::Compound` of `r`/`i`) must match
`Datatype::Complex` / `DType::Complex`, or take the compound view with
`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a
`Complex` arm. Files are unchanged; nothing needs rewriting.
## Chunk dimensions of 2^32 or more were refused
**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). An error, never wrong
data. Nothing for users to do but upgrade; code that matches on
`DataLayout::Chunked::chunk_dimensions` gets `u64`s now.
libhdf5 2.x writes a chunk larger than 4 GiB with layout message version
5, whose dimensions take up to 8 bytes each, so a chunk dimension can be
2^32 or more (a 4 GiB chunk of 1-byte elements needs one).
`DataLayout::Chunked::chunk_dimensions` was `Vec<u32>`: such a file was
refused when opened for reading (`InvalidChunkDimensions`, "larger than
2^32 - 1"), and the writer refused such a chunk. Chunk dimensions are
`u64` throughout (format crate, facade, editor, `h5rs`, Python bindings);
a version-3 layout, whose dimensions are 4 bytes, is never written for
them (a chunk that large always takes version 5, as in libhdf5). Tests:
`huge_chunks_interop` (`huge_chunk_dims_list` always, over
`fixtures/huge_chunk_dims.h5`, which libhdf5 2.0.0 wrote: `u8` chunks of
2^32 + 7 in a Single Chunk and an Extensible Array index; with
`CLAWHDF5_HUGE_CHUNKS=1`, reads of it and of an unfiltered sparse file
h5py writes, and clawhdf5's own such chunks read back by h5py 3.16 and
h5dump 2.2.0), unit tests of the 5-byte encoding, and the wasm package
test (the file lists; reading such a chunk on wasm32 is
`FormatError::Overflow`).
## Writing a 4 GiB unfiltered chunk held several copies of it
**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). Memory only. To write
large data with the least memory, hand the builder the data
(`DatasetBuilder::with_u8_data_owned`) and write with
`FileBuilder::write` (or `FileWriter::finish_with`).
`FileWriter::finish` copied every chunk into a per-chunk buffer, laid the
chunks out into one buffer per pass (two passes, the first kept), then
copied everything into the file's buffer, and `FileBuilder::write` wrote
that buffer out: about five times a dataset's size for one unfiltered
chunk, six with `with_u8_data` (which copies its argument). Measured
before and after on tank, 2026-09-29, peak resident memory of
`write_one_chunk 1073741824 OUT <mode>` (`crates/clawhdf5/examples`, one
unfiltered 1 GiB chunk of `u8`): `owned` 5.0 GiB before (the writer at
`4260af4`), 1.0 GiB after (commit `0065c6b`); `slice` 6.0 GiB before,
2.0 GiB after. With 4 GiB + 7 bytes and `owned`: 4.00 GiB after. The
files written are byte for byte the same. Now a chunk that is a contiguous
run of the dataset's data is borrowed (and a filtered one compressed
straight from it), the layout refers to chunks instead of copying them,
and `FileBuilder::write` streams the file to disk
(`FileWriter::finish_with`); contiguous datasets are not copied either.
Test: `writer_unfiltered_huge_chunk_round_trip` (with
`CLAWHDF5_HUGE_CHUNKS=1`: 4 GiB + 7 bytes written and read back by
clawhdf5, h5py 3.16 and h5dump 2.2.0; the test process peaked at 4.00 GiB).
## NetCDF-4: variables' dimensions are guessed from sizes
@@ -590,7 +808,8 @@ dimension scales, and h5netcdf 1.8.1 and xarray files (tank, 2026-09-28,
`CLAWHDF5_PYTHON=<venv with h5netcdf> cargo test -p clawhdf5-netcdf4`).
Files without dimension scales still get dimensions by size (netCDF-C
gives them `phony_dim_<n>`), as before; see `CHANGELOG.md` for a
comparison over the conformance corpus's netCDF-readable files.
comparison over the conformance corpus's netCDF-readable files. (Fixed
2026-09-29: see the entry above.)
## HDF5 1.8 could not read the files we wrote
+6 -2
View File
@@ -123,8 +123,12 @@ else throws an `Error` naming the type.
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
Every response body is cut off past the length asked for. More in
`docs/known-issues.md`.
- Compound, reference, opaque and variable-length-sequence datasets are
refused with an error. Attributes of those types are listed with
- HDF5 2.0 native complex datasets read as `[re, im]` pairs: `dtype`
`complex<f64>`, `elementShape` ending in `2`, the parts interleaved in
the typed array.
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
reference, opaque and variable-length-sequence datasets are refused with
an error. Attributes of those types are listed with
`value: null` and their `dtype`.
- No Zstd or SZIP filters (they link C): such a dataset fails with
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
+6
View File
@@ -93,6 +93,12 @@ expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]&lt;i32&gt;</dd>'
# 3-D: leading dimension held at 0, window over the last two.
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
# HDF5 2.0 native complex (written when h5py's libhdf5 is 2.0 or later):
# each cell is the [re, im] pair.
if "$PY" -c 'import sys, h5py; sys.exit(not getattr(h5py.get_config(), "has_native_complex", False))'; then
expect $H5 "/native_c128" '<dd>complex&lt;f64&gt;</dd>' '<dd>(3, 2)</dd>' '<td>[1, -1.5]</td>' \
'<td>[5, -7.5]</td>'
fi
# Unsupported type: an error, not values.
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
+22 -3
View File
@@ -80,6 +80,15 @@ with h5py.File(h5, "w") as f:
**hdf5plugin.Zstd())
comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")])
f.create_dataset("table", data=comp)
if getattr(h5py.get_config(), "has_native_complex", False):
# HDF5 2.0 native complex (class 11): read as [re, im] pairs.
from h5py import h5d, h5s, h5t
for name, t, dt in [(b"native_c128", h5t.COMPLEX_IEEE_F64LE, "<c16"),
(b"native_c64_be", h5t.COMPLEX_IEEE_F32BE, ">c8")]:
z = (np.arange(6) - 1.5j * np.arange(6)).reshape(3, 2).astype(dt)
d = h5d.create(f.id, name, t, h5s.create_simple(z.shape))
d.write(h5s.ALL, h5s.ALL, z, mtype=t)
g = f.create_group("sensors")
g.attrs["location"] = "lab"
g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4"))
@@ -105,6 +114,8 @@ def kind(dt):
return "strings"
if dt.subdtype is not None:
return kind(dt.subdtype[0])
if dt.kind == "c":
return "f64" if dt.itemsize == 16 else "f32"
if dt.kind == "f":
return "f64" if dt.itemsize == 8 else "f32"
if dt.kind in "iu":
@@ -120,20 +131,25 @@ def flat(a, k):
return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()]
if k.startswith(("i", "u")):
return [str(int(x)) for x in a.ravel()]
if a.dtype.kind == "c":
# [re, im] pairs, in order.
a = np.stack([a.real, a.imag], axis=-1)
return [float(x) for x in a.ravel()]
def entry(ds, slab=None):
k = kind(ds.dtype)
data = ds[()]
e = {"kind": k, "shape": list(np.shape(data)), "values": flat(data, k)}
# A complex element reads as its two parts: one more dimension.
pair = [2] if ds.dtype.kind == "c" else []
e = {"kind": k, "shape": list(np.shape(data)) + pair, "values": flat(data, k)}
if slab:
start, count, stride = slab
idx = tuple(slice(s, s + (c - 1) * st + 1, st)
for s, c, st in zip(start, count, stride))
part = ds[idx]
e["slab"] = {"start": start, "count": count, "stride": stride,
"shape": list(part.shape), "values": flat(part, k)}
"shape": list(part.shape) + pair, "values": flat(part, k)}
return e
@@ -176,7 +192,10 @@ def describe(path):
"datasets": sorted(sets)}
for n, o in members.items():
walk(key.rstrip("/") + "/" + n, o)
elif obj.dtype.names:
elif obj.dtype.names or (
obj.dtype.kind == "c"
and obj.id.get_type().get_class() != getattr(h5py.h5t, "COMPLEX", None)):
# h5py's own complex numbers are a compound {r, i}: refused.
expected["errors"][key] = "compound"
else:
expected["datasets"][key] = entry(obj, slab_for(obj))
+11 -2
View File
@@ -399,7 +399,7 @@ async function remoteTests() {
await fails(async () => {
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
blockSize: 512,
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": `bytes 0-511/${statSync(join(fixDir, "fixture.h5")).size}` })),
});
await f.read("/grid");
}, /the server sent 0-511/, "wrong range");
@@ -514,6 +514,15 @@ async function limitTests() {
`huge chunk: ${name}`);
}
hc.free();
// Likewise chunks whose dimension is 2^32 or more (u8 chunks of 2^32 + 7).
const dims = pkg.open(new Uint8Array(readFileSync(join(import.meta.dirname, "..", "..", "..",
"crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5"))));
eq(dims.list("/").map((e) => e.name), ["earray", "single"], "huge chunk dims: list");
for (const name of ["single", "earray"]) {
await fails(() => dims.readHyperslab(`/${name}`, [0], [4]), /exceeds this platform.s address space/,
`huge chunk dims: ${name}`);
}
dims.free();
}
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
@@ -550,7 +559,7 @@ async function floodTests() {
const range = new Headers(init.headers).get("Range");
if (range === "bytes=0-511") return fetch(url, init);
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/${statSync(join(fixDir, "fixture.h5")).size}` } });
},
}).catch((e) => e);
// The open itself may need a second range: flooded either way.