Write files HDF5 1.8 can read: libver_bounds #28

Open
osobh wants to merge 3 commits from feat/libver-v18 into main
26 changed files with 3154 additions and 260 deletions
+1 -1
View File
@@ -40,7 +40,7 @@ jobs:
python3 -m venv /opt/interop python3 -m venv /opt/interop
# maturin + pytest: ci-test.sh builds the Python package # maturin + pytest: ci-test.sh builds the Python package
# (crates/clawhdf5-py) and runs its tests against h5py. # (crates/clawhdf5-py) and runs its tests against h5py.
/opt/interop/bin/pip install --no-cache-dir h5py numpy netCDF4 xarray hdf5plugin maturin pytest /opt/interop/bin/pip install --no-cache-dir h5py numpy netCDF4 xarray h5netcdf hdf5plugin maturin pytest
echo "/opt/interop/bin" >> "$GITHUB_PATH" echo "/opt/interop/bin" >> "$GITHUB_PATH"
- name: Show interop library versions - name: Show interop library versions
# h5dump's version too: the h5rs dump test requires its exact output # h5dump's version too: the h5rs dump test requires its exact output
+62
View File
@@ -536,6 +536,68 @@ The rows and columns of the uncompressed layouts are within 20% (chunked
column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not column 0.45 -> 0.49 ms, contiguous column 2.55 -> 2.61 ms). This run does not
explain the slower windows. explain the slower windows.
### HDF5 1.8 format: version-1 B-tree chunk indexes (2026-09-28, tank)
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), branch `feat/libver-v18`
at `3c61635`, to decide whether the writer's default should become the
HDF5 1.8 format (`libver_bounds(LibVer::V18, LibVer::V18)`: version-1 B-tree
chunk indexes) instead of the 1.10 format (Fixed Array indexes for these
datasets). **Not an idle machine:** two other agents were building; the
1-minute load average was 7.0 to 7.6 throughout (the rule is below 2), so
treat differences under about 20% as noise. Default and `--v18` runs
alternated, three of each per chunk size; medians of the three.
> **Run:** `cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk N [--v18]`
> with N = 256 (128 chunks per dataset) and N = 32 (8192 chunks per dataset)
Write (the whole 3-dataset file) and file size:
| chunks | 1.10 write ms | 1.8 write ms | 1.10 file bytes | 1.8 file bytes | size |
|---|---:|---:|---:|---:|---:|
| 256 x 256 | 467-500 | 467-483 | 134 916 080 | 134 933 848 | +0.013% |
| 32 x 32 | 490-495 | 489-506 | 138 879 336 | 139 465 112 | +0.42% |
Reads (chunked datasets; the contiguous one does not change), ms. The
window labels are the harness's, which count 256 x 256 chunks: with 32 x 32
chunks the 64 x 64 window covers 9 chunks and the 512 x 512 one 289.
| chunks | layout | read | 1.10 | 1.8 | 1.8 / 1.10 |
|---|---|---|---:|---:|---:|
| 256 | chunked + deflate | full (first) | 6.70 | 7.30 | 1.09x |
| 256 | chunked + deflate | full (repeat) | 4.80 | 3.80 | 0.79x |
| 256 | chunked + deflate | 64 x 64 window (1 chunk) | 0.15 | 0.16 | 1.07x |
| 256 | chunked + deflate | 512 x 512 window (4-9 chunks) | 2.40 | 1.25 | 0.52x |
| 256 | chunked + deflate | one row | 0.95 | 0.95 | 1.00x |
| 256 | chunked + deflate | one column | 1.93 | 1.94 | 1.01x |
| 256 | chunked | full (first) | 11.30 | 11.80 | 1.04x |
| 256 | chunked | full (repeat) | 6.20 | 5.60 | 0.90x |
| 256 | chunked | 64 x 64 window (1 chunk) | 0.04 | 0.04 | 1.00x |
| 256 | chunked | 512 x 512 window (4-9 chunks) | 1.50 | 0.37 | 0.25x |
| 256 | chunked | one row | 0.05 | 0.05 | 1.00x |
| 256 | chunked | one column | 0.44 | 0.45 | 1.02x |
| 32 | chunked + deflate | full (first) | 11.50 | 12.20 | 1.06x |
| 32 | chunked + deflate | full (repeat) | 9.30 | 9.30 | 1.00x |
| 32 | chunked + deflate | 64 x 64 window (1 chunk) | 0.45 | 0.89 | 1.98x |
| 32 | chunked + deflate | 512 x 512 window (4-9 chunks) | 1.97 | 2.41 | 1.22x |
| 32 | chunked + deflate | one row | 0.69 | 1.13 | 1.64x |
| 32 | chunked + deflate | one column | 1.26 | 1.70 | 1.35x |
| 32 | chunked | full (first) | 11.80 | 13.00 | 1.10x |
| 32 | chunked | full (repeat) | 6.30 | 6.20 | 0.98x |
| 32 | chunked | 64 x 64 window (1 chunk) | 0.36 | 0.83 | 2.31x |
| 32 | chunked | 512 x 512 window (4-9 chunks) | 0.70 | 1.16 | 1.66x |
| 32 | chunked | one row | 0.38 | 0.84 | 2.21x |
| 32 | chunked | one column | 1.05 | 1.21 | 1.15x |
With 128 chunks per dataset the two formats read and write alike (the 512 x
512 windows' 0.25x and 0.52x are not explained by the index and are likely
the load). With 8192 chunks, full reads stay within 10%, but a selection on
a freshly opened file costs about 0.4 to 0.5 ms more through the version-1
B-tree (1.2x to 2.3x). Each timed selection opens the file anew, so the
likely cause (not profiled) is walking the B-tree's nodes (2.6 KB each,
about 150 per dataset here) against a Fixed Array's few blocks. Writing costs the
same; files grow by about 36 bytes per chunk. The default therefore stays
the 1.10 format; the 1.8 format is opt-in.
## Local file speed after range reads ## Local file speed after range reads
### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank) ### `ObjectHeader::parse` back at 8f59b2e's speed (2026-09-27, tank)
+109
View File
@@ -2,6 +2,115 @@
## Unreleased ## Unreleased
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
dimension it had written fewer records of got `dim_<n>`, and dimensions
of one size could be swapped. It now resolves them as netCDF-C does
(`libhdf5/hdf5open.c`): the ids in the variable's `_Netcdf4Coordinates`
(each scale's `_Netcdf4Dimid`), else the scales its `DIMENSION_LIST`
references (object references, also the revised `H5T_STD_REF`; of
several scales attached to one axis, the last, as netCDF-C's
`dimscale_visitor` keeps), looked up in its group and then each parent
group; a coordinate variable is on
its own scale. Only an axis with neither (a file not written by a netCDF
library) is still matched by size. A variable may use one dimension
twice (`(p, p)`).
- `variables()` and `variable_names()` leave out the dimension scales that
are only dimensions (`NAME` "This is a netCDF dimension but not a netCDF
variable."), as netCDF-C does, and `variable()` refuses them
(`VariableNotFound`). A variable stored as `_nc4_non_coord_<name>`
(netCDF-C's name for a variable sharing a dimension's name without being
its coordinate variable) is listed and found as `<name>`.
`NetCDF4File::variable_names` is new; `NetCDF4Group::variable_names` used
to list every dataset.
- A variable along an unlimited dimension has the dimension's length, as in
netCDF: `Variable::shape` is that length and the reads return that many
values, the records it has not written as the fill value (`_FillValue`,
else netCDF's `NC_FILL_*` for the type; `""` for strings; NaN from
`read_f64`). It was the HDF5 extent. `Variable::stored_shape` (new) is the
HDF5 extent. Where the unlimited dimension is not a variable's first,
unwritten values are placed per row, as netCDF-C's element and row reads
return them; a whole-variable read through netCDF-C 4.9.3
(netCDF4-python 1.7.4) instead returns the written values first, then the
fill.
- A variable's attributes are read when it is opened (they were read on
first use).
- Tests, compared with netCDF4-python 1.7.4 (netCDF-C 4.9.3) variable by
variable (dimensions, shape, every value): the known-issues reproducer;
two dimensions of one size in either order, one dimension used twice, a
scalar, a non-coordinate variable named like a dimension, a subgroup and
a sub-subgroup on their ancestors' dimensions; unwritten records of
`i1`/`i4`/`u8`/`f4`/`f8`/string variables, with and without
`_FillValue`; a file of h5py dimension scales (no netCDF attributes);
h5netcdf 1.8.1 and xarray 2026.7.0 files (engines netcdf4 and h5netcdf).
The h5netcdf cases skip when h5netcdf is not installed; CI now installs
it. A one-off comparison over the 78 conformance-corpus files
netCDF4-python opens (tank, 2026-09-28; a throwaway test, not committed)
found 52 files with the same variables, dimensions and shapes, and the
same values in every numeric variable of up to 5000 elements (enum
variables were not compared). The other 26 have no dimension scales
(netCDF-C names their axes `phony_dim_<n>`; this crate still matches by
size or names them `dim_<size>`) or hold datasets of types netCDF-C
skips (opaque, references). Affected v2.1.0 to v2.7.0.
`docs/known-issues.md`.
### Writing files HDF5 1.8 can read (2026-09-28)
- New `LibVer` (`V18`, `V110`, `V112`, `V114`, `V200`, `Latest`; re-exported
as `clawhdf5::LibVer`) and `FileBuilder::libver_bounds(low, high)`
(`FileWriter::libver_bounds` in the format crate), as libhdf5's
`H5Pset_libver_bounds` and h5py's `libver=(low, high)`. The default,
`(V110, Latest)`, writes exactly what was written before (the HDF5 1.10
format, which HDF5 1.8.23 refuses to open).
- A low bound of `V18` writes what libhdf5 2.x writes for h5py's
`libver=('v108', 'latest')`: a version-2 superblock (version 3 only for a
paged file), version-3 layout messages for contiguous, compact and chunked
datasets, and a version-1 B-tree chunk index for every chunked dataset,
resizable and multi-unlimited ones included, instead of the single chunk,
Fixed Array, Extensible Array and version-2 B-tree indexes. The B-tree
writer (`btree_v1_write`) replays `H5B_insert` with the chunk callbacks of
`H5Dbtree.c` for chunks arriving in row-major order: libhdf5's split
ratios, right keys moved exactly when `H5D__btree_cmp3` moves them, the
root kept at its address. Its trees are libhdf5's node for node (levels,
child counts, keys) for 1-D, 2-D and 3-D datasets with two- and
three-level trees, deflated or not, against libhdf5 2.0 writing without a
chunk cache (`chunk_btrees_match_libhdf5`). An empty chunked dataset has
no tree (undefined address), as in libhdf5.
- A high bound refuses, with the new `FormatError::LibverBound` and before
anything is written, what needs a newer format: virtual datasets and the
paged file-space strategy (1.10), the 1.12 reference types (datatype
version 4), native complex numbers (datatype version 5, HDF5 2.0), and a
low bound above the high one. `(V18, V18)` therefore writes a file HDF5
1.8 reads, or fails; `(V18, Latest)` writes such objects in their newer
format, as libhdf5 does. `Datatype::max_encoded_version` reports the
version a type needs.
- Checked against a real HDF5 1.8: `scripts/build-hdf5-1.8.sh` builds
1.8.23 (the last 1.8 release) with its tools. `clawhdf5-tools`'
`tests/libver_v18.rs` writes every writer feature under `(V18, V18)`
(contiguous, empty, scalar, compact, f16/f32/f64/i64, chunked with
deflate + shuffle + Fletcher-32 and edge chunks, resizable with one and
two unlimited dimensions, finite maxshape, 100 000 one-element chunks for
a three-level tree, a 100 x 100 grid of chunks, fill values, fixed-length
strings, compound, enum, array, compact/dense/creation-ordered groups,
soft, hard and external links, dense attributes); HDF5 1.8.23's h5dump
dumps the whole file exactly as h5dump 1.14 does and returns our bytes
for every numeric dataset (`-b LE`), and h5py, clawhdf5 and
`h5rs check --data` read it. Then `FileEditor` appends to the resizable
datasets 300 times (splitting B-tree nodes), grows the three-level one,
rewrites a deflated row, grows a 2-D dataset and sets compact and dense
attributes; h5py appends too; and every reader checks again. The 1.8
checks are skipped where no HDF5 1.8 is found (`CLAWHDF5_H5DUMP18` or
`~/.cache/hdf5-1.8.23`), as in CI.
- The default stays the 1.10 format: with 8192 chunks per dataset, small
selections on a freshly opened file read 1.2x to 2.3x slower through a
version-1 B-tree (read harness, tank, 2026-09-28, under load; full reads
and writes within 10%, files about 36 bytes per chunk larger).
`read_harness` gains `--v18` and `--chunk N`. `BENCHMARKS.md`, "HDF5 1.8
format".
- Not yet: the Python bindings' `'w'` mode has no `libver` argument, and
the pre-1.8 format (version-0 superblock, symbol-table groups) cannot be
written. `docs/known-issues.md`.
### A dropped `FileEditor` releases its lock at once (2026-09-28) ### A dropped `FileEditor` releases its lock at once (2026-09-28)
- `FileEditor`'s `flock` could outlive the editor for a moment when - `FileEditor`'s `flock` could outlive the editor for a moment when
another thread forked to spawn a process: the child shared the locked another thread forked to spawn a process: the child shared the locked
+6 -4
View File
@@ -13,7 +13,8 @@ It reads superblocks v0–3, every group and chunk-index structure libhdf5
writes, the standard filters and the common plugin filters, writes, the standard filters and the common plugin filters,
variable-length data and virtual datasets, and follows files a SWMR writer variable-length data and virtual datasets, and follows files a SWMR writer
is appending to. It reads files from libhdf5, h5py and netCDF-4 and writes is appending to. It reads files from libhdf5, h5py and netCDF-4 and writes
files they read. The same library opens files over HTTP and in object files they read, in the HDF5 1.10 format or, on request, in one HDF5 1.8
reads. The same library opens files over HTTP and in object
stores by range requests, runs in the browser as WebAssembly, and has stores by range requests, runs in the browser as WebAssembly, and has
Python bindings with an h5py-shaped API. Python bindings with an h5py-shaped API.
@@ -112,10 +113,10 @@ Limits and open issues, with dates, are in
| Area | Supported | Read only | Not supported | | Area | Supported | Read only | Not supported |
|---|---|---|---| |---|---|---|---|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers | Metadata cache images | Writing files HDF5 1.8 can read | | **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped | | **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 | | **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does); fill values; resizable datasets; virtual datasets (read limits in known-issues) | Chunk indexes v1 B-tree and implicit (the editor also changes them) | External raw data files (explicit error) | | **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error) |
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) | | **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; filters this build cannot encode (refused before any write) | | **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; filters this build cannot encode (refused before any write) |
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) | | **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
@@ -411,7 +412,8 @@ scripts/ci-test.sh # what CI runs: fmt, clippy matrix, tests,
conformance/run.sh # the conformance report (needs h5py, hdf5plugin, h5dump) conformance/run.sh # the conformance report (needs h5py, hdf5plugin, h5dump)
``` ```
The interop suites need a Python with h5py (and netCDF4, xarray); on a The interop suites need a Python with h5py (and netCDF4, xarray; the
NetCDF-4 tests of h5netcdf-written files skip without h5netcdf); on a
PEP 668 system that has to be a virtualenv, which `ci-test.sh` finds as PEP 668 system that has to be a virtualenv, which `ci-test.sh` finds as
`.venv` or through `CLAWHDF5_PYTHON`. Without one they skip; set `.venv` or through `CLAWHDF5_PYTHON`. Without one they skip; set
`CLAWHDF5_REQUIRE_INTEROP=1` to make that a failure, as CI does: `CLAWHDF5_REQUIRE_INTEROP=1` to make that a failure, as CI does:
+32 -9
View File
@@ -7,15 +7,19 @@
//! ```text //! ```text
//! cargo run --release -p clawhdf5-bench --bin read_harness //! cargo run --release -p clawhdf5-bench --bin read_harness
//! cargo run --release -p clawhdf5-bench --bin read_harness -- --large # 512 MB //! cargo run --release -p clawhdf5-bench --bin read_harness -- --large # 512 MB
//! cargo run --release -p clawhdf5-bench --bin read_harness -- --v18 # HDF5 1.8 format
//! cargo run --release -p clawhdf5-bench --bin read_harness -- --chunk 32 # 32 x 32 chunks
//! ``` //! ```
//!
//! `--v18` writes the file with `libver_bounds(V18, V18)` (version-1 B-tree
//! chunk indexes) instead of the default 1.10 format (Fixed Array indexes
//! here), to compare the two.
use std::time::{Duration, Instant}; use std::time::{Duration, Instant};
use clawhdf5::{File, FileBuilder}; use clawhdf5::{File, FileBuilder, LibVer};
use clawhdf5_format::selection::Selection; use clawhdf5_format::selection::Selection;
const CHUNK: u64 = 256;
struct Layout { struct Layout {
name: &'static str, name: &'static str,
chunked: bool, chunked: bool,
@@ -46,16 +50,19 @@ fn value(row: u64, col: u64) -> f64 {
(row * 100_003 + col) as f64 * 0.5 (row * 100_003 + col) as f64 * 0.5
} }
fn write_file(path: &std::path::Path, rows: u64, cols: u64) { fn write_file(path: &std::path::Path, rows: u64, cols: u64, chunk: u64, v18: bool) {
let data: Vec<f64> = (0..rows) let data: Vec<f64> = (0..rows)
.flat_map(|r| (0..cols).map(move |c| value(r, c))) .flat_map(|r| (0..cols).map(move |c| value(r, c)))
.collect(); .collect();
let mut builder = FileBuilder::new(); let mut builder = FileBuilder::new();
if v18 {
builder.libver_bounds(LibVer::V18, LibVer::V18);
}
for (i, layout) in LAYOUTS.iter().enumerate() { for (i, layout) in LAYOUTS.iter().enumerate() {
let ds = builder.create_dataset(&format!("d{i}")); let ds = builder.create_dataset(&format!("d{i}"));
ds.with_f64_data(&data).with_shape(&[rows, cols]); ds.with_f64_data(&data).with_shape(&[rows, cols]);
if layout.chunked { if layout.chunked {
ds.with_chunks(&[CHUNK, CHUNK]); ds.with_chunks(&[chunk, chunk]);
} }
if layout.deflate { if layout.deflate {
ds.with_deflate(4); ds.with_deflate(4);
@@ -91,7 +98,14 @@ fn slab(start: [u64; 2], count: [u64; 2]) -> Selection {
} }
fn main() { fn main() {
let large = std::env::args().any(|a| a == "--large"); let args: Vec<String> = std::env::args().collect();
let large = args.iter().any(|a| a == "--large");
let v18 = args.iter().any(|a| a == "--v18");
let chunk: u64 = args
.iter()
.position(|a| a == "--chunk")
.and_then(|i| args.get(i + 1))
.map_or(256, |c| c.parse().expect("--chunk N"));
let (rows, cols) = if large { (8192, 8192) } else { (4096, 2048) }; let (rows, cols) = if large { (8192, 8192) } else { (4096, 2048) };
let total_mb = (rows * cols * 8) as f64 / (1 << 20) as f64; let total_mb = (rows * cols * 8) as f64 / (1 << 20) as f64;
if cfg!(debug_assertions) { if cfg!(debug_assertions) {
@@ -100,12 +114,21 @@ fn main() {
let dir = tempfile::TempDir::new().unwrap(); let dir = tempfile::TempDir::new().unwrap();
let path = dir.path().join("read_harness.h5"); let path = dir.path().join("read_harness.h5");
write_file(&path, rows, cols); let t = Instant::now();
let file_mb = std::fs::metadata(&path).unwrap().len() as f64 / (1 << 20) as f64; write_file(&path, rows, cols, chunk, v18);
let write_ms = t.elapsed().as_secs_f64() * 1e3;
let file_bytes = std::fs::metadata(&path).unwrap().len();
let file_mb = file_bytes as f64 / (1 << 20) as f64;
println!("## Read harness"); println!("## Read harness");
println!( println!(
"\n{rows} x {cols} f64 ({total_mb:.0} MB per dataset), chunks {CHUNK} x {CHUNK}, file {file_mb:.0} MB\n" "\n{rows} x {cols} f64 ({total_mb:.0} MB per dataset), chunks {chunk} x {chunk}, \
format {}, file {file_mb:.0} MB ({file_bytes} bytes), written in {write_ms:.0} ms\n",
if v18 {
"1.8 (v1 B-tree)"
} else {
"1.10 (default)"
}
); );
// (label, selection, elements selected) // (label, selection, elements selected)
+3 -1
View File
@@ -36,7 +36,9 @@ clawhdf5-format = { git = "https://git.redclaw.dev/quantumclaw/clawhdf5" }
`type_builders` (datasets, groups, attributes, compound and enum types, `type_builders` (datasets, groups, attributes, compound and enum types,
links, virtual datasets, creation-order tracking); chunk indexes and links, virtual datasets, creation-order tracking); chunk indexes and
dense-storage B-trees of any size (`chunked_write`, `btree_v2_write`, dense-storage B-trees of any size (`chunked_write`, `btree_v2_write`,
`ea_writer`). Output is read by h5py and h5dump. `ea_writer`, and the version-1 chunk B-tree of `btree_v1_write`). Output
is read by h5py and h5dump; `FileWriter::libver_bounds` (`libver`) picks
the format: HDF5 1.10 by default, or one HDF5 1.8 reads.
- **Filters:** `filter_pipeline` and `filter_registry` (look up by ID; other - **Filters:** `filter_pipeline` and `filter_registry` (look up by ID; other
IDs can be registered at run time with `register_filter`). Built in: IDs can be registered at run time with `register_filter`). Built in:
deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4, deflate, shuffle, Fletcher-32, N-Bit, scale-offset; behind features LZ4,
@@ -0,0 +1,470 @@
//! Writing a version-1 B-tree chunk index (node type 1): the chunk index of
//! layout message versions 1-3, and the only one HDF5 1.8 reads.
//!
//! The tree is built the way libhdf5 builds it when the chunks reach it one
//! after another in row-major order (a whole-dataset `H5Dwrite` of a 1-D
//! dataset, or of any dataset without a chunk cache; with one, libhdf5
//! inserts the small chunks of a multi-dimensional dataset in the order its
//! cache evicts them, which fills the nodes differently): each
//! chunk goes through the same steps as `H5B_insert` (`H5B.c`) with the
//! chunk callbacks of `H5Dbtree.c`, so nodes split where libhdf5's split,
//! with its default split ratios (a full right-most node keeps 90% of its
//! children, a left-most one 10%, any other half), and keys hold what
//! libhdf5's hold:
//!
//! - a chunk's key is its size in the file, its filter mask and its offsets
//! (the element-size coordinate 0);
//! - a node's final key is the zero-size key one chunk past the chunk that
//! last moved it (every scaled coordinate plus one, `H5D__btree_new_node`),
//! which libhdf5 moves only when a new chunk is not below it
//! (`H5D__btree_cmp3`) — so after an even number of appends in one
//! dimension it lies on the last chunk itself;
//! - a full root is copied to a new node and becomes the parent of the copy
//! and its new sibling, so the root's address (the layout message's) never
//! changes.
//!
//! Nodes are laid out in the order libhdf5 allocates them (the root first,
//! then each new node as a split creates it), all of the full node size, the
//! unused slots zero.
#[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec};
use core::cmp::Ordering;
use crate::error::FormatError;
/// libhdf5's default chunk B-tree K (`HDF5_BTREE_CHUNK_IK_DEF`): nodes hold
/// up to 2K = 64 children. Superblocks of version 2 cannot record another
/// value without a superblock extension, which this writer does not emit.
pub(crate) const CHUNK_BTREE_K: u16 = 32;
/// libhdf5's default split ratios (`H5D_XFER_BTREE_SPLIT_RATIO_DEF`) for a
/// left-most, middle and right-most node.
const SPLIT_RATIOS: [f64; 3] = [0.1, 0.5, 0.9];
/// A chunk to index: scaled coordinates (offset / chunk dimension) in each
/// dataset dimension, stored size, filter mask and address.
pub(crate) struct ChunkEntry {
pub(crate) scaled: Vec<u64>,
pub(crate) nbytes: u64,
pub(crate) filter_mask: u32,
pub(crate) address: u64,
}
#[derive(Debug, Clone, PartialEq, Eq)]
struct Key {
nbytes: u32,
mask: u32,
/// Scaled coordinates, the element-size one (0 or 1) last.
scaled: Vec<u64>,
}
impl Key {
/// `H5D__btree_new_node`'s right key: one chunk past `self` in every
/// dimension, with no storage.
fn right_of(&self) -> Key {
Key {
nbytes: 0,
mask: 0,
scaled: self.scaled.iter().map(|s| s + 1).collect(),
}
}
fn cmp_scaled(&self, other: &Key) -> Ordering {
self.scaled.cmp(&other.scaled)
}
}
#[derive(Debug, Clone)]
struct Node {
level: u8,
left: Option<usize>,
right: Option<usize>,
/// `children.len() + 1` keys once the node holds a child.
keys: Vec<Key>,
/// Chunk addresses in a leaf, node indexes above.
children: Vec<u64>,
}
/// What an insertion below a node did (`H5B__insert_helper`'s outputs).
#[derive(Default)]
struct Ret {
/// The node's new left key (`lt_key_changed`).
lt: Option<Key>,
/// The node's new right key (`rt_key_changed`).
rt: Option<Key>,
/// The node split: the key shared by the halves and the new right node.
split: Option<(Key, usize)>,
}
struct Tree {
nodes: Vec<Node>,
two_k: usize,
}
fn bad(why: &str) -> FormatError {
FormatError::SerializationError(format!("version-1 B-tree chunk index: {why}"))
}
impl Tree {
fn new(k: u16) -> Self {
Self {
nodes: vec![Node {
level: 0,
left: None,
right: None,
keys: Vec::new(),
children: Vec::new(),
}],
two_k: 2 * usize::from(k),
}
}
/// `H5B_insert` of `key` (a chunk after every chunk already inserted).
fn insert(&mut self, key: &Key, addr: u64) -> Result<(), FormatError> {
let r = self.insert_helper(0, key, addr, 64)?;
let Some((md, split)) = r.split else {
return Ok(());
};
// The root split: copy it to a new node and make the root the
// parent of the copy and its new right sibling.
let lt = r.lt.unwrap_or_else(|| self.nodes[0].keys[0].clone());
let rt = match r.rt {
Some(rt) => rt,
None => self.nodes[split]
.keys
.last()
.cloned()
.ok_or_else(|| bad("empty node"))?,
};
let moved = self.nodes[0].clone();
let level = moved.level;
let moved_id = self.nodes.len();
self.nodes.push(moved);
self.nodes[split].left = Some(moved_id);
self.nodes[0] = Node {
level: level + 1,
left: None,
right: None,
keys: vec![lt, md, rt],
children: vec![moved_id as u64, split as u64],
};
Ok(())
}
fn insert_helper(
&mut self,
id: usize,
key: &Key,
addr: u64,
depth: u8,
) -> Result<Ret, FormatError> {
if depth == 0 {
return Err(bad("tree too deep"));
}
let n = self.nodes[id].children.len();
let level = self.nodes[id].level;
let mut ret = Ret::default();
if n == 0 {
// The first chunk (H5B_INS_FIRST): its key and the right key.
let node = &mut self.nodes[id];
node.keys = vec![key.clone(), key.right_of()];
node.children = vec![addr];
return Ok(ret);
}
// Binary search with H5D__btree_cmp3: 1 when the chunk is not below
// the right key, -1 when below the left key, else 0.
let (mut lo, mut hi, mut idx) = (0usize, n, 0usize);
let mut cmp = Ordering::Less;
while lo < hi && cmp != Ordering::Equal {
idx = (lo + hi) / 2;
let node = &self.nodes[id];
cmp = if key.cmp_scaled(&node.keys[idx + 1]) != Ordering::Less {
Ordering::Greater
} else if key.cmp_scaled(&node.keys[idx]) == Ordering::Less {
Ordering::Less
} else {
Ordering::Equal
};
if cmp == Ordering::Less {
hi = idx;
} else {
lo = idx + 1;
}
}
let (mut lt_changed, mut rt_changed) = (false, false);
// The child to add after child `idx`, with its left key.
let mut new_child: Option<(Key, u64)> = None;
match cmp {
Ordering::Less => return Err(bad("chunks out of order")),
Ordering::Greater if idx + 1 < n => {
return Err(bad("cannot place chunk"));
}
Ordering::Greater if level == 0 => {
// Past every chunk of the right-most leaf: a new maximum
// (H5B_INS_RIGHT through `new_node`), which moves the right
// key one chunk past it.
idx = n - 1;
self.nodes[id].keys[idx + 1] = key.right_of();
rt_changed = true;
new_child = Some((key.clone(), addr));
}
Ordering::Equal if level == 0 => {
// Inside the last chunk's range: H5D__btree_insert adds it
// to the right of that chunk; the right key stays.
if key.scaled == self.nodes[id].keys[idx].scaled {
return Err(bad("duplicate chunk"));
}
new_child = Some((key.clone(), addr));
}
_ => {
if cmp == Ordering::Greater {
idx = n - 1;
}
let child = usize::try_from(self.nodes[id].children[idx])
.map_err(|_| bad("bad node index"))?;
let r = self.insert_helper(child, key, addr, depth - 1)?;
if let Some(lt) = r.lt {
self.nodes[id].keys[idx] = lt;
lt_changed = true;
}
if let Some(rt) = r.rt {
self.nodes[id].keys[idx + 1] = rt;
rt_changed = true;
}
if let Some((md, split)) = r.split {
new_child = Some((md, split as u64));
}
}
}
// Pass the node's changed end keys up, as H5B__insert_helper does.
if lt_changed && idx == 0 {
ret.lt = Some(self.nodes[id].keys[0].clone());
}
if rt_changed && idx + 1 >= n {
ret.rt = Some(self.nodes[id].keys[idx + 1].clone());
}
if let Some((md, child)) = new_child {
// A full node splits first; the child goes to the half that
// holds child `idx`.
let (mut target, mut split) = (id, None);
if n == self.two_k {
let s = self.split(id, idx);
let nleft = self.nodes[id].children.len();
if idx >= nleft {
idx -= nleft;
target = s;
}
split = Some(s);
}
// H5B__insert_child (H5B_INS_RIGHT): the new child after child
// `idx`, its left key after that child's.
let node = &mut self.nodes[target];
node.keys.insert(idx + 1, md);
node.children.insert(idx + 1, child);
ret.split = split.map(|s| (self.nodes[s].keys[0].clone(), s));
}
Ok(ret)
}
/// `H5B__split` of the full node `id`, the insertion going after child
/// `idx`; returns the new right node.
fn split(&mut self, id: usize, idx: usize) -> usize {
let node = &self.nodes[id];
let ratio = if node.right.is_none() {
SPLIT_RATIOS[2]
} else if node.left.is_none() {
SPLIT_RATIOS[0]
} else {
SPLIT_RATIOS[1]
};
let mut nleft = (self.two_k as f64 * ratio) as usize;
if idx < nleft && nleft == self.two_k {
nleft -= 1;
} else if idx >= nleft && nleft == 0 {
nleft += 1;
}
let new_id = self.nodes.len();
let right = Node {
level: node.level,
left: Some(id),
right: node.right,
keys: node.keys[nleft..].to_vec(),
children: node.children[nleft..].to_vec(),
};
let old_right = node.right;
self.nodes.push(right);
if let Some(r) = old_right {
self.nodes[r].left = Some(new_id);
}
let node = &mut self.nodes[id];
node.keys.truncate(nleft + 1);
node.children.truncate(nleft);
node.right = Some(new_id);
new_id
}
}
/// Bytes of one node of a chunk B-tree with `ndims` key dimensions (the
/// dataset's rank plus the element-size one).
fn node_size(two_k: usize, ndims: usize, offset_size: usize) -> usize {
let key = 8 + 8 * ndims;
8 + 2 * offset_size + (two_k + 1) * key + two_k * offset_size
}
/// Build the chunk B-tree for `chunks`, given in row-major order of their
/// scaled coordinates, with nodes laid out from `base_address`. `chunk_dims`
/// are the chunk's dimensions (the dataset's rank of them) and `elem_size`
/// the element size, the key's last dimension. Returns the nodes' bytes; the
/// root is at `base_address`. `chunks` must not be empty: an index without
/// chunks has no tree (its address is undefined).
pub(crate) fn build_chunk_btree_v1_at(
chunks: &[ChunkEntry],
chunk_dims: &[u64],
elem_size: u32,
base_address: u64,
offset_size: u8,
) -> Result<Vec<u8>, FormatError> {
if chunks.is_empty() {
return Err(bad("no chunks"));
}
let rank = chunk_dims.len();
let mut tree = Tree::new(CHUNK_BTREE_K);
for c in chunks {
if c.scaled.len() != rank {
return Err(bad("chunk rank differs from the dataset's"));
}
let nbytes = u32::try_from(c.nbytes).map_err(|_| {
FormatError::SerializationError(format!(
"a chunk of {} bytes cannot be indexed by a version-1 B-tree \
(HDF5 1.8 chunks are under 4 GiB)",
c.nbytes
))
})?;
let mut scaled = c.scaled.clone();
scaled.push(0);
let key = Key {
nbytes,
mask: c.filter_mask,
scaled,
};
tree.insert(&key, c.address)?;
}
let os = usize::from(offset_size);
let ndims = rank + 1;
let nsize = node_size(tree.two_k, ndims, os);
let addr_of = |id: usize| base_address + (id * nsize) as u64;
let mut dims: Vec<u64> = chunk_dims.to_vec();
dims.push(u64::from(elem_size));
let mut out = vec![0u8; tree.nodes.len() * nsize];
for (i, node) in tree.nodes.iter().enumerate() {
let d = &mut out[i * nsize..(i + 1) * nsize];
d[0..4].copy_from_slice(b"TREE");
d[4] = 1; // node type: raw data chunks
d[5] = node.level;
let n = u16::try_from(node.children.len()).map_err(|_| bad("node too large"))?;
d[6..8].copy_from_slice(&n.to_le_bytes());
let undef = u64::MAX;
put_addr(&mut d[8..], node.left.map_or(undef, addr_of), os);
put_addr(&mut d[8 + os..], node.right.map_or(undef, addr_of), os);
let mut p = 8 + 2 * os;
for (k, key) in node.keys.iter().enumerate() {
d[p..p + 4].copy_from_slice(&key.nbytes.to_le_bytes());
d[p + 4..p + 8].copy_from_slice(&key.mask.to_le_bytes());
for (j, (&s, &dim)) in key.scaled.iter().zip(&dims).enumerate() {
let off = s
.checked_mul(dim)
.ok_or_else(|| FormatError::Overflow("chunk key offset".into()))?;
d[p + 8 + 8 * j..p + 16 + 8 * j].copy_from_slice(&off.to_le_bytes());
}
p += 8 + 8 * ndims;
if let Some(&child) = node.children.get(k) {
let a = if node.level == 0 {
child
} else {
addr_of(usize::try_from(child).map_err(|_| bad("bad node index"))?)
};
put_addr(&mut d[p..], a, os);
p += os;
}
}
}
Ok(out)
}
fn put_addr(d: &mut [u8], v: u64, os: usize) {
d[..os].copy_from_slice(&v.to_le_bytes()[..os]);
}
#[cfg(test)]
mod tests {
use super::*;
fn build(n: u64) -> Tree {
let mut t = Tree::new(CHUNK_BTREE_K);
for i in 0..n {
let key = Key {
nbytes: 80,
mask: 0,
scaled: vec![i, 0],
};
t.insert(&key, 1000 + i).unwrap();
}
t
}
/// Leaves in order from the root, with their child counts.
fn leaves(t: &Tree, id: usize, out: &mut Vec<usize>) {
let n = &t.nodes[id];
if n.level == 0 {
out.push(n.children.len());
} else {
for &c in &n.children {
leaves(t, c as usize, out);
}
}
}
#[test]
fn sequential_appends_split_as_libhdf5_does() {
// libhdf5 2.0 (h5py, libver=('v108', 'latest')) writes 1000 chunks
// as a root over 17 leaves of 57 chunks and one of 31, with the
// root's right key on the last chunk (9990, 8 for 10-element f8
// chunks).
let t = build(1000);
assert_eq!(t.nodes[0].level, 1);
let mut l = Vec::new();
leaves(&t, 0, &mut l);
let mut want = vec![57; 17];
want.push(31);
assert_eq!(l, want);
assert_eq!(t.nodes[0].keys.last().unwrap().scaled, vec![999, 1]);
// 100 000 chunks: three levels, a root of 31 children.
let t = build(100_000);
assert_eq!(t.nodes[0].level, 2);
assert_eq!(t.nodes[0].children.len(), 31);
}
#[test]
fn right_key_moves_every_other_append() {
let t = build(5);
assert_eq!(t.nodes[0].keys.last().unwrap().scaled, vec![5, 1]);
let t = build(6);
assert_eq!(t.nodes[0].keys.last().unwrap().scaled, vec![5, 1]);
}
#[test]
fn keys_and_siblings_are_consistent() {
let t = build(5000);
for (i, n) in t.nodes.iter().enumerate() {
assert!(n.children.len() <= t.two_k);
assert_eq!(n.keys.len(), n.children.len() + 1);
if let Some(r) = n.right {
assert_eq!(t.nodes[r].left, Some(i));
assert_eq!(n.keys.last(), t.nodes[r].keys.first());
}
}
}
}
+105
View File
@@ -7,6 +7,7 @@ use crate::addr::saturating_usize;
#[cfg(not(feature = "std"))] #[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec}; use alloc::{format, vec, vec::Vec};
use crate::btree_v1_write;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2}; use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::checksum::jenkins_lookup3; use crate::checksum::jenkins_lookup3;
use crate::chunk_cache::{CACHE_LINE_SIZE, align_to_cache_line}; use crate::chunk_cache::{CACHE_LINE_SIZE, align_to_cache_line};
@@ -19,6 +20,7 @@ use crate::filter_pipeline::{
FilterPipeline, FilterPipeline,
}; };
use crate::filters::compress_chunk_masked; use crate::filters::compress_chunk_masked;
use crate::libver::LibVer;
/// Round a file offset up to the next cache-line boundary. /// Round a file offset up to the next cache-line boundary.
/// ///
/// This ensures chunk data starts at an address that is a multiple of the /// This ensures chunk data starts at an address that is a multiple of the
@@ -866,6 +868,23 @@ pub fn build_chunked_data_from_precompressed(
base_address: u64, base_address: u64,
maxshape: Option<&[u64]>, maxshape: Option<&[u64]>,
) -> Result<ChunkedDataResult, FormatError> { ) -> Result<ChunkedDataResult, FormatError> {
build_chunked_data_from_precompressed_libver(pre, base_address, maxshape, LibVer::Latest)
}
/// [`build_chunked_data_from_precompressed`] for a file whose low library
/// version bound is `low`: below [`LibVer::V110`] (that is, for HDF5 1.8)
/// every chunked dataset gets a version-3 layout message and a version-1
/// B-tree chunk index, whatever its maximum shape, as libhdf5 writes it;
/// otherwise the version-4 layout and the index libhdf5 picks for it.
pub fn build_chunked_data_from_precompressed_libver(
pre: &PrecompressedChunks,
base_address: u64,
maxshape: Option<&[u64]>,
low: LibVer,
) -> Result<ChunkedDataResult, FormatError> {
if low < LibVer::V110 {
return build_btree_v1_chunked_data(pre, base_address, maxshape);
}
let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?; let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?;
let offset_size: u8 = 8; let offset_size: u8 = 8;
let length_size: u8 = 8; let length_size: u8 = 8;
@@ -992,6 +1011,92 @@ pub fn build_chunked_data_from_precompressed(
}) })
} }
/// Lay out precompressed chunks at `base_address` followed by a version-1
/// B-tree chunk index, with a version-3 layout message: what libhdf5 writes
/// for a chunked dataset under a low bound of 1.8.
fn build_btree_v1_chunked_data(
pre: &PrecompressedChunks,
base_address: u64,
maxshape: Option<&[u64]>,
) -> Result<ChunkedDataResult, FormatError> {
if let Some(ms) = maxshape {
let bad = |what: &str| FormatError::ChunkedReadError(format!("maxshape: {what}"));
if ms.len() != pre.shape.len() {
return Err(bad("rank differs from the shape"));
}
if ms.iter().zip(&pre.shape).any(|(&m, &s)| m < s) {
return Err(bad("smaller than the shape"));
}
}
let offset_size: u8 = 8;
let mut data_buf = Vec::new();
let mut entries = Vec::with_capacity(pre.chunks.len());
for (i, (_raw_size, stored, filter_mask)) in pre.chunks.iter().enumerate() {
let aligned_offset = align_to_cache_line(data_buf.len());
if aligned_offset > data_buf.len() {
data_buf.resize(aligned_offset, 0u8);
}
entries.push(btree_v1_write::ChunkEntry {
scaled: scaled_coords(&pre.shape, &pre.chunk_dims, i),
nbytes: stored.len() as u64,
filter_mask: *filter_mask,
address: base_address + data_buf.len() as u64,
});
data_buf.extend_from_slice(stored);
}
let element_size = u32::try_from(pre.element_size)
.map_err(|_| FormatError::Overflow("element size".into()))?;
// A dataset with no chunks has no tree: its address is undefined, as
// libhdf5 leaves it until the first chunk is written.
let btree_address = if entries.is_empty() {
u64::MAX
} else {
let aligned_idx = align_to_cache_line(data_buf.len());
if aligned_idx > data_buf.len() {
data_buf.resize(aligned_idx, 0u8);
}
let addr = base_address + data_buf.len() as u64;
let tree = btree_v1_write::build_chunk_btree_v1_at(
&entries,
&pre.chunk_dims,
element_size,
addr,
offset_size,
)?;
data_buf.extend_from_slice(&tree);
addr
};
let layout_message =
serialize_v3_chunked(&pre.chunk_dims, btree_address, offset_size, element_size)?;
Ok(ChunkedDataResult {
data_bytes: data_buf,
layout_message,
pipeline_message: pre.pipeline_message.clone(),
})
}
/// A version-3 layout message for a chunked dataset: dimensionality (the
/// rank plus one), the B-tree's address, then each chunk dimension and the
/// element size, four bytes each.
fn serialize_v3_chunked(
chunk_dims: &[u64],
btree_address: u64,
offset_size: u8,
element_size: u32,
) -> Result<Vec<u8>, FormatError> {
let ndims = u8::try_from(chunk_dims.len() + 1)
.map_err(|_| FormatError::Overflow("chunked layout rank".into()))?;
let mut buf = vec![3u8, 2, ndims];
push_addr(&mut buf, btree_address, offset_size);
for &d in chunk_dims {
let d =
u32::try_from(d).map_err(|_| FormatError::Overflow(format!("chunk dimension {d}")))?;
buf.extend_from_slice(&d.to_le_bytes());
}
buf.extend_from_slice(&element_size.to_le_bytes());
Ok(buf)
}
/// Most slots a Fixed Array index may have before we refuse to build it: its /// Most slots a Fixed Array index may have before we refuse to build it: its
/// data block holds one element per chunk of the *maximum* extent, so a huge /// data block holds one element per chunk of the *maximum* extent, so a huge
/// finite maxshape with small chunks would otherwise exhaust memory. /// finite maxshape with small chunks would otherwise exhaust memory.
+20
View File
@@ -1216,6 +1216,26 @@ impl Datatype {
} }
} }
/// The highest datatype message version in this type's encoding, its
/// members' and base types' included (the version decides which HDF5
/// releases can read it: 1-3 HDF5 1.8, 4 HDF5 1.12, 5 HDF5 2.0).
pub fn max_encoded_version(&self) -> u8 {
let own = self.serialize().first().map_or(0, |b| b >> 4);
let inner = match self {
Datatype::Compound { members, .. } => members
.iter()
.map(|m| m.datatype.max_encoded_version())
.max()
.unwrap_or(0),
Datatype::Enumeration { base_type, .. }
| Datatype::VariableLength { base_type, .. }
| Datatype::Array { base_type, .. }
| Datatype::Complex { base_type, .. } => base_type.max_encoded_version(),
_ => 0,
};
own.max(inner)
}
/// Check that this datatype can be written: every part of it has an /// Check that this datatype can be written: every part of it has an
/// on-disk encoding, and the encoding is one the reader (and libhdf5) /// on-disk encoding, and the encoding is one the reader (and libhdf5)
/// accepts. [`Self::serialize`] cannot report errors, so the writer calls /// accepts. [`Self::serialize`] cannot report errors, so the writer calls
+20
View File
@@ -167,6 +167,19 @@ pub enum FormatError {
VlDataError(String), VlDataError(String),
/// Serialization error. /// Serialization error.
SerializationError(String), SerializationError(String),
/// The file's library version bounds
/// ([`FileWriter::libver_bounds`](crate::file_writer::FileWriter::libver_bounds))
/// do not allow what was asked for: `what` needs the format of HDF5
/// `needs` or later, and the high bound is `high` (or the low bound is
/// above the high one, with `needs` the low bound).
LibverBound {
/// What cannot be written.
what: String,
/// The oldest release whose format holds it.
needs: crate::libver::LibVer,
/// The file's high bound.
high: crate::libver::LibVer,
},
/// Dataset is missing data. /// Dataset is missing data.
DatasetMissingData, DatasetMissingData,
/// Dataset is missing shape. /// Dataset is missing shape.
@@ -450,6 +463,13 @@ impl fmt::Display for FormatError {
FormatError::SerializationError(msg) => { FormatError::SerializationError(msg) => {
write!(f, "serialization error: {msg}") write!(f, "serialization error: {msg}")
} }
FormatError::LibverBound { what, needs, high } => {
write!(
f,
"{what} needs the HDF5 {needs} file format, above the high \
library version bound ({high})"
)
}
FormatError::DatasetMissingData => { FormatError::DatasetMissingData => {
write!(f, "dataset is missing data") write!(f, "dataset is missing data")
} }
+236 -8
View File
@@ -5,12 +5,13 @@
use crate::addr::saturating_usize; use crate::addr::saturating_usize;
#[cfg(not(feature = "std"))] #[cfg(not(feature = "std"))]
use alloc::{format, vec, vec::Vec}; use alloc::{format, string::String, vec, vec::Vec};
use crate::attribute::AttributeMessage; use crate::attribute::AttributeMessage;
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2}; use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
use crate::chunked_write::{ use crate::chunked_write::{
ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed, precompress_chunks, ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed_libver,
precompress_chunks,
}; };
use crate::data_layout::VdsMapping; use crate::data_layout::VdsMapping;
use crate::dataspace::{Dataspace, DataspaceType}; use crate::dataspace::{Dataspace, DataspaceType};
@@ -31,6 +32,7 @@ pub use crate::type_builders::ProvenanceConfig;
pub use crate::type_builders::{AttrValue, CompoundTypeBuilder, EnumTypeBuilder}; pub use crate::type_builders::{AttrValue, CompoundTypeBuilder, EnumTypeBuilder};
use crate::datatype::{CharacterSet, Datatype}; use crate::datatype::{CharacterSet, Datatype};
use crate::libver::LibVer;
pub(crate) const OFFSET_SIZE: u8 = 8; pub(crate) const OFFSET_SIZE: u8 = 8;
pub(crate) const LENGTH_SIZE: u8 = 8; pub(crate) const LENGTH_SIZE: u8 = 8;
@@ -168,13 +170,15 @@ pub(crate) fn build_dataset_oh(
attrs: AttrStorage<'_>, attrs: AttrStorage<'_>,
fill_message: &[u8], fill_message: &[u8],
refcount: u32, refcount: u32,
layout_version: u8,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
let mut w = ObjectHeaderWriter::new(); let mut w = ObjectHeaderWriter::new();
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01); w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE)); w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
w.add_message_with_flags(MessageType::FillValue, fill_message.to_vec(), 0x01); w.add_message_with_flags(MessageType::FillValue, fill_message.to_vec(), 0x01);
// Versions 3 and 4 encode a contiguous layout the same way.
let mut dl = Vec::new(); let mut dl = Vec::new();
dl.push(4); // version dl.push(layout_version);
dl.push(1); // class = contiguous dl.push(1); // class = contiguous
// An empty dataset has no storage: its address must be the undefined // An empty dataset has no storage: its address must be the undefined
// address, as libhdf5 writes it. A real address with size 0 trips // address, as libhdf5 writes it. A real address with size 0 trips
@@ -198,14 +202,16 @@ pub(crate) fn build_compact_dataset_oh(
attrs: AttrStorage<'_>, attrs: AttrStorage<'_>,
fill_message: &[u8], fill_message: &[u8],
refcount: u32, refcount: u32,
layout_version: u8,
) -> Result<Vec<u8>, FormatError> { ) -> Result<Vec<u8>, FormatError> {
let mut w = ObjectHeaderWriter::new(); let mut w = ObjectHeaderWriter::new();
w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01); w.add_message_with_flags(MessageType::Datatype, dt.serialize(), 0x01);
w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE)); w.add_message(MessageType::Dataspace, ds.serialize(LENGTH_SIZE));
w.add_message_with_flags(MessageType::FillValue, fill_message.to_vec(), 0x01); w.add_message_with_flags(MessageType::FillValue, fill_message.to_vec(), 0x01);
// Compact layout message: version=4, class=0, u16 size, inline data // Compact layout message: version (3 and 4 are the same here), class=0,
// u16 size, inline data
let mut dl = Vec::new(); let mut dl = Vec::new();
dl.push(4); // version dl.push(layout_version);
dl.push(0); // class = compact dl.push(0); // class = compact
dl.extend_from_slice(&(data.len() as u16).to_le_bytes()); dl.extend_from_slice(&(data.len() as u16).to_le_bytes());
dl.extend_from_slice(data); dl.extend_from_slice(data);
@@ -1371,6 +1377,10 @@ pub struct FileWriter {
/// file-space strategy (a File Space Info message in the superblock /// file-space strategy (a File Space Info message in the superblock
/// extension). /// extension).
page_size: Option<u32>, page_size: Option<u32>,
/// Library version bounds: the low bound picks the format versions
/// written, the high bound limits the features allowed.
low: LibVer,
high: LibVer,
} }
impl Default for FileWriter { impl Default for FileWriter {
@@ -1485,9 +1495,37 @@ impl FileWriter {
alignment_threshold: 0, alignment_threshold: 0,
alignment_bytes: 0, alignment_bytes: 0,
page_size: None, page_size: None,
low: LibVer::V110,
high: LibVer::Latest,
} }
} }
/// Set the library version bounds, as libhdf5's `H5Pset_libver_bounds`
/// (h5py's `libver=(low, high)`): the oldest HDF5 release whose format
/// the file uses (`low`), and the newest whose features it may use
/// (`high`). See [`crate::libver`] for what each bound changes.
///
/// The default, `(LibVer::V110, LibVer::Latest)`, is what clawhdf5 has
/// always written: the HDF5 1.10 format (version-3 superblock, version-4
/// layouts with the 1.10 chunk indexes), readable by HDF5 1.10 and later.
///
/// `(LibVer::V18, LibVer::V18)` writes a file HDF5 1.8 can read — the
/// low bound libhdf5 2.0 uses by default: a version-2 superblock,
/// version-3 layouts, and a version-1 B-tree for every chunked dataset,
/// resizable ones included; [`Self::finish`] then fails with
/// [`FormatError::LibverBound`] for anything HDF5 1.8 cannot read
/// (virtual datasets, a paged file, the 1.12 reference types, native
/// complex numbers). With a low bound of 1.8 and a later high bound
/// such objects are written in the newer format, as libhdf5 writes them;
/// the rest of the file stays readable by 1.8.
///
/// A low bound above the high bound makes [`Self::finish`] fail.
pub fn libver_bounds(&mut self, low: LibVer, high: LibVer) -> &mut Self {
self.low = low;
self.high = high;
self
}
/// Set global file alignment: datasets with raw data >= `threshold` bytes /// Set global file alignment: datasets with raw data >= `threshold` bytes
/// will have their data aligned to `bytes` boundary. /// will have their data aligned to `bytes` boundary.
/// ///
@@ -1583,6 +1621,33 @@ impl FileWriter {
))); )));
} }
let (low, high) = (self.low, self.high);
let within_bounds = |what: &dyn Fn() -> String, needs: LibVer| {
if needs > high {
Err(FormatError::LibverBound {
what: what(),
needs,
high,
})
} else {
Ok(())
}
};
within_bounds(&|| format!("a low library version bound of {low}"), low)?;
if page_size.is_some() {
within_bounds(&|| "the paged file-space strategy".into(), LibVer::V110)?;
}
// Versions 3 of the layout message and 2 of the superblock are what
// HDF5 1.8 reads; 1.10 added version 4 (with its chunk indexes) and
// version 3. A paged file needs the version-3 superblock whatever
// the low bound (libhdf5 raises it as far as the high bound allows).
let layout_version: u8 = if low < LibVer::V110 { 3 } else { 4 };
let superblock_version: u8 = if low < LibVer::V110 && page_size.is_none() {
2
} else {
3
};
// The group tree, in layout order: groups depth-first from the root, // The group tree, in layout order: groups depth-first from the root,
// then every group's datasets in the same order. // then every group's datasets in the same order.
let tree = writer_tree::build(self.root, self.track_order)?; let tree = writer_tree::build(self.root, self.track_order)?;
@@ -1621,9 +1686,20 @@ impl FileWriter {
let ds_attrs = all_ds.iter().flat_map(|d| &d.attrs); let ds_attrs = all_ds.iter().flat_map(|d| &d.attrs);
for a in group_attrs.chain(ds_attrs) { for a in group_attrs.chain(ds_attrs) {
a.datatype.check_encodable()?; a.datatype.check_encodable()?;
within_bounds(
&|| format!("the datatype of attribute {:?}", a.name),
LibVer::for_datatype_version(a.datatype.max_encoded_version()),
)?;
} }
for d in &all_ds { for d in &all_ds {
d.dt.check_encodable()?; d.dt.check_encodable()?;
within_bounds(
&|| "a dataset's datatype".into(),
LibVer::for_datatype_version(d.dt.max_encoded_version()),
)?;
if d.virtual_sources.is_some() {
within_bounds(&|| "a virtual dataset".into(), LibVer::V110)?;
}
} }
let is_vds: Vec<bool> = all_ds.iter().map(|d| d.virtual_sources.is_some()).collect(); let is_vds: Vec<bool> = all_ds.iter().map(|d| d.virtual_sources.is_some()).collect();
@@ -1749,10 +1825,11 @@ impl FileWriter {
elem_size, elem_size,
&d.chunk_options, &d.chunk_options,
)?; )?;
let result = build_chunked_data_from_precompressed( let result = build_chunked_data_from_precompressed_libver(
&pre, &pre,
dummy_cursor, dummy_cursor,
d.maxshape.as_deref(), d.maxshape.as_deref(),
low,
)?; )?;
dummy_cursor += result.data_bytes.len() as u64; dummy_cursor += result.data_bytes.len() as u64;
let oh = build_chunked_dataset_oh( let oh = build_chunked_dataset_oh(
@@ -1785,6 +1862,7 @@ impl FileWriter {
}, },
&d.fill_message, &d.fill_message,
d.refcount, d.refcount,
layout_version,
)?; )?;
dummy_blobs.push(DataBlob { dummy_blobs.push(DataBlob {
data: vec![], data: vec![],
@@ -1804,6 +1882,7 @@ impl FileWriter {
}, },
&d.fill_message, &d.fill_message,
d.refcount, d.refcount,
layout_version,
)?; )?;
dummy_blobs.push(DataBlob { dummy_blobs.push(DataBlob {
data: vec![], data: vec![],
@@ -1904,13 +1983,14 @@ impl FileWriter {
let base_address = cursor2 as u64; let base_address = cursor2 as u64;
// Reuse precompressed chunks from Pass 1 — avoids re-compressing // Reuse precompressed chunks from Pass 1 — avoids re-compressing
// the same data a second time. // the same data a second time.
let result = build_chunked_data_from_precompressed( let result = build_chunked_data_from_precompressed_libver(
dummy_blobs[i] dummy_blobs[i]
.precompressed .precompressed
.as_ref() .as_ref()
.expect("chunked dataset missing precompressed cache"), .expect("chunked dataset missing precompressed cache"),
base_address, base_address,
d.maxshape.as_deref(), d.maxshape.as_deref(),
low,
)?; )?;
cursor2 += result.data_bytes.len(); cursor2 += result.data_bytes.len();
let oh = build_chunked_dataset_oh( let oh = build_chunked_dataset_oh(
@@ -1944,6 +2024,7 @@ impl FileWriter {
}, },
&d.fill_message, &d.fill_message,
d.refcount, d.refcount,
layout_version,
)?; )?;
ds_blobs2.push(DataBlob { ds_blobs2.push(DataBlob {
data: vec![], data: vec![],
@@ -1973,6 +2054,7 @@ impl FileWriter {
}, },
&d.fill_message, &d.fill_message,
d.refcount, d.refcount,
layout_version,
)?; )?;
let mut data = vec![0u8; padding]; let mut data = vec![0u8; padding];
data.extend_from_slice(&d.raw); data.extend_from_slice(&d.raw);
@@ -1997,7 +2079,7 @@ impl FileWriter {
let mut buf = Vec::with_capacity(cursor2); let mut buf = Vec::with_capacity(cursor2);
let sb = Superblock { let sb = Superblock {
version: 3, version: superblock_version,
offset_size: OFFSET_SIZE, offset_size: OFFSET_SIZE,
length_size: LENGTH_SIZE, length_size: LENGTH_SIZE,
base_address: 0, base_address: 0,
@@ -2783,4 +2865,150 @@ mod tests {
assert_eq!(sb.version, 3); assert_eq!(sb.version, 3);
assert_eq!(sb.page_size, None); assert_eq!(sb.page_size, None);
} }
fn layout_of(bytes: &[u8], name: &str) -> Vec<u8> {
let sb = Superblock::parse(bytes, 0).unwrap();
let addr = resolve_path_any(bytes, &sb, name).unwrap();
let hdr = ObjectHeader::parse(bytes, addr as usize, 8, 8).unwrap();
hdr.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
.unwrap()
.data
.clone()
}
#[test]
fn libver_v18_writes_the_1_8_format() {
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("contig").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("compact").with_f64_data(&[3.0]).compact();
fw.create_dataset("grow")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
fw.create_dataset("none")
.with_f64_data(&[])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
let bytes = fw.finish().unwrap();
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 2);
assert_eq!(layout_of(&bytes, "contig")[..2], [3, 1]);
assert_eq!(layout_of(&bytes, "compact")[..2], [3, 0]);
let grow = layout_of(&bytes, "grow");
// Version 3, chunked, 2 dimensions (the element size is the last),
// B-tree address, chunk dims 2 and 8.
assert_eq!(grow[..3], [3, 2, 2]);
assert_eq!(grow[11..], [2, 0, 0, 0, 8, 0, 0, 0]);
let root = u64::from_le_bytes(grow[3..11].try_into().unwrap()) as usize;
assert_eq!(&bytes[root..root + 5], b"TREE\x01");
// No chunks, no tree.
assert_eq!(layout_of(&bytes, "none")[3..11], [0xff; 8]);
assert_eq!(read_dataset_f64(&bytes, "grow"), vec![1.0, 2.0, 3.0]);
assert_eq!(read_dataset_f64(&bytes, "contig"), vec![1.0, 2.0]);
assert_eq!(read_dataset_f64(&bytes, "compact"), vec![3.0]);
}
#[test]
fn default_libver_bounds_keep_the_1_10_format() {
let mut fw = FileWriter::new();
fw.create_dataset("contig").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("grow")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
let default = fw.finish().unwrap();
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V110, LibVer::Latest);
fw.create_dataset("contig").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("grow")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_maxshape(&[u64::MAX])
.with_chunks(&[2]);
assert_eq!(fw.finish().unwrap(), default);
assert_eq!(layout_of(&default, "contig")[0], 4);
assert_eq!(layout_of(&default, "grow")[..2], [4, 2]);
}
#[test]
fn libver_high_bound_refuses_newer_features() {
let bound = |r: Result<Vec<u8>, FormatError>, needs: LibVer| match r {
Err(FormatError::LibverBound { needs: n, high, .. }) => {
assert_eq!((n, high), (needs, LibVer::V18));
}
other => panic!("expected a bound error, got {other:?}"),
};
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("z")
.with_native_complex_f64_data(&[[1.0, 2.0]]);
bound(fw.finish(), LibVer::V200);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("x").with_f64_data(&[1.0]).set_attr(
"z",
AttrValue::Raw {
datatype: crate::type_builders::make_native_complex_f64_type(),
shape: vec![],
data: vec![0; 16],
},
);
bound(fw.finish(), LibVer::V200);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("r").with_compound_data(
Datatype::Reference {
size: 16,
ref_type: crate::datatype::ReferenceType::Object2,
},
vec![0; 16],
1,
);
bound(fw.finish(), LibVer::V112);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18);
fw.create_dataset("src").with_f64_data(&[1.0, 2.0]);
fw.create_dataset("vds")
.with_shape(&[2])
.with_f64_data(&[])
.with_virtual_sources(vec![VdsMapping {
source_file: ".".into(),
source_dataset: "src".into(),
source_selection: sel_all(),
virtual_selection: sel_hyper_1d(0, 2),
}]);
bound(fw.finish(), LibVer::V110);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::V18)
.with_page_size(4096);
bound(fw.finish(), LibVer::V110);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V110, LibVer::V18);
bound(fw.finish(), LibVer::V110);
}
#[test]
fn libver_low_v18_high_latest_allows_newer_objects() {
// As libhdf5 does: the object that needs a newer format gets it,
// the rest of the file keeps the 1.8 format.
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::Latest);
fw.create_dataset("z")
.with_native_complex_f64_data(&[[1.0, 2.0]]);
let bytes = fw.finish().unwrap();
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 2);
let mut fw = FileWriter::new();
fw.libver_bounds(LibVer::V18, LibVer::Latest)
.with_page_size(4096);
fw.create_dataset("x").with_f64_data(&[1.0]);
let bytes = fw.finish().unwrap();
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 3);
assert_eq!(layout_of(&bytes, "x")[0], 3);
}
} }
+2
View File
@@ -61,6 +61,7 @@ pub mod addr;
pub mod attribute; pub mod attribute;
pub mod attribute_info; pub mod attribute_info;
pub mod btree_v1; pub mod btree_v1;
mod btree_v1_write;
pub mod btree_v2; pub mod btree_v2;
mod btree_v2_write; mod btree_v2_write;
mod bulk_alloc; mod bulk_alloc;
@@ -107,6 +108,7 @@ pub mod group_v1;
pub mod group_v2; pub mod group_v2;
#[cfg(feature = "parallel")] #[cfg(feature = "parallel")]
pub mod lane_partition; pub mod lane_partition;
pub mod libver;
pub mod link_info; pub mod link_info;
pub mod link_message; pub mod link_message;
pub mod local_heap; pub mod local_heap;
+93
View File
@@ -0,0 +1,93 @@
//! Library version bounds for writing: which HDF5 releases can read a file.
//!
//! libhdf5 picks the version of every object it writes from the file's
//! *low* bound (`H5Pset_libver_bounds`; h5py's `libver=`): the oldest
//! format version that holds the object, but never older than the one the
//! low bound names. The *high* bound caps it: a feature that needs a newer
//! format than the high bound is an error. [`LibVer`] names the same
//! releases, and [`crate::file_writer::FileWriter::libver_bounds`] sets them.
//!
//! What the low bound changes in what clawhdf5 writes:
//!
//! | | low [`LibVer::V18`] | low [`LibVer::V110`] or later (the default) |
//! |---|---|---|
//! | superblock | version 2 | version 3 |
//! | data layout message | version 3 | version 4 |
//! | chunk index | version-1 B-tree (every chunked dataset) | single chunk, Fixed Array, Extensible Array or version-2 B-tree, as libhdf5 picks |
//!
//! Everything else (version-2 object headers, link and group-info messages,
//! dense storage in fractal heaps with version-2 B-trees, filter pipeline
//! version 2, fill value version 3, datatype versions up to 3) is the same
//! and already readable by HDF5 1.8.
//!
//! What the high bound refuses: anything that needs 1.10 (virtual datasets,
//! the paged file-space strategy) above [`LibVer::V18`], the 1.12 reference
//! types (datatype version 4) above [`LibVer::V110`], and HDF5 2.0's native
//! complex numbers (datatype version 5) above [`LibVer::V114`].
//! `libver_bounds(LibVer::V18, LibVer::V18)` therefore writes a file HDF5
//! 1.8 can read, or fails.
use core::fmt;
/// An HDF5 library release, as a bound on the file format versions a writer
/// may use (libhdf5's `H5F_libver_t`). Ordered oldest first.
///
/// There is no `Earliest`: clawhdf5 cannot write the pre-1.8 format
/// (symbol-table groups, version-1 object headers).
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash)]
#[non_exhaustive]
pub enum LibVer {
/// HDF5 1.8 (`H5F_LIBVER_V18`, h5py `'v108'`).
V18,
/// HDF5 1.10 (`H5F_LIBVER_V110`, h5py `'v110'`).
V110,
/// HDF5 1.12 (`H5F_LIBVER_V112`, h5py `'v112'`).
V112,
/// HDF5 1.14 (`H5F_LIBVER_V114`, h5py `'v114'`).
V114,
/// HDF5 2.0 (`H5F_LIBVER_V200`).
V200,
/// The newest format this build of clawhdf5 writes
/// (`H5F_LIBVER_LATEST`, h5py `'latest'`).
Latest,
}
impl LibVer {
/// The release a datatype message of this version first appeared in:
/// versions 1-3 are readable by HDF5 1.8, 4 needs 1.12, 5 needs 2.0.
pub(crate) fn for_datatype_version(version: u8) -> Self {
match version {
0..=3 => LibVer::V18,
4 => LibVer::V112,
_ => LibVer::V200,
}
}
}
impl fmt::Display for LibVer {
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
f.write_str(match self {
LibVer::V18 => "1.8",
LibVer::V110 => "1.10",
LibVer::V112 => "1.12",
LibVer::V114 => "1.14",
LibVer::V200 => "2.0",
LibVer::Latest => "latest",
})
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn ordered_oldest_first() {
assert!(LibVer::V18 < LibVer::V110);
assert!(LibVer::V114 < LibVer::V200);
assert!(LibVer::V200 < LibVer::Latest);
assert_eq!(LibVer::for_datatype_version(3), LibVer::V18);
assert_eq!(LibVer::for_datatype_version(4), LibVer::V112);
assert_eq!(LibVer::for_datatype_version(5), LibVer::V200);
}
}
+25 -7
View File
@@ -36,17 +36,35 @@ let values: Vec<f64> = temp.read_f64()?;
| Item | What | | Item | What |
|---|---| |---|---|
| `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` | | `NetCDF4File` | `open`, `from_bytes`, `dimensions`, `variables`, `variable_names`, `variable`, `global_attrs`, `group`, `group_names`, `nc_properties`, and `hdf5_file` for the underlying `clawhdf5::File` |
| `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `attrs`, nested `group`) | | `NetCDF4Group` | the same for a sub-group (`dimensions`, `variables`, `variable_names`, `attrs`, nested `group`) |
| `Variable` | `name`, `shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` | | `Variable` | `name`, `shape`, `stored_shape`, `dimensions`, `nc_type`, `is_coordinate`, `attrs`, `cf_attributes`; `read_f64` (CF scale/offset and fill applied), `read_raw_f32`/`_f64`/`_i32`/`_i64`/`_u64`, `read_string`, `read_raw` |
| `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) | | `Dimension` | `name`, `size`, `is_unlimited` (an unlimited dimension's `size` is its current length as netCDF-C reports it: the largest extent of the variables using it) |
| `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` | | `CfAttributes` | CF convention attributes: `units`, `long_name`, `standard_name`, `fill_value` (`_FillValue`), `missing_value`, `scale_factor`, `add_offset`, `valid_range`, `calendar`, `axis` |
| `NcType` | the NetCDF type of a variable | | `NcType` | the NetCDF type of a variable |
No cargo features. Tests compare against files written by netCDF4-python Variables and dimensions follow netCDF-C:
(`tests/interop_tests.rs`; the CI job requires them with
`CLAWHDF5_REQUIRE_INTEROP=1`). What the HDF5 reader underneath cannot - A variable's dimensions are the ones the file names: the ids in its
read is listed in [`docs/known-issues.md`](../../docs/known-issues.md). `_Netcdf4Coordinates` attribute, else the dimension scales its
`DIMENSION_LIST` references, found in its group or a parent group. Only
an axis the file names no dimension for (an HDF5 file not written by a
netCDF library) gets the first dimension of the group of the same size,
else an anonymous `dim_<size>`.
- Dimension scales that are only dimensions are not variables; a dataset
`_nc4_non_coord_<name>` is the variable `<name>`.
- A variable along an unlimited dimension has the dimension's length:
`shape` is that length, and the reads return that many values, the
records the variable has not written as its fill value (`_FillValue`,
else netCDF's default for the type; NaN from `read_f64`).
`stored_shape` is the HDF5 dataset's extent.
No cargo features. Tests compare against files written by netCDF4-python,
h5py dimension scales, h5netcdf and xarray, variable by variable with what
netCDF4-python reads (`tests/interop_tests.rs`; the CI job requires them
with `CLAWHDF5_REQUIRE_INTEROP=1`; the h5netcdf cases skip when h5netcdf is
not installed). What the HDF5 reader underneath cannot read is listed in
[`docs/known-issues.md`](../../docs/known-issues.md).
## License ## License
+144 -94
View File
@@ -22,73 +22,123 @@ pub struct Dimension {
pub is_unlimited: bool, pub is_unlimited: bool,
} }
/// Extract dimensions from an HDF5 group (root or subgroup). /// A dimension scale of one group: the dataset that defines a dimension.
#[derive(Debug, Clone)]
pub(crate) struct Scale {
/// Object header address of the scale's dataset (what a variable's
/// `DIMENSION_LIST` references).
pub address: u64,
/// Its `_Netcdf4Dimid` (what a variable's `_Netcdf4Coordinates` lists).
pub dimid: Option<i64>,
/// Index of its dimension in [`GroupDims::dims`].
pub dim: usize,
}
/// The dimensions a group defines, with the scales that define them.
#[derive(Debug, Clone, Default)]
pub(crate) struct GroupDims {
/// The group's dimensions, in `_Netcdf4Dimid` order (then discovery order).
pub dims: Vec<Dimension>,
/// The dimension scales behind `dims`; empty when the group has no
/// dimension scale and `dims` were inferred from 1-D datasets.
pub scales: Vec<Scale>,
}
impl GroupDims {
/// The dimension defined by the scale at `address`.
pub fn by_address(&self, address: u64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.address == address)
.map(|s| &self.dims[s.dim])
}
/// The dimension whose scale has `_Netcdf4Dimid` `id`.
pub fn by_dimid(&self, id: i64) -> Option<&Dimension> {
self.scales
.iter()
.find(|s| s.dimid == Some(id))
.map(|s| &self.dims[s.dim])
}
}
/// The dimensions of an HDF5 group (root or subgroup).
/// ///
/// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed /// NetCDF-4 stores dimensions as datasets with `CLASS=DIMENSION_SCALE`. A fixed
/// dimension's size is the dataset's first (and typically only) shape extent. /// dimension's size is the dataset's first (and typically only) shape extent.
/// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace; /// Unlimited dimensions have `max_dimensions[0] == u64::MAX` in the HDF5 dataspace;
/// their size is computed by `unlimited_len`. /// their size is computed by `unlimited_len`. A group with no dimension
pub(crate) fn extract_dimensions( /// scale at all (not written by a netCDF library) gets one dimension per
/// 1-D dataset instead.
pub(crate) fn group_dims(
file: &clawhdf5::File, file: &clawhdf5::File,
group: &clawhdf5::Group<'_>, group: &clawhdf5::Group<'_>,
) -> Result<Vec<Dimension>, Error> { ) -> Result<GroupDims, Error> {
let addresses: HashMap<String, u64> = group.entries()?.into_iter().collect();
let dataset_names = group.datasets()?; let dataset_names = group.datasets()?;
let mut dims = Vec::new(); // (dimid, dimension, scale address), in discovery order.
let mut seen_dimids: HashMap<i64, usize> = HashMap::new(); let mut found: Vec<(Option<i64>, Dimension, u64)> = Vec::new();
for ds_name in &dataset_names { for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?; let ds = group.dataset(ds_name)?;
let attrs = ds.attrs()?; let attrs = ds.attrs()?;
// Check if this is a dimension scale
if !is_dimension_scale(&attrs) { if !is_dimension_scale(&attrs) {
continue; continue;
} }
let Some(&address) = addresses.get(ds_name) else {
continue;
};
let shape = ds.shape()?; let shape = ds.shape()?;
let is_unlimited = check_unlimited(file, group, ds_name); let is_unlimited = is_unlimited(&ds);
let size = if is_unlimited { let size = if is_unlimited {
unlimited_len(file, &attrs, &shape) unlimited_len(file, &attrs, &shape)
} else { } else {
shape.first().copied().unwrap_or(0) shape.first().copied().unwrap_or(0)
}; };
let dimid = get_dimid(&attrs);
let dim = Dimension { let dim = Dimension {
name: ds_name.clone(), name: ds_name.clone(),
size, size,
is_unlimited, is_unlimited,
}; };
found.push((get_dimid(&attrs), dim, address));
if let Some(id) = dimid {
seen_dimids.insert(id, dims.len());
}
dims.push(dim);
} }
// Sort by dimid if available, otherwise keep discovery order if found.is_empty() {
if !seen_dimids.is_empty() { // Fallback: infer dimensions from dataset shapes and names.
let mut pairs: Vec<(i64, Dimension)> = Vec::new(); // In NetCDF-4, coordinate variables are datasets whose name matches
let mut unordered = Vec::new(); // a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
for (i, dim) in dims.into_iter().enumerate() { let mut dims = Vec::new();
let id = seen_dimids for ds_name in &dataset_names {
.iter() let ds = group.dataset(ds_name)?;
.find(|(_, idx)| **idx == i) let shape = ds.shape()?;
.map(|(k, _)| *k); if shape.len() == 1 {
if let Some(id) = id { dims.push(Dimension {
pairs.push((id, dim)); name: ds_name.clone(),
} else { size: shape[0],
unordered.push(dim); is_unlimited: is_unlimited(&ds),
});
} }
} }
pairs.sort_by_key(|(id, _)| *id); return Ok(GroupDims {
dims = pairs.into_iter().map(|(_, d)| d).collect(); dims,
dims.extend(unordered); scales: Vec::new(),
});
} }
Ok(dims) // By dimid; scales without one keep their discovery order after those
// with one (the sort is stable).
found.sort_by_key(|(id, ..)| (id.is_none(), id.unwrap_or(0)));
let mut out = GroupDims::default();
for (i, (dimid, dim, address)) in found.into_iter().enumerate() {
out.dims.push(dim);
out.scales.push(Scale {
address,
dimid,
dim: i,
});
}
Ok(out)
} }
/// The start of the `NAME` attribute netCDF-C gives a dimension scale that /// The start of the `NAME` attribute netCDF-C gives a dimension scale that
@@ -107,10 +157,7 @@ const PURE_DIMENSION_NAME: &str = "This is a netCDF dimension but not a netCDF v
/// scale's own extent, as before. /// scale's own extent, as before.
fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 { fn unlimited_len(file: &clawhdf5::File, attrs: &HashMap<String, AttrValue>, shape: &[u64]) -> u64 {
let own = shape.first().copied().unwrap_or(0); let own = shape.first().copied().unwrap_or(0);
let is_variable = !matches!( let is_variable = !is_pure_dimension(attrs);
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
);
let Some(refs) = reference_list(file, attrs) else { let Some(refs) = reference_list(file, attrs) else {
return own; return own;
}; };
@@ -171,8 +218,59 @@ fn reference_list(
Some(addresses.into_iter().map(|r| r.address).zip(axes).collect()) Some(addresses.into_iter().map(|r| r.address).zip(axes).collect())
} }
/// The dimension scale attached to each axis of a variable, from its
/// `DIMENSION_LIST` attribute (HDF5 dimension scales: one variable-length
/// sequence of object references per axis) — the address of the scale
/// netCDF-C takes for the axis, or `None` for an axis with none. netCDF-C's
/// `dimscale_visitor` lets `H5DSiterate_scales` visit every scale attached
/// to the axis and keeps the last, so with several (h5py's `attach_scale`
/// twice) the last one is the axis's dimension. `None` overall when the
/// attribute is missing or not in that form.
pub(crate) fn dimension_list(
file: &clawhdf5::File,
attrs: &HashMap<String, AttrValue>,
) -> Option<Vec<Option<u64>>> {
use clawhdf5_format::data_read::read_object_references;
use clawhdf5_format::datatype::Datatype;
use clawhdf5_format::vl_data::VlResolver;
let Some(AttrValue::Raw { datatype, data, .. }) = attrs.get("DIMENSION_LIST") else {
return None;
};
let Datatype::VariableLength {
is_string: false,
base_type,
..
} = datatype
else {
return None;
};
let sb = file.superblock();
let base_size = usize::try_from(base_type.type_size()).ok()?;
let sequences = VlResolver::new_in(file.storage(), sb.offset_size, sb.length_size)
.sequences(data, base_size)
.ok()?;
sequences
.iter()
.map(|refs| {
let refs = read_object_references(refs, base_type, sb.offset_size).ok()?;
Some(refs.iter().rev().find(|r| !r.is_null()).map(|r| r.address))
})
.collect()
}
/// Whether a dataset is a dimension scale that is only a dimension, not a
/// netCDF variable: netCDF-C and h5netcdf give it this `NAME`, and netCDF-C
/// does not list it among the variables.
pub(crate) fn is_pure_dimension(attrs: &HashMap<String, AttrValue>) -> bool {
is_dimension_scale(attrs)
&& matches!(
attrs.get("NAME"),
Some(AttrValue::String(n)) if n.starts_with(PURE_DIMENSION_NAME)
)
}
/// Check if a dataset's attributes mark it as a dimension scale. /// Check if a dataset's attributes mark it as a dimension scale.
fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool { pub(crate) fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
if let Some(AttrValue::String(class)) = attrs.get("CLASS") { if let Some(AttrValue::String(class)) = attrs.get("CLASS") {
return class == "DIMENSION_SCALE"; return class == "DIMENSION_SCALE";
} }
@@ -180,7 +278,7 @@ fn is_dimension_scale(attrs: &HashMap<String, AttrValue>) -> bool {
} }
/// Get the _Netcdf4Dimid attribute value if present. /// Get the _Netcdf4Dimid attribute value if present.
fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> { pub(crate) fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
match attrs.get("_Netcdf4Dimid") { match attrs.get("_Netcdf4Dimid") {
Some(AttrValue::I64(id)) => Some(*id), Some(AttrValue::I64(id)) => Some(*id),
Some(AttrValue::U64(id)) => Some(*id as i64), Some(AttrValue::U64(id)) => Some(*id as i64),
@@ -188,56 +286,8 @@ fn get_dimid(attrs: &HashMap<String, AttrValue>) -> Option<i64> {
} }
} }
/// Check if a dimension is unlimited by inspecting the HDF5 dataspace max_dimensions. /// Whether a dataset's first axis is unlimited (`max_dimensions[0] ==
/// /// u64::MAX` in its dataspace).
/// A dimension is unlimited when `max_dimensions[0] == u64::MAX` in the HDF5 dataspace. fn is_unlimited(ds: &clawhdf5::Dataset<'_>) -> bool {
fn check_unlimited(_file: &clawhdf5::File, group: &clawhdf5::Group<'_>, ds_name: &str) -> bool { matches!(ds.max_dimensions(), Ok(Some(max_dims)) if max_dims.first() == Some(&u64::MAX))
let ds = match group.dataset(ds_name) {
Ok(ds) => ds,
Err(_) => return false,
};
match ds.max_dimensions() {
Ok(Some(max_dims)) => max_dims.first().copied() == Some(u64::MAX),
_ => false,
}
}
/// Extract dimensions from an HDF5 group using both dimension scale attributes
/// and variable DIMENSION_LIST references.
///
/// This is a more robust approach that also discovers dimensions from variables
/// that reference them, even when dimension scales aren't explicitly set.
pub(crate) fn extract_dimensions_from_datasets(
group: &clawhdf5::Group<'_>,
file: &clawhdf5::File,
) -> Result<Vec<Dimension>, Error> {
// First try the standard approach with DIMENSION_SCALE
let mut dims = extract_dimensions(file, group)?;
// If we found dimensions, return them
if !dims.is_empty() {
return Ok(dims);
}
// Fallback: infer dimensions from dataset shapes and names.
// In NetCDF-4, coordinate variables are datasets whose name matches
// a dimension name. If there are no explicit DIMENSION_SCALE attributes,
// we look for 1-D datasets that might be coordinate variables.
let dataset_names = group.datasets()?;
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
if shape.len() == 1 {
// This 1-D dataset could be a coordinate variable / dimension
let is_unlimited = check_unlimited(file, group, ds_name);
dims.push(Dimension {
name: ds_name.clone(),
size: shape[0],
is_unlimited,
});
}
}
Ok(dims)
} }
+23 -17
View File
@@ -9,12 +9,15 @@ use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension}; use crate::dimension::{self, Dimension};
use crate::error::Error; use crate::error::Error;
use crate::variable::{self, Variable}; use crate::scope::{self, Scope};
use crate::variable::Variable;
/// A NetCDF-4 group corresponding to an HDF5 group. /// A NetCDF-4 group corresponding to an HDF5 group.
pub struct NetCDF4Group<'f> { pub struct NetCDF4Group<'f> {
/// Group name. /// Group name.
name: String, name: String,
/// Path of the group from the root (`/`-separated).
path: String,
/// Underlying HDF5 file. /// Underlying HDF5 file.
file: &'f clawhdf5::File, file: &'f clawhdf5::File,
/// Underlying HDF5 group. /// Underlying HDF5 group.
@@ -25,11 +28,13 @@ impl<'f> NetCDF4Group<'f> {
/// Create a new NetCDF4Group from an HDF5 group. /// Create a new NetCDF4Group from an HDF5 group.
pub(crate) fn new( pub(crate) fn new(
name: String, name: String,
path: String,
file: &'f clawhdf5::File, file: &'f clawhdf5::File,
hdf5_group: clawhdf5::Group<'f>, hdf5_group: clawhdf5::Group<'f>,
) -> Self { ) -> Self {
Self { Self {
name, name,
path,
file, file,
hdf5_group, hdf5_group,
} }
@@ -40,27 +45,22 @@ impl<'f> NetCDF4Group<'f> {
&self.name &self.name
} }
/// List dimensions defined in this group. /// List dimensions defined in this group (not those of its parent
/// groups, which its variables can also use).
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> { pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
dimension::extract_dimensions_from_datasets(&self.hdf5_group, self.file) Ok(dimension::group_dims(self.file, &self.hdf5_group)?.dims)
} }
/// List variables in this group. /// List variables in this group: its datasets, except the dimension
/// scales that are only dimensions. Their dimensions can be defined in
/// this group or a parent group.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> { pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let dims = self.dimensions()?; Scope::new(self.file, &self.path)?.variables()
variable::build_variables(&self.hdf5_group, &dims)
} }
/// Get a specific variable by name. /// Get a specific variable by name.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> { pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let dims = self.dimensions()?; scope::variable_at(self.file, &self.path, name)
let ds = self
.hdf5_group
.dataset(name)
.map_err(|_| Error::VariableNotFound(name.to_string()))?;
let shape = ds.shape()?;
let var_dims = crate::variable::match_dimensions_to_variable(&shape, &dims);
Ok(Variable::new(name.to_string(), ds, var_dims))
} }
/// Read all attributes of this group. /// Read all attributes of this group.
@@ -79,12 +79,18 @@ impl<'f> NetCDF4Group<'f> {
.hdf5_group .hdf5_group
.group(name) .group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?; .map_err(|_| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(name.to_string(), self.file, hdf5_group)) Ok(NetCDF4Group::new(
name.to_string(),
format!("{}/{name}", self.path),
self.file,
hdf5_group,
))
} }
/// List dataset (variable) names in this group. /// The names of this group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> { pub fn variable_names(&self) -> Result<Vec<String>, Error> {
Ok(self.hdf5_group.datasets()?) Scope::new(self.file, &self.path)?.variable_names()
} }
} }
+19 -13
View File
@@ -27,6 +27,7 @@ pub mod cf;
pub mod dimension; pub mod dimension;
pub mod error; pub mod error;
pub mod group; pub mod group;
mod scope;
pub mod types; pub mod types;
pub mod variable; pub mod variable;
@@ -75,25 +76,25 @@ impl NetCDF4File {
/// List dimensions defined in the root group. /// List dimensions defined in the root group.
pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> { pub fn dimensions(&self) -> Result<Vec<Dimension>, Error> {
dimension::extract_dimensions_from_datasets(&self.hdf5.root(), &self.hdf5) Ok(dimension::group_dims(&self.hdf5, &self.hdf5.root())?.dims)
} }
/// List all variables in the root group. /// List all variables in the root group: its datasets, except the
/// dimension scales that are only dimensions (netCDF-C does not list
/// them either).
pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> { pub fn variables(&self) -> Result<Vec<Variable<'_>>, Error> {
let dims = self.dimensions()?; scope::Scope::new(&self.hdf5, "/")?.variables()
variable::build_variables(&self.hdf5.root(), &dims) }
/// The names of the root group's variables (see
/// [`variables`](Self::variables)).
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
scope::Scope::new(&self.hdf5, "/")?.variable_names()
} }
/// Get a specific variable by name from the root group. /// Get a specific variable by name from the root group.
pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> { pub fn variable(&self, name: &str) -> Result<Variable<'_>, Error> {
let dims = self.dimensions()?; scope::variable_at(&self.hdf5, "", name)
let ds = self
.hdf5
.dataset(name)
.map_err(|_| Error::VariableNotFound(name.to_string()))?;
let shape = ds.shape()?;
let var_dims = variable::match_dimensions_to_variable(&shape, &dims);
Ok(Variable::new(name.to_string(), ds, var_dims))
} }
/// Read all global (root group) attributes. /// Read all global (root group) attributes.
@@ -112,7 +113,12 @@ impl NetCDF4File {
.hdf5 .hdf5
.group(name) .group(name)
.map_err(|_| Error::GroupNotFound(name.to_string()))?; .map_err(|_| Error::GroupNotFound(name.to_string()))?;
Ok(NetCDF4Group::new(name.to_string(), &self.hdf5, hdf5_group)) Ok(NetCDF4Group::new(
name.to_string(),
name.to_string(),
&self.hdf5,
hdf5_group,
))
} }
/// Access the underlying HDF5 file for advanced operations. /// Access the underlying HDF5 file for advanced operations.
+231
View File
@@ -0,0 +1,231 @@
//! A group's variables and the dimensions they are defined on.
//!
//! netCDF-C (`libhdf5/hdf5open.c`) gives a variable its dimensions from the
//! file, never by size: the dimension ids in its `_Netcdf4Coordinates`
//! attribute (each dimension scale's `_Netcdf4Dimid`), else the dimension
//! scales its `DIMENSION_LIST` attribute references, looked up in the
//! variable's group and then each parent group up to the root. Only an axis
//! with neither (a file not written by a netCDF library) gets a dimension
//! by size. Dimension scales that are only dimensions are not variables, and
//! a variable stored as `_nc4_non_coord_<name>` (a variable sharing a
//! dimension's name without being its coordinate variable) is `<name>`.
use std::collections::{HashMap, HashSet};
use clawhdf5::AttrValue;
use crate::dimension::{self, Dimension, GroupDims};
use crate::error::Error;
use crate::variable::Variable;
/// The prefix netCDF-C gives the dataset of a variable that has a
/// dimension's name but is not that dimension's coordinate variable (the
/// dimension's scale holds the name).
const NON_COORD_PREFIX: &str = "_nc4_non_coord_";
/// A group, with the dimensions visible from it.
pub(crate) struct Scope<'f> {
file: &'f clawhdf5::File,
group: clawhdf5::Group<'f>,
/// This group's dimensions, then its parent's, and so on to the root's.
levels: Vec<GroupDims>,
}
impl<'f> Scope<'f> {
/// The group at `path` (`/`-separated from the root; `""` or `"/"` is
/// the root).
pub fn new(file: &'f clawhdf5::File, path: &str) -> Result<Self, Error> {
let parts: Vec<&str> = path.split('/').filter(|p| !p.is_empty()).collect();
let mut levels = Vec::with_capacity(parts.len() + 1);
for n in (0..=parts.len()).rev() {
let group = file.group(&parts[..n].join("/"))?;
levels.push(dimension::group_dims(file, &group)?);
}
let group = file.group(&parts.join("/"))?;
Ok(Self {
file,
group,
levels,
})
}
/// The group's datasets, as `(dataset name, object header address)` in
/// listing order.
fn datasets(&self) -> Result<Vec<(String, u64)>, Error> {
let datasets: HashSet<String> = self.group.datasets()?.into_iter().collect();
Ok(self
.group
.entries()?
.into_iter()
.filter(|(name, _)| datasets.contains(name))
.collect())
}
/// The group's variables: every dataset but the dimension scales that
/// are only dimensions.
pub fn variables(&self) -> Result<Vec<Variable<'f>>, Error> {
let mut variables = Vec::new();
for (ds_name, address) in self.datasets()? {
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
continue;
}
variables.push(self.variable_from(nc_name(&ds_name), address, ds, attrs)?);
}
Ok(variables)
}
/// The names of the group's variables.
pub fn variable_names(&self) -> Result<Vec<String>, Error> {
let mut names = Vec::new();
for (ds_name, address) in self.datasets()? {
let attrs = self.file.dataset_at(address)?.attrs()?;
if !dimension::is_pure_dimension(&attrs) {
names.push(nc_name(&ds_name));
}
}
Ok(names)
}
/// The variable called `name`: the dataset `_nc4_non_coord_<name>` if
/// there is one, else the dataset `<name>` unless it is only a
/// dimension.
pub fn variable(&self, name: &str) -> Result<Variable<'f>, Error> {
let not_found = || Error::VariableNotFound(name.to_string());
let datasets = self.datasets()?;
let prefixed = format!("{NON_COORD_PREFIX}{name}");
let address = datasets
.iter()
.find(|(n, _)| *n == prefixed)
.or_else(|| datasets.iter().find(|(n, _)| n == name))
.map(|&(_, address)| address)
.ok_or_else(not_found)?;
let ds = self.file.dataset_at(address)?;
let attrs = ds.attrs()?;
if dimension::is_pure_dimension(&attrs) {
return Err(not_found());
}
self.variable_from(nc_name(name), address, ds, attrs)
}
fn variable_from(
&self,
name: String,
address: u64,
ds: clawhdf5::Dataset<'f>,
attrs: HashMap<String, AttrValue>,
) -> Result<Variable<'f>, Error> {
let shape = ds.shape()?;
let dims = self.variable_dims(address, &attrs, &shape);
Ok(Variable::new(name, ds, dims, attrs))
}
/// The first dimension, searching this group and then its ancestors,
/// that `find` picks.
fn find<'a>(
&'a self,
find: impl Fn(&'a GroupDims) -> Option<&'a Dimension>,
) -> Option<Dimension> {
self.levels.iter().find_map(find).cloned()
}
/// The dimensions of the dataset at `address`, one per axis of `shape`,
/// as netCDF-C resolves them (see the module docs).
fn variable_dims(
&self,
address: u64,
attrs: &HashMap<String, AttrValue>,
shape: &[u64],
) -> Vec<Dimension> {
let rank = shape.len();
let mut dims: Vec<Option<Dimension>> = vec![None; rank];
if rank == 0 {
return Vec::new();
}
// A coordinate variable is the scale of its (first) dimension.
dims[0] = self.levels[0].by_address(address).cloned();
if let Some(ids) = coordinates(attrs).filter(|ids| ids.len() == rank) {
for (slot, id) in dims.iter_mut().zip(ids) {
if slot.is_none() {
*slot = self.find(|level| level.by_dimid(id));
}
}
}
if dims.iter().any(Option::is_none)
&& let Some(scales) =
dimension::dimension_list(self.file, attrs).filter(|s| s.len() == rank)
{
for (slot, scale) in dims.iter_mut().zip(scales) {
if slot.is_none()
&& let Some(scale) = scale
{
*slot = self.find(|level| level.by_address(scale));
}
}
}
// Neither: the first dimension of this group of the same size not
// already taken by another such axis, else an anonymous one.
let own = &self.levels[0].dims;
let mut used = vec![false; own.len()];
dims.into_iter()
.zip(shape)
.map(|(dim, &size)| {
dim.unwrap_or_else(|| {
match own
.iter()
.enumerate()
.find(|&(i, d)| !used[i] && d.size == size)
{
Some((i, d)) => {
used[i] = true;
d.clone()
}
None => Dimension {
name: format!("dim_{size}"),
size,
is_unlimited: false,
},
}
})
})
.collect()
}
}
/// The variable `name` of the group at `group_path`; `name` may itself be
/// a path (`"sub/var"`), relative to that group.
pub(crate) fn variable_at<'f>(
file: &'f clawhdf5::File,
group_path: &str,
name: &str,
) -> Result<Variable<'f>, Error> {
match name.trim_start_matches('/').rsplit_once('/') {
Some((dir, leaf)) => Scope::new(file, &format!("{group_path}/{dir}"))
.map_err(|_| Error::VariableNotFound(name.to_string()))?
.variable(leaf),
None => Scope::new(file, group_path)?.variable(name.trim_start_matches('/')),
}
}
/// The netCDF name of the dataset `ds_name`.
fn nc_name(ds_name: &str) -> String {
ds_name
.strip_prefix(NON_COORD_PREFIX)
.unwrap_or(ds_name)
.to_string()
}
/// A variable's `_Netcdf4Coordinates`: the `_Netcdf4Dimid` of the dimension
/// of each axis.
fn coordinates(attrs: &HashMap<String, AttrValue>) -> Option<Vec<i64>> {
match attrs.get("_Netcdf4Coordinates")? {
AttrValue::I64Array(ids) => Some(ids.clone()),
AttrValue::I64(id) => Some(vec![*id]),
AttrValue::U64Array(ids) => ids.iter().map(|&id| i64::try_from(id).ok()).collect(),
AttrValue::U64(id) => Some(vec![i64::try_from(*id).ok()?]),
_ => None,
}
}
+208 -81
View File
@@ -2,12 +2,18 @@
//! //!
//! Variables in NetCDF-4 are HDF5 datasets. This module wraps them with //! Variables in NetCDF-4 are HDF5 datasets. This module wraps them with
//! dimension associations and CF attribute support. //! dimension associations and CF attribute support.
//!
//! A variable along an unlimited dimension has that dimension's length in
//! netCDF, even when fewer records of it have been written (its HDF5 dataset
//! is shorter): [`Variable::shape`] is the netCDF shape and the reads return
//! that many values, the unwritten ones as the fill value, as netCDF-C does.
//! [`Variable::stored_shape`] is the dataset's extent.
use std::collections::HashMap; use std::collections::HashMap;
use clawhdf5::AttrValue; use clawhdf5::AttrValue;
use crate::cf::{self, CfAttributes}; use crate::cf::{self, CfAttributes, FillValue};
use crate::dimension::Dimension; use crate::dimension::Dimension;
use crate::error::Error; use crate::error::Error;
use crate::types::{NcType, dtype_to_nctype}; use crate::types::{NcType, dtype_to_nctype};
@@ -20,18 +26,23 @@ pub struct Variable<'f> {
dataset: clawhdf5::Dataset<'f>, dataset: clawhdf5::Dataset<'f>,
/// Dimensions associated with this variable. /// Dimensions associated with this variable.
dims: Vec<Dimension>, dims: Vec<Dimension>,
/// Cached attributes. /// The dataset's attributes.
attrs_cache: Option<HashMap<String, AttrValue>>, attrs: HashMap<String, AttrValue>,
} }
impl<'f> Variable<'f> { impl<'f> Variable<'f> {
/// Create a new Variable wrapping an HDF5 dataset. /// Create a new Variable wrapping an HDF5 dataset.
pub(crate) fn new(name: String, dataset: clawhdf5::Dataset<'f>, dims: Vec<Dimension>) -> Self { pub(crate) fn new(
name: String,
dataset: clawhdf5::Dataset<'f>,
dims: Vec<Dimension>,
attrs: HashMap<String, AttrValue>,
) -> Self {
Self { Self {
name, name,
dataset, dataset,
dims, dims,
attrs_cache: None, attrs,
} }
} }
@@ -40,13 +51,27 @@ impl<'f> Variable<'f> {
&self.name &self.name
} }
/// The dimensions of this variable. /// The dimensions of this variable, one per axis: the ones the file
/// gives it (`_Netcdf4Coordinates`, else `DIMENSION_LIST`), found in its
/// group or a parent group. An axis the file gives no dimension (a file
/// not written by a netCDF library) gets the first dimension of the
/// variable's group of the same size, else an anonymous `dim_<size>`.
pub fn dimensions(&self) -> &[Dimension] { pub fn dimensions(&self) -> &[Dimension] {
&self.dims &self.dims
} }
/// The shape of this variable (dimension sizes). /// The shape of this variable as netCDF reports it: along an unlimited
/// dimension, the dimension's current length (the longest variable on
/// it), even if fewer records of this variable have been written;
/// otherwise the dataset's extent. The reads return this many values.
pub fn shape(&self) -> Result<Vec<u64>, Error> { pub fn shape(&self) -> Result<Vec<u64>, Error> {
Ok(nc_shape(&self.dataset.shape()?, &self.dims))
}
/// The extent of the HDF5 dataset: what has been written. It differs
/// from [`shape`](Self::shape) only along an unlimited dimension that
/// another variable has more records of.
pub fn stored_shape(&self) -> Result<Vec<u64>, Error> {
Ok(self.dataset.shape()?) Ok(self.dataset.shape()?)
} }
@@ -58,59 +83,88 @@ impl<'f> Variable<'f> {
/// Read all attributes as a HashMap. /// Read all attributes as a HashMap.
pub fn attrs(&mut self) -> Result<&HashMap<String, AttrValue>, Error> { pub fn attrs(&mut self) -> Result<&HashMap<String, AttrValue>, Error> {
if self.attrs_cache.is_none() { Ok(&self.attrs)
self.attrs_cache = Some(self.dataset.attrs()?);
}
Ok(self
.attrs_cache
.as_ref()
.expect("invariant: attrs_cache is Some after initialization"))
} }
/// Extract CF convention attributes. /// Extract CF convention attributes.
pub fn cf_attributes(&mut self) -> Result<CfAttributes, Error> { pub fn cf_attributes(&mut self) -> Result<CfAttributes, Error> {
let attrs = self.attrs()?; Ok(cf::extract_cf_attributes(&self.attrs))
Ok(cf::extract_cf_attributes(attrs))
} }
/// Read data as f64 with scale_factor/add_offset applied. /// Read data as f64 with scale_factor/add_offset applied.
/// ///
/// Missing values (matching `_FillValue` or `missing_value`) become NaN. /// Missing values (matching `_FillValue` or `missing_value`) become NaN.
/// If no scale_factor or add_offset attributes exist, returns the raw f64 data. /// If no scale_factor or add_offset attributes exist, returns the raw f64 data.
/// Records along an unlimited dimension that this variable has not
/// written (see [`shape`](Self::shape)) are NaN.
pub fn read_f64(&mut self) -> Result<Vec<f64>, Error> { pub fn read_f64(&mut self) -> Result<Vec<f64>, Error> {
let raw = self.dataset.read_f64()?; let raw = self.dataset.read_f64()?;
let cf = self.cf_attributes()?; let cf = cf::extract_cf_attributes(&self.attrs);
Ok(cf::apply_scale_offset(&raw, &cf)) self.padded(cf::apply_scale_offset(&raw, &cf), || Ok(f64::NAN))
} }
/// Read raw data as f64 without any scale/offset transformation. /// Read raw data as f64 without any scale/offset transformation.
///
/// Unwritten records along an unlimited dimension read as the fill
/// value (`_FillValue`, else netCDF's default for the type), as in the
/// other `read_raw_*` methods and [`read_string`](Self::read_string).
pub fn read_raw_f64(&self) -> Result<Vec<f64>, Error> { pub fn read_raw_f64(&self) -> Result<Vec<f64>, Error> {
Ok(self.dataset.read_f64()?) self.padded_read(self.dataset.read_f64()?)
} }
/// Read raw data as f32 without any scale/offset transformation. /// Read raw data as f32 without any scale/offset transformation.
pub fn read_raw_f32(&self) -> Result<Vec<f32>, Error> { pub fn read_raw_f32(&self) -> Result<Vec<f32>, Error> {
Ok(self.dataset.read_f32()?) self.padded_read(self.dataset.read_f32()?)
} }
/// Read raw data as i32 without any scale/offset transformation. /// Read raw data as i32 without any scale/offset transformation.
pub fn read_raw_i32(&self) -> Result<Vec<i32>, Error> { pub fn read_raw_i32(&self) -> Result<Vec<i32>, Error> {
Ok(self.dataset.read_i32()?) self.padded_read(self.dataset.read_i32()?)
} }
/// Read raw data as i64 without any scale/offset transformation. /// Read raw data as i64 without any scale/offset transformation.
pub fn read_raw_i64(&self) -> Result<Vec<i64>, Error> { pub fn read_raw_i64(&self) -> Result<Vec<i64>, Error> {
Ok(self.dataset.read_i64()?) self.padded_read(self.dataset.read_i64()?)
} }
/// Read raw data as u64 without any scale/offset transformation. /// Read raw data as u64 without any scale/offset transformation.
pub fn read_raw_u64(&self) -> Result<Vec<u64>, Error> { pub fn read_raw_u64(&self) -> Result<Vec<u64>, Error> {
Ok(self.dataset.read_u64()?) self.padded_read(self.dataset.read_u64()?)
} }
/// Read raw data as strings. /// Read raw data as strings.
pub fn read_string(&self) -> Result<Vec<String>, Error> { pub fn read_string(&self) -> Result<Vec<String>, Error> {
Ok(self.dataset.read_string()?) self.padded_read(self.dataset.read_string()?)
}
/// The fill value netCDF-C gives the variable's unwritten values: its
/// `_FillValue`, else the default fill value of its type (`NC_FILL_*`).
fn fill_value(&self) -> Result<FillValue, Error> {
if let Some(fill) = cf::extract_cf_attributes(&self.attrs).fill_value {
return Ok(fill);
}
Ok(default_fill(self.nc_type()?))
}
/// `data`, read in the dataset's extent, laid out in the variable's
/// netCDF shape with the fill value in the positions not written.
fn padded_read<T: Clone + FromFill>(&self, data: Vec<T>) -> Result<Vec<T>, Error> {
self.padded(data, || Ok(T::from_fill(&self.fill_value()?)))
}
/// Like [`padded_read`](Self::padded_read), padding with what `fill`
/// returns (called only when there is something to pad).
fn padded<T: Clone>(
&self,
data: Vec<T>,
fill: impl FnOnce() -> Result<T, Error>,
) -> Result<Vec<T>, Error> {
let extent = self.dataset.shape()?;
let shape = nc_shape(&extent, &self.dims);
if shape == extent {
return Ok(data);
}
pad(data, &extent, &shape, fill()?)
} }
/// Read raw bytes without any type conversion. /// Read raw bytes without any type conversion.
@@ -122,23 +176,23 @@ impl<'f> Variable<'f> {
let dtype = self.dataset.dtype()?; let dtype = self.dataset.dtype()?;
match dtype { match dtype {
clawhdf5::DType::F64 => { clawhdf5::DType::F64 => {
let vals = self.dataset.read_f64()?; let vals = self.read_raw_f64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::F32 => { clawhdf5::DType::F32 => {
let vals = self.dataset.read_f32()?; let vals = self.read_raw_f32()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::I32 => { clawhdf5::DType::I32 => {
let vals = self.dataset.read_i32()?; let vals = self.read_raw_i32()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::I64 => { clawhdf5::DType::I64 => {
let vals = self.dataset.read_i64()?; let vals = self.read_raw_i64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
clawhdf5::DType::U64 => { clawhdf5::DType::U64 => {
let vals = self.dataset.read_u64()?; let vals = self.read_raw_u64()?;
Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect()) Ok(vals.iter().flat_map(|v| v.to_le_bytes()).collect())
} }
other => { other => {
@@ -169,64 +223,137 @@ impl std::fmt::Debug for Variable<'_> {
} }
} }
/// Build variables from a group's datasets and associated dimensions. /// The netCDF shape of a variable whose dataset has `extent`: along an
pub(crate) fn build_variables<'f>( /// unlimited dimension the dimension's length, which is at least the extent.
group: &clawhdf5::Group<'f>, fn nc_shape(extent: &[u64], dims: &[Dimension]) -> Vec<u64> {
available_dims: &[Dimension], if dims.len() != extent.len() {
) -> Result<Vec<Variable<'f>>, Error> { return extent.to_vec();
let dataset_names = group.datasets()?;
let mut variables = Vec::new();
for ds_name in &dataset_names {
let ds = group.dataset(ds_name)?;
let shape = ds.shape()?;
// Associate dimensions with this variable.
// First try DIMENSION_LIST attribute, then fall back to shape matching.
let var_dims = match_dimensions_to_variable(&shape, available_dims);
variables.push(Variable::new(ds_name.clone(), ds, var_dims));
} }
extent
Ok(variables) .iter()
.zip(dims)
.map(|(&e, d)| if d.is_unlimited { e.max(d.size) } else { e })
.collect()
} }
/// Match dimensions to a variable based on shape. /// `data`, row-major in `extent`, placed in a row-major array of `shape`
/// /// (as many axes, each at least as long) filled with `fill`.
/// For each axis of the variable, find a dimension with matching size. fn pad<T: Clone>(data: Vec<T>, extent: &[u64], shape: &[u64], fill: T) -> Result<Vec<T>, Error> {
/// If multiple dimensions have the same size, prefer exact name matching let too_big = || Error::TypeError(format!("variable of shape {shape:?} is too large"));
/// from the convention order. let to_usize = |dims: &[u64]| -> Result<Vec<usize>, Error> {
pub(crate) fn match_dimensions_to_variable( dims.iter()
shape: &[u64], .map(|&d| usize::try_from(d).map_err(|_| too_big()))
available_dims: &[Dimension], .collect()
) -> Vec<Dimension> { };
let mut result = Vec::with_capacity(shape.len()); let (extent, shape) = (to_usize(extent)?, to_usize(shape)?);
let total = shape
// Track which dimensions have been used to avoid duplicates .iter()
let mut used = vec![false; available_dims.len()]; .try_fold(1usize, |n, &d| n.checked_mul(d))
.ok_or_else(too_big)?;
for &dim_size in shape { if extent.len() != shape.len()
let mut matched = false; || extent.iter().zip(&shape).any(|(e, s)| e > s)
|| extent.iter().product::<usize>() != data.len()
// Find a dimension with matching size that hasn't been used yet {
for (i, dim) in available_dims.iter().enumerate() { return Err(Error::TypeError(format!(
if !used[i] && dim.size == dim_size { "{} values of extent {extent:?} do not fit shape {shape:?}",
result.push(dim.clone()); data.len()
used[i] = true; )));
matched = true; }
let (Some((&row, outer)), Some(&row_stride)) = (extent.split_last(), shape.last()) else {
return Ok(data);
};
let mut out = vec![fill; total];
if row == 0 {
return Ok(out);
}
// The position of the current row along each outer axis.
let mut index = vec![0usize; outer.len()];
for chunk in data.chunks_exact(row) {
let offset = index.iter().zip(&shape).fold(0, |o, (&i, &n)| o * n + i);
out[offset * row_stride..][..row].clone_from_slice(chunk);
for (i, &n) in index.iter_mut().zip(outer).rev() {
*i += 1;
if *i < n {
break; break;
} }
} *i = 0;
if !matched {
// Create an anonymous dimension for unmatched sizes
result.push(Dimension {
name: format!("dim_{dim_size}"),
size: dim_size,
is_unlimited: false,
});
} }
} }
Ok(out)
}
result /// netCDF's default fill value for a type (`NC_FILL_*` in `netcdf.h`).
fn default_fill(nc_type: NcType) -> FillValue {
match nc_type {
NcType::Byte => FillValue::Int(-127),
NcType::UByte => FillValue::UInt(255),
NcType::Short => FillValue::Int(-32767),
NcType::UShort => FillValue::UInt(65535),
NcType::Int => FillValue::Int(-2_147_483_647),
NcType::UInt => FillValue::UInt(4_294_967_295),
NcType::Int64 => FillValue::Int(-9_223_372_036_854_775_806),
NcType::UInt64 => FillValue::UInt(18_446_744_073_709_551_614),
NcType::Float => FillValue::Float(f64::from(9.969_21e36_f32)),
NcType::Double => FillValue::Float(9.969_209_968_386_869e36),
NcType::String => FillValue::String(String::new()),
NcType::Char => FillValue::Int(0),
}
}
/// A fill value converted to the element type of a read, as the read
/// converts the stored values.
trait FromFill {
fn from_fill(fill: &FillValue) -> Self;
}
macro_rules! numeric_from_fill {
($($t:ty),*) => {$(
impl FromFill for $t {
fn from_fill(fill: &FillValue) -> Self {
match fill {
FillValue::Float(v) => *v as $t,
FillValue::Int(v) => *v as $t,
FillValue::UInt(v) => *v as $t,
FillValue::String(_) => <$t>::default(),
}
}
}
)*};
}
numeric_from_fill!(f64, f32, i32, i64, u64);
impl FromFill for String {
fn from_fill(fill: &FillValue) -> Self {
match fill {
FillValue::String(s) => s.clone(),
_ => String::new(),
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn pad_places_rows() {
// (2, 1) written of (2, 4): each row padded, not the tail.
let out = pad(vec![1, 2], &[2, 1], &[2, 4], 0).unwrap();
assert_eq!(out, vec![1, 0, 0, 0, 2, 0, 0, 0]);
// Leading axis short.
let out = pad(vec![1, 2, 3, 4], &[2, 2], &[3, 2], -1).unwrap();
assert_eq!(out, vec![1, 2, 3, 4, -1, -1]);
// Nothing written.
let out = pad(Vec::<i32>::new(), &[0, 3], &[2, 3], 7).unwrap();
assert_eq!(out, vec![7; 6]);
// 3-D, middle axis short.
let out = pad(vec![1, 2, 3, 4], &[2, 1, 2], &[2, 2, 2], 0).unwrap();
assert_eq!(out, vec![1, 2, 0, 0, 3, 4, 0, 0]);
}
#[test]
fn pad_rejects_wrong_length() {
assert!(pad(vec![1, 2, 3], &[2, 2], &[3, 2], 0).is_err());
assert!(pad(vec![1, 2, 3, 4], &[2, 2], &[1, 4], 0).is_err());
}
} }
+382 -1
View File
@@ -4,7 +4,7 @@
use std::process::Command; use std::process::Command;
use clawhdf5_netcdf4::{AttrValue, NetCDF4File}; use clawhdf5_netcdf4::{AttrValue, NcType, NetCDF4File};
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Helpers // Helpers
@@ -461,3 +461,384 @@ with nc.Dataset({path:?}) as f:
} }
assert_eq!(got, expected); assert_eq!(got, expected);
} }
// ===========================================================================
// Variables' dimensions, shapes and values as netCDF4-python reports them
// ===========================================================================
/// Whether python can import `module`.
fn python_has(module: &str) -> bool {
Command::new(python())
.args(["-c", &format!("import {module}")])
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
/// h5netcdf is not in every interop environment (CI installs it; a local
/// `.venv` may not have it), so its tests skip without it even under
/// `CLAWHDF5_REQUIRE_INTEROP=1`.
macro_rules! skip_if_no_h5netcdf {
() => {
if !python_has("h5netcdf") {
eprintln!("SKIP: python3 with h5netcdf not available");
return;
}
};
}
/// Every variable of the file at `path`, in every group, as netCDF4-python
/// reports it: `"<group path> <name> (<dims>) (<shape>)"` and its values
/// (numeric variables; element by element with masking off, so unwritten
/// records are the fill value), sorted by the description.
///
/// Values are read one element at a time because netCDF-C 4.9.3 lays out a
/// whole-variable read of a variable shorter than an unlimited dimension
/// that is not its first wrongly (the written values first, then the fill);
/// element reads, and reads of one index of the leading axis, are right.
fn netcdf4_view(path: &std::path::Path) -> Vec<(String, Vec<f64>)> {
let script = r#"
import sys
import numpy as np
import netCDF4 as nc
def walk(g):
for name, v in g.variables.items():
v.set_auto_mask(False)
head = "%s %s (%s) (%s)" % (g.path, name, ",".join(v.dimensions), ",".join(map(str, v.shape)))
vals = []
if v.dtype != str and v.dtype.kind in "iuf":
vals = [repr(float(v[i])) for i in np.ndindex(v.shape)]
print(head + "|" + " ".join(vals))
for sub in g.groups.values():
walk(sub)
with nc.Dataset(sys.argv[1]) as f:
walk(f)
"#;
let out = Command::new(python())
.args(["-c", script, &path.display().to_string()])
.output()
.expect("failed to run python3");
assert!(
out.status.success(),
"{}",
String::from_utf8_lossy(&out.stderr)
);
let mut view: Vec<(String, Vec<f64>)> = String::from_utf8(out.stdout)
.unwrap()
.lines()
.map(|line| {
let (head, vals) = line.split_once('|').unwrap();
let vals = vals
.split_whitespace()
.map(|v| v.parse().unwrap())
.collect();
(head.to_string(), vals)
})
.collect();
view.sort_by(|a, b| a.0.cmp(&b.0));
view
}
/// The same view of the file through clawhdf5-netcdf4.
fn clawhdf5_view(path: &std::path::Path) -> Vec<(String, Vec<f64>)> {
fn describe(
group_path: &str,
vars: Vec<clawhdf5_netcdf4::Variable<'_>>,
) -> Vec<(String, Vec<f64>)> {
vars.into_iter()
.map(|v| {
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
let shape: Vec<String> = v.shape().unwrap().iter().map(u64::to_string).collect();
let head = format!(
"{group_path} {} ({}) ({})",
v.name(),
dims.join(","),
shape.join(",")
);
let vals = match v.nc_type().unwrap() {
NcType::String | NcType::Char => Vec::new(),
_ => v.read_raw_f64().unwrap(),
};
(head, vals)
})
.collect()
}
fn walk(
group_path: &str,
group: &clawhdf5_netcdf4::NetCDF4Group<'_>,
out: &mut Vec<(String, Vec<f64>)>,
) {
out.extend(describe(group_path, group.variables().unwrap()));
for name in group.group_names().unwrap() {
walk(
&format!("{group_path}/{name}"),
&group.group(&name).unwrap(),
out,
);
}
}
let file = NetCDF4File::open(path).unwrap();
let mut view = describe("/", file.variables().unwrap());
for name in file.group_names().unwrap() {
walk(&format!("/{name}"), &file.group(&name).unwrap(), &mut view);
}
view.sort_by(|a, b| a.0.cmp(&b.0));
view
}
/// clawhdf5-netcdf4 reports the same variables, dimensions, shapes and
/// values (bit for bit, NaN equal to NaN) as netCDF4-python.
fn assert_same_view(path: &std::path::Path) {
let want = netcdf4_view(path);
let got = clawhdf5_view(path);
let heads = |v: &[(String, Vec<f64>)]| v.iter().map(|(h, _)| h.clone()).collect::<Vec<_>>();
assert_eq!(heads(&got), heads(&want), "variables differ from netCDF4's");
for ((head, got), (_, want)) in got.iter().zip(&want) {
let same = got.len() == want.len()
&& got
.iter()
.zip(want)
.all(|(a, b)| a.to_bits() == b.to_bits() || (a.is_nan() && b.is_nan()));
assert!(same, "{head}: got {got:?}, netCDF4 reads {want:?}");
}
}
/// The reproducer of the known-issues entry: `a` is on the unlimited `time`
/// (5 long through `b`) with 2 records, not on an anonymous `dim_2`; the
/// pure dimension scales `time` and `empty` are not variables; `a` has
/// shape (5,) and reads its 3 unwritten records as the fill value.
#[test]
fn variable_dimensions_come_from_the_file() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("repro.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("time", None)
f.createDimension("empty", None)
f.createDimension("x", 3)
f.createVariable("a", "i4", ("time",))[0:2] = [1, 2]
f.createVariable("b", "f4", ("time", "x"))[0:5, :] = np.arange(15).reshape(5, 3)
f.createVariable("e", "i4", ("empty",))
f.createVariable("c", "i4", ("x",))[:] = [7, 8, 9]
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let mut names = file.variable_names().unwrap();
names.sort();
assert_eq!(names, ["a", "b", "c", "e"]);
assert!(matches!(
file.variable("time"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
let a = file.variable("a").unwrap();
assert_eq!(a.dimensions()[0].name, "time");
assert_eq!(a.shape().unwrap(), [5]);
assert_eq!(a.stored_shape().unwrap(), [2]);
assert_eq!(
a.read_raw_i32().unwrap(),
[1, 2, -2_147_483_647, -2_147_483_647, -2_147_483_647]
);
}
/// Dimensions of one size are told apart by the file, not by order: `p`
/// and `q` are both 2 long, and `v(q, p)`, `same(p, p)` (one dimension
/// twice), a scalar, `q`'s coordinate variable, a variable called `p` that
/// is not `p`'s coordinate variable (stored as `_nc4_non_coord_p`), and
/// variables in a subgroup and a sub-subgroup on dimensions of their
/// ancestors.
#[test]
fn equal_size_and_inherited_dimensions_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("dims.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("p", 2)
f.createDimension("q", 2)
f.createVariable("v", "i4", ("q", "p"))[:] = np.array([[1, 2], [3, 4]])
f.createVariable("same", "i4", ("p", "p"))[:] = np.array([[5, 6], [7, 8]])
f.createVariable("s", "f8", ())[...] = 3.5
f.createVariable("q", "f4", ("q",))[:] = [0, 1]
f.createVariable("p", "f4", ("q", "p"))[:] = np.array([[0, 1], [2, 3]])
g = f.createGroup("g")
g.createDimension("r", 2)
g.createVariable("w", "i4", ("r", "q", "p"))[:] = np.arange(8).reshape(2, 2, 2)
h = g.createGroup("h")
h.createVariable("z", "i4", ("p", "r"))[:] = np.array([[1, 2], [3, 4]])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let v = file.variable("v").unwrap();
let dims: Vec<&str> = v.dimensions().iter().map(|d| d.name.as_str()).collect();
assert_eq!(dims, ["q", "p"]);
let p = file.variable("p").unwrap();
assert!(!p.is_coordinate());
assert!(file.variable("q").unwrap().is_coordinate());
let s = file.variable("s").unwrap();
assert!(s.dimensions().is_empty());
assert_eq!(s.shape().unwrap(), Vec::<u64>::new());
let z = file
.group("g")
.unwrap()
.group("h")
.unwrap()
.variable("z")
.unwrap();
let dims: Vec<&str> = z.dimensions().iter().map(|d| d.name.as_str()).collect();
assert_eq!(dims, ["p", "r"]);
}
/// Variables shorter than their unlimited dimension have its length and
/// read the fill value (`_FillValue`, else netCDF's default for the type)
/// where nothing was written — also when the unlimited dimension is not
/// the first; `read_f64` gives NaN there.
#[test]
fn unwritten_records_read_as_fill_like_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("pad.nc");
run_python(&format!(
r#"
import netCDF4 as nc
import numpy as np
with nc.Dataset({path:?}, "w") as f:
f.createDimension("t", None)
f.createDimension("x", 2)
f.createVariable("a", "i4", ("t",))[0:2] = [1, 2]
f.createVariable("f", "f4", ("x", "t"), fill_value=-5.0)[:, 0:1] = np.array([[1], [2]])
f.createVariable("d", "f8", ("t",))[0:4] = [1, 2, 3, 4]
f.createVariable("u", "u8", ("t",))[0:1] = [1]
f.createVariable("b", "i1", ("t", "x"))[0:3, :] = np.ones((3, 2))
f.createVariable("st", str, ("t",))[0] = "hi"
g = f.createGroup("g")
g.createVariable("k", "f4", ("t",))[0:1] = [9]
"#,
path = path.display().to_string()
));
assert_same_view(&path);
let file = NetCDF4File::open(&path).unwrap();
let mut f = file.variable("f").unwrap();
assert_eq!(f.shape().unwrap(), [2, 4]);
assert_eq!(f.stored_shape().unwrap(), [2, 1]);
assert_eq!(
f.read_raw_f32().unwrap(),
[1.0, -5.0, -5.0, -5.0, 2.0, -5.0, -5.0, -5.0]
);
let read = f.read_f64().unwrap();
assert_eq!(read[0], 1.0);
assert!(read[1].is_nan() && read[7].is_nan());
let st = file.variable("st").unwrap();
assert_eq!(st.read_string().unwrap(), ["hi", "", "", ""]);
assert_eq!(st.shape().unwrap(), [4]);
}
/// A file with HDF5 dimension scales but none of netCDF's own attributes
/// (h5py's `dims` API): the dimensions come from `DIMENSION_LIST`, so
/// `v(q, p)` is not `v(p, q)` although both are 2 long; with two scales
/// attached to one axis (`w`), netCDF-C takes the last.
#[test]
fn h5py_dimension_scales_match_netcdf4_python() {
skip_if_no_netcdf4!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("scales.h5");
run_python(&format!(
r#"
import h5py
import numpy as np
with h5py.File({path:?}, "w") as f:
f["p"] = np.arange(2.0)
f["q"] = np.arange(2.0) + 10
f["p"].make_scale("p")
f["q"].make_scale("q")
f["v"] = np.arange(4).reshape(2, 2)
f["v"].dims[0].attach_scale(f["q"])
f["v"].dims[1].attach_scale(f["p"])
f["w"] = np.arange(2)
f["w"].dims[0].attach_scale(f["p"])
f["w"].dims[0].attach_scale(f["q"])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
/// Files h5netcdf writes (its own implementation of the netCDF-4
/// conventions over h5py): an unlimited dimension, equal sizes, a subgroup
/// on inherited dimensions, a scalar.
#[test]
fn h5netcdf_file_matches_netcdf4_python() {
skip_if_no_netcdf4!();
skip_if_no_h5netcdf!();
let dir = tempfile::tempdir().unwrap();
let path = dir.path().join("h5netcdf.nc");
run_python(&format!(
r#"
import h5netcdf
import numpy as np
with h5netcdf.File({path:?}, "w") as f:
f.dimensions = {{"p": 2, "q": 2, "t": None}}
f.create_variable("v", ("q", "p"), "i4")[...] = np.array([[1, 2], [3, 4]])
f.create_variable("q", ("q",), "f4")[...] = [0, 1]
f.create_variable("same", ("p", "p"), "i4")[...] = np.array([[5, 6], [7, 8]])
a = f.create_variable("a", ("t", "p"), "f8")
f.resize_dimension("t", 3)
a[...] = np.ones((3, 2))
f.create_variable("short", ("t",), "i4")
g = f.create_group("g")
g.dimensions = {{"r": 2}}
g.create_variable("w", ("r", "q", "p"), "i4")[...] = np.arange(8).reshape(2, 2, 2)
g.create_variable("s", (), "f8")[...] = 2.5
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
/// Files xarray writes, through netCDF4 and (when installed) h5netcdf:
/// coordinates, two dimensions of one size, an unlimited dimension.
#[test]
fn xarray_files_match_netcdf4_python() {
skip_if_no_netcdf4!();
skip_if_no_xarray!();
let dir = tempfile::tempdir().unwrap();
let mut engines = vec!["netcdf4"];
if python_has("h5netcdf") {
engines.push("h5netcdf");
} else {
eprintln!("SKIP: xarray with engine h5netcdf (h5netcdf not available)");
}
for engine in engines {
let path = dir.path().join(format!("xarray_{engine}.nc"));
run_python(&format!(
r#"
import numpy as np
import xarray as xr
ds = xr.Dataset(
{{
"temp": (("time", "lat", "lon"), np.arange(12.0).reshape(3, 2, 2)),
"grid": (("lon", "lat"), np.array([[1, 2], [3, 4]], dtype="i4")),
"scalar": ((), 1.5),
}},
coords={{"time": [0.0, 6.0, 12.0], "lat": [10.0, 20.0], "lon": [5.0, 6.0]}},
)
ds.to_netcdf({path:?}, engine={engine:?}, unlimited_dims=["time"])
"#,
path = path.display().to_string()
));
assert_same_view(&path);
}
}
@@ -754,3 +754,53 @@ fn test_dimension_struct_equality() {
}; };
assert_ne!(d1, d3); assert_ne!(d1, d3);
} }
/// A dimension scale that is only a dimension (netCDF-C's `NAME`) is not a
/// variable, and `_nc4_non_coord_<name>` is the variable `<name>`, found in
/// place of the scale of the same name.
#[test]
fn test_pure_dimensions_hidden_and_non_coord_names() {
let pure = "This is a netCDF dimension but not a netCDF variable. 2";
let mut b = FileBuilder::new();
b.create_dataset("x")
.with_f32_data(&[0.0, 0.0])
.with_shape(&[2])
.set_attr("CLASS", AttrValue::String("DIMENSION_SCALE".into()))
.set_attr("NAME", AttrValue::String(pure.into()))
.set_attr("_Netcdf4Dimid", AttrValue::I64(0));
b.create_dataset("_nc4_non_coord_x")
.with_f64_data(&[1.0, 2.0, 3.0])
.with_shape(&[3]);
b.create_dataset("v")
.with_f64_data(&[5.0, 6.0])
.with_shape(&[2]);
let file = NetCDF4File::from_bytes(b.finish().unwrap()).unwrap();
let dims = file.dimensions().unwrap();
assert_eq!(dims.len(), 1);
assert_eq!(dims[0].name, "x");
let mut names = file.variable_names().unwrap();
names.sort();
assert_eq!(names, ["v", "x"]);
let x = file.variable("x").unwrap();
assert_eq!(x.name(), "x");
assert_eq!(x.read_raw_f64().unwrap(), [1.0, 2.0, 3.0]);
assert!(!x.is_coordinate());
// No DIMENSION_LIST: `v` gets `x` by size, as before.
assert_eq!(file.variable("v").unwrap().dimensions()[0].name, "x");
}
/// `variable` still takes a path relative to the group, as it did when it
/// opened the dataset by path.
#[test]
fn test_variable_by_path() {
let file = NetCDF4File::from_bytes(make_grouped_netcdf4()).unwrap();
let pressure = file.variable("surface/pressure").unwrap();
assert_eq!(pressure.name(), "pressure");
assert_eq!(pressure.read_raw_f64().unwrap(), [1013.25, 1012.0, 1011.5]);
assert!(file.variable("/time").is_ok());
assert!(matches!(
file.variable("nowhere/pressure"),
Err(clawhdf5_netcdf4::Error::VariableNotFound(_))
));
}
+736
View File
@@ -0,0 +1,736 @@
//! Files written with `FileBuilder::libver_bounds(LibVer::V18, LibVer::V18)`
//! must be readable by HDF5 1.8. Every writer feature is written under that
//! bound and read back by HDF5 1.8.23's h5dump (values dumped in binary and
//! compared), by h5py (libhdf5 2.x), by h5dump 1.14, by clawhdf5 and by
//! `h5rs check --data`; then `FileEditor` grows, appends to and annotates
//! the file (splitting version-1 B-tree nodes) and every reader checks it
//! again. The version-1 B-trees we write are compared node by node with
//! the ones libhdf5 writes for the same data under h5py's
//! `libver=('v108', 'latest')`.
//!
//! HDF5 1.8 is found through `CLAWHDF5_H5DUMP18` (the path to its h5dump)
//! or at `~/.cache/hdf5-1.8.23/bin/h5dump`, where
//! `scripts/build-hdf5-1.8.sh` builds it; without it the 1.8 checks are
//! skipped (CI has no HDF5 1.8), even with `CLAWHDF5_REQUIRE_INTEROP=1`.
//! h5py/numpy (`CLAWHDF5_PYTHON`) and h5dump are needed otherwise; they
//! skip when missing unless `CLAWHDF5_REQUIRE_INTEROP=1`.
use std::path::{Path, PathBuf};
use std::process::{Command, Output};
use clawhdf5::{AttrValue, File, FileBuilder, FileEditor, LibVer, Selection};
use clawhdf5_format::datatype::{CharacterSet, Datatype, StringPadding};
use clawhdf5_format::file_writer::{CompoundTypeBuilder, EnumTypeBuilder};
use clawhdf5_format::type_builders::make_i32_type;
fn python() -> String {
std::env::var("CLAWHDF5_PYTHON").unwrap_or_else(|_| "python3".to_string())
}
fn interop_required() -> bool {
std::env::var("CLAWHDF5_REQUIRE_INTEROP").is_ok_and(|v| v == "1")
}
fn available(cmd: &str, args: &[&str]) -> bool {
Command::new(cmd)
.args(args)
.output()
.map(|o| o.status.success())
.unwrap_or(false)
}
fn tools_ok() -> bool {
let ok =
available(&python(), &["-c", "import h5py, numpy"]) && available("h5dump", &["--version"]);
if !ok {
assert!(
!interop_required(),
"CLAWHDF5_REQUIRE_INTEROP=1 but h5py/numpy or h5dump is not available"
);
eprintln!("SKIP: h5py/numpy or h5dump not available");
}
ok
}
/// HDF5 1.8's h5dump, when there is one.
fn h5dump18() -> Option<PathBuf> {
let p = match std::env::var_os("CLAWHDF5_H5DUMP18") {
Some(p) => PathBuf::from(p),
None => PathBuf::from(std::env::var_os("HOME")?).join(".cache/hdf5-1.8.23/bin/h5dump"),
};
let o = Command::new(&p).arg("--version").output().ok()?;
let v = String::from_utf8_lossy(&o.stdout).to_string();
if !v.contains("1.8.") {
eprintln!("SKIP 1.8 checks: {} is not HDF5 1.8 ({v})", p.display());
return None;
}
Some(p)
}
fn py(script: &str) -> String {
let o = Command::new(python())
.args(["-c", script])
.output()
.expect("run python");
assert!(
o.status.success(),
"python failed:\n{script}\nSTDOUT: {}\nSTDERR: {}",
String::from_utf8_lossy(&o.stdout),
String::from_utf8_lossy(&o.stderr)
);
String::from_utf8_lossy(&o.stdout).trim().to_string()
}
fn text(o: &Output) -> String {
format!(
"{}{}",
String::from_utf8_lossy(&o.stdout),
String::from_utf8_lossy(&o.stderr)
)
}
fn tmpdir() -> tempfile::TempDir {
tempfile::TempDir::new_in(env!("CARGO_TARGET_TMPDIR")).unwrap()
}
/// The hyperslab of `count` elements from `start`.
fn block(start: &[u64], count: &[u64]) -> Selection {
Selection::Hyperslab {
start: start.to_vec(),
stride: vec![1; start.len()],
count: count.to_vec(),
block: vec![1; start.len()],
}
}
fn le<T: Copy, const N: usize>(v: &[T], f: impl Fn(T) -> [u8; N]) -> Vec<u8> {
v.iter().flat_map(|&x| f(x)).collect()
}
/// The datasets of the test file and the bytes each holds (little-endian,
/// row-major, as h5py's `tobytes()` and h5dump's `-b LE` give them).
struct Expect {
datasets: Vec<(String, Vec<u8>)>,
}
impl Expect {
fn set(&mut self, name: &str, bytes: Vec<u8>) {
match self.datasets.iter_mut().find(|(n, _)| n == name) {
Some(e) => e.1 = bytes,
None => self.datasets.push((name.to_string(), bytes)),
}
}
}
const MANY: usize = 100_000;
/// Write every feature under the 1.8 bound.
fn write_file(path: &Path) -> Expect {
let mut e = Expect {
datasets: Vec::new(),
};
let mut b = FileBuilder::new();
b.libver_bounds(LibVer::V18, LibVer::V18);
// Contiguous, with dense attributes (more than 8).
let v: Vec<f64> = (0..1000).map(|i| i as f64 * 0.5).collect();
let d = b.create_dataset("contig").with_f64_data(&v);
for i in 0..12 {
d.set_attr(&format!("a{i:02}"), AttrValue::I64(i));
}
e.set("/contig", le(&v, f64::to_le_bytes));
b.create_dataset("empty").with_f64_data(&[]);
e.set("/empty", vec![]);
b.create_dataset("scalar")
.with_f64_data(&[2.5])
.with_shape(&[]);
e.set("/scalar", 2.5f64.to_le_bytes().to_vec());
let v: Vec<i32> = (0..16).map(|i| i * 3 - 7).collect();
b.create_dataset("compact").with_i32_data(&v).compact();
e.set("/compact", le(&v, i32::to_le_bytes));
let v: Vec<f32> = (0..20).map(|i| i as f32 / 3.0).collect();
b.create_dataset("f32").with_f32_data(&v);
e.set("/f32", le(&v, f32::to_le_bytes));
let v: Vec<f32> = vec![0.5, -2.0, 1024.0, 0.0];
b.create_dataset("f16").with_f16_data(&v);
e.set(
"/f16",
le(&v, |x| {
clawhdf5_format::float16::f32_to_f16_bits(x).to_le_bytes()
}),
);
// Chunked with every built-in filter HDF5 1.8 has, and edge chunks.
let v: Vec<f32> = (0..60 * 70).map(|i| (i % 97) as f32 * 1.25).collect();
b.create_dataset("chunked")
.with_f32_data(&v)
.with_shape(&[60, 70])
.with_chunks(&[16, 16])
.with_deflate(6)
.with_shuffle()
.with_fletcher32();
e.set("/chunked", le(&v, f32::to_le_bytes));
// What the 1.10 indexes would be: Extensible Array, version-2 B-tree,
// Fixed Array, single chunk. All become version-1 B-trees.
let v: Vec<i32> = (0..25).collect();
b.create_dataset("resizable")
.with_i32_data(&v)
.with_maxshape(&[u64::MAX])
.with_chunks(&[4]);
e.set("/resizable", le(&v, i32::to_le_bytes));
let v: Vec<i64> = (0..30).map(|i| i * 1_000_000_007).collect();
b.create_dataset("resizable2")
.with_i64_data(&v)
.with_shape(&[5, 6])
.with_maxshape(&[u64::MAX, u64::MAX])
.with_chunks(&[2, 4]);
e.set("/resizable2", le(&v, i64::to_le_bytes));
let v: Vec<i32> = (0..10).map(|i| -i).collect();
b.create_dataset("fixedmax")
.with_i32_data(&v)
.with_maxshape(&[100])
.with_chunks(&[3])
.with_deflate(1);
e.set("/fixedmax", le(&v, i32::to_le_bytes));
let v: Vec<i32> = (0..8).map(|i| i * i).collect();
b.create_dataset("single")
.with_i32_data(&v)
.with_chunks(&[8]);
e.set("/single", le(&v, i32::to_le_bytes));
// Enough chunks for a three-level tree, and a 2-D two-level one.
let v: Vec<i32> = (0..MANY as i32).map(|i| i ^ 0x5a5a).collect();
b.create_dataset("many")
.with_i32_data(&v)
.with_maxshape(&[u64::MAX])
.with_chunks(&[1]);
e.set("/many", le(&v, i32::to_le_bytes));
let v: Vec<f64> = (0..100 * 100).map(|i| i as f64).collect();
b.create_dataset("grid")
.with_f64_data(&v)
.with_shape(&[100, 100])
.with_chunks(&[1, 1]);
e.set("/grid", le(&v, f64::to_le_bytes));
b.create_dataset("empty_chunked")
.with_f64_data(&[])
.with_maxshape(&[u64::MAX])
.with_chunks(&[10]);
e.set("/empty_chunked", vec![]);
let v: Vec<i32> = (0..12).collect();
b.create_dataset("filled")
.with_i32_data(&v)
.with_maxshape(&[u64::MAX])
.with_chunks(&[5])
.with_fill_value(&(-9i32).to_le_bytes());
e.set("/filled", le(&v, i32::to_le_bytes));
// Datatypes.
let raw = b"abcdehello\0\0\0\0\0".to_vec();
b.create_dataset("strings").with_compound_data(
Datatype::String {
size: 5,
padding: StringPadding::NullPad,
charset: CharacterSet::Ascii,
},
raw.clone(),
3,
);
e.set("/strings", raw);
let ct = CompoundTypeBuilder::new()
.i32_field("a")
.f64_field("b")
.build();
let mut raw = Vec::new();
for i in 0..5i32 {
raw.extend_from_slice(&i.to_le_bytes());
raw.extend_from_slice(&(f64::from(i) * 1.5).to_le_bytes());
}
b.create_dataset("compound")
.with_compound_data(ct, raw.clone(), 5);
e.set("/compound", raw);
let et = EnumTypeBuilder::i32_based()
.value("RED", 0)
.value("GREEN", 1)
.value("BLUE", 7)
.build();
let v = [0, 7, 1, 1, 0];
b.create_dataset("enum").with_enum_i32_data(et, &v);
e.set("/enum", le(&v, i32::to_le_bytes));
let v: Vec<i32> = (0..24).collect();
b.create_dataset("array").with_array_data(
make_i32_type(),
&[2, 3],
le(&v, i32::to_le_bytes),
4,
);
e.set("/array", le(&v, i32::to_le_bytes));
let v: Vec<i64> = vec![-1, 0, i64::MAX];
b.create_dataset("i64").with_i64_data(&v);
e.set("/i64", le(&v, i64::to_le_bytes));
// Groups: compact with links of every kind, dense (links and
// attributes), and tracking creation order.
let mut g = b.create_group("g_compact");
g.create_dataset("x").with_i32_data(&[1, 2, 3]);
e.set("/g_compact/x", le(&[1i32, 2, 3], i32::to_le_bytes));
g.add_soft_link("soft", "/contig");
g.add_external_link("ext", "other.h5", "/data");
g.set_attr("title", AttrValue::String("compact group".into()));
b.add_group(g.finish());
b.add_hard_link("hard", "/g_compact/x");
let mut g = b.create_group("g_dense");
for i in 0..20 {
let v = [i, i + 1];
g.create_dataset(&format!("d{i:02}")).with_i32_data(&v);
e.set(&format!("/g_dense/d{i:02}"), le(&v, i32::to_le_bytes));
}
for i in 0..12 {
g.set_attr(&format!("attr{i:02}"), AttrValue::F64(f64::from(i) / 4.0));
}
b.add_group(g.finish());
let mut g = b.create_group("g_order");
g.track_order(true);
for i in 0..10 {
let n = format!("z{}", 9 - i);
g.create_dataset(&n).with_i32_data(&[i]);
e.set(&format!("/g_order/{n}"), i.to_le_bytes().to_vec());
}
for i in 0..10 {
g.set_attr(&format!("b{}", 9 - i), AttrValue::I64(i));
}
b.add_group(g.finish());
b.create_dataset("deep/er/path").with_i32_data(&[42]);
e.set("/deep/er/path", 42i32.to_le_bytes().to_vec());
b.set_attr("version", AttrValue::F64(1.8));
b.set_attr("ints", AttrValue::I64Array(vec![1, -2, 3]));
b.set_attr("text", AttrValue::String("readable by 1.8".into()));
b.set_attr(
"texts",
AttrValue::StringArray(vec!["a".into(), "bb".into(), "ccc".into()]),
);
b.set_attr("big", AttrValue::U64(u64::MAX));
b.write(path).unwrap();
e
}
/// Structural facts of a file: superblock and layout message versions.
fn check_versions(path: &Path) {
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
let bytes = std::fs::read(path).unwrap();
assert_eq!(bytes[8], 2, "superblock version");
let f = File::open(path).unwrap();
let sb = f.superblock().clone();
for name in [
"contig",
"compact",
"chunked",
"resizable",
"resizable2",
"many",
] {
let addr = clawhdf5_format::group_v2::resolve_path_any(&bytes, &sb, name).unwrap();
let oh = ObjectHeader::parse(&bytes, addr as usize, 8, 8).unwrap();
let layout = oh
.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
.unwrap();
assert_eq!(layout.data[0], 3, "layout version of {name}");
}
}
/// clawhdf5 reads every dataset's bytes back.
fn check_ours(path: &Path, e: &Expect) {
let f = File::open(path).unwrap();
for (name, want) in &e.datasets {
let got = f
.dataset(name)
.unwrap()
.read_selection(&Selection::All)
.unwrap();
assert!(got == *want, "our read of {name} differs");
}
let attrs = f.dataset("contig").unwrap().attrs().unwrap();
assert!((12..=13).contains(&attrs.len()), "{attrs:?}");
}
/// h5py (libhdf5 2.x) reads every dataset's bytes back.
fn check_h5py(path: &Path, e: &Expect, dir: &Path) {
let mut script = format!(
"import h5py, numpy as np\nf = h5py.File({p:?}, 'r')\n",
p = path.to_str().unwrap()
);
for (i, (name, want)) in e.datasets.iter().enumerate() {
let exp = dir.join(format!("expect{i}.bin"));
std::fs::write(&exp, want).unwrap();
script.push_str(&format!(
"a = np.ascontiguousarray(f[{name:?}][()]).tobytes()\n\
assert a == open({x:?}, 'rb').read(), {name:?}\n",
x = exp.to_str().unwrap()
));
}
script.push_str(
"assert f.attrs['text'] in (b'readable by 1.8', 'readable by 1.8')\n\
assert list(f.attrs['ints']) == [1, -2, 3]\n\
assert len(f['g_dense'].attrs) == 12 and len(f['g_dense']) == 20\n\
assert list(f['g_order']) == ['z9', 'z8', 'z7', 'z6', 'z5', 'z4', 'z3', 'z2', 'z1', 'z0']\n\
assert f['g_compact/soft'].shape == (1000,)\n\
assert f['resizable'].maxshape == (None,)\n\
assert f['filled'].fillvalue == -9\n\
print('ok')\n",
);
assert_eq!(py(&script), "ok");
}
/// h5dump (1.14) and `h5rs check --data` accept the file.
fn check_tools(path: &Path) {
let p = path.to_str().unwrap();
let o = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["check", "--data", "-q", p])
.output()
.unwrap();
assert!(o.status.success(), "h5rs check --data {p}:\n{}", text(&o));
let o = Command::new("h5dump")
.args(["-o", "/dev/null", p])
.output()
.unwrap();
assert!(
o.status.success() && o.stderr.is_empty(),
"h5dump {p}:\n{}",
text(&o)
);
}
/// HDF5 1.8's h5dump reads the whole file exactly as h5dump 1.14 does, and
/// dumps each numeric dataset's values as the bytes we wrote.
fn check_18(h5dump: &Path, path: &Path, e: &Expect, dir: &Path) {
let p = path.to_str().unwrap();
let o = Command::new(h5dump).arg(p).output().unwrap();
assert!(
o.status.success() && o.stderr.is_empty(),
"h5dump 1.8 {p}:\n{}",
text(&o)
);
let out18 = String::from_utf8_lossy(&o.stdout).to_string();
assert!(out18.contains("EXTERNAL_LINK \"ext\""), "{out18}");
let o = Command::new("h5dump").arg(p).output().unwrap();
assert!(o.status.success(), "h5dump {p}:\n{}", text(&o));
// HDF5 1.8 has no name for IEEE half floats.
let out = String::from_utf8_lossy(&o.stdout).replace(
"H5T_IEEE_F16LE",
"16-bit little-endian floating-point 16-bit precision",
);
if let Some((n, (a, b))) = out18
.lines()
.zip(out.lines())
.enumerate()
.find(|(_, (a, b))| a != b)
{
panic!("h5dump 1.8 and h5dump differ at line {}:\n{a}\n{b}", n + 1);
}
assert_eq!(out18.lines().count(), out.lines().count());
// `-b` writes nothing for compounds, enums, arrays, strings and half
// floats (HDF5 1.8
// has no native half float), for h5py's files too: those are covered
// by the text above.
let skip = ["/f16", "/strings", "/compound", "/enum", "/array"];
for (i, (name, want)) in e.datasets.iter().enumerate() {
if want.is_empty() || skip.contains(&name.as_str()) {
continue;
}
let bin = dir.join(format!("dump18_{i}.bin"));
let o = Command::new(h5dump)
.args(["-d", name, "-b", "LE", "-o"])
.arg(&bin)
.arg(p)
.output()
.unwrap();
assert!(o.status.success(), "h5dump 1.8 -d {name}:\n{}", text(&o));
let got = std::fs::read(&bin).unwrap();
assert!(got == *want, "HDF5 1.8 read of {name} differs");
}
}
fn check_all(path: &Path, e: &Expect, dir: &Path, h5dump: Option<&Path>) {
check_ours(path, e);
check_h5py(path, e, dir);
check_tools(path);
if let Some(h) = h5dump {
check_18(h, path, e, dir);
}
}
#[test]
fn every_feature_reads_in_hdf5_1_8() {
if !tools_ok() {
return;
}
let h18 = h5dump18();
let dir = tmpdir();
let path = dir.path().join("v18.h5");
let mut e = write_file(&path);
check_versions(&path);
check_all(&path, &e, dir.path(), h18.as_deref());
// FileEditor: grow the unlimited datasets (splitting B-tree nodes),
// overwrite values in filtered and unfiltered chunks, set attributes
// (compact and dense).
let mut ed = FileEditor::open(&path).unwrap();
let mut res: Vec<i32> = (0..25).collect();
for round in 0..300usize {
let n = res.len() as u64;
let add = 1 + (round % 9) as u64;
ed.resize("resizable", &[n + add]).unwrap();
let vals: Vec<i32> = (0..add).map(|k| (n + k) as i32 * 7 - 3).collect();
ed.write_values("resizable", &block(&[n], &[add]), &vals)
.unwrap();
res.extend(&vals);
}
e.set("/resizable", le(&res, i32::to_le_bytes));
let mut many: Vec<i32> = (0..MANY as i32).map(|i| i ^ 0x5a5a).collect();
ed.resize("many", &[MANY as u64 + 500]).unwrap();
let vals: Vec<i32> = (0..500).collect();
ed.write_values("many", &block(&[MANY as u64], &[500]), &vals)
.unwrap();
many.extend(&vals);
many[12_345] = -1;
ed.write_values("many", &block(&[12_345], &[1]), &[-1i32])
.unwrap();
e.set("/many", le(&many, i32::to_le_bytes));
let mut chunked: Vec<f32> = (0..60 * 70).map(|i| (i % 97) as f32 * 1.25).collect();
let row: Vec<f32> = (0..70).map(|i| -(i as f32)).collect();
ed.write_values("chunked", &block(&[31, 0], &[1, 70]), &row)
.unwrap();
chunked[31 * 70..32 * 70].copy_from_slice(&row);
e.set("/chunked", le(&chunked, f32::to_le_bytes));
ed.resize("resizable2", &[9, 6]).unwrap();
let mut r2: Vec<i64> = (0..30).map(|i| i * 1_000_000_007).collect();
r2.extend(std::iter::repeat_n(0, 24));
e.set("/resizable2", le(&r2, i64::to_le_bytes));
ed.set_attr("contig", "added", &AttrValue::F64(3.25))
.unwrap();
ed.set_attr("g_dense", "attr03", &AttrValue::F64(-1.0))
.unwrap();
ed.set_attr("/", "text", &AttrValue::String("edited".into()))
.unwrap();
drop(ed);
check_ours(&path, &e);
let f = File::open(&path).unwrap();
assert!(matches!(
f.dataset("contig").unwrap().attr("added").unwrap(),
Some(AttrValue::F64(v)) if v == 3.25
));
drop(f);
// Everything but the root's "text" attribute is as check_h5py expects.
py(&format!(
"import h5py\nf = h5py.File({p:?}, 'r')\n\
assert f.attrs['text'] in (b'edited', 'edited')\n\
assert f['contig'].attrs['added'] == 3.25\n\
assert f['g_dense'].attrs['attr03'] == -1.0\n",
p = path.to_str().unwrap()
));
let o = Command::new(env!("CARGO_BIN_EXE_h5rs"))
.args(["check", "--data", "-q", path.to_str().unwrap()])
.output()
.unwrap();
assert!(o.status.success(), "h5rs check after edits:\n{}", text(&o));
if let Some(h) = h18.as_deref() {
check_18(h, &path, &e, dir.path());
}
// h5py (libhdf5 2.x) goes on appending to the 1.8-format file, and 1.8
// still reads it.
py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'r+') as f:\n\
\x20 d = f['resizable']\n\
\x20 n = d.shape[0]\n\
\x20 d.resize((n + 100,))\n\
\x20 d[n:] = np.arange(100, dtype='<i4') + 5000\n",
p = path.to_str().unwrap()
));
res.extend((0..100).map(|k| 5000 + k));
e.set("/resizable", le(&res, i32::to_le_bytes));
check_ours(&path, &e);
if let Some(h) = h18.as_deref() {
check_18(h, &path, &e, dir.path());
}
}
// ---- The trees against libhdf5's ----
/// One node of a version-1 chunk B-tree: level, and per key its stored
/// size and offsets; children below.
#[derive(Debug, PartialEq)]
struct TreeNode {
level: u8,
keys: Vec<(u32, Vec<u64>)>,
children: Vec<TreeNode>,
}
fn read_tree(bytes: &[u8], addr: u64, ndims: usize, sizes: bool) -> TreeNode {
let a = addr as usize;
assert_eq!(&bytes[a..a + 4], b"TREE");
let level = bytes[a + 5];
let n = u16::from_le_bytes([bytes[a + 6], bytes[a + 7]]) as usize;
let mut p = a + 24;
let mut keys = Vec::new();
let mut kids = Vec::new();
for i in 0..=n {
let size = u32::from_le_bytes(bytes[p..p + 4].try_into().unwrap());
let offs = (0..ndims)
.map(|d| u64::from_le_bytes(bytes[p + 8 + 8 * d..p + 16 + 8 * d].try_into().unwrap()))
.collect();
keys.push((if sizes { size } else { 0 }, offs));
p += 8 + 8 * ndims;
if i < n {
kids.push(u64::from_le_bytes(bytes[p..p + 8].try_into().unwrap()));
p += 8;
}
}
let children = if level > 0 {
kids.iter()
.map(|&c| read_tree(bytes, c, ndims, sizes))
.collect()
} else {
Vec::new()
};
TreeNode {
level,
keys,
children,
}
}
/// Where two trees first differ (libhdf5's first).
fn first_difference(a: &TreeNode, b: &TreeNode, at: &str) -> Option<String> {
if a.level != b.level || a.keys.len() != b.keys.len() {
return Some(format!(
"{at}: level {} with {} keys vs level {} with {} keys",
a.level,
a.keys.len(),
b.level,
b.keys.len()
));
}
if let Some(i) = (0..a.keys.len()).find(|&i| a.keys[i] != b.keys[i]) {
return Some(format!("{at} key {i}: {:?} vs {:?}", a.keys[i], b.keys[i]));
}
a.children
.iter()
.zip(&b.children)
.enumerate()
.find_map(|(i, (x, y))| first_difference(x, y, &format!("{at}/{i}")))
}
/// The chunk B-tree of dataset `name`: (tree, ndims) from its layout.
fn tree_of(path: &Path, name: &str, sizes: bool) -> TreeNode {
use clawhdf5_format::message_type::MessageType;
use clawhdf5_format::object_header::ObjectHeader;
let bytes = std::fs::read(path).unwrap();
let f = File::open(path).unwrap();
let sb = f.superblock().clone();
let addr = clawhdf5_format::group_v2::resolve_path_any(&bytes, &sb, name).unwrap();
let oh = ObjectHeader::parse(&bytes, addr as usize, 8, 8).unwrap();
let l = &oh
.messages
.iter()
.find(|m| m.msg_type == MessageType::DataLayout)
.unwrap()
.data;
assert_eq!((l[0], l[1]), (3, 2), "{name}: layout v3, chunked");
let ndims = l[2] as usize;
let root = u64::from_le_bytes(l[3..11].try_into().unwrap());
read_tree(&bytes, root, ndims, sizes)
}
/// Our version-1 chunk B-trees are libhdf5's, node for node (levels, child
/// counts, every key's offsets, and its chunk size where the chunks are
/// the same bytes), for 1-D, 2-D and 3-D datasets with two- and three-level
/// trees, filtered or not.
///
/// libhdf5 inserts each chunk into the tree when it leaves its chunk cache.
/// A whole-dataset write with no cache (`rdcc_nbytes=0`, or chunks larger
/// than the cache) inserts them in row-major order, as we build the tree;
/// with the default cache small chunks of a 1-D dataset still arrive in
/// order, but those of a multi-dimensional one arrive in the order the
/// cache's hash evicts them, which gives the same keys in differently
/// filled nodes. Both trees index the same chunks; we do not model the
/// cache.
#[test]
fn chunk_btrees_match_libhdf5() {
if !tools_ok() {
return;
}
let dir = tmpdir();
let theirs = dir.path().join("libhdf5.h5");
py(&format!(
"import h5py, numpy as np\n\
with h5py.File({p:?}, 'w', libver=('v108', 'latest'), rdcc_nbytes=0) as f:\n\
\x20 f.create_dataset('d1000', data=np.arange(10000.0), chunks=(10,))\n\
\x20 f.create_dataset('d999', data=np.arange(9990.0), chunks=(10,))\n\
\x20 f.create_dataset('big', data=np.arange(100000, dtype='<i4'), chunks=(1,), maxshape=(None,))\n\
\x20 f.create_dataset('grid', data=np.arange(10000.0).reshape(100, 100), chunks=(1, 1))\n\
\x20 f.create_dataset('cube', data=np.arange(27000, dtype='<i2').reshape(30, 30, 30), chunks=(2, 3, 5))\n\
\x20 f.create_dataset('gz', data=np.arange(20000, dtype='<i8') % 13, chunks=(7,), compression='gzip')\n",
p = theirs.to_str().unwrap()
));
let ours = dir.path().join("ours.h5");
let mut b = FileBuilder::new();
b.libver_bounds(LibVer::V18, LibVer::V18);
let v: Vec<f64> = (0..10000).map(f64::from).collect();
b.create_dataset("d1000")
.with_f64_data(&v)
.with_chunks(&[10]);
b.create_dataset("d999")
.with_f64_data(&v[..9990])
.with_chunks(&[10]);
let v: Vec<i32> = (0..100_000).collect();
b.create_dataset("big")
.with_i32_data(&v)
.with_chunks(&[1])
.with_maxshape(&[u64::MAX]);
let v: Vec<f64> = (0..10000).map(f64::from).collect();
b.create_dataset("grid")
.with_f64_data(&v)
.with_shape(&[100, 100])
.with_chunks(&[1, 1]);
let raw: Vec<u8> = (0..27000i16).flat_map(|x| x.to_le_bytes()).collect();
b.create_dataset("cube")
.with_compound_data(
clawhdf5_format::datatype::Datatype::FixedPoint {
size: 2,
byte_order: clawhdf5_format::datatype::DatatypeByteOrder::LittleEndian,
signed: true,
bit_offset: 0,
bit_precision: 16,
},
raw,
27000,
)
.with_shape(&[30, 30, 30])
.with_chunks(&[2, 3, 5]);
let v: Vec<i64> = (0..20000).map(|i| i % 13).collect();
b.create_dataset("gz")
.with_i64_data(&v)
.with_chunks(&[7])
.with_deflate(4);
b.write(&ours).unwrap();
for (name, sizes) in [
("d1000", true),
("d999", true),
("big", true),
("grid", true),
("cube", true),
// Compressed sizes differ between zlib-rs and zlib.
("gz", false),
] {
let a = tree_of(&theirs, name, sizes);
let b = tree_of(&ours, name, sizes);
if let Some(d) = first_difference(&a, &b, "root") {
panic!("{name}: our chunk B-tree differs from libhdf5's: {d}");
}
}
}
+1
View File
@@ -68,6 +68,7 @@ pub use writer::{DatasetSpec, create_datasets_parallel};
// Re-export useful types from clawhdf5-format for advanced users // Re-export useful types from clawhdf5-format for advanced users
pub use clawhdf5_format::data_layout::VdsMapping; pub use clawhdf5_format::data_layout::VdsMapping;
pub use clawhdf5_format::dict_encoding::{DictEncoded, DictionaryEncoder}; pub use clawhdf5_format::dict_encoding::{DictEncoded, DictionaryEncoder};
pub use clawhdf5_format::libver::LibVer;
pub use clawhdf5_format::property_list::{ pub use clawhdf5_format::property_list::{
DatasetCreateProps, FileAccessProps, FileCreateProps, lib_version, DatasetCreateProps, FileAccessProps, FileCreateProps, lib_version,
}; };
+29
View File
@@ -1,6 +1,7 @@
//! Writing API: FileBuilder and GroupBuilder for creating HDF5 files. //! Writing API: FileBuilder and GroupBuilder for creating HDF5 files.
use clawhdf5_format::file_writer::FileWriter as FormatWriter; use clawhdf5_format::file_writer::FileWriter as FormatWriter;
use clawhdf5_format::libver::LibVer;
use clawhdf5_format::type_builders::{ use clawhdf5_format::type_builders::{
AttrValue, DatasetBuilder as FormatDatasetBuilder, FinishedGroup, AttrValue, DatasetBuilder as FormatDatasetBuilder, FinishedGroup,
GroupBuilder as FormatGroupBuilder, GroupBuilder as FormatGroupBuilder,
@@ -98,6 +99,34 @@ impl FileBuilder {
self self
} }
/// Set the library version bounds, as h5py's `libver=(low, high)`: the
/// oldest HDF5 release whose file format is used, and the newest whose
/// features are allowed. The default is `(LibVer::V110,
/// LibVer::Latest)`, the HDF5 1.10 format clawhdf5 has always written.
///
/// `libver_bounds(LibVer::V18, LibVer::V18)` writes a file HDF5 1.8 can
/// read (version-1 B-tree chunk indexes, version-2 superblock, the low
/// bound libhdf5 2.0 uses by default), or fails with
/// [`Error::Format`] (`FormatError::LibverBound`) for anything 1.8
/// cannot read: virtual datasets, the 1.12 reference types, native
/// complex numbers. See `clawhdf5_format::libver` for the details.
///
/// ```
/// use clawhdf5::{FileBuilder, LibVer};
///
/// let mut b = FileBuilder::new();
/// b.libver_bounds(LibVer::V18, LibVer::V18);
/// b.create_dataset("x")
/// .with_f64_data(&[1.0, 2.0, 3.0])
/// .with_maxshape(&[u64::MAX]);
/// let bytes = b.finish().unwrap();
/// assert_eq!(bytes[8], 2); // superblock version 2
/// ```
pub fn libver_bounds(&mut self, low: LibVer, high: LibVer) -> &mut Self {
self.writer.libver_bounds(low, high);
self
}
/// Set an attribute on the root group. /// Set an attribute on the root group.
pub fn set_attr(&mut self, name: &str, value: AttrValue) { pub fn set_attr(&mut self, name: &str, value: AttrValue) {
self.writer.set_root_attr(name, value); self.writer.set_root_attr(name, value);
+104 -24
View File
@@ -15,9 +15,6 @@ Checked against `main` at `9b5803f` on 2026-09-28.
| Issue | Kind | Since | | Issue | Kind | Since |
|---|---|---| |---|---|---|
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 | | [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 | | [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 | | [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
@@ -200,8 +197,16 @@ wrong data.
`filter_registry::register_filter`. `filter_registry::register_filter`.
- **Writer:** in dense storage (more than 8 attributes on an object, or - **Writer:** in dense storage (more than 8 attributes on an object, or
more than 8 links in a group) one attribute or link message over 65 515 more than 8 links in a group) one attribute or link message over 65 515
bytes is an error (no huge fractal-heap objects). The writer does not bytes is an error (no huge fractal-heap objects). Files HDF5 1.8 reads
produce output that HDF5 1.8 can read. are opt-in (`libver_bounds(LibVer::V18, LibVer::V18)`, since 2026-09-28,
[history](#hdf5-18-could-not-read-the-files-we-wrote)): by default the
writer uses the HDF5 1.10 format. The pre-1.8 format (version-0
superblock, symbol-table groups; h5py's `libver='earliest'`) cannot be
written, and the Python bindings' `'w'` mode has no `libver` argument.
Under the 1.8 bound, chunk B-trees equal libhdf5's node for node when
libhdf5 inserts the chunks in row-major order; libhdf5 with its chunk
cache inserts the small chunks of a multi-dimensional dataset in eviction
order, which fills the nodes differently (same chunks, same keys).
- **Checks we deliberately do not make:** - **Checks we deliberately do not make:**
- a float sign bit position outside the type, and a size-0 string type: - a float sign bit position outside the type, and a size-0 string type:
clawhdf5 up to v2.7.0 wrote them; clawhdf5 up to v2.7.0 wrote them;
@@ -413,25 +418,6 @@ always expose it. A fix belongs in `conformance/ref_bugs.py` (more or more
varied reads for this object) or in documenting the file as a known varied reads for this object) or in documenting the file as a known
refusal; neither is done. refusal; neither is done.
## NetCDF-4: variables' dimensions are guessed from sizes
**Status:** open (found 2026-09-28 while fixing unlimited dimension sizes).
`clawhdf5-netcdf4` gives each variable the dimensions it finds by size
(`match_dimensions_to_variable`: the first unused dimension of equal size,
else an anonymous `dim_<n>`), not the ones its `DIMENSION_LIST` names, and
`variables()` also lists the dimension scales that are only dimensions
(netCDF-C hides them). With netCDF4-python: unlimited `time` and `empty`,
`x` (3), `a(time)` with 2 records, `b(time, x)` with 5, `e(empty)` — netCDF4
reports `time` = 5, `a` on `time`, and variables `a`, `b`, `e` and a `c(x)`;
clawhdf5-netcdf4 reports `time` = 5 (right) but `a` on `dim_2`, and also
variables `time` and `empty` (the scales, put on `empty`). Two dimensions of
one size can be swapped the same way. Related: netCDF4 gives `a` the shape
(5,) (a variable along an unlimited dimension has the dimension's length,
unwritten records read as fill); `Variable::shape` is the HDF5 extent,
`[2]`, and reads return those 2 values. `dimensions()` is right.
Workaround: read the variable's `_Netcdf4Coordinates` attribute (dimension
ids, matching each scale's `_Netcdf4Dimid`).
## Small floats decode as libhdf5 does, not as the OCP MX specification ## Small floats decode as libhdf5 does, not as the OCP MX specification
**Status:** open, deliberate (documented 2026-09-28). HDF5 2.x predefines **Status:** open, deliberate (documented 2026-09-28). HDF5 2.x predefines
@@ -484,6 +470,100 @@ Newest first. "Before any release" means no tagged release (v2.7.0 and
earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date earlier) contains the bug. Full detail is in `CHANGELOG.md` under the date
given. given.
## NetCDF-4: variables' dimensions are guessed from sizes
**Status:** fixed 2026-09-28 (branch `fix/netcdf-dimension-list`). Affected
every release (v2.1.0 to v2.7.0: size matching dates from the crate's
first version). Wrong metadata only: stored values were always read right.
Users who read `_Netcdf4Coordinates` themselves can use
`Variable::dimensions` again; code that relied on `Variable::shape` being
the HDF5 extent, or on the reads returning only the written records, should
use `Variable::stored_shape` (new) — `shape` and the reads now follow
netCDF (below). `variables()` no longer lists pure dimension scales.
Found 2026-09-28 while fixing unlimited dimension sizes.
`clawhdf5-netcdf4` gave each variable the dimensions it found by size
(`match_dimensions_to_variable`: the first unused dimension of equal size,
else an anonymous `dim_<n>`), not the ones its `DIMENSION_LIST` names, and
`variables()` also listed the dimension scales that are only dimensions
(netCDF-C hides them). With netCDF4-python: unlimited `time` and `empty`,
`x` (3), `a(time)` with 2 records, `b(time, x)` with 5, `e(empty)` — netCDF4
reports `time` = 5, `a` on `time`, and variables `a`, `b`, `e` and a `c(x)`;
clawhdf5-netcdf4 reported `time` = 5 (right) but `a` on `dim_2`, and also
variables `time` and `empty` (the scales, put on `empty`). Two dimensions of
one size could be swapped the same way. Related: netCDF4 gives `a` the shape
(5,) (a variable along an unlimited dimension has the dimension's length,
unwritten records read as fill); `Variable::shape` was the HDF5 extent,
`[2]`, and reads returned those 2 values. `dimensions()` was right. The
workaround was to read the variable's `_Netcdf4Coordinates` attribute
(dimension ids, matching each scale's `_Netcdf4Dimid`).
Variables now get their dimensions as netCDF-C resolves them
(`libhdf5/hdf5open.c`): `_Netcdf4Coordinates` ids, else the scales
`DIMENSION_LIST` references, looked up in the variable's group and its
parents; size matching remains only for axes with neither (files not
written by a netCDF library). Pure dimension scales are not variables, and
`_nc4_non_coord_<name>` datasets are the variables `<name>`. A variable
along an unlimited dimension has the dimension's length, and its unwritten
records read as the fill value. One difference from netCDF-C 4.9.3 is
deliberate: when the unlimited dimension is not a variable's first
(`f(x, t)` with 1 of 4 records), a whole-variable read through netCDF-C
returns the written values first and then the fill
(`[[1, 2, fill, fill], [fill, ...]]` for rows `[1]` and `[2]`), while its
element and row reads — and clawhdf5-netcdf4 — place each row's values in
their row (`[[1, fill, fill, fill], [2, fill, fill, fill]]`).
`interop_tests` compares variables, dimensions, shapes and every value
with netCDF4-python 1.7.4 for the reproducer, dimensions of equal size,
one dimension used twice, a scalar, subgroups on their parents'
dimensions, unwritten records with and without `_FillValue`, h5py
dimension scales, and h5netcdf 1.8.1 and xarray files (tank, 2026-09-28,
`CLAWHDF5_PYTHON=<venv with h5netcdf> cargo test -p clawhdf5-netcdf4`).
Files without dimension scales still get dimensions by size (netCDF-C
gives them `phony_dim_<n>`), as before; see `CHANGELOG.md` for a
comparison over the conformance corpus's netCDF-readable files.
## HDF5 1.8 could not read the files we wrote
**Status:** fixed 2026-09-28 (branch `feat/libver-v18`), as an opt-in.
Affected every release (v2.1.0 to v2.7.0): the writer only ever wrote the
HDF5 1.10 format. Users who need HDF5 1.8 to read their files call
`FileBuilder::libver_bounds(LibVer::V18, LibVer::V18)` (format crate:
`FileWriter::libver_bounds`) and write them again; the default output is
unchanged. Listed until now under
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
HDF5 1.8.23's h5dump refused a default clawhdf5 file outright ("unable to
open file": its superblock is version 3; tank, 2026-09-28). The files
also use version-4 layout messages and the 1.10 chunk indexes (single
chunk, Fixed Array, Extensible Array, version-2 B-tree). libhdf5 2.0 made
the 1.8 format its default low bound (`H5F_LIBVER_V18`).
With a low bound of 1.8 the writer now writes what libhdf5 2.x writes for
h5py's `libver=('v108', 'latest')`: a version-2 superblock, version-3
layout messages, and a version-1 B-tree for every chunked dataset,
resizable ones included, built the way `H5B_insert` builds it (checked node
for node against libhdf5's trees). Everything else the writer emits was
already 1.8's (version-2 object headers, link messages, dense storage in
fractal heaps and version-2 B-trees, filter pipeline version 2, fill value
version 3, datatypes up to version 3). With a high bound of 1.8, what 1.8
cannot read is `FormatError::LibverBound` before anything is written:
virtual datasets, the paged file-space strategy, the 1.12 reference types,
native complex numbers.
`crates/clawhdf5-tools/tests/libver_v18.rs` writes every writer feature
under the 1.8 bound; HDF5 1.8.23's h5dump (built by
`scripts/build-hdf5-1.8.sh`; skipped where it is missing, as in CI) dumps
the whole file exactly as h5dump 1.14 does and returns our bytes for each
numeric dataset, and h5py, clawhdf5 and `h5rs check --data` agree; again
after `FileEditor` grows and appends to it (splitting B-tree nodes) and
sets attributes, and after h5py appends.
The default stays the 1.10 format: on the read harness (tank, 2026-09-28,
under load from other builds) a freshly opened file with 8192 chunks
per dataset reads small selections 1.2x to 2.3x slower through a version-1
B-tree (see `BENCHMARKS.md`, "HDF5 1.8 format").
## A dropped `FileEditor` could keep its file locked for a moment ## A dropped `FileEditor` could keep its file locked for a moment
**Status:** fixed 2026-09-28 (#23), before any release **Status:** fixed 2026-09-28 (#23), before any release
+43
View File
@@ -0,0 +1,43 @@
#!/usr/bin/env bash
# Build libhdf5 1.8.23 (the last 1.8 release) with its command-line tools, as
# the oracle for files written with `LibVer::V18` (the `libver_v18` test in
# clawhdf5-tools finds h5dump through CLAWHDF5_H5DUMP18 or this default prefix).
#
# bash scripts/build-hdf5-1.8.sh [PREFIX]
#
# PREFIX defaults to ~/.cache/hdf5-1.8.23; sources go to PREFIX-src and the
# build tree to PREFIX-build. Reuses an existing install. Needs git, cmake, a C
# compiler and zlib headers.
set -euo pipefail
PREFIX="${1:-$HOME/.cache/hdf5-1.8.23}"
SRC="$PREFIX-src"
BUILD="$PREFIX-build"
if [ -x "$PREFIX/bin/h5dump" ]; then
echo "reusing $PREFIX/bin/h5dump"
"$PREFIX/bin/h5dump" --version
exit 0
fi
if [ ! -d "$SRC" ]; then
git clone --depth 1 --branch hdf5-1_8_23 https://github.com/HDFGroup/hdf5.git "$SRC"
fi
# 1.8 predates current compilers (GCC 14 turns its pointer-type mismatches in
# the tools into errors): pin gnu99 and demote those errors to warnings.
CFLAGS18="-w -std=gnu99 -Wno-error=incompatible-pointer-types"
CFLAGS18="$CFLAGS18 -Wno-error=implicit-function-declaration -Wno-error=int-conversion"
cmake -S "$SRC" -B "$BUILD" \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_INSTALL_PREFIX="$PREFIX" \
-DCMAKE_C_FLAGS="$CFLAGS18" \
-DBUILD_SHARED_LIBS=ON \
-DBUILD_TESTING=OFF \
-DHDF5_BUILD_TOOLS=ON \
-DHDF5_BUILD_EXAMPLES=OFF \
-DHDF5_BUILD_CPP_LIB=OFF \
-DHDF5_BUILD_FORTRAN=OFF \
-DHDF5_BUILD_HL_LIB=OFF \
-DHDF5_BUILD_JAVA=OFF \
-DHDF5_ENABLE_Z_LIB_SUPPORT=ON \
-DHDF5_ENABLE_SZIP_SUPPORT=OFF
cmake --build "$BUILD" -j "${JOBS:-6}"
cmake --install "$BUILD"
"$PREFIX/bin/h5dump" --version