Compare commits
6
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4cc3111151 | ||
|
|
7fba2f3996 | ||
|
|
549e442aff | ||
|
|
728ffeef16 | ||
|
|
e4b3a76dc7 | ||
|
|
e4ba09946f |
@@ -601,6 +601,42 @@ about 150 per dataset here) against a Fixed Array's few blocks. Writing costs th
|
|||||||
same; files grow by about 36 bytes per chunk. The default therefore stays
|
same; files grow by about 36 bytes per chunk. The default therefore stays
|
||||||
the 1.10 format; the 1.8 format is opt-in.
|
the 1.10 format; the 1.8 format is opt-in.
|
||||||
|
|
||||||
|
### Streaming `FileBuilder::write` (2026-09-29, tank)
|
||||||
|
|
||||||
|
`FileBuilder::write` now streams the file through a buffered writer
|
||||||
|
instead of building it in memory first (`feat/huge-chunk-dims`, the change
|
||||||
|
that removes copies of 4 GiB unfiltered chunks). Measured 2026-09-29 on
|
||||||
|
tank (AMD Ryzen 7 7800X3D), idle (1-minute load average 1.66 to 1.83 at the
|
||||||
|
start of each run): `main` at `4260af4` against the stacked branch at
|
||||||
|
`549e442`, alternating, three full runs each; median (min-max) of
|
||||||
|
criterion's estimate, ms.
|
||||||
|
|
||||||
|
> **Run:** `cargo bench -p clawhdf5-bench --bench h5bench_write -- --noplot --warm-up-time 1 --measurement-time 3`
|
||||||
|
|
||||||
|
The first candidate used a 1 MiB write buffer and was 1.35x to 1.83x slower
|
||||||
|
on every 512 x 512 (1 MiB) chunked case when the full suite ran (e.g.
|
||||||
|
`write_2d_chunked/512x512` 0.454 -> 0.754 ms): the buffer is above glibc's
|
||||||
|
128 KiB mmap threshold, and after the smaller cases it was mapped afresh on
|
||||||
|
each write (1 629 368 minor page faults over the `write_2d_chunked` group
|
||||||
|
against 12 888 on `main`). Run alone, the case showed no difference. With a
|
||||||
|
64 KiB buffer (`549e442`):
|
||||||
|
|
||||||
|
| benchmark | main | candidate | ratio |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| write_1d_contiguous/100000 | 0.216 (0.215-0.218) | 0.206 (0.204-0.207) | 0.96x |
|
||||||
|
| write_2d_chunked/32x32 | 0.062 (0.061-0.063) | 0.061 (0.061-0.064) | 0.99x |
|
||||||
|
| write_2d_chunked/128x128 | 0.102 (0.100-0.103) | 0.101 (0.099-0.102) | 0.98x |
|
||||||
|
| write_2d_chunked/512x512 | 0.460 (0.457-0.463) | 0.466 (0.463-0.477) | 1.01x |
|
||||||
|
| write_2d_chunked_zstd/deflate-6/512x512 | 0.457 (0.450-0.465) | 0.465 (0.458-0.477) | 1.02x |
|
||||||
|
| write_2d_chunked_zstd/zstd-3/512x512 | 0.376 (0.369-0.381) | 0.387 (0.379-0.390) | 1.03x |
|
||||||
|
| write_2d_chunked_pcodec/pcodec/512x512 | 0.906 (0.895-0.908) | 0.901 (0.898-0.914) | 0.99x |
|
||||||
|
| write_multi_dataset/64 | 0.136 (0.135-0.138) | 0.136 (0.135-0.136) | 1.00x |
|
||||||
|
| write_with_attrs/64 | 0.063 (0.062-0.063) | 0.063 (0.062-0.063) | 1.00x |
|
||||||
|
|
||||||
|
The other 18 cases are within 2%. Minor page faults per full run: 59 602 to
|
||||||
|
64 486 on both sides. The write path is as fast as before; what changes is
|
||||||
|
peak memory for large unfiltered chunks (see CHANGELOG, 2026-09-29).
|
||||||
|
|
||||||
### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank)
|
### HDF5 1.8 format: version-1 B-tree chunk indexes, idle re-run (2026-09-28, tank)
|
||||||
|
|
||||||
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load
|
Measured 2026-09-28 on tank (AMD Ryzen 7 7800X3D), idle: 1-minute load
|
||||||
|
|||||||
+122
@@ -56,6 +56,128 @@
|
|||||||
and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4).
|
and `docs/known-issues.md` (tank, 2026-09-29, netCDF4-python 1.7.4).
|
||||||
Affected v2.1.0 to v2.7.0.
|
Affected v2.1.0 to v2.7.0.
|
||||||
|
|
||||||
|
### HDF5 2.0 native complex is its own type on read; Python `libver=` (2026-09-29)
|
||||||
|
- **`Datatype::parse` returns `Datatype::Complex { size, base_type }`** for
|
||||||
|
a class-11 message, also as a compound member, array base or
|
||||||
|
variable-length base. It used to return the equivalent `{r, i}`
|
||||||
|
compound (`docs/known-issues.md`, "HDF5 2.0 native complex numbers read
|
||||||
|
as a `{r, i}` compound"). **Breaking** for code that matched the
|
||||||
|
compound view of a native complex type: `Dataset::raw_datatype()` and
|
||||||
|
attribute datatypes are now `Datatype::Complex`;
|
||||||
|
`Datatype::complex_as_compound(size, base)` still gives the `{r, i}` view,
|
||||||
|
and `data_read::read_compound_fields` accepts a `Complex` directly.
|
||||||
|
Re-serializing a parsed type (e.g. copying it to another file) now writes
|
||||||
|
class 11 again instead of a compound.
|
||||||
|
- **Facade:** `DType::Complex(Box<DType>)` (new variant, **breaking** for
|
||||||
|
exhaustive `match`es over `DType`): `Complex(F32)` is numpy `complex64`,
|
||||||
|
`Complex(F64)` `complex128`, a binary16 part `Complex(Other("float16"))`
|
||||||
|
(the same `Other` name small floats already use); `Display` prints
|
||||||
|
`complex<f64>`. h5py's own complex encoding, the compound `{r, i}`, is
|
||||||
|
still `DType::Compound`. `read_complex_f64`/`read_complex_f32` read both.
|
||||||
|
- **`h5rs`:** `dump` prints native complex as h5dump 2.2.0 does:
|
||||||
|
`H5T_COMPLEX_IEEE_F{16,32,64}{LE,BE}` (else `H5T_COMPLEX { <base> }`), in
|
||||||
|
arrays and compounds too, and values as `1.5-2i` (`%g%+gi`; for binary16
|
||||||
|
parts, which h5dump has no C type for, `1+-2i` as it prints them). `ls`
|
||||||
|
shows `complex64`, `complex128-be`, `complex32` (and `complex<part>` for
|
||||||
|
non-IEEE parts); `ls -v` shows h5ls 2.2.0's `complex number of` /
|
||||||
|
`IEEE 64-bit big-endian float`. `diff` compares complex values part by
|
||||||
|
part and prints the difference as h5diff 2.2.0 does (`1+0i`); a native
|
||||||
|
complex and an `{r, i}` compound are no longer comparable (different
|
||||||
|
classes, as in h5diff). `dump --json` keeps the `{r, i}` compound and
|
||||||
|
`[re, im]` values: hdf5-json (h5json 2.0.0) has no complex class.
|
||||||
|
- **Browser (`clawhdf5-wasm`):** native complex datasets are readable, as
|
||||||
|
`[re, im]` pairs: `info().elementShape` ends in `2`, `read()` returns the
|
||||||
|
parts interleaved in a `Float32Array`/`Float64Array` (binary16 parts
|
||||||
|
widened to f32), and `dtype` is `complex<f64>` (`complex<f32
|
||||||
|
(big-endian)>`, `array[2]<complex<f32>>`). The viewer shows each element
|
||||||
|
as `[re, im]` with no change to the page. h5py's `{r, i}` compound is
|
||||||
|
still refused, as every compound is.
|
||||||
|
- **Python:** `clawhdf5.File(path, 'w', libver=...)` takes h5py's values:
|
||||||
|
`'v108'`, `'v110'`, `'v112'`, `'v114'`, `'v200'`, `'latest'` (the low
|
||||||
|
bound; high `'latest'`, as h5py) or a `(low, high)` tuple, mapped to
|
||||||
|
`FileBuilder::libver_bounds`. `'earliest'` as the low bound writes the
|
||||||
|
1.8 format with a `UserWarning` (clawhdf5 cannot write the pre-1.8
|
||||||
|
format); as the high bound it is a `ValueError`, as are unknown names and
|
||||||
|
a low bound above the high one. It is ignored for `'r'` and refused
|
||||||
|
(`NotImplementedError`) for `'r+'`/`'a'`. Native complex already read as
|
||||||
|
numpy `complex64`/`complex128` and still does
|
||||||
|
(`test_read_native_complex_from_h5py`).
|
||||||
|
- Tests (tank, 2026-09-29): `crates/clawhdf5-tools/tests/native_complex_dump.rs`
|
||||||
|
compares `h5rs dump` with h5dump 2.2.0's output stored next to a fixture
|
||||||
|
written by h5py 3.16 / libhdf5 2.0.0
|
||||||
|
(`crates/clawhdf5/tests/fixtures/gen_native_complex.py`:
|
||||||
|
`F32LE`/`F64LE`/`F64BE`/`F16LE`, scalar, compound member, array, root
|
||||||
|
attribute), line for line outside the values and value for value within
|
||||||
|
them, and `ls`/`diff` with h5ls/h5diff 2.2.0;
|
||||||
|
`integration_tests::complex_datasets_and_attributes_round_trip`
|
||||||
|
(`DType::Complex`, `raw_datatype()`); `clawhdf5-wasm` unit tests on the
|
||||||
|
fixture, `h5py_interop` and `examples/wasm-viewer/test` (Node:
|
||||||
|
`test.mjs`, 274 + 1324 checks; the page in headless Chromium for
|
||||||
|
`/native_c128`, `/native_c64_be`, `/pairs`) with two native complex
|
||||||
|
datasets added to `make_fixture.py` (when h5py's libhdf5 is 2.0+);
|
||||||
|
`crates/clawhdf5-py/tests/test_libver.py` (every bound; `'v108'` output
|
||||||
|
opens in HDF5 1.8.23's h5dump and in h5py). `test.mjs` no longer
|
||||||
|
hard-codes the fixture's length.
|
||||||
|
|
||||||
|
### Chunk dimensions of 2^32 or more; 4 GiB chunks written without copies (2026-09-29)
|
||||||
|
- **Chunk dimensions of 2^32 or more** (libhdf5 2.x writes them with
|
||||||
|
layout message version 5, up to 8 bytes per dimension; a 4 GiB chunk of
|
||||||
|
1-byte elements needs one) are read and written. They were refused
|
||||||
|
(`InvalidChunkDimensions`). **Public API changes:**
|
||||||
|
`DataLayout::Chunked::chunk_dimensions` is `Vec<u64>` (was `Vec<u32>`),
|
||||||
|
and the `chunk_dimensions` argument of `chunked_read::{collect_chunk_info_checked,
|
||||||
|
collect_chunk_info_checked_in, generate_implicit_chunks_in_grid}`,
|
||||||
|
`fixed_array::{read_fixed_array_chunks, read_fixed_array_chunks_in}`,
|
||||||
|
`extensible_array::{read_extensible_array_chunks,
|
||||||
|
read_extensible_array_chunks_in}`, and the `chunk_dims` argument of
|
||||||
|
`chunked_write::serialize_v4_single_chunk_pub`, are `&[u64]` (were
|
||||||
|
`&[u32]`). A chunk whose size overflows 64 bits is refused when a dataset is read
|
||||||
|
(`InvalidChunkDimensions`), and the writer refuses a chunk dimension of
|
||||||
|
0 up front. A version-3 layout (4-byte dimensions) is never written for
|
||||||
|
such a chunk: over 4 GiB, it always takes version 5, as in libhdf5.
|
||||||
|
- **Writing without copies:** `FileWriter` no longer copies chunks into a
|
||||||
|
buffer per pass and then into the file's buffer. A chunk that is a
|
||||||
|
contiguous run of the dataset's data (a dataset stored as one chunk of
|
||||||
|
its shape, blocks of whole rows) is borrowed when unfiltered and
|
||||||
|
compressed straight from it when filtered, and contiguous datasets are
|
||||||
|
not copied. New `FileWriter::finish_with(put)` hands the file out piece
|
||||||
|
by piece; `FileBuilder::write` uses it (through a 1 MiB `BufWriter`,
|
||||||
|
still to a temporary file renamed into place), so a file is never
|
||||||
|
assembled in memory. New `DatasetBuilder::with_u8_data_owned(Vec<u8>)`
|
||||||
|
takes the data without the copy `with_u8_data` makes. Peak resident
|
||||||
|
memory of one unfiltered 1 GiB chunk of `u8`
|
||||||
|
(`cargo run --release -p clawhdf5 --example write_one_chunk --
|
||||||
|
1073741824 OUT owned|slice`, tank, 2026-09-29): 5.0 GiB before (the
|
||||||
|
writer of `4260af4`), 1.0 GiB after with `owned`; 6.0 and 2.0 GiB with
|
||||||
|
`slice`; 4 GiB + 7 bytes `owned`, 4.00 GiB. Same bytes written. Write
|
||||||
|
speed was not re-measured (the machine was not idle): the A/B of
|
||||||
|
`FileBuilder::write` and `finish` on small and large files is still to
|
||||||
|
be run.
|
||||||
|
- `chunked_write::extract_chunk` (behind `split_into_chunks`) no longer
|
||||||
|
panics when the data is shorter than the shape; the missing part is
|
||||||
|
zeros, as it was meant to be.
|
||||||
|
- Tests: `crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5` (31 KB,
|
||||||
|
libhdf5 2.0.0 via h5py 3.16, `gen_huge_chunks.py dims`: `u8` chunks of
|
||||||
|
2^32 + 7 deflated twice, Single Chunk and Extensible Array), and in
|
||||||
|
`huge_chunks_interop`: `huge_chunk_dims_list` (always), and with
|
||||||
|
`CLAWHDF5_HUGE_CHUNKS=1` `huge_chunk_dims_read`,
|
||||||
|
`huge_chunk_dims_unfiltered_read` (a sparse file h5py writes, Single
|
||||||
|
Chunk and Fixed Array, mmap and positioned reads),
|
||||||
|
`writer_huge_chunk_dims_round_trip` (deflated, read back by clawhdf5,
|
||||||
|
h5py 3.16 and — layouts — h5dump 2.2.0),
|
||||||
|
`writer_unfiltered_huge_chunk_round_trip` (an unfiltered 4 GiB + 7 byte
|
||||||
|
chunk: clawhdf5, h5py, and h5dump 2.2.0 printing values past 2^32) and
|
||||||
|
`writer_huge_chunks_lz4_zstd_round_trip` (with the `lz4`/`zstd`
|
||||||
|
features; h5py through hdf5plugin 7.1.0). The whole opt-in suite
|
||||||
|
(`cargo test --release -p clawhdf5 --features lz4,zstd --test
|
||||||
|
huge_chunks_interop -- --test-threads=1`, h5py and h5dump 2.2.0
|
||||||
|
included) passed with a peak of 10.04 GiB resident (tank, 2026-09-29,
|
||||||
|
commit `0065c6b`). Unit tests: 5-byte dimension encoding, contiguous
|
||||||
|
chunk detection, `finish_with` against `finish`. The wasm package test
|
||||||
|
lists the new fixture and gets `FormatError::Overflow` reading it.
|
||||||
|
`docs/known-issues.md`: two fixed entries; the open "Chunks of 4 GiB or
|
||||||
|
more: limits" keeps the refused filters, the editor and memory.
|
||||||
|
|
||||||
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
|
### NetCDF-4: variables' dimensions come from the file (2026-09-28)
|
||||||
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
|
- `clawhdf5-netcdf4` gave each variable the first unused dimension of
|
||||||
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
|
equal size (else an anonymous `dim_<n>`), so a variable on an unlimited
|
||||||
|
|||||||
@@ -102,7 +102,8 @@ writes 85.2 µs vs 877 µs (10.3x); 64 group creates 130 µs vs 1.37 ms
|
|||||||
(10.6x); a 512×512 `f32` chunked deflate-6 write 1.44 ms vs 65.0 ms
|
(10.6x); a 512×512 `f32` chunked deflate-6 write 1.44 ms vs 65.0 ms
|
||||||
(re-measured 2026-09-23 with the pure-Rust deflate: 1.46 ms vs 51.4 ms,
|
(re-measured 2026-09-23 with the pure-Rust deflate: 1.46 ms vs 51.4 ms,
|
||||||
35x); a 100K `f32` sequential write is a tie. The writer (`FileBuilder`)
|
35x); a 100K `f32` sequential write is a tie. The writer (`FileBuilder`)
|
||||||
assembles a file in memory and writes it once, which is part of that
|
lays out the whole file before writing it (when these were measured, in
|
||||||
|
one buffer in memory), which is part of that
|
||||||
difference; read the caveats in [BENCHMARKS.md](BENCHMARKS.md#caveats)
|
difference; read the caveats in [BENCHMARKS.md](BENCHMARKS.md#caveats)
|
||||||
before quoting these.
|
before quoting these.
|
||||||
|
|
||||||
@@ -115,12 +116,12 @@ Limits and open issues, with dates, are in
|
|||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
|
| **File format** | Superblock v0–v3, user blocks, v1/v2 object headers; writing the HDF5 1.10 format (default) or, with `libver_bounds(LibVer::V18, LibVer::V18)`, files HDF5 1.8 reads (checked with HDF5 1.8.23) | Metadata cache images | Writing the pre-1.8 format (version-0 superblock, symbol-table groups) |
|
||||||
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
|
| **Groups and links** | Symbol-table, compact and dense groups (tested to 100 000 links), creation order, soft and hard links; writing external links | | Following external links (explicit error); user-defined links are skipped |
|
||||||
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; reads surface it as `{r, i}`) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
|
| **Datatypes** | Integers and IEEE floats of every width and byte order (incl. `f16`), enums, compounds (every version, incl. HDF5 2.0's v5), arrays, fixed-length strings, opaque, complex: h5py's `{r, i}` compound (`with_complex_f64_data`) and HDF5 2.0's native class 11 (`with_native_complex_f64_data`, opt-in: only libhdf5 2.0+ reads it; read as its own type, `DType::Complex`, and printed by `h5rs` as h5dump 2.x prints it) | Variable-length strings and sequences, object references; HDF5 2.x's small floats (bfloat16, FP8 E4M3/E5M2, FP6 E2M3/E3M2, FP4 E2M1: every bit pattern decoded as libhdf5 2.2.0 decodes it) and other non-IEEE floats up to 64 bits | Writing variable-length data; writing non-IEEE floats; decoding region and attribute references; x87 long double and binary128 |
|
||||||
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error); chunk dimensions of 2^32 or more |
|
| **Layouts and chunk indexes** | Compact, contiguous and chunked; chunk indexes single chunk, Fixed Array, Extensible Array and v2 B-tree (the writer picks one as libhdf5 does), and v1 B-tree (every chunked dataset under the 1.8 bound, split as libhdf5 splits it); chunks of 4 GiB or more and chunk dimensions of 2^32 or more (HDF5 2.0's layout message version 5); fill values; resizable datasets; virtual datasets (read limits in known-issues) | The implicit chunk index (the editor also changes it) | External raw data files (explicit error) |
|
||||||
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
|
| **Filters** | deflate (pure-Rust zlib-rs), shuffle, Fletcher-32, LZ4 (opt-in), Zstd (C, opt-in); plugins LZF, bitshuffle, bzip2, Blosc 1 | N-Bit, scale-offset, SZIP (C, opt-in); plugins Blosc2 and ZFP | Other filter IDs, unless you register a codec (`filter_registry::register_filter`) |
|
||||||
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
|
| **Editing in place** | `FileEditor`: overwrite values, grow and shrink chunked datasets (every index), set attributes (compact and dense), in files from h5py or clawhdf5 | | Creating or deleting objects in an existing file; deleting attributes; new chunks in implicit indexes; VL data; rewriting chunks of 4 GiB or more; filters this build cannot encode (refused before any write) |
|
||||||
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
|
| **Access** | Local files (mmap or buffered), bytes in memory, any `Storage` backend, HTTP(S) and S3/GCS/Azure via `clawhdf5-remote`, SWMR reading (`File::open_swmr`, `Dataset::refresh`) | Remote files and the browser are read-only | SWMR writing; remote SWMR; MPI collective I/O (`clawhdf5-io`'s `mpi-io` reads on one rank and broadcasts) |
|
||||||
| **Bindings** | Python (read, `'w'` for numeric and complex arrays, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`); no Zstd/SZIP/pcodec, no compound, reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
|
| **Bindings** | Python (read, `'w'` for numeric and complex arrays with h5py's `libver=`, `'r+'` editing, URLs), NetCDF-4 (CF scale/offset/fill) | WebAssembly (`open(bytes)`, `openUrl`; HDF5 2.0 native complex as `[re, im]` pairs); no Zstd/SZIP/pcodec, no compound (h5py's `{r, i}` complex included), reference, opaque, bitfield, time or VL-sequence datasets | Node.js (the package does not work; see known-issues) |
|
||||||
|
|
||||||
Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`,
|
Plugin filters other than LZF are cargo features (`bitshuffle`, `bzip2`,
|
||||||
`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure
|
`blosc`, `blosc2`, `zfp`, or `plugin-filters` for all of them), all pure
|
||||||
|
|||||||
@@ -555,7 +555,7 @@ pub(crate) fn checked_byte_len(elements: u64, elem_size: usize) -> Result<usize,
|
|||||||
/// HDF5 2.0 writes them with layout version 5. A zero chunk dimension used
|
/// HDF5 2.0 writes them with layout version 5. A zero chunk dimension used
|
||||||
/// to read as all fill values, and a huge one to hang the reader.
|
/// to read as all fill values, and a huge one to hang the reader.
|
||||||
pub(crate) fn chunk_geometry(
|
pub(crate) fn chunk_geometry(
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
layout_version: u8,
|
layout_version: u8,
|
||||||
dataspace: &Dataspace,
|
dataspace: &Dataspace,
|
||||||
elem_size: usize,
|
elem_size: usize,
|
||||||
@@ -576,15 +576,26 @@ pub(crate) fn chunk_geometry(
|
|||||||
"chunk size must be > 0, dim = {d}"
|
"chunk size must be > 0, dim = {d}"
|
||||||
)));
|
)));
|
||||||
}
|
}
|
||||||
let bytes = spatial
|
// Saturating: 33 dimensions of up to 64 bits each can overflow u128.
|
||||||
.iter()
|
let bytes = spatial.iter().fold(elem_size as u128, |acc, &c| {
|
||||||
.fold(elem_size as u128, |acc, &c| acc * u128::from(c));
|
acc.saturating_mul(u128::from(c))
|
||||||
|
});
|
||||||
|
// So the indexes can compute a chunk's size in 64 bits.
|
||||||
|
if bytes > u128::from(u64::MAX) {
|
||||||
|
return Err(FormatError::InvalidChunkDimensions(format!(
|
||||||
|
"chunk size overflows 64 bits (chunk {spatial:?} of {elem_size}-byte elements)"
|
||||||
|
)));
|
||||||
|
}
|
||||||
if layout_version < 4 && bytes > u128::from(u32::MAX) {
|
if layout_version < 4 && bytes > u128::from(u32::MAX) {
|
||||||
return Err(FormatError::InvalidChunkDimensions(format!(
|
return Err(FormatError::InvalidChunkDimensions(format!(
|
||||||
"chunk size must be < 4GB with v1 b-tree index (chunk {spatial:?} of {elem_size}-byte elements)"
|
"chunk size must be < 4GB with v1 b-tree index (chunk {spatial:?} of {elem_size}-byte elements)"
|
||||||
)));
|
)));
|
||||||
}
|
}
|
||||||
Ok((rank, spatial.iter().map(|&c| c as usize).collect()))
|
let spatial = spatial
|
||||||
|
.iter()
|
||||||
|
.map(|&c| to_usize(c))
|
||||||
|
.collect::<Result<Vec<usize>, _>>()?;
|
||||||
|
Ok((rank, spatial))
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The size of one element of `dt` as stored in the file: a
|
/// The size of one element of `dt` as stored in the file: a
|
||||||
@@ -625,7 +636,7 @@ pub fn check_chunk_element_size(
|
|||||||
return Ok(());
|
return Ok(());
|
||||||
};
|
};
|
||||||
let expected = stored_element_size(datatype, offset_size);
|
let expected = stored_element_size(datatype, offset_size);
|
||||||
if u64::from(stored) != expected {
|
if stored != expected {
|
||||||
return Err(FormatError::InvalidChunkDimensions(format!(
|
return Err(FormatError::InvalidChunkDimensions(format!(
|
||||||
"stored datatype size in chunk layout does not match datatype description \
|
"stored datatype size in chunk layout does not match datatype description \
|
||||||
(layout {stored} bytes, datatype {expected})"
|
(layout {stored} bytes, datatype {expected})"
|
||||||
@@ -759,7 +770,7 @@ pub fn collect_chunk_info_in<S: Storage + ?Sized>(
|
|||||||
pub fn collect_chunk_info_checked(
|
pub fn collect_chunk_info_checked(
|
||||||
file_data: &[u8],
|
file_data: &[u8],
|
||||||
btree_address: u64,
|
btree_address: u64,
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
length_size: u8,
|
length_size: u8,
|
||||||
) -> Result<Vec<ChunkInfo>, FormatError> {
|
) -> Result<Vec<ChunkInfo>, FormatError> {
|
||||||
@@ -776,7 +787,7 @@ pub fn collect_chunk_info_checked(
|
|||||||
pub fn collect_chunk_info_checked_in<S: Storage + ?Sized>(
|
pub fn collect_chunk_info_checked_in<S: Storage + ?Sized>(
|
||||||
file_data: &S,
|
file_data: &S,
|
||||||
btree_address: u64,
|
btree_address: u64,
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
length_size: u8,
|
length_size: u8,
|
||||||
) -> Result<Vec<ChunkInfo>, FormatError> {
|
) -> Result<Vec<ChunkInfo>, FormatError> {
|
||||||
@@ -808,7 +819,7 @@ pub fn collect_chunk_info_checked_in<S: Storage + ?Sized>(
|
|||||||
.iter_mut()
|
.iter_mut()
|
||||||
.zip(chunk.offsets.iter().zip(chunk_dimensions))
|
.zip(chunk.offsets.iter().zip(chunk_dimensions))
|
||||||
{
|
{
|
||||||
*w = o / u64::from(d);
|
*w = o / d;
|
||||||
}
|
}
|
||||||
wanted[ndims - 1] = 0;
|
wanted[ndims - 1] = 0;
|
||||||
if let Some(i) = root.find(&wanted) {
|
if let Some(i) = root.find(&wanted) {
|
||||||
@@ -849,9 +860,9 @@ impl ChunkNode {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn scale_keys(&mut self, dims: &[u32]) {
|
fn scale_keys(&mut self, dims: &[u64]) {
|
||||||
for (k, &d) in self.keys.iter_mut().zip(dims.iter().cycle()) {
|
for (k, &d) in self.keys.iter_mut().zip(dims.iter().cycle()) {
|
||||||
*k /= u64::from(d);
|
*k /= d;
|
||||||
}
|
}
|
||||||
if let ChunkChildren::Nodes(nodes) = &mut self.children {
|
if let ChunkChildren::Nodes(nodes) = &mut self.children {
|
||||||
for n in nodes {
|
for n in nodes {
|
||||||
@@ -919,9 +930,9 @@ fn btree_cmp3(lt: &[u64], scaled: &[u64], rt: &[u64]) -> core::cmp::Ordering {
|
|||||||
|
|
||||||
/// Check one v1 B-tree chunk key's offsets (see
|
/// Check one v1 B-tree chunk key's offsets (see
|
||||||
/// [`collect_chunk_info_checked`]).
|
/// [`collect_chunk_info_checked`]).
|
||||||
fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u32]) -> Result<(), FormatError> {
|
fn check_key_offsets(offsets: &[u64], chunk_dimensions: &[u64]) -> Result<(), FormatError> {
|
||||||
for (&offset, &dim) in offsets.iter().zip(chunk_dimensions) {
|
for (&offset, &dim) in offsets.iter().zip(chunk_dimensions) {
|
||||||
if dim == 0 || offset % u64::from(dim) != 0 {
|
if dim == 0 || offset % dim != 0 {
|
||||||
return Err(FormatError::ChunkedReadError(format!(
|
return Err(FormatError::ChunkedReadError(format!(
|
||||||
"bad coordinate offset {offsets:?} for chunk dimensions {chunk_dimensions:?}"
|
"bad coordinate offset {offsets:?} for chunk dimensions {chunk_dimensions:?}"
|
||||||
)));
|
)));
|
||||||
@@ -937,7 +948,7 @@ fn read_key_offsets(
|
|||||||
file_data: &[u8],
|
file_data: &[u8],
|
||||||
pos: usize,
|
pos: usize,
|
||||||
ndims: usize,
|
ndims: usize,
|
||||||
chunk_dimensions: Option<&[u32]>,
|
chunk_dimensions: Option<&[u64]>,
|
||||||
out: &mut Vec<u64>,
|
out: &mut Vec<u64>,
|
||||||
) -> Result<(), FormatError> {
|
) -> Result<(), FormatError> {
|
||||||
let start = out.len();
|
let start = out.len();
|
||||||
@@ -966,7 +977,7 @@ fn parse_chunk_node<S: Storage + ?Sized>(
|
|||||||
file_data: &S,
|
file_data: &S,
|
||||||
btree_address: u64,
|
btree_address: u64,
|
||||||
ndims: usize,
|
ndims: usize,
|
||||||
chunk_dimensions: Option<&[u32]>,
|
chunk_dimensions: Option<&[u64]>,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
depth: usize,
|
depth: usize,
|
||||||
stored: &mut Vec<ChunkInfo>,
|
stored: &mut Vec<ChunkInfo>,
|
||||||
@@ -1075,7 +1086,7 @@ fn parse_chunk_node<S: Storage + ?Sized>(
|
|||||||
pub fn generate_implicit_chunks(
|
pub fn generate_implicit_chunks(
|
||||||
base_address: u64,
|
base_address: u64,
|
||||||
dataset_dims: &[u64],
|
dataset_dims: &[u64],
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
) -> Vec<ChunkInfo> {
|
) -> Vec<ChunkInfo> {
|
||||||
generate_implicit_chunks_in_grid(
|
generate_implicit_chunks_in_grid(
|
||||||
@@ -1097,17 +1108,18 @@ pub fn generate_implicit_chunks_in_grid(
|
|||||||
base_address: u64,
|
base_address: u64,
|
||||||
dataset_dims: &[u64],
|
dataset_dims: &[u64],
|
||||||
max_dims: &[u64],
|
max_dims: &[u64],
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
) -> Vec<ChunkInfo> {
|
) -> Vec<ChunkInfo> {
|
||||||
let rank = chunk_dimensions.len();
|
let rank = chunk_dimensions.len();
|
||||||
let chunk_byte_size: u64 =
|
let chunk_byte_size: u64 = chunk_dimensions
|
||||||
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
|
.iter()
|
||||||
|
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
|
||||||
|
|
||||||
let mut num_chunks_per_dim = Vec::with_capacity(rank);
|
let mut num_chunks_per_dim = Vec::with_capacity(rank);
|
||||||
let mut grid_per_dim = Vec::with_capacity(rank);
|
let mut grid_per_dim = Vec::with_capacity(rank);
|
||||||
for d in 0..rank {
|
for d in 0..rank {
|
||||||
let ch = chunk_dimensions[d] as u64;
|
let ch = chunk_dimensions[d];
|
||||||
let n = dataset_dims[d].div_ceil(ch);
|
let n = dataset_dims[d].div_ceil(ch);
|
||||||
num_chunks_per_dim.push(n);
|
num_chunks_per_dim.push(n);
|
||||||
grid_per_dim.push(max_dims.get(d).map_or(n, |m| m.div_ceil(ch)).max(n));
|
grid_per_dim.push(max_dims.get(d).map_or(n, |m| m.div_ceil(ch)).max(n));
|
||||||
@@ -1125,7 +1137,7 @@ pub fn generate_implicit_chunks_in_grid(
|
|||||||
let nchunks = num_chunks_per_dim[d];
|
let nchunks = num_chunks_per_dim[d];
|
||||||
let chunk_idx = remaining % nchunks;
|
let chunk_idx = remaining % nchunks;
|
||||||
remaining /= nchunks;
|
remaining /= nchunks;
|
||||||
offsets[d] = chunk_idx * chunk_dimensions[d] as u64;
|
offsets[d] = chunk_idx * chunk_dimensions[d];
|
||||||
grid_idx = grid_idx.saturating_add(chunk_idx.saturating_mul(down));
|
grid_idx = grid_idx.saturating_add(chunk_idx.saturating_mul(down));
|
||||||
down = down.saturating_mul(grid_per_dim[d]);
|
down = down.saturating_mul(grid_per_dim[d]);
|
||||||
}
|
}
|
||||||
@@ -1350,7 +1362,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
|
|||||||
}
|
}
|
||||||
(4, Some(2)) => {
|
(4, Some(2)) => {
|
||||||
// Implicit index — use spatial chunk dims only
|
// Implicit index — use spatial chunk dims only
|
||||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
|
||||||
generate_implicit_chunks_in_grid(
|
generate_implicit_chunks_in_grid(
|
||||||
addr,
|
addr,
|
||||||
&dataspace.dimensions,
|
&dataspace.dimensions,
|
||||||
@@ -1364,7 +1376,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
|
|||||||
}
|
}
|
||||||
(4, Some(3)) => {
|
(4, Some(3)) => {
|
||||||
// Fixed Array — use spatial chunk dims only
|
// Fixed Array — use spatial chunk dims only
|
||||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
|
||||||
let header = FixedArrayHeader::parse_in(
|
let header = FixedArrayHeader::parse_in(
|
||||||
file_data,
|
file_data,
|
||||||
checked_addr(addr)?,
|
checked_addr(addr)?,
|
||||||
@@ -1384,7 +1396,7 @@ pub fn list_chunks_in<S: Storage + ?Sized>(
|
|||||||
}
|
}
|
||||||
(4, Some(4)) => {
|
(4, Some(4)) => {
|
||||||
// Extensible Array — use spatial chunk dims only
|
// Extensible Array — use spatial chunk dims only
|
||||||
let spatial_chunk_dims: &[u32] = &chunk_dimensions[..rank];
|
let spatial_chunk_dims: &[u64] = &chunk_dimensions[..rank];
|
||||||
let header = ExtensibleArrayHeader::parse_in(
|
let header = ExtensibleArrayHeader::parse_in(
|
||||||
file_data,
|
file_data,
|
||||||
checked_addr(addr)?,
|
checked_addr(addr)?,
|
||||||
@@ -2435,12 +2447,12 @@ mod tests {
|
|||||||
/// A leaf whose final key is the one libhdf5 writes past the last chunk:
|
/// A leaf whose final key is the one libhdf5 writes past the last chunk:
|
||||||
/// each coordinate of the last chunk plus its chunk dimension (the
|
/// each coordinate of the last chunk plus its chunk dimension (the
|
||||||
/// element size last).
|
/// element size last).
|
||||||
fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u32], offset_size: u8) -> Vec<u8> {
|
fn build_chunk_btree_leaf_dims(chunks: &[ChunkInfo], dims: &[u64], offset_size: u8) -> Vec<u8> {
|
||||||
let last = &chunks.last().expect("a chunk").offsets;
|
let last = &chunks.last().expect("a chunk").offsets;
|
||||||
let end: Vec<u64> = dims
|
let end: Vec<u64> = dims
|
||||||
.iter()
|
.iter()
|
||||||
.enumerate()
|
.enumerate()
|
||||||
.map(|(d, &c)| last.get(d).copied().unwrap_or(0) + u64::from(c))
|
.map(|(d, &c)| last.get(d).copied().unwrap_or(0) + c)
|
||||||
.collect();
|
.collect();
|
||||||
build_chunk_btree_leaf_to(chunks, &end, offset_size)
|
build_chunk_btree_leaf_to(chunks, &end, offset_size)
|
||||||
}
|
}
|
||||||
@@ -2771,7 +2783,7 @@ mod tests {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// Build B-tree at offset 0x100
|
// Build B-tree at offset 0x100
|
||||||
let dims = [chunk_size_elems as u32, elem_size as u32];
|
let dims = [chunk_size_elems as u64, elem_size as u64];
|
||||||
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
|
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
|
||||||
let btree_addr = 0x100usize;
|
let btree_addr = 0x100usize;
|
||||||
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
|
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
|
||||||
@@ -3145,7 +3157,7 @@ mod tests {
|
|||||||
data_offset += compressed.len() + 16; // some padding
|
data_offset += compressed.len() + 16; // some padding
|
||||||
}
|
}
|
||||||
|
|
||||||
let dims = [chunk_elems as u32, elem_size as u32];
|
let dims = [chunk_elems as u64, elem_size as u64];
|
||||||
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
|
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
|
||||||
let btree_addr = 0x100usize;
|
let btree_addr = 0x100usize;
|
||||||
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
|
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
|
||||||
@@ -3228,7 +3240,7 @@ mod tests {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
let dims = [chunk_dims[0] as u32, chunk_dims[1] as u32, elem_size as u32];
|
let dims = [chunk_dims[0] as u64, chunk_dims[1] as u64, elem_size as u64];
|
||||||
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
|
let btree = build_chunk_btree_leaf_dims(&chunk_infos, &dims, os);
|
||||||
let btree_addr = 0x100usize;
|
let btree_addr = 0x100usize;
|
||||||
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
|
file_data[btree_addr..btree_addr + btree.len()].copy_from_slice(&btree);
|
||||||
@@ -3408,7 +3420,7 @@ mod tests {
|
|||||||
}
|
}
|
||||||
|
|
||||||
let layout = DataLayout::Chunked {
|
let layout = DataLayout::Chunked {
|
||||||
chunk_dimensions: vec![chunk_elems as u32, elem_size as u32],
|
chunk_dimensions: vec![chunk_elems as u64, elem_size as u64],
|
||||||
btree_address: Some(data_addr as u64),
|
btree_address: Some(data_addr as u64),
|
||||||
version: 4,
|
version: 4,
|
||||||
chunk_index_type: Some(1),
|
chunk_index_type: Some(1),
|
||||||
|
|||||||
@@ -5,12 +5,14 @@ extern crate alloc;
|
|||||||
|
|
||||||
use crate::addr::saturating_usize;
|
use crate::addr::saturating_usize;
|
||||||
#[cfg(not(feature = "std"))]
|
#[cfg(not(feature = "std"))]
|
||||||
use alloc::{format, vec, vec::Vec};
|
use alloc::{borrow::Cow, format, vec, vec::Vec};
|
||||||
|
#[cfg(feature = "std")]
|
||||||
|
use std::borrow::Cow;
|
||||||
|
|
||||||
use crate::btree_v1_write;
|
use crate::btree_v1_write;
|
||||||
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
|
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
|
||||||
use crate::checksum::jenkins_lookup3;
|
use crate::checksum::jenkins_lookup3;
|
||||||
use crate::chunk_cache::{CACHE_LINE_SIZE, align_to_cache_line};
|
use crate::chunk_cache::CACHE_LINE_SIZE;
|
||||||
use crate::chunk_grid::ChunkGrid;
|
use crate::chunk_grid::ChunkGrid;
|
||||||
use crate::ea_writer;
|
use crate::ea_writer;
|
||||||
use crate::error::FormatError;
|
use crate::error::FormatError;
|
||||||
@@ -464,7 +466,9 @@ fn extract_chunk(
|
|||||||
let dst: usize = (0..rank).map(|d| idx[d] * chunk_strides[d]).sum::<usize>() * element_size;
|
let dst: usize = (0..rank).map(|d| idx[d] * chunk_strides[d]).sum::<usize>() * element_size;
|
||||||
// Whole elements only, as far as `raw_data` reaches.
|
// Whole elements only, as far as `raw_data` reaches.
|
||||||
let n = row.min(raw_data.len().saturating_sub(src)) / element_size * element_size;
|
let n = row.min(raw_data.len().saturating_sub(src)) / element_size * element_size;
|
||||||
chunk[dst..dst + n].copy_from_slice(&raw_data[src..src + n]);
|
if n > 0 {
|
||||||
|
chunk[dst..dst + n].copy_from_slice(&raw_data[src..src + n]);
|
||||||
|
}
|
||||||
// Next row: advance every dimension but the last.
|
// Next row: advance every dimension but the last.
|
||||||
let mut d = rank - 1;
|
let mut d = rank - 1;
|
||||||
loop {
|
loop {
|
||||||
@@ -481,6 +485,66 @@ fn extract_chunk(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The bytes of the `linear_idx`-th chunk (see [`extract_chunk`]),
|
||||||
|
/// borrowed from `raw_data` when the chunk is one contiguous run of it: a
|
||||||
|
/// chunk inside the dataset (not past its edge) whose dimensions after the
|
||||||
|
/// first one above 1 equal the dataset's. A dataset stored as one chunk of
|
||||||
|
/// its own shape is such a run, so its chunk is never copied. Other chunks
|
||||||
|
/// are copied out, as [`extract_chunk`] does.
|
||||||
|
fn chunk_bytes_of<'a>(
|
||||||
|
raw_data: &'a [u8],
|
||||||
|
shape: &[u64],
|
||||||
|
chunk_dims: &[u64],
|
||||||
|
element_size: usize,
|
||||||
|
linear_idx: u64,
|
||||||
|
) -> Cow<'a, [u8]> {
|
||||||
|
if shape.is_empty() {
|
||||||
|
return Cow::Borrowed(raw_data);
|
||||||
|
}
|
||||||
|
if let Some(range) = contiguous_chunk(shape, chunk_dims, element_size, linear_idx)
|
||||||
|
&& let (Ok(start), Ok(end)) = (usize::try_from(range.start), usize::try_from(range.end))
|
||||||
|
&& end <= raw_data.len()
|
||||||
|
{
|
||||||
|
return Cow::Borrowed(&raw_data[start..end]);
|
||||||
|
}
|
||||||
|
Cow::Owned(extract_chunk(raw_data, shape, chunk_dims, element_size, linear_idx).1)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The byte range of the `linear_idx`-th chunk in the row-major dataset,
|
||||||
|
/// when the chunk is a contiguous run of it (see [`chunk_bytes_of`]).
|
||||||
|
fn contiguous_chunk(
|
||||||
|
shape: &[u64],
|
||||||
|
chunk_dims: &[u64],
|
||||||
|
element_size: usize,
|
||||||
|
linear_idx: u64,
|
||||||
|
) -> Option<core::ops::Range<u64>> {
|
||||||
|
let rank = shape.len();
|
||||||
|
// Past the first chunk dimension above 1, every chunk dimension must
|
||||||
|
// span the dataset's.
|
||||||
|
let first_wide = chunk_dims.iter().position(|&c| c > 1).unwrap_or(rank);
|
||||||
|
if (first_wide + 1..rank).any(|d| chunk_dims[d] != shape[d]) {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
let mut remaining = linear_idx;
|
||||||
|
let mut start = 0u64;
|
||||||
|
let mut stride = element_size as u64;
|
||||||
|
for d in (0..rank).rev() {
|
||||||
|
let n = shape[d].div_ceil(chunk_dims[d]);
|
||||||
|
let offset = (remaining % n).checked_mul(chunk_dims[d])?;
|
||||||
|
remaining /= n;
|
||||||
|
// A chunk past the dataset's edge is padded: not a run of it.
|
||||||
|
if offset.checked_add(chunk_dims[d])? > shape[d] {
|
||||||
|
return None;
|
||||||
|
}
|
||||||
|
start = start.checked_add(offset.checked_mul(stride)?)?;
|
||||||
|
stride = stride.checked_mul(shape[d])?;
|
||||||
|
}
|
||||||
|
let len = chunk_dims
|
||||||
|
.iter()
|
||||||
|
.try_fold(element_size as u64, |acc, &c| acc.checked_mul(c))?;
|
||||||
|
Some(start..start.checked_add(len)?)
|
||||||
|
}
|
||||||
|
|
||||||
/// Parallel compression threshold: use rayon when chunk count exceeds this.
|
/// Parallel compression threshold: use rayon when chunk count exceeds this.
|
||||||
///
|
///
|
||||||
/// Lowered to 2 to enable parallel compression for typical 4-chunk workloads
|
/// Lowered to 2 to enable parallel compression for typical 4-chunk workloads
|
||||||
@@ -508,25 +572,21 @@ const PARALLEL_COMPRESS_MAX_CHUNK_BYTES: u64 = 64 << 20;
|
|||||||
/// at most [`PARALLEL_COMPRESS_MAX_CHUNK_BYTES`], compression runs across
|
/// at most [`PARALLEL_COMPRESS_MAX_CHUNK_BYTES`], compression runs across
|
||||||
/// rayon threads; otherwise it is sequential. Output order matches chunk
|
/// rayon threads; otherwise it is sequential. Output order matches chunk
|
||||||
/// order, so per-chunk bytes are identical to the sequential path.
|
/// order, so per-chunk bytes are identical to the sequential path.
|
||||||
fn compress_all_chunks(
|
fn compress_all_chunks<'a>(
|
||||||
raw_data: &[u8],
|
raw_data: &'a [u8],
|
||||||
shape: &[u64],
|
shape: &[u64],
|
||||||
chunk_dims: &[u64],
|
chunk_dims: &[u64],
|
||||||
element_size: usize,
|
element_size: usize,
|
||||||
chunk_bytes: u64,
|
chunk_bytes: u64,
|
||||||
pipeline: &Option<FilterPipeline>,
|
pipeline: &Option<FilterPipeline>,
|
||||||
) -> Result<Vec<(u64, Vec<u8>, u32)>, FormatError> {
|
) -> Result<Vec<PreparedChunk<'a>>, FormatError> {
|
||||||
let one = |i: u64| -> Result<(u64, Vec<u8>, u32), FormatError> {
|
let one = |i: u64| -> Result<PreparedChunk<'a>, FormatError> {
|
||||||
let (_, raw) = if shape.is_empty() {
|
let raw = chunk_bytes_of(raw_data, shape, chunk_dims, element_size, i);
|
||||||
(Vec::new(), raw_data.to_vec())
|
|
||||||
} else {
|
|
||||||
extract_chunk(raw_data, shape, chunk_dims, element_size, i)
|
|
||||||
};
|
|
||||||
let raw_size = raw.len() as u64;
|
let raw_size = raw.len() as u64;
|
||||||
match pipeline {
|
match pipeline {
|
||||||
Some(pl) => {
|
Some(pl) => {
|
||||||
let (stored, mask) = compress_chunk_masked(&raw, pl, element_size as u32)?;
|
let (stored, mask) = compress_chunk_masked(&raw, pl, element_size as u32)?;
|
||||||
Ok((raw_size, stored, mask))
|
Ok((raw_size, Cow::Owned(stored), mask))
|
||||||
}
|
}
|
||||||
None => Ok((raw_size, raw, 0)),
|
None => Ok((raw_size, raw, 0)),
|
||||||
}
|
}
|
||||||
@@ -557,7 +617,7 @@ fn compress_all_chunks(
|
|||||||
/// layout/pipeline messages. `base_address` is where the blob will be placed in the file.
|
/// layout/pipeline messages. `base_address` is where the blob will be placed in the file.
|
||||||
/// Serialize a v4 single chunk layout message (public for OH size estimation).
|
/// Serialize a v4 single chunk layout message (public for OH size estimation).
|
||||||
pub fn serialize_v4_single_chunk_pub(
|
pub fn serialize_v4_single_chunk_pub(
|
||||||
chunk_dims: &[u32],
|
chunk_dims: &[u64],
|
||||||
chunk_address: u64,
|
chunk_address: u64,
|
||||||
filtered_size: Option<u64>,
|
filtered_size: Option<u64>,
|
||||||
filter_mask: Option<u32>,
|
filter_mask: Option<u32>,
|
||||||
@@ -578,7 +638,7 @@ pub fn serialize_v4_single_chunk_pub(
|
|||||||
/// Serialize a v4 (or, for a chunk of 4 GiB or more, v5) single chunk
|
/// Serialize a v4 (or, for a chunk of 4 GiB or more, v5) single chunk
|
||||||
/// layout message.
|
/// layout message.
|
||||||
fn serialize_v4_single_chunk(
|
fn serialize_v4_single_chunk(
|
||||||
chunk_dims: &[u32],
|
chunk_dims: &[u64],
|
||||||
chunk_address: u64,
|
chunk_address: u64,
|
||||||
filtered_size: Option<u64>,
|
filtered_size: Option<u64>,
|
||||||
filter_mask: Option<u32>,
|
filter_mask: Option<u32>,
|
||||||
@@ -622,7 +682,7 @@ fn serialize_v4_single_chunk(
|
|||||||
|
|
||||||
/// Serialize a v4 Fixed Array layout message.
|
/// Serialize a v4 Fixed Array layout message.
|
||||||
fn serialize_v4_fixed_array(
|
fn serialize_v4_fixed_array(
|
||||||
chunk_dims: &[u32],
|
chunk_dims: &[u64],
|
||||||
fixed_array_address: u64,
|
fixed_array_address: u64,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
@@ -653,7 +713,8 @@ fn serialize_v4_fixed_array(
|
|||||||
/// dimensions, then the element size). Each takes the fewest bytes that hold
|
/// dimensions, then the element size). Each takes the fewest bytes that hold
|
||||||
/// the largest, as libhdf5 computes it (`H5D__chunk_set_sizes`:
|
/// the largest, as libhdf5 computes it (`H5D__chunk_set_sizes`:
|
||||||
/// `(log2(dim) + 8) / 8`); HDF5 2.0.0 refuses any other width.
|
/// `(log2(dim) + 8) / 8`); HDF5 2.0.0 refuses any other width.
|
||||||
pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u32], element_size: u32) {
|
pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u64], element_size: u32) {
|
||||||
|
let element_size = u64::from(element_size);
|
||||||
let max_dim = chunk_dims
|
let max_dim = chunk_dims
|
||||||
.iter()
|
.iter()
|
||||||
.copied()
|
.copied()
|
||||||
@@ -661,14 +722,14 @@ pub(crate) fn push_v4_chunk_dims(buf: &mut Vec<u8>, chunk_dims: &[u32], element_
|
|||||||
.max()
|
.max()
|
||||||
.unwrap_or(1)
|
.unwrap_or(1)
|
||||||
.max(1);
|
.max(1);
|
||||||
let width = (32 - max_dim.leading_zeros()).div_ceil(8) as usize;
|
let width = (64 - max_dim.leading_zeros()).div_ceil(8) as usize;
|
||||||
buf.push(width as u8);
|
buf.push(width as u8);
|
||||||
for &d in chunk_dims.iter().chain(core::iter::once(&element_size)) {
|
for &d in chunk_dims.iter().chain(core::iter::once(&element_size)) {
|
||||||
buf.extend_from_slice(&d.to_le_bytes()[..width]);
|
buf.extend_from_slice(&d.to_le_bytes()[..width]);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
fn layout_v4_chunked_prefix(chunk_dims: &[u32], element_size: u32, version: u8) -> Vec<u8> {
|
fn layout_v4_chunked_prefix(chunk_dims: &[u64], element_size: u32, version: u8) -> Vec<u8> {
|
||||||
let mut buf = Vec::new();
|
let mut buf = Vec::new();
|
||||||
buf.push(version);
|
buf.push(version);
|
||||||
buf.push(2); // class = chunked
|
buf.push(2); // class = chunked
|
||||||
@@ -884,7 +945,66 @@ pub fn precompress_chunks(
|
|||||||
element_size: usize,
|
element_size: usize,
|
||||||
options: &ChunkOptions,
|
options: &ChunkOptions,
|
||||||
) -> Result<PrecompressedChunks, FormatError> {
|
) -> Result<PrecompressedChunks, FormatError> {
|
||||||
let (_, chunk_bytes) = checked_chunk_dims(chunk_dims, element_size)?;
|
let p = prepare_chunks(raw_data, shape, chunk_dims, element_size, options)?;
|
||||||
|
Ok(PrecompressedChunks {
|
||||||
|
chunks: p
|
||||||
|
.chunks
|
||||||
|
.into_iter()
|
||||||
|
.map(|(raw, stored, mask)| (raw, stored.into_owned(), mask))
|
||||||
|
.collect(),
|
||||||
|
has_filters: p.has_filters,
|
||||||
|
element_size: p.element_size,
|
||||||
|
shape: p.shape,
|
||||||
|
chunk_dims: p.chunk_dims,
|
||||||
|
pipeline_message: p.pipeline_message,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/// One chunk of [`PreparedChunks`]: its raw size, its stored bytes and its
|
||||||
|
/// filter mask.
|
||||||
|
pub(crate) type PreparedChunk<'a> = (u64, Cow<'a, [u8]>, u32);
|
||||||
|
|
||||||
|
/// [`PrecompressedChunks`] whose stored bytes may borrow the dataset's data:
|
||||||
|
/// an unfiltered chunk that is a contiguous run of it (a dataset stored as
|
||||||
|
/// one chunk, say) is not copied. What the file writer lays out.
|
||||||
|
pub(crate) struct PreparedChunks<'a> {
|
||||||
|
pub(crate) chunks: Vec<PreparedChunk<'a>>,
|
||||||
|
pub(crate) has_filters: bool,
|
||||||
|
pub(crate) element_size: usize,
|
||||||
|
pub(crate) shape: Vec<u64>,
|
||||||
|
pub(crate) chunk_dims: Vec<u64>,
|
||||||
|
pub(crate) pipeline_message: Option<Vec<u8>>,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl PrecompressedChunks {
|
||||||
|
/// A view of these chunks as [`PreparedChunks`] (their bytes borrowed).
|
||||||
|
fn prepared(&self) -> PreparedChunks<'_> {
|
||||||
|
PreparedChunks {
|
||||||
|
chunks: self
|
||||||
|
.chunks
|
||||||
|
.iter()
|
||||||
|
.map(|(raw, stored, mask)| (*raw, Cow::Borrowed(&stored[..]), *mask))
|
||||||
|
.collect(),
|
||||||
|
has_filters: self.has_filters,
|
||||||
|
element_size: self.element_size,
|
||||||
|
shape: self.shape.clone(),
|
||||||
|
chunk_dims: self.chunk_dims.clone(),
|
||||||
|
pipeline_message: self.pipeline_message.clone(),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// [`precompress_chunks`], borrowing what it can (see [`PreparedChunks`]).
|
||||||
|
/// A filtered chunk that is a contiguous run of `raw_data` is compressed
|
||||||
|
/// straight from it, without a copy of the raw chunk.
|
||||||
|
pub(crate) fn prepare_chunks<'a>(
|
||||||
|
raw_data: &'a [u8],
|
||||||
|
shape: &[u64],
|
||||||
|
chunk_dims: &[u64],
|
||||||
|
element_size: usize,
|
||||||
|
options: &ChunkOptions,
|
||||||
|
) -> Result<PreparedChunks<'a>, FormatError> {
|
||||||
|
let chunk_bytes = checked_chunk_dims(chunk_dims, element_size)?;
|
||||||
if chunk_bytes > MAX_V4_CHUNK_BYTES {
|
if chunk_bytes > MAX_V4_CHUNK_BYTES {
|
||||||
check_huge_chunk_filters(options, chunk_bytes)?;
|
check_huge_chunk_filters(options, chunk_bytes)?;
|
||||||
}
|
}
|
||||||
@@ -902,7 +1022,7 @@ pub fn precompress_chunks(
|
|||||||
&pipeline,
|
&pipeline,
|
||||||
)?;
|
)?;
|
||||||
|
|
||||||
Ok(PrecompressedChunks {
|
Ok(PreparedChunks {
|
||||||
chunks,
|
chunks,
|
||||||
has_filters,
|
has_filters,
|
||||||
element_size,
|
element_size,
|
||||||
@@ -912,26 +1032,87 @@ pub fn precompress_chunks(
|
|||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The chunk dimensions as the layout message stores them (each below
|
/// A piece of a chunked dataset's data blob as laid out by
|
||||||
/// 2^32), and one chunk's size in bytes. A chunk dimension of 2^32 or more
|
/// [`place_chunked_data`]: zero padding, a chunk's stored bytes (by index
|
||||||
/// (which HDF5 2.0 can store, in wider fields) is refused, as it is when
|
/// into [`PreparedChunks::chunks`]), or index structures.
|
||||||
/// read: clawhdf5 holds chunk dimensions as `u32`. So is a chunk whose size
|
#[derive(Debug)]
|
||||||
/// overflows 64 bits, or that this platform cannot hold in memory (a chunk
|
pub(crate) enum Piece {
|
||||||
/// of 4 GiB or more on a 32-bit target).
|
Zeros(usize),
|
||||||
fn checked_chunk_dims(
|
Chunk(usize),
|
||||||
chunk_dims: &[u64],
|
Bytes(Vec<u8>),
|
||||||
element_size: usize,
|
}
|
||||||
) -> Result<(Vec<u32>, u64), FormatError> {
|
|
||||||
let dims = chunk_dims
|
/// A chunked dataset laid out at an address, without its chunk bytes
|
||||||
.iter()
|
/// copied: the pieces of its data blob, in order, and its messages.
|
||||||
.map(|&d| {
|
pub(crate) struct ChunkedPlacement {
|
||||||
u32::try_from(d).map_err(|_| {
|
pub(crate) pieces: Vec<Piece>,
|
||||||
FormatError::InvalidChunkDimensions(format!(
|
/// Length of the data blob in bytes.
|
||||||
"chunk dimension {d} is 2^32 or more, which clawhdf5 does not support"
|
pub(crate) len: u64,
|
||||||
))
|
pub(crate) layout_message: Vec<u8>,
|
||||||
})
|
pub(crate) pipeline_message: Option<Vec<u8>>,
|
||||||
})
|
}
|
||||||
.collect::<Result<Vec<u32>, _>>()?;
|
|
||||||
|
impl ChunkedPlacement {
|
||||||
|
/// The data blob as one buffer (the chunks copied in).
|
||||||
|
fn into_result(self, pre: &PreparedChunks<'_>) -> ChunkedDataResult {
|
||||||
|
let mut data_bytes = Vec::with_capacity(saturating_usize(self.len));
|
||||||
|
for piece in &self.pieces {
|
||||||
|
match piece {
|
||||||
|
Piece::Zeros(n) => data_bytes.resize(data_bytes.len() + n, 0),
|
||||||
|
Piece::Chunk(i) => data_bytes.extend_from_slice(&pre.chunks[*i].1),
|
||||||
|
Piece::Bytes(b) => data_bytes.extend_from_slice(b),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
ChunkedDataResult {
|
||||||
|
data_bytes,
|
||||||
|
layout_message: self.layout_message,
|
||||||
|
pipeline_message: self.pipeline_message,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Builds the pieces of a data blob, tracking its length.
|
||||||
|
#[derive(Default)]
|
||||||
|
struct Pieces {
|
||||||
|
pieces: Vec<Piece>,
|
||||||
|
len: u64,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Pieces {
|
||||||
|
/// Pad to the next cache line.
|
||||||
|
fn align(&mut self) {
|
||||||
|
let aligned = align_chunk_offset(self.len);
|
||||||
|
if aligned > self.len {
|
||||||
|
self.pieces
|
||||||
|
.push(Piece::Zeros(saturating_usize(aligned - self.len)));
|
||||||
|
self.len = aligned;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn chunk(&mut self, i: usize, len: usize) {
|
||||||
|
self.pieces.push(Piece::Chunk(i));
|
||||||
|
self.len += len as u64;
|
||||||
|
}
|
||||||
|
|
||||||
|
fn bytes(&mut self, b: Vec<u8>) {
|
||||||
|
self.len += b.len() as u64;
|
||||||
|
self.pieces.push(Piece::Bytes(b));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// One chunk's size in bytes. A chunk dimension of 0 is refused, and so is
|
||||||
|
/// a chunk whose size overflows 64 bits, or that this platform cannot hold
|
||||||
|
/// in memory (a chunk of 4 GiB or more on a 32-bit target). Chunk
|
||||||
|
/// dimensions of 2^32 or more are allowed, as in libhdf5 2.x: such a chunk
|
||||||
|
/// is over 4 GiB, so it takes layout message version 5, whose dimension
|
||||||
|
/// fields are up to 8 bytes wide (a version-3 layout, with 4-byte fields,
|
||||||
|
/// is never written for it).
|
||||||
|
fn checked_chunk_dims(chunk_dims: &[u64], element_size: usize) -> Result<u64, FormatError> {
|
||||||
|
if chunk_dims.contains(&0) {
|
||||||
|
return Err(FormatError::InvalidChunkDimensions(format!(
|
||||||
|
"chunk dimensions {chunk_dims:?} include 0"
|
||||||
|
)));
|
||||||
|
}
|
||||||
let bytes = chunk_dims
|
let bytes = chunk_dims
|
||||||
.iter()
|
.iter()
|
||||||
.try_fold(element_size as u64, |acc, &d| acc.checked_mul(d))
|
.try_fold(element_size as u64, |acc, &d| acc.checked_mul(d))
|
||||||
@@ -941,7 +1122,7 @@ fn checked_chunk_dims(
|
|||||||
"a chunk of {chunk_dims:?} x {element_size} bytes exceeds this platform's address space"
|
"a chunk of {chunk_dims:?} x {element_size} bytes exceeds this platform's address space"
|
||||||
))
|
))
|
||||||
})?;
|
})?;
|
||||||
Ok((dims, bytes))
|
Ok(bytes)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// The filters clawhdf5 can apply to a chunk of more than `u32::MAX` bytes:
|
/// The filters clawhdf5 can apply to a chunk of more than `u32::MAX` bytes:
|
||||||
@@ -1010,7 +1191,22 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
low: LibVer,
|
low: LibVer,
|
||||||
high: LibVer,
|
high: LibVer,
|
||||||
) -> Result<ChunkedDataResult, FormatError> {
|
) -> Result<ChunkedDataResult, FormatError> {
|
||||||
let (_, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?;
|
let pre = pre.prepared();
|
||||||
|
let placement = place_chunked_data(&pre, base_address, maxshape, low, high)?;
|
||||||
|
Ok(placement.into_result(&pre))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// [`build_chunked_data_from_precompressed_libver`] without copying the
|
||||||
|
/// chunks into one buffer: the data blob as [`Piece`]s referring to
|
||||||
|
/// `pre`'s chunks. The file writer writes the pieces out one by one.
|
||||||
|
pub(crate) fn place_chunked_data(
|
||||||
|
pre: &PreparedChunks<'_>,
|
||||||
|
base_address: u64,
|
||||||
|
maxshape: Option<&[u64]>,
|
||||||
|
low: LibVer,
|
||||||
|
high: LibVer,
|
||||||
|
) -> Result<ChunkedPlacement, FormatError> {
|
||||||
|
let chunk_bytes = checked_chunk_dims(&pre.chunk_dims, pre.element_size)?;
|
||||||
if chunk_bytes > MAX_V4_CHUNK_BYTES {
|
if chunk_bytes > MAX_V4_CHUNK_BYTES {
|
||||||
if high < LibVer::V200 {
|
if high < LibVer::V200 {
|
||||||
return Err(FormatError::LibverBound {
|
return Err(FormatError::LibverBound {
|
||||||
@@ -1020,7 +1216,7 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
} else if low < LibVer::V110 {
|
} else if low < LibVer::V110 {
|
||||||
return build_btree_v1_chunked_data(pre, base_address, maxshape);
|
return place_btree_v1_chunked_data(pre, base_address, maxshape);
|
||||||
}
|
}
|
||||||
let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?;
|
let index = ChunkIndexPlan::new(&pre.shape, maxshape, &pre.chunk_dims)?;
|
||||||
let offset_size: u8 = 8;
|
let offset_size: u8 = 8;
|
||||||
@@ -1028,17 +1224,14 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
let num_chunks = pre.chunks.len();
|
let num_chunks = pre.chunks.len();
|
||||||
let element_size = pre.element_size;
|
let element_size = pre.element_size;
|
||||||
|
|
||||||
let mut data_buf = Vec::new();
|
let mut blob = Pieces::default();
|
||||||
let mut written_chunks = Vec::with_capacity(num_chunks);
|
let mut written_chunks = Vec::with_capacity(num_chunks);
|
||||||
|
|
||||||
for (raw_size, compressed, filter_mask) in &pre.chunks {
|
for (i, (raw_size, compressed, filter_mask)) in pre.chunks.iter().enumerate() {
|
||||||
let aligned_offset = align_to_cache_line(data_buf.len());
|
blob.align();
|
||||||
if aligned_offset > data_buf.len() {
|
let address = base_address + blob.len;
|
||||||
data_buf.resize(aligned_offset, 0u8);
|
|
||||||
}
|
|
||||||
let address = base_address + data_buf.len() as u64;
|
|
||||||
let compressed_size = compressed.len() as u64;
|
let compressed_size = compressed.len() as u64;
|
||||||
data_buf.extend_from_slice(compressed);
|
blob.chunk(i, compressed.len());
|
||||||
written_chunks.push(WrittenChunk {
|
written_chunks.push(WrittenChunk {
|
||||||
address,
|
address,
|
||||||
compressed_size,
|
compressed_size,
|
||||||
@@ -1047,28 +1240,22 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
let (chunk_dims_u32, chunk_bytes) = checked_chunk_dims(&pre.chunk_dims, element_size)?;
|
|
||||||
let version = layout_version_for(chunk_bytes);
|
let version = layout_version_for(chunk_bytes);
|
||||||
|
blob.align();
|
||||||
let aligned_idx = align_to_cache_line(data_buf.len());
|
|
||||||
if aligned_idx > data_buf.len() {
|
|
||||||
data_buf.resize(aligned_idx, 0u8);
|
|
||||||
}
|
|
||||||
|
|
||||||
let layout_message = match &index {
|
let layout_message = match &index {
|
||||||
ChunkIndexPlan::ExtensibleArray(grid) => {
|
ChunkIndexPlan::ExtensibleArray(grid) => {
|
||||||
let ea_address = base_address + data_buf.len() as u64;
|
let ea_address = base_address + blob.len;
|
||||||
let slots = index_slots(grid, &pre.shape, &pre.chunk_dims, &written_chunks, None)?;
|
let slots = index_slots(grid, &pre.shape, &pre.chunk_dims, &written_chunks, None)?;
|
||||||
let ea_bytes = ea_writer::build_extensible_array_at(
|
blob.bytes(ea_writer::build_extensible_array_at(
|
||||||
&slots,
|
&slots,
|
||||||
offset_size,
|
offset_size,
|
||||||
length_size,
|
length_size,
|
||||||
pre.has_filters,
|
pre.has_filters,
|
||||||
ea_address,
|
ea_address,
|
||||||
);
|
));
|
||||||
data_buf.extend_from_slice(&ea_bytes);
|
|
||||||
ea_writer::serialize_v4_extensible_array(
|
ea_writer::serialize_v4_extensible_array(
|
||||||
&chunk_dims_u32,
|
&pre.chunk_dims,
|
||||||
ea_address,
|
ea_address,
|
||||||
offset_size,
|
offset_size,
|
||||||
element_size as u32,
|
element_size as u32,
|
||||||
@@ -1084,7 +1271,7 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
};
|
};
|
||||||
let filter_mask = pre.has_filters.then_some(written_chunks[0].filter_mask);
|
let filter_mask = pre.has_filters.then_some(written_chunks[0].filter_mask);
|
||||||
serialize_v4_single_chunk(
|
serialize_v4_single_chunk(
|
||||||
&chunk_dims_u32,
|
&pre.chunk_dims,
|
||||||
chunk_addr,
|
chunk_addr,
|
||||||
filtered_size,
|
filtered_size,
|
||||||
filter_mask,
|
filter_mask,
|
||||||
@@ -1094,7 +1281,7 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
)
|
)
|
||||||
}
|
}
|
||||||
ChunkIndexPlan::FixedArray(grid, nslots) => {
|
ChunkIndexPlan::FixedArray(grid, nslots) => {
|
||||||
let fa_address = base_address + data_buf.len() as u64;
|
let fa_address = base_address + blob.len;
|
||||||
let slots = index_slots(
|
let slots = index_slots(
|
||||||
grid,
|
grid,
|
||||||
&pre.shape,
|
&pre.shape,
|
||||||
@@ -1102,16 +1289,15 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
&written_chunks,
|
&written_chunks,
|
||||||
Some(*nslots),
|
Some(*nslots),
|
||||||
)?;
|
)?;
|
||||||
let fa_bytes = build_fixed_array_at(
|
blob.bytes(build_fixed_array_at(
|
||||||
&slots,
|
&slots,
|
||||||
offset_size,
|
offset_size,
|
||||||
length_size,
|
length_size,
|
||||||
pre.has_filters,
|
pre.has_filters,
|
||||||
fa_address,
|
fa_address,
|
||||||
);
|
));
|
||||||
data_buf.extend_from_slice(&fa_bytes);
|
|
||||||
serialize_v4_fixed_array(
|
serialize_v4_fixed_array(
|
||||||
&chunk_dims_u32,
|
&pre.chunk_dims,
|
||||||
fa_address,
|
fa_address,
|
||||||
offset_size,
|
offset_size,
|
||||||
element_size as u32,
|
element_size as u32,
|
||||||
@@ -1120,7 +1306,7 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
)
|
)
|
||||||
}
|
}
|
||||||
ChunkIndexPlan::BTreeV2 => {
|
ChunkIndexPlan::BTreeV2 => {
|
||||||
let bt_address = base_address + data_buf.len() as u64;
|
let bt_address = base_address + blob.len;
|
||||||
let records: Vec<(Vec<u64>, &WrittenChunk)> = written_chunks
|
let records: Vec<(Vec<u64>, &WrittenChunk)> = written_chunks
|
||||||
.iter()
|
.iter()
|
||||||
.enumerate()
|
.enumerate()
|
||||||
@@ -1134,9 +1320,9 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
pre.has_filters,
|
pre.has_filters,
|
||||||
bt_address,
|
bt_address,
|
||||||
)?;
|
)?;
|
||||||
data_buf.extend_from_slice(&bt_bytes);
|
blob.bytes(bt_bytes);
|
||||||
serialize_v4_btree_v2(
|
serialize_v4_btree_v2(
|
||||||
&chunk_dims_u32,
|
&pre.chunk_dims,
|
||||||
bt_address,
|
bt_address,
|
||||||
offset_size,
|
offset_size,
|
||||||
element_size as u32,
|
element_size as u32,
|
||||||
@@ -1146,8 +1332,9 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
}
|
}
|
||||||
};
|
};
|
||||||
|
|
||||||
Ok(ChunkedDataResult {
|
Ok(ChunkedPlacement {
|
||||||
data_bytes: data_buf,
|
pieces: blob.pieces,
|
||||||
|
len: blob.len,
|
||||||
layout_message,
|
layout_message,
|
||||||
pipeline_message: pre.pipeline_message.clone(),
|
pipeline_message: pre.pipeline_message.clone(),
|
||||||
})
|
})
|
||||||
@@ -1156,11 +1343,11 @@ pub fn build_chunked_data_from_precompressed_libver(
|
|||||||
/// Lay out precompressed chunks at `base_address` followed by a version-1
|
/// Lay out precompressed chunks at `base_address` followed by a version-1
|
||||||
/// B-tree chunk index, with a version-3 layout message: what libhdf5 writes
|
/// B-tree chunk index, with a version-3 layout message: what libhdf5 writes
|
||||||
/// for a chunked dataset under a low bound of 1.8.
|
/// for a chunked dataset under a low bound of 1.8.
|
||||||
fn build_btree_v1_chunked_data(
|
fn place_btree_v1_chunked_data(
|
||||||
pre: &PrecompressedChunks,
|
pre: &PreparedChunks<'_>,
|
||||||
base_address: u64,
|
base_address: u64,
|
||||||
maxshape: Option<&[u64]>,
|
maxshape: Option<&[u64]>,
|
||||||
) -> Result<ChunkedDataResult, FormatError> {
|
) -> Result<ChunkedPlacement, FormatError> {
|
||||||
if let Some(ms) = maxshape {
|
if let Some(ms) = maxshape {
|
||||||
let bad = |what: &str| FormatError::ChunkedReadError(format!("maxshape: {what}"));
|
let bad = |what: &str| FormatError::ChunkedReadError(format!("maxshape: {what}"));
|
||||||
if ms.len() != pre.shape.len() {
|
if ms.len() != pre.shape.len() {
|
||||||
@@ -1171,20 +1358,17 @@ fn build_btree_v1_chunked_data(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
let offset_size: u8 = 8;
|
let offset_size: u8 = 8;
|
||||||
let mut data_buf = Vec::new();
|
let mut blob = Pieces::default();
|
||||||
let mut entries = Vec::with_capacity(pre.chunks.len());
|
let mut entries = Vec::with_capacity(pre.chunks.len());
|
||||||
for (i, (_raw_size, stored, filter_mask)) in pre.chunks.iter().enumerate() {
|
for (i, (_raw_size, stored, filter_mask)) in pre.chunks.iter().enumerate() {
|
||||||
let aligned_offset = align_to_cache_line(data_buf.len());
|
blob.align();
|
||||||
if aligned_offset > data_buf.len() {
|
|
||||||
data_buf.resize(aligned_offset, 0u8);
|
|
||||||
}
|
|
||||||
entries.push(btree_v1_write::ChunkEntry {
|
entries.push(btree_v1_write::ChunkEntry {
|
||||||
scaled: scaled_coords(&pre.shape, &pre.chunk_dims, i),
|
scaled: scaled_coords(&pre.shape, &pre.chunk_dims, i),
|
||||||
nbytes: stored.len() as u64,
|
nbytes: stored.len() as u64,
|
||||||
filter_mask: *filter_mask,
|
filter_mask: *filter_mask,
|
||||||
address: base_address + data_buf.len() as u64,
|
address: base_address + blob.len,
|
||||||
});
|
});
|
||||||
data_buf.extend_from_slice(stored);
|
blob.chunk(i, stored.len());
|
||||||
}
|
}
|
||||||
let element_size = u32::try_from(pre.element_size)
|
let element_size = u32::try_from(pre.element_size)
|
||||||
.map_err(|_| FormatError::Overflow("element size".into()))?;
|
.map_err(|_| FormatError::Overflow("element size".into()))?;
|
||||||
@@ -1193,11 +1377,8 @@ fn build_btree_v1_chunked_data(
|
|||||||
let btree_address = if entries.is_empty() {
|
let btree_address = if entries.is_empty() {
|
||||||
u64::MAX
|
u64::MAX
|
||||||
} else {
|
} else {
|
||||||
let aligned_idx = align_to_cache_line(data_buf.len());
|
blob.align();
|
||||||
if aligned_idx > data_buf.len() {
|
let addr = base_address + blob.len;
|
||||||
data_buf.resize(aligned_idx, 0u8);
|
|
||||||
}
|
|
||||||
let addr = base_address + data_buf.len() as u64;
|
|
||||||
let tree = btree_v1_write::build_chunk_btree_v1_at(
|
let tree = btree_v1_write::build_chunk_btree_v1_at(
|
||||||
&entries,
|
&entries,
|
||||||
&pre.chunk_dims,
|
&pre.chunk_dims,
|
||||||
@@ -1205,13 +1386,14 @@ fn build_btree_v1_chunked_data(
|
|||||||
addr,
|
addr,
|
||||||
offset_size,
|
offset_size,
|
||||||
)?;
|
)?;
|
||||||
data_buf.extend_from_slice(&tree);
|
blob.bytes(tree);
|
||||||
addr
|
addr
|
||||||
};
|
};
|
||||||
let layout_message =
|
let layout_message =
|
||||||
serialize_v3_chunked(&pre.chunk_dims, btree_address, offset_size, element_size)?;
|
serialize_v3_chunked(&pre.chunk_dims, btree_address, offset_size, element_size)?;
|
||||||
Ok(ChunkedDataResult {
|
Ok(ChunkedPlacement {
|
||||||
data_bytes: data_buf,
|
pieces: blob.pieces,
|
||||||
|
len: blob.len,
|
||||||
layout_message,
|
layout_message,
|
||||||
pipeline_message: pre.pipeline_message.clone(),
|
pipeline_message: pre.pipeline_message.clone(),
|
||||||
})
|
})
|
||||||
@@ -1424,7 +1606,7 @@ fn build_btree_v2_chunk_index_at(
|
|||||||
|
|
||||||
/// Serialize a v4 layout message for a version-2 B-tree chunk index.
|
/// Serialize a v4 layout message for a version-2 B-tree chunk index.
|
||||||
fn serialize_v4_btree_v2(
|
fn serialize_v4_btree_v2(
|
||||||
chunk_dims: &[u32],
|
chunk_dims: &[u64],
|
||||||
btree_address: u64,
|
btree_address: u64,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
@@ -2184,7 +2366,7 @@ mod tests {
|
|||||||
/// filtered index element stores the chunk's size in 8 bytes.
|
/// filtered index element stores the chunk's size in 8 bytes.
|
||||||
#[test]
|
#[test]
|
||||||
fn huge_chunk_layout_messages_and_index_elements() {
|
fn huge_chunk_layout_messages_and_index_elements() {
|
||||||
let dims = [HUGE_DIM as u32];
|
let dims = [HUGE_DIM];
|
||||||
let parsed = |msg: &[u8]| {
|
let parsed = |msg: &[u8]| {
|
||||||
assert_eq!(msg[0], 5, "layout message version");
|
assert_eq!(msg[0], 5, "layout message version");
|
||||||
// Class chunked, then (after the flags) 2 dimensions of 4 bytes.
|
// Class chunked, then (after the flags) 2 dimensions of 4 bytes.
|
||||||
@@ -2197,7 +2379,7 @@ mod tests {
|
|||||||
chunk_index_type,
|
chunk_index_type,
|
||||||
..
|
..
|
||||||
} => {
|
} => {
|
||||||
assert_eq!(chunk_dimensions, vec![HUGE_DIM as u32, 8]);
|
assert_eq!(chunk_dimensions, vec![HUGE_DIM, 8]);
|
||||||
chunk_index_type.unwrap()
|
chunk_index_type.unwrap()
|
||||||
}
|
}
|
||||||
other => panic!("{other:?}"),
|
other => panic!("{other:?}"),
|
||||||
@@ -2245,22 +2427,89 @@ mod tests {
|
|||||||
assert_eq!(u16::from_le_bytes([bt[10], bt[11]]), 28);
|
assert_eq!(u16::from_le_bytes([bt[10], bt[11]]), 28);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// A chunk dimension of 2^32 or more (a 1-byte element, 2^32 + 7 per
|
||||||
|
/// chunk): the dimensions take 5 bytes each, as libhdf5 2.x writes
|
||||||
|
/// them (`(log2(dim) + 8) / 8`), and parse back whole.
|
||||||
|
#[test]
|
||||||
|
fn chunk_dimensions_of_2_pow_32_or_more_take_5_bytes() {
|
||||||
|
const DIM: u64 = (1 << 32) + 7;
|
||||||
|
let msg = serialize_v4_single_chunk(&[DIM], 0x800, None, None, 8, 1, 5);
|
||||||
|
assert_eq!(&msg[..5], &[5, 2, 0, 2, 5]);
|
||||||
|
assert_eq!(&msg[5..10], &[7, 0, 0, 0, 1]);
|
||||||
|
assert_eq!(&msg[10..15], &[1, 0, 0, 0, 0]);
|
||||||
|
assert_eq!(msg[15], 1, "Single Chunk index");
|
||||||
|
match DataLayout::parse(&msg, 8, 8).unwrap() {
|
||||||
|
DataLayout::Chunked {
|
||||||
|
chunk_dimensions, ..
|
||||||
|
} => assert_eq!(chunk_dimensions, [DIM, 1]),
|
||||||
|
other => panic!("{other:?}"),
|
||||||
|
}
|
||||||
|
// The largest dimension takes all 8 bytes.
|
||||||
|
let mut buf = Vec::new();
|
||||||
|
push_v4_chunk_dims(&mut buf, &[u64::MAX, 3], 4);
|
||||||
|
assert_eq!(buf[0], 8);
|
||||||
|
assert_eq!(buf.len(), 1 + 3 * 8);
|
||||||
|
// A version-3 layout has 4-byte dimensions: refused, not truncated.
|
||||||
|
assert!(serialize_v3_chunked(&[DIM], 0x800, 8, 1).is_err());
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Which chunks are contiguous runs of the dataset's data (borrowed,
|
||||||
|
/// not copied, when unfiltered).
|
||||||
|
#[test]
|
||||||
|
fn contiguous_chunks_are_borrowed() {
|
||||||
|
// Rows of a 2-D dataset: chunk (2, 5) of shape (5, 5).
|
||||||
|
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 0), Some(0..40));
|
||||||
|
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 1), Some(40..80));
|
||||||
|
// The last row chunk runs past the edge: padded, so copied.
|
||||||
|
assert_eq!(contiguous_chunk(&[5, 5], &[2, 5], 4, 2), None);
|
||||||
|
// Columns split: not contiguous.
|
||||||
|
assert_eq!(contiguous_chunk(&[4, 6], &[2, 3], 1, 0), None);
|
||||||
|
// Leading dimensions of 1, then a partial row: contiguous.
|
||||||
|
assert_eq!(contiguous_chunk(&[2, 3, 8], &[1, 1, 4], 2, 5), Some(40..48));
|
||||||
|
// The whole dataset as one chunk.
|
||||||
|
assert_eq!(contiguous_chunk(&[3, 4], &[3, 4], 8, 0), Some(0..96));
|
||||||
|
|
||||||
|
let raw: Vec<u8> = (0..100).collect();
|
||||||
|
let whole = chunk_bytes_of(&raw, &[10, 10], &[10, 10], 1, 0);
|
||||||
|
assert!(matches!(whole, Cow::Borrowed(b) if b.as_ptr() == raw.as_ptr()));
|
||||||
|
let rows = chunk_bytes_of(&raw, &[10, 10], &[3, 10], 1, 1);
|
||||||
|
assert!(matches!(rows, Cow::Borrowed(b) if b == &raw[30..60]));
|
||||||
|
// Data shorter than the shape is copied (and zero-padded) as before.
|
||||||
|
let short = chunk_bytes_of(&raw[..50], &[10, 10], &[10, 10], 1, 0);
|
||||||
|
assert!(matches!(&short, Cow::Owned(v) if v.len() == 100));
|
||||||
|
// Borrowed or copied, the bytes are what extract_chunk gives.
|
||||||
|
for (shape, chunk) in [
|
||||||
|
(&[10u64, 10][..], &[3u64, 10][..]),
|
||||||
|
(&[10, 10], &[3, 4]),
|
||||||
|
(&[4, 5, 5], &[1, 5, 5]),
|
||||||
|
(&[4, 5, 5], &[1, 2, 5]),
|
||||||
|
] {
|
||||||
|
for i in 0..chunk_count(shape, chunk) {
|
||||||
|
let got = chunk_bytes_of(&raw, shape, chunk, 1, i);
|
||||||
|
assert_eq!(&got[..], &extract_chunk(&raw, shape, chunk, 1, i).1[..]);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn huge_chunks_refused_where_unsupported() {
|
fn huge_chunks_refused_where_unsupported() {
|
||||||
// A chunk dimension of 2^32 or more.
|
// A chunk dimension of 0.
|
||||||
assert!(matches!(
|
assert!(matches!(
|
||||||
checked_chunk_dims(&[1 << 32], 1),
|
checked_chunk_dims(&[4, 0], 1),
|
||||||
Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32")
|
Err(FormatError::InvalidChunkDimensions(_))
|
||||||
));
|
));
|
||||||
|
// A chunk dimension of 2^32 or more is fine (on a 64-bit target).
|
||||||
|
#[cfg(target_pointer_width = "64")]
|
||||||
|
assert_eq!(
|
||||||
|
checked_chunk_dims(&[(1 << 32) + 5], 1).unwrap(),
|
||||||
|
(1 << 32) + 5
|
||||||
|
);
|
||||||
// A chunk size that overflows u64.
|
// A chunk size that overflows u64.
|
||||||
assert!(matches!(
|
assert!(matches!(
|
||||||
checked_chunk_dims(&[u32::MAX.into(), u32::MAX.into(), 2], 8),
|
checked_chunk_dims(&[u32::MAX.into(), u32::MAX.into(), 2], 8),
|
||||||
Err(FormatError::Overflow(_))
|
Err(FormatError::Overflow(_))
|
||||||
));
|
));
|
||||||
assert_eq!(
|
assert_eq!(checked_chunk_dims(&[HUGE_DIM], 8).unwrap(), HUGE_BYTES);
|
||||||
checked_chunk_dims(&[HUGE_DIM], 8).unwrap(),
|
|
||||||
(vec![HUGE_DIM as u32], HUGE_BYTES)
|
|
||||||
);
|
|
||||||
// Filters that cannot take a chunk that large.
|
// Filters that cannot take a chunk that large.
|
||||||
for plugin in [
|
for plugin in [
|
||||||
PluginFilter::Lzf,
|
PluginFilter::Lzf,
|
||||||
|
|||||||
@@ -34,7 +34,7 @@ const MAX_LAYOUT_NDIMS: usize = 33;
|
|||||||
/// (`H5O__layout_decode`): at most [`MAX_LAYOUT_NDIMS`], no dimension 0, and
|
/// (`H5O__layout_decode`): at most [`MAX_LAYOUT_NDIMS`], no dimension 0, and
|
||||||
/// before version 4 at least one dataspace dimension plus the element size.
|
/// before version 4 at least one dataspace dimension plus the element size.
|
||||||
/// A zero chunk dimension used to read the dataset as all fill values.
|
/// A zero chunk dimension used to read the dataset as all fill values.
|
||||||
fn check_chunk_dims(dims: Vec<u32>, layout_version: u8) -> Result<Vec<u32>, FormatError> {
|
fn check_chunk_dims(dims: Vec<u64>, layout_version: u8) -> Result<Vec<u64>, FormatError> {
|
||||||
if dims.len() > MAX_LAYOUT_NDIMS {
|
if dims.len() > MAX_LAYOUT_NDIMS {
|
||||||
return Err(FormatError::InvalidChunkDimensions(
|
return Err(FormatError::InvalidChunkDimensions(
|
||||||
"dimensionality is too large".into(),
|
"dimensionality is too large".into(),
|
||||||
@@ -72,7 +72,7 @@ pub enum DataLayout {
|
|||||||
/// Chunked: data stored in chunks via a B-tree.
|
/// Chunked: data stored in chunks via a B-tree.
|
||||||
Chunked {
|
Chunked {
|
||||||
/// Chunk dimension sizes.
|
/// Chunk dimension sizes.
|
||||||
chunk_dimensions: Vec<u32>,
|
chunk_dimensions: Vec<u64>,
|
||||||
/// B-tree address, or `None` if undefined.
|
/// B-tree address, or `None` if undefined.
|
||||||
btree_address: Option<u64>,
|
btree_address: Option<u64>,
|
||||||
/// Layout version (3 or 4). Version 1/2 messages (HDF5 1.4/1.6-era)
|
/// Layout version (3 or 4). Version 1/2 messages (HDF5 1.4/1.6-era)
|
||||||
@@ -404,11 +404,11 @@ impl DataLayout {
|
|||||||
_ => return Err(FormatError::InvalidLayoutClass(layout_class)),
|
_ => return Err(FormatError::InvalidLayoutClass(layout_class)),
|
||||||
};
|
};
|
||||||
ensure_len(data, p, dimensionality * 4)?;
|
ensure_len(data, p, dimensionality * 4)?;
|
||||||
let dims: Vec<u32> = data[p..p + dimensionality * 4]
|
let dims: Vec<u64> = data[p..p + dimensionality * 4]
|
||||||
.as_chunks::<4>()
|
.as_chunks::<4>()
|
||||||
.0
|
.0
|
||||||
.iter()
|
.iter()
|
||||||
.map(|c| u32::from_le_bytes(*c))
|
.map(|c| u64::from(u32::from_le_bytes(*c)))
|
||||||
.collect();
|
.collect();
|
||||||
p += dimensionality * 4;
|
p += dimensionality * 4;
|
||||||
match layout_class {
|
match layout_class {
|
||||||
@@ -424,7 +424,7 @@ impl DataLayout {
|
|||||||
1 => {
|
1 => {
|
||||||
let size = dims
|
let size = dims
|
||||||
.iter()
|
.iter()
|
||||||
.try_fold(1u64, |acc, &d| acc.checked_mul(d as u64))
|
.try_fold(1u64, |acc, &d| acc.checked_mul(d))
|
||||||
.ok_or_else(|| {
|
.ok_or_else(|| {
|
||||||
FormatError::Overflow(format!("contiguous layout size {dims:?}"))
|
FormatError::Overflow(format!("contiguous layout size {dims:?}"))
|
||||||
})?;
|
})?;
|
||||||
@@ -489,7 +489,12 @@ impl DataLayout {
|
|||||||
ensure_len(data, p, dimensionality * 4)?;
|
ensure_len(data, p, dimensionality * 4)?;
|
||||||
let mut chunk_dimensions = Vec::with_capacity(dimensionality);
|
let mut chunk_dimensions = Vec::with_capacity(dimensionality);
|
||||||
for _ in 0..dimensionality {
|
for _ in 0..dimensionality {
|
||||||
let dim = u32::from_le_bytes([data[p], data[p + 1], data[p + 2], data[p + 3]]);
|
let dim = u64::from(u32::from_le_bytes([
|
||||||
|
data[p],
|
||||||
|
data[p + 1],
|
||||||
|
data[p + 2],
|
||||||
|
data[p + 3],
|
||||||
|
]));
|
||||||
chunk_dimensions.push(dim);
|
chunk_dimensions.push(dim);
|
||||||
p += 4;
|
p += 4;
|
||||||
}
|
}
|
||||||
@@ -565,14 +570,8 @@ impl DataLayout {
|
|||||||
.iter()
|
.iter()
|
||||||
.rev()
|
.rev()
|
||||||
.fold(0u64, |acc, &b| (acc << 8) | u64::from(b));
|
.fold(0u64, |acc, &b| (acc << 8) | u64::from(b));
|
||||||
// Chunk dimensions are held as u32; HDF5 2.0 can write
|
// HDF5 2.0 writes dimensions of 2^32 or more (layout
|
||||||
// larger ones (layout version 5), which are refused
|
// version 5, chunks over 4 GiB) in 5 to 8 bytes.
|
||||||
// rather than truncated.
|
|
||||||
let val = u32::try_from(val).map_err(|_| {
|
|
||||||
FormatError::InvalidChunkDimensions(format!(
|
|
||||||
"chunk dimension {val} is larger than 2^32 - 1, which is not supported"
|
|
||||||
))
|
|
||||||
})?;
|
|
||||||
chunk_dimensions.push(val);
|
chunk_dimensions.push(val);
|
||||||
p += dim_size_encoded_length;
|
p += dim_size_encoded_length;
|
||||||
}
|
}
|
||||||
@@ -858,7 +857,7 @@ mod tests {
|
|||||||
.unwrap_or_else(|e| panic!("width {width}: {e:?}"));
|
.unwrap_or_else(|e| panic!("width {width}: {e:?}"));
|
||||||
assert!(
|
assert!(
|
||||||
matches!(&layout, DataLayout::Chunked { chunk_dimensions, .. }
|
matches!(&layout, DataLayout::Chunked { chunk_dimensions, .. }
|
||||||
if chunk_dimensions.iter().map(|&d| u64::from(d)).eq(dims)),
|
if chunk_dimensions[..] == dims),
|
||||||
"width {width}: {layout:?}"
|
"width {width}: {layout:?}"
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
@@ -871,12 +870,13 @@ mod tests {
|
|||||||
)
|
)
|
||||||
);
|
);
|
||||||
}
|
}
|
||||||
// A dimension past u32 cannot be represented and is refused, not
|
// HDF5 2.0 writes dimensions of 2^32 or more (up to 8 bytes each).
|
||||||
// truncated.
|
for (width, dim) in [(5u8, (1u64 << 32) + 3), (8, u64::MAX)] {
|
||||||
assert!(matches!(
|
assert!(matches!(
|
||||||
DataLayout::parse(&v4_chunked_msg(5, &[1 << 32, 8]), 8, 8),
|
DataLayout::parse(&v4_chunked_msg(width, &[dim, 1]), 8, 8),
|
||||||
Err(FormatError::InvalidChunkDimensions(m)) if m.contains("2^32")
|
Ok(DataLayout::Chunked { chunk_dimensions, .. }) if chunk_dimensions == [dim, 1]
|
||||||
));
|
));
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
|
|||||||
@@ -145,13 +145,16 @@ pub enum Datatype {
|
|||||||
/// imaginary, in rectangular form. `size` is twice the base size and the
|
/// imaginary, in rectangular form. `size` is twice the base size and the
|
||||||
/// base is an IEEE float.
|
/// base is an IEEE float.
|
||||||
///
|
///
|
||||||
/// This variant exists for **writing** (see
|
/// [`Datatype::parse`] returns this variant for every class-11 message,
|
||||||
|
/// also inside compounds, arrays and variable-length types. Readers that
|
||||||
|
/// want the member view (`{r, i}`, as h5py writes complex numbers by
|
||||||
|
/// default) use [`Datatype::complex_as_compound`]; `read_compound_fields`
|
||||||
|
/// accepts this variant directly.
|
||||||
|
///
|
||||||
|
/// Writing it is opt-in (see
|
||||||
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
|
/// `type_builders::make_native_complex_f64_type`): only libhdf5 2.0 and
|
||||||
/// newer can read class 11, so it is opt-in and h5py's compound `{r, i}`
|
/// newer can read class 11, so h5py's compound `{r, i}` stays the
|
||||||
/// stays the default complex encoding. [`Datatype::parse`] still
|
/// default complex encoding.
|
||||||
/// surfaces a class-11 message as that equivalent `{r, i}` compound, so
|
|
||||||
/// every compound reader handles both encodings; parsing what this
|
|
||||||
/// variant serializes therefore yields a `Compound`, not a `Complex`.
|
|
||||||
Complex { size: u32, base_type: Box<Datatype> },
|
Complex { size: u32, base_type: Box<Datatype> },
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -857,9 +860,7 @@ impl Datatype {
|
|||||||
// Complex number (HDF5 2.0, datatype version 5). The properties
|
// Complex number (HDF5 2.0, datatype version 5). The properties
|
||||||
// are a single base floating-point datatype message; an element
|
// are a single base floating-point datatype message; an element
|
||||||
// is two consecutive base-type values (real, imaginary). There
|
// is two consecutive base-type values (real, imaginary). There
|
||||||
// is no member list. Surface it as the equivalent two-member
|
// is no member list.
|
||||||
// compound `{r, i}` — the same shape h5py writes for numpy
|
|
||||||
// complex dtypes — so downstream compound readers work as-is.
|
|
||||||
if version != 5 {
|
if version != 5 {
|
||||||
return Err(FormatError::InvalidDatatypeVersion {
|
return Err(FormatError::InvalidDatatypeVersion {
|
||||||
class: class_id,
|
class: class_id,
|
||||||
@@ -875,7 +876,13 @@ impl Datatype {
|
|||||||
actual: size as usize,
|
actual: size as usize,
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
Ok((Self::complex_as_compound(size, &base_type), pos))
|
Ok((
|
||||||
|
Datatype::Complex {
|
||||||
|
size,
|
||||||
|
base_type: Box::new(base_type),
|
||||||
|
},
|
||||||
|
pos,
|
||||||
|
))
|
||||||
}
|
}
|
||||||
_ => Err(FormatError::InvalidDatatypeClass(class_id)),
|
_ => Err(FormatError::InvalidDatatypeClass(class_id)),
|
||||||
}
|
}
|
||||||
@@ -1174,9 +1181,8 @@ impl Datatype {
|
|||||||
|
|
||||||
/// The `{r, i}` compound equivalent to a native complex type of `size`
|
/// The `{r, i}` compound equivalent to a native complex type of `size`
|
||||||
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the
|
/// bytes over `base_type`: `r` at offset 0, `i` right after it — the
|
||||||
/// shape h5py writes for numpy complex dtypes, and what [`Self::parse`]
|
/// shape h5py writes for numpy complex dtypes. Readers that want a
|
||||||
/// returns for a class-11 message. Readers that meet a
|
/// [`Datatype::Complex`] as members handle it through this view.
|
||||||
/// [`Datatype::Complex`] handle it through this view.
|
|
||||||
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
|
pub fn complex_as_compound(size: u32, base_type: &Datatype) -> Datatype {
|
||||||
let base_size = base_type.type_size();
|
let base_size = base_type.type_size();
|
||||||
Datatype::Compound {
|
Datatype::Compound {
|
||||||
@@ -1849,21 +1855,24 @@ mod tests {
|
|||||||
fn test_complex_v5_from_hdf5_2_0() {
|
fn test_complex_v5_from_hdf5_2_0() {
|
||||||
let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap();
|
let (dt, consumed) = Datatype::parse(&COMPLEX_F64_HDF5_2_0).unwrap();
|
||||||
assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len());
|
assert_eq!(consumed, COMPLEX_F64_HDF5_2_0.len());
|
||||||
match dt {
|
match &dt {
|
||||||
Datatype::Compound { size, members } => {
|
Datatype::Complex { size, base_type } => {
|
||||||
assert_eq!(size, 16);
|
assert_eq!(*size, 16);
|
||||||
assert_eq!(members.len(), 2);
|
assert!(matches!(
|
||||||
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
|
base_type.as_ref(),
|
||||||
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
|
Datatype::FloatingPoint { size: 8, .. }
|
||||||
for m in &members {
|
));
|
||||||
assert!(matches!(
|
|
||||||
m.datatype,
|
|
||||||
Datatype::FloatingPoint { size: 8, .. }
|
|
||||||
));
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
other => panic!("expected Compound, got {other:?}"),
|
other => panic!("expected Complex, got {other:?}"),
|
||||||
}
|
}
|
||||||
|
// The member view is h5py's `{r, i}` compound.
|
||||||
|
let Datatype::Compound { members, .. } =
|
||||||
|
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
|
||||||
|
else {
|
||||||
|
unreachable!()
|
||||||
|
};
|
||||||
|
assert_eq!((members[0].name.as_str(), members[0].byte_offset), ("r", 0));
|
||||||
|
assert_eq!((members[1].name.as_str(), members[1].byte_offset), ("i", 8));
|
||||||
}
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
@@ -1887,7 +1896,7 @@ mod tests {
|
|||||||
assert_eq!(members.len(), 2);
|
assert_eq!(members.len(), 2);
|
||||||
assert!(matches!(
|
assert!(matches!(
|
||||||
&members[0].datatype,
|
&members[0].datatype,
|
||||||
Datatype::Compound { size: 16, members } if members.len() == 2
|
Datatype::Complex { size: 16, .. }
|
||||||
));
|
));
|
||||||
assert_eq!(
|
assert_eq!(
|
||||||
(members[1].name.as_str(), members[1].byte_offset),
|
(members[1].name.as_str(), members[1].byte_offset),
|
||||||
@@ -1918,12 +1927,9 @@ mod tests {
|
|||||||
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
|
assert_eq!(dt.serialize(), COMPLEX_F64_HDF5_2_0);
|
||||||
assert_eq!(dt.type_size(), 16);
|
assert_eq!(dt.type_size(), 16);
|
||||||
dt.check_encodable().unwrap();
|
dt.check_encodable().unwrap();
|
||||||
// Parsing surfaces class 11 as the equivalent `{r, i}` compound.
|
// Parsing returns the native complex type unchanged.
|
||||||
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
|
let (parsed, _) = Datatype::parse(&dt.serialize()).unwrap();
|
||||||
assert_eq!(
|
assert_eq!(parsed, dt);
|
||||||
parsed,
|
|
||||||
Datatype::complex_as_compound(16, &crate::type_builders::make_f64_type())
|
|
||||||
);
|
|
||||||
|
|
||||||
let f32c = make_native_complex_f32_type().serialize();
|
let f32c = make_native_complex_f32_type().serialize();
|
||||||
assert_eq!(
|
assert_eq!(
|
||||||
|
|||||||
@@ -14,7 +14,7 @@ use crate::chunked_write::{
|
|||||||
|
|
||||||
/// Serialize a v4 Extensible Array layout message.
|
/// Serialize a v4 Extensible Array layout message.
|
||||||
pub(crate) fn serialize_v4_extensible_array(
|
pub(crate) fn serialize_v4_extensible_array(
|
||||||
chunk_dims: &[u32],
|
chunk_dims: &[u64],
|
||||||
ea_address: u64,
|
ea_address: u64,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
|
|||||||
@@ -454,7 +454,7 @@ pub fn read_extensible_array_chunks(
|
|||||||
header: &ExtensibleArrayHeader,
|
header: &ExtensibleArrayHeader,
|
||||||
dataset_dims: &[u64],
|
dataset_dims: &[u64],
|
||||||
max_dims: Option<&[u64]>,
|
max_dims: Option<&[u64]>,
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
length_size: u8,
|
length_size: u8,
|
||||||
@@ -480,7 +480,7 @@ pub fn read_extensible_array_chunks_in<S: Storage + ?Sized>(
|
|||||||
header: &ExtensibleArrayHeader,
|
header: &ExtensibleArrayHeader,
|
||||||
dataset_dims: &[u64],
|
dataset_dims: &[u64],
|
||||||
max_dims: Option<&[u64]>,
|
max_dims: Option<&[u64]>,
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
_length_size: u8,
|
_length_size: u8,
|
||||||
@@ -489,12 +489,12 @@ pub fn read_extensible_array_chunks_in<S: Storage + ?Sized>(
|
|||||||
|
|
||||||
// Linear indexes follow the maximum dimensions, with the unlimited
|
// Linear indexes follow the maximum dimensions, with the unlimited
|
||||||
// dimension swizzled to the slowest position (see `chunk_grid`).
|
// dimension swizzled to the slowest position (see `chunk_grid`).
|
||||||
let dims_u64: Vec<u64> = chunk_dimensions.iter().map(|&d| d as u64).collect();
|
let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, chunk_dimensions)?;
|
||||||
let grid = ChunkGrid::extensible_array(dataset_dims, max_dims, &dims_u64)?;
|
|
||||||
let grid = &grid;
|
let grid = &grid;
|
||||||
|
|
||||||
let chunk_byte_size: u64 =
|
let chunk_byte_size: u64 = chunk_dimensions
|
||||||
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
|
.iter()
|
||||||
|
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
|
||||||
|
|
||||||
// Parse index block (EAIB): signature(4) + version(1) + client_id(1)
|
// Parse index block (EAIB): signature(4) + version(1) + client_id(1)
|
||||||
// + header address(offset_size), then the inline elements, then the
|
// + header address(offset_size), then the inline elements, then the
|
||||||
@@ -930,7 +930,7 @@ mod tests {
|
|||||||
|
|
||||||
let header = ExtensibleArrayHeader::parse(&file_data, aehd_offset, os, ls).unwrap();
|
let header = ExtensibleArrayHeader::parse(&file_data, aehd_offset, os, ls).unwrap();
|
||||||
let ds_dims = vec![40u64]; // 2 chunks × 20 elements
|
let ds_dims = vec![40u64]; // 2 chunks × 20 elements
|
||||||
let chunk_dims = vec![20u32];
|
let chunk_dims = vec![20u64];
|
||||||
let chunks = read_extensible_array_chunks(
|
let chunks = read_extensible_array_chunks(
|
||||||
&file_data,
|
&file_data,
|
||||||
&header,
|
&header,
|
||||||
@@ -1057,7 +1057,7 @@ mod tests {
|
|||||||
let file_data = build_inline_plus_data_blocks();
|
let file_data = build_inline_plus_data_blocks();
|
||||||
let header = ExtensibleArrayHeader::parse(&file_data, 0x100, os, ls).unwrap();
|
let header = ExtensibleArrayHeader::parse(&file_data, 0x100, os, ls).unwrap();
|
||||||
let ds_dims = vec![40u64];
|
let ds_dims = vec![40u64];
|
||||||
let chunk_dims = vec![10u32];
|
let chunk_dims = vec![10u64];
|
||||||
let chunks = read_extensible_array_chunks(
|
let chunks = read_extensible_array_chunks(
|
||||||
&file_data,
|
&file_data,
|
||||||
&header,
|
&header,
|
||||||
|
|||||||
@@ -3,15 +3,14 @@
|
|||||||
//! Produces valid HDF5 files with v3 superblock, v2 object headers,
|
//! Produces valid HDF5 files with v3 superblock, v2 object headers,
|
||||||
//! link messages, contiguous datasets, inline and dense attributes.
|
//! link messages, contiguous datasets, inline and dense attributes.
|
||||||
|
|
||||||
use crate::addr::saturating_usize;
|
use crate::addr::{saturating_usize, to_usize};
|
||||||
#[cfg(not(feature = "std"))]
|
#[cfg(not(feature = "std"))]
|
||||||
use alloc::{format, string::String, vec, vec::Vec};
|
use alloc::{format, string::String, vec, vec::Vec};
|
||||||
|
|
||||||
use crate::attribute::AttributeMessage;
|
use crate::attribute::AttributeMessage;
|
||||||
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
|
use crate::btree_v2_write::{BTreeV2Params, build_btree_v2};
|
||||||
use crate::chunked_write::{
|
use crate::chunked_write::{
|
||||||
ChunkOptions, PrecompressedChunks, build_chunked_data_from_precompressed_libver,
|
ChunkOptions, Piece, PreparedChunks, place_chunked_data, prepare_chunks,
|
||||||
precompress_chunks,
|
|
||||||
};
|
};
|
||||||
use crate::data_layout::VdsMapping;
|
use crate::data_layout::VdsMapping;
|
||||||
use crate::dataspace::{Dataspace, DataspaceType};
|
use crate::dataspace::{Dataspace, DataspaceType};
|
||||||
@@ -1612,7 +1611,30 @@ impl FileWriter {
|
|||||||
self
|
self
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// The file, in memory. Its datasets' data is copied into the buffer
|
||||||
|
/// once (a dataset stored as one chunk of its own shape, or unfiltered
|
||||||
|
/// chunks that are contiguous runs of its data, are not copied before);
|
||||||
|
/// to write a file without holding it all, use [`Self::finish_with`].
|
||||||
pub fn finish(self) -> Result<Vec<u8>, FormatError> {
|
pub fn finish(self) -> Result<Vec<u8>, FormatError> {
|
||||||
|
let mut buf = Vec::new();
|
||||||
|
self.finish_into(&mut buf)?;
|
||||||
|
Ok(buf)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// The file, handed to `put` in order, a piece at a time, without
|
||||||
|
/// assembling it in memory: an unfiltered chunk or contiguous dataset
|
||||||
|
/// is passed straight from the data the dataset was given. Writing a
|
||||||
|
/// dataset stored as one unfiltered chunk this way needs little memory
|
||||||
|
/// beyond the dataset's own data. `put`'s first error stops the write
|
||||||
|
/// and is returned.
|
||||||
|
pub fn finish_with<E: From<FormatError>>(
|
||||||
|
self,
|
||||||
|
put: impl FnMut(&[u8]) -> Result<(), E>,
|
||||||
|
) -> Result<(), E> {
|
||||||
|
self.finish_into(&mut FnSink(put))
|
||||||
|
}
|
||||||
|
|
||||||
|
fn finish_into<S: Sink>(self, sink: &mut S) -> Result<(), S::Error> {
|
||||||
let page_size = self.page_size;
|
let page_size = self.page_size;
|
||||||
if let Some(ps) = page_size
|
if let Some(ps) = page_size
|
||||||
&& !(MIN_FILE_SPACE_PAGE_SIZE..=MAX_FILE_SPACE_PAGE_SIZE).contains(&ps)
|
&& !(MIN_FILE_SPACE_PAGE_SIZE..=MAX_FILE_SPACE_PAGE_SIZE).contains(&ps)
|
||||||
@@ -1620,7 +1642,8 @@ impl FileWriter {
|
|||||||
return Err(FormatError::SerializationError(format!(
|
return Err(FormatError::SerializationError(format!(
|
||||||
"file space page size {ps} is outside libhdf5's \
|
"file space page size {ps} is outside libhdf5's \
|
||||||
{MIN_FILE_SPACE_PAGE_SIZE}..={MAX_FILE_SPACE_PAGE_SIZE} bytes"
|
{MIN_FILE_SPACE_PAGE_SIZE}..={MAX_FILE_SPACE_PAGE_SIZE} bytes"
|
||||||
)));
|
))
|
||||||
|
.into());
|
||||||
}
|
}
|
||||||
|
|
||||||
let (low, high) = (self.low, self.high);
|
let (low, high) = (self.low, self.high);
|
||||||
@@ -1774,15 +1797,16 @@ impl FileWriter {
|
|||||||
})
|
})
|
||||||
.collect::<Result<_, _>>()?;
|
.collect::<Result<_, _>>()?;
|
||||||
|
|
||||||
struct DataBlob {
|
struct DataBlob<'a> {
|
||||||
data: Vec<u8>,
|
data: Vec<u8>,
|
||||||
oh_bytes: Vec<u8>,
|
oh_bytes: Vec<u8>,
|
||||||
/// Cached compressed chunks for chunked datasets; reused in Pass 2
|
/// Cached compressed chunks for chunked datasets; reused in Pass 2
|
||||||
/// to avoid re-compressing the same data.
|
/// to avoid re-compressing the same data. Unfiltered chunks
|
||||||
precompressed: Option<PrecompressedChunks>,
|
/// borrow the dataset's data where they can.
|
||||||
|
precompressed: Option<PreparedChunks<'a>>,
|
||||||
}
|
}
|
||||||
|
|
||||||
let mut dummy_blobs: Vec<DataBlob> = Vec::new();
|
let mut dummy_blobs: Vec<DataBlob<'_>> = Vec::new();
|
||||||
let mut dummy_cursor = 0u64;
|
let mut dummy_cursor = 0u64;
|
||||||
for (i, d) in all_ds.iter().enumerate() {
|
for (i, d) in all_ds.iter().enumerate() {
|
||||||
let dense_blob = ds_dense[i]
|
let dense_blob = ds_dense[i]
|
||||||
@@ -1820,21 +1844,16 @@ impl FileWriter {
|
|||||||
.resolve_chunk_dims_for(&d.ds.dimensions, elem_size);
|
.resolve_chunk_dims_for(&d.ds.dimensions, elem_size);
|
||||||
// Compress once in Pass 1; cache the result so Pass 2 can skip
|
// Compress once in Pass 1; cache the result so Pass 2 can skip
|
||||||
// re-compression and just rebuild the index with real addresses.
|
// re-compression and just rebuild the index with real addresses.
|
||||||
let pre = precompress_chunks(
|
let pre = prepare_chunks(
|
||||||
&d.raw,
|
&d.raw,
|
||||||
&d.ds.dimensions,
|
&d.ds.dimensions,
|
||||||
&chunk_dims,
|
&chunk_dims,
|
||||||
elem_size,
|
elem_size,
|
||||||
&d.chunk_options,
|
&d.chunk_options,
|
||||||
)?;
|
)?;
|
||||||
let result = build_chunked_data_from_precompressed_libver(
|
let result =
|
||||||
&pre,
|
place_chunked_data(&pre, dummy_cursor, d.maxshape.as_deref(), low, high)?;
|
||||||
dummy_cursor,
|
dummy_cursor += result.len;
|
||||||
d.maxshape.as_deref(),
|
|
||||||
low,
|
|
||||||
high,
|
|
||||||
)?;
|
|
||||||
dummy_cursor += result.data_bytes.len() as u64;
|
|
||||||
let oh = build_chunked_dataset_oh(
|
let oh = build_chunked_dataset_oh(
|
||||||
&d.dt,
|
&d.dt,
|
||||||
&d.ds,
|
&d.ds,
|
||||||
@@ -1849,7 +1868,7 @@ impl FileWriter {
|
|||||||
d.refcount,
|
d.refcount,
|
||||||
)?;
|
)?;
|
||||||
dummy_blobs.push(DataBlob {
|
dummy_blobs.push(DataBlob {
|
||||||
data: result.data_bytes,
|
data: Vec::new(),
|
||||||
oh_bytes: oh,
|
oh_bytes: oh,
|
||||||
precompressed: Some(pre),
|
precompressed: Some(pre),
|
||||||
});
|
});
|
||||||
@@ -1955,7 +1974,7 @@ impl FileWriter {
|
|||||||
})
|
})
|
||||||
.collect::<Result<_, FormatError>>()?;
|
.collect::<Result<_, FormatError>>()?;
|
||||||
|
|
||||||
let mut ds_blobs2: Vec<DataBlob> = Vec::new();
|
let mut ds_blobs2: Vec<OutBlob<'_>> = Vec::new();
|
||||||
let global_align_threshold = self.alignment_threshold;
|
let global_align_threshold = self.alignment_threshold;
|
||||||
let global_align_bytes = self.alignment_bytes;
|
let global_align_bytes = self.alignment_bytes;
|
||||||
for (i, d) in all_ds.iter().enumerate() {
|
for (i, d) in all_ds.iter().enumerate() {
|
||||||
@@ -1977,26 +1996,21 @@ impl FileWriter {
|
|||||||
&d.fill_message,
|
&d.fill_message,
|
||||||
d.refcount,
|
d.refcount,
|
||||||
)?;
|
)?;
|
||||||
ds_blobs2.push(DataBlob {
|
ds_blobs2.push(OutBlob {
|
||||||
data: gcol_bytes.clone(),
|
data: vec![Out::Slice(gcol_bytes)],
|
||||||
oh_bytes: oh,
|
oh_bytes: oh,
|
||||||
precompressed: None,
|
|
||||||
});
|
});
|
||||||
} else if is_chunked[i] {
|
} else if is_chunked[i] {
|
||||||
let base_address = cursor2 as u64;
|
let base_address = cursor2 as u64;
|
||||||
// Reuse precompressed chunks from Pass 1 — avoids re-compressing
|
// Reuse precompressed chunks from Pass 1 — avoids re-compressing
|
||||||
// the same data a second time.
|
// the same data a second time.
|
||||||
let result = build_chunked_data_from_precompressed_libver(
|
let pre = dummy_blobs[i]
|
||||||
dummy_blobs[i]
|
.precompressed
|
||||||
.precompressed
|
.as_ref()
|
||||||
.as_ref()
|
.expect("chunked dataset missing precompressed cache");
|
||||||
.expect("chunked dataset missing precompressed cache"),
|
let result =
|
||||||
base_address,
|
place_chunked_data(pre, base_address, d.maxshape.as_deref(), low, high)?;
|
||||||
d.maxshape.as_deref(),
|
cursor2 += to_usize(result.len)?;
|
||||||
low,
|
|
||||||
high,
|
|
||||||
)?;
|
|
||||||
cursor2 += result.data_bytes.len();
|
|
||||||
let oh = build_chunked_dataset_oh(
|
let oh = build_chunked_dataset_oh(
|
||||||
&d.dt,
|
&d.dt,
|
||||||
&d.ds,
|
&d.ds,
|
||||||
@@ -2010,11 +2024,16 @@ impl FileWriter {
|
|||||||
&d.fill_message,
|
&d.fill_message,
|
||||||
d.refcount,
|
d.refcount,
|
||||||
)?;
|
)?;
|
||||||
ds_blobs2.push(DataBlob {
|
let data = result
|
||||||
data: result.data_bytes,
|
.pieces
|
||||||
oh_bytes: oh,
|
.into_iter()
|
||||||
precompressed: None,
|
.map(|p| match p {
|
||||||
});
|
Piece::Zeros(n) => Out::Zeros(n),
|
||||||
|
Piece::Chunk(c) => Out::Slice(&pre.chunks[c].1),
|
||||||
|
Piece::Bytes(b) => Out::Bytes(b),
|
||||||
|
})
|
||||||
|
.collect();
|
||||||
|
ds_blobs2.push(OutBlob { data, oh_bytes: oh });
|
||||||
} else if is_compact[i] {
|
} else if is_compact[i] {
|
||||||
// Compact: data is inline in the object header, no external blob
|
// Compact: data is inline in the object header, no external blob
|
||||||
let oh = build_compact_dataset_oh(
|
let oh = build_compact_dataset_oh(
|
||||||
@@ -2030,10 +2049,9 @@ impl FileWriter {
|
|||||||
d.refcount,
|
d.refcount,
|
||||||
layout_version,
|
layout_version,
|
||||||
)?;
|
)?;
|
||||||
ds_blobs2.push(DataBlob {
|
ds_blobs2.push(OutBlob {
|
||||||
data: vec![],
|
data: vec![],
|
||||||
oh_bytes: oh,
|
oh_bytes: oh,
|
||||||
precompressed: None,
|
|
||||||
});
|
});
|
||||||
} else {
|
} else {
|
||||||
// Determine alignment: per-dataset overrides global
|
// Determine alignment: per-dataset overrides global
|
||||||
@@ -2060,13 +2078,10 @@ impl FileWriter {
|
|||||||
d.refcount,
|
d.refcount,
|
||||||
layout_version,
|
layout_version,
|
||||||
)?;
|
)?;
|
||||||
let mut data = vec![0u8; padding];
|
|
||||||
data.extend_from_slice(&d.raw);
|
|
||||||
cursor2 += d.raw.len();
|
cursor2 += d.raw.len();
|
||||||
ds_blobs2.push(DataBlob {
|
ds_blobs2.push(OutBlob {
|
||||||
data,
|
data: vec![Out::Zeros(padding), Out::Slice(&d.raw)],
|
||||||
oh_bytes: oh,
|
oh_bytes: oh,
|
||||||
precompressed: None,
|
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -2080,7 +2095,8 @@ impl FileWriter {
|
|||||||
cursor2 = cursor2.next_multiple_of(ps as usize);
|
cursor2 = cursor2.next_multiple_of(ps as usize);
|
||||||
}
|
}
|
||||||
let eof_addr2 = cursor2 as u64;
|
let eof_addr2 = cursor2 as u64;
|
||||||
let mut buf = Vec::with_capacity(cursor2);
|
sink.reserve(cursor2);
|
||||||
|
let mut out = Counted { sink, written: 0 };
|
||||||
|
|
||||||
let sb = Superblock {
|
let sb = Superblock {
|
||||||
version: superblock_version,
|
version: superblock_version,
|
||||||
@@ -2103,9 +2119,9 @@ impl FileWriter {
|
|||||||
checksum: None,
|
checksum: None,
|
||||||
page_size: None,
|
page_size: None,
|
||||||
};
|
};
|
||||||
buf.extend_from_slice(&sb.serialize());
|
out.put(&sb.serialize())?;
|
||||||
if let Some(ref ext) = sb_ext {
|
if let Some(ref ext) = sb_ext {
|
||||||
buf.extend_from_slice(ext);
|
out.put(ext)?;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Group OHs + dense blobs (link blob, then attr blob, matching pass 2)
|
// Group OHs + dense blobs (link blob, then attr blob, matching pass 2)
|
||||||
@@ -2132,32 +2148,120 @@ impl FileWriter {
|
|||||||
g.refcount,
|
g.refcount,
|
||||||
)?;
|
)?;
|
||||||
debug_assert_eq!(oh.len(), group_oh_sizes[gi]);
|
debug_assert_eq!(oh.len(), group_oh_sizes[gi]);
|
||||||
debug_assert_eq!(buf.len() as u64, group_addrs2[gi]);
|
debug_assert_eq!(out.written as u64, group_addrs2[gi]);
|
||||||
buf.extend_from_slice(&oh);
|
out.put(&oh)?;
|
||||||
if let Some(ref b) = link_blob {
|
if let Some(ref b) = link_blob {
|
||||||
buf.extend_from_slice(&b.blob);
|
out.put(&b.blob)?;
|
||||||
}
|
}
|
||||||
if let Some(ref blob) = group_dense_blobs[gi] {
|
if let Some(ref blob) = group_dense_blobs[gi] {
|
||||||
buf.extend_from_slice(&blob.blob);
|
out.put(&blob.blob)?;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// Dataset OHs + dense blobs
|
// Dataset OHs + dense blobs
|
||||||
for (i, blob) in ds_blobs2.iter().enumerate() {
|
for (i, blob) in ds_blobs2.iter().enumerate() {
|
||||||
buf.extend_from_slice(&blob.oh_bytes);
|
out.put(&blob.oh_bytes)?;
|
||||||
if let Some(ref dense) = ds_dense_blobs[i] {
|
if let Some(ref dense) = ds_dense_blobs[i] {
|
||||||
buf.extend_from_slice(&dense.blob);
|
out.put(&dense.blob)?;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// Data
|
// Data
|
||||||
for blob in &ds_blobs2 {
|
for blob in &ds_blobs2 {
|
||||||
buf.extend_from_slice(&blob.data);
|
for piece in &blob.data {
|
||||||
|
match piece {
|
||||||
|
Out::Bytes(b) => out.put(b)?,
|
||||||
|
Out::Slice(s) => out.put(s)?,
|
||||||
|
Out::Zeros(n) => out.zeros(*n)?,
|
||||||
|
}
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
debug_assert_eq!(buf.len(), data_end);
|
debug_assert_eq!(out.written, data_end);
|
||||||
buf.resize(cursor2, 0);
|
out.zeros(cursor2 - data_end)?;
|
||||||
Ok(buf)
|
Ok(())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A piece of a dataset's data as [`FileWriter`] writes it out.
|
||||||
|
enum Out<'a> {
|
||||||
|
Bytes(Vec<u8>),
|
||||||
|
/// Borrowed from the dataset's data or its prepared chunks.
|
||||||
|
Slice(&'a [u8]),
|
||||||
|
Zeros(usize),
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A dataset's object header and the pieces of its data.
|
||||||
|
struct OutBlob<'a> {
|
||||||
|
oh_bytes: Vec<u8>,
|
||||||
|
data: Vec<Out<'a>>,
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Where [`FileWriter`] puts a file's bytes, in order.
|
||||||
|
trait Sink {
|
||||||
|
type Error: From<FormatError>;
|
||||||
|
|
||||||
|
/// The file will be `_len` bytes long.
|
||||||
|
fn reserve(&mut self, _len: usize) {}
|
||||||
|
|
||||||
|
fn put(&mut self, bytes: &[u8]) -> Result<(), Self::Error>;
|
||||||
|
|
||||||
|
fn zeros(&mut self, n: usize) -> Result<(), Self::Error> {
|
||||||
|
const ZEROS: [u8; 4096] = [0; 4096];
|
||||||
|
let mut left = n;
|
||||||
|
while left > 0 {
|
||||||
|
let k = left.min(ZEROS.len());
|
||||||
|
self.put(&ZEROS[..k])?;
|
||||||
|
left -= k;
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
impl Sink for Vec<u8> {
|
||||||
|
type Error = FormatError;
|
||||||
|
|
||||||
|
fn reserve(&mut self, len: usize) {
|
||||||
|
self.reserve_exact(len.saturating_sub(self.len()));
|
||||||
|
}
|
||||||
|
|
||||||
|
fn put(&mut self, bytes: &[u8]) -> Result<(), FormatError> {
|
||||||
|
self.extend_from_slice(bytes);
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn zeros(&mut self, n: usize) -> Result<(), FormatError> {
|
||||||
|
self.resize(self.len() + n, 0);
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A [`Sink`] calling a function ([`FileWriter::finish_with`]).
|
||||||
|
struct FnSink<F>(F);
|
||||||
|
|
||||||
|
impl<E: From<FormatError>, F: FnMut(&[u8]) -> Result<(), E>> Sink for FnSink<F> {
|
||||||
|
type Error = E;
|
||||||
|
|
||||||
|
fn put(&mut self, bytes: &[u8]) -> Result<(), E> {
|
||||||
|
(self.0)(bytes)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// A [`Sink`] and the number of bytes put into it.
|
||||||
|
struct Counted<'s, S> {
|
||||||
|
sink: &'s mut S,
|
||||||
|
written: usize,
|
||||||
|
}
|
||||||
|
|
||||||
|
impl<S: Sink> Counted<'_, S> {
|
||||||
|
fn put(&mut self, bytes: &[u8]) -> Result<(), S::Error> {
|
||||||
|
self.written += bytes.len();
|
||||||
|
self.sink.put(bytes)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn zeros(&mut self, n: usize) -> Result<(), S::Error> {
|
||||||
|
self.written += n;
|
||||||
|
self.sink.zeros(n)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -3015,4 +3119,75 @@ mod tests {
|
|||||||
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 3);
|
assert_eq!(Superblock::parse(&bytes, 0).unwrap().version, 3);
|
||||||
assert_eq!(layout_of(&bytes, "x")[0], 3);
|
assert_eq!(layout_of(&bytes, "x")[0], 3);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// `finish_with` hands out exactly the bytes `finish` returns, for every
|
||||||
|
/// storage (contiguous, compact, chunked filtered and not, chunks
|
||||||
|
/// borrowed and copied, a version-1 B-tree, a paged file), and stops at
|
||||||
|
/// the sink's first error.
|
||||||
|
#[test]
|
||||||
|
fn finish_with_matches_finish() {
|
||||||
|
let build = |low: LibVer, paged: bool| {
|
||||||
|
let mut fw = FileWriter::new();
|
||||||
|
fw.libver_bounds(low, LibVer::Latest);
|
||||||
|
if paged {
|
||||||
|
fw.with_page_size(4096);
|
||||||
|
}
|
||||||
|
let v: Vec<f64> = (0..300).map(f64::from).collect();
|
||||||
|
fw.create_dataset("contig").with_f64_data(&v);
|
||||||
|
fw.create_dataset("compact")
|
||||||
|
.with_f64_data(&[1.0, 2.0])
|
||||||
|
.compact();
|
||||||
|
fw.create_dataset("one_chunk")
|
||||||
|
.with_f64_data(&v)
|
||||||
|
.with_shape(&[30, 10])
|
||||||
|
.with_chunks(&[30, 10]);
|
||||||
|
fw.create_dataset("rows")
|
||||||
|
.with_f64_data(&v)
|
||||||
|
.with_shape(&[30, 10])
|
||||||
|
.with_chunks(&[7, 10]);
|
||||||
|
fw.create_dataset("tiles")
|
||||||
|
.with_f64_data(&v)
|
||||||
|
.with_shape(&[30, 10])
|
||||||
|
.with_chunks(&[7, 3]);
|
||||||
|
fw.create_dataset("deflated")
|
||||||
|
.with_f64_data(&v)
|
||||||
|
.with_chunks(&[64])
|
||||||
|
.with_shuffle()
|
||||||
|
.with_deflate(4);
|
||||||
|
fw.create_dataset("grows")
|
||||||
|
.with_u8_data_owned((0..=255).collect())
|
||||||
|
.with_chunks(&[100])
|
||||||
|
.with_maxshape(&[u64::MAX]);
|
||||||
|
fw
|
||||||
|
};
|
||||||
|
for (low, paged) in [
|
||||||
|
(LibVer::V18, false),
|
||||||
|
(LibVer::V110, false),
|
||||||
|
(LibVer::Latest, true),
|
||||||
|
] {
|
||||||
|
let bytes = build(low, paged).finish().unwrap();
|
||||||
|
let mut streamed = Vec::new();
|
||||||
|
build(low, paged)
|
||||||
|
.finish_with(|b: &[u8]| {
|
||||||
|
streamed.extend_from_slice(b);
|
||||||
|
Ok::<(), FormatError>(())
|
||||||
|
})
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(streamed, bytes, "{low:?} paged {paged}");
|
||||||
|
}
|
||||||
|
|
||||||
|
let mut calls = 0;
|
||||||
|
let err = build(LibVer::V110, false)
|
||||||
|
.finish_with(|_| {
|
||||||
|
calls += 1;
|
||||||
|
if calls == 3 {
|
||||||
|
Err(FormatError::SerializationError("full".into()))
|
||||||
|
} else {
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
})
|
||||||
|
.unwrap_err();
|
||||||
|
assert_eq!(err, FormatError::SerializationError("full".into()));
|
||||||
|
assert_eq!(calls, 3);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -155,7 +155,7 @@ pub fn read_fixed_array_chunks(
|
|||||||
header: &FixedArrayHeader,
|
header: &FixedArrayHeader,
|
||||||
dataset_dims: &[u64],
|
dataset_dims: &[u64],
|
||||||
max_dims: Option<&[u64]>,
|
max_dims: Option<&[u64]>,
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
length_size: u8,
|
length_size: u8,
|
||||||
@@ -180,7 +180,7 @@ pub fn read_fixed_array_chunks_in<S: Storage + ?Sized>(
|
|||||||
header: &FixedArrayHeader,
|
header: &FixedArrayHeader,
|
||||||
dataset_dims: &[u64],
|
dataset_dims: &[u64],
|
||||||
max_dims: Option<&[u64]>,
|
max_dims: Option<&[u64]>,
|
||||||
chunk_dimensions: &[u32],
|
chunk_dimensions: &[u64],
|
||||||
element_size: u32,
|
element_size: u32,
|
||||||
offset_size: u8,
|
offset_size: u8,
|
||||||
_length_size: u8,
|
_length_size: u8,
|
||||||
@@ -227,11 +227,11 @@ pub fn read_fixed_array_chunks_in<S: Storage + ?Sized>(
|
|||||||
|
|
||||||
// The index is laid out over the chunk grid of the *maximum* dimensions
|
// The index is laid out over the chunk grid of the *maximum* dimensions
|
||||||
// (row-major), so a dataset smaller than its maxshape has gaps.
|
// (row-major), so a dataset smaller than its maxshape has gaps.
|
||||||
let dims_u64: Vec<u64> = chunk_dimensions.iter().map(|&d| d as u64).collect();
|
let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, chunk_dimensions)?;
|
||||||
let grid = ChunkGrid::fixed_array(dataset_dims, max_dims, &dims_u64)?;
|
|
||||||
|
|
||||||
let chunk_byte_size: u64 =
|
let chunk_byte_size: u64 = chunk_dimensions
|
||||||
chunk_dimensions.iter().map(|&d| d as u64).product::<u64>() * element_size as u64;
|
.iter()
|
||||||
|
.fold(u64::from(element_size), |acc, &d| acc.saturating_mul(d));
|
||||||
|
|
||||||
let mut chunks = Vec::new();
|
let mut chunks = Vec::new();
|
||||||
// `rel` is relative to the data block, whose bytes are in `w`.
|
// `rel` is relative to the data block, whose bytes are in `w`.
|
||||||
@@ -670,7 +670,7 @@ mod tests {
|
|||||||
let header =
|
let header =
|
||||||
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
|
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
|
||||||
let ds_dims = vec![100u64];
|
let ds_dims = vec![100u64];
|
||||||
let chunk_dims = vec![20u32];
|
let chunk_dims = vec![20u64];
|
||||||
let chunks = read_fixed_array_chunks(
|
let chunks = read_fixed_array_chunks(
|
||||||
&file_data,
|
&file_data,
|
||||||
&header,
|
&header,
|
||||||
@@ -747,7 +747,7 @@ mod tests {
|
|||||||
let header =
|
let header =
|
||||||
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
|
FixedArrayHeader::parse(&file_data, fahd_offset, offset_size, length_size).unwrap();
|
||||||
let ds_dims = vec![60u64];
|
let ds_dims = vec![60u64];
|
||||||
let chunk_dims = vec![20u32];
|
let chunk_dims = vec![20u64];
|
||||||
let chunks = read_fixed_array_chunks(
|
let chunks = read_fixed_array_chunks(
|
||||||
&file_data,
|
&file_data,
|
||||||
&header,
|
&header,
|
||||||
@@ -848,7 +848,7 @@ mod tests {
|
|||||||
assert_eq!(header.num_elements, 11);
|
assert_eq!(header.num_elements, 11);
|
||||||
|
|
||||||
let ds_dims = vec![11u64 * 20];
|
let ds_dims = vec![11u64 * 20];
|
||||||
let chunk_dims = vec![20u32];
|
let chunk_dims = vec![20u64];
|
||||||
let chunks = read_fixed_array_chunks(
|
let chunks = read_fixed_array_chunks(
|
||||||
&file_data,
|
&file_data,
|
||||||
&header,
|
&header,
|
||||||
|
|||||||
@@ -663,11 +663,17 @@ impl DatasetBuilder {
|
|||||||
}
|
}
|
||||||
|
|
||||||
pub fn with_u8_data(&mut self, data: &[u8]) -> &mut Self {
|
pub fn with_u8_data(&mut self, data: &[u8]) -> &mut Self {
|
||||||
|
self.with_u8_data_owned(data.to_vec())
|
||||||
|
}
|
||||||
|
|
||||||
|
/// [`Self::with_u8_data`] without the copy: the builder takes the
|
||||||
|
/// vector. For large data this halves the memory a write needs.
|
||||||
|
pub fn with_u8_data_owned(&mut self, data: Vec<u8>) -> &mut Self {
|
||||||
self.datatype = Some(make_u8_type());
|
self.datatype = Some(make_u8_type());
|
||||||
self.data = Some(data.to_vec());
|
|
||||||
if self.shape.is_none() {
|
if self.shape.is_none() {
|
||||||
self.shape = Some(vec![data.len() as u64]);
|
self.shape = Some(vec![data.len() as u64]);
|
||||||
}
|
}
|
||||||
|
self.data = Some(data);
|
||||||
self
|
self
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -104,7 +104,22 @@ writes `float64`, `float32`, `int64`, `int32`, `uint8`, `complex64` and
|
|||||||
`complex128` arrays; the file is written on `close()`. Complex arrays are
|
`complex128` arrays; the file is written on `close()`. Complex arrays are
|
||||||
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
|
stored as h5py stores them, a compound `{r, i}` that every libhdf5 reads
|
||||||
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
|
(not HDF5 2.0's native complex type, which only libhdf5 2.0+ reads; the
|
||||||
Rust API writes that on request).
|
Rust API writes that on request). Both forms read back as numpy
|
||||||
|
`complex64`/`complex128`.
|
||||||
|
|
||||||
|
`libver=` sets the library version bounds as in h5py: `'v108'`,
|
||||||
|
`'v110'`, `'v112'`, `'v114'`, `'v200'` or `'latest'` (the low bound; the
|
||||||
|
high bound is then `'latest'`), or a `(low, high)` tuple. With `'v108'`
|
||||||
|
the file is written in the HDF5 1.8 format (version-2 superblock,
|
||||||
|
version-1 B-tree chunk indexes), which HDF5 1.8 reads; the default is the
|
||||||
|
HDF5 1.10 format. `'earliest'` as the low bound writes the 1.8 format too,
|
||||||
|
with a `UserWarning`: clawhdf5 cannot write the pre-1.8 format. The
|
||||||
|
argument is ignored when reading and refused for `'r+'`/`'a'`.
|
||||||
|
|
||||||
|
```python
|
||||||
|
with clawhdf5.File("old.h5", "w", libver="v108") as f: # HDF5 1.8 reads it
|
||||||
|
f.create_dataset("x", data=np.arange(10.0), chunks=(5,), compression="gzip")
|
||||||
|
```
|
||||||
|
|
||||||
## Editing a file in place
|
## Editing a file in place
|
||||||
|
|
||||||
|
|||||||
@@ -13,10 +13,13 @@ use crate::attrs::PyAttrs;
|
|||||||
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
|
use crate::group::{PyGroup, ReadGroup, WriteGroupState, finalize_write_group};
|
||||||
use crate::handle::Handle;
|
use crate::handle::Handle;
|
||||||
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
|
use crate::{DatasetSpec, OwnedAttrValue, apply_dataset_spec, extract_numpy_data, to_py_err};
|
||||||
|
use clawhdf5_rs::LibVer;
|
||||||
|
|
||||||
/// Internal state for write mode.
|
/// Internal state for write mode.
|
||||||
struct WriteState {
|
struct WriteState {
|
||||||
path: PathBuf,
|
path: PathBuf,
|
||||||
|
/// `libver=` as (low, high); `None` keeps the writer's default.
|
||||||
|
libver: Option<(LibVer, LibVer)>,
|
||||||
root_datasets: Vec<DatasetSpec>,
|
root_datasets: Vec<DatasetSpec>,
|
||||||
root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
|
root_attrs: Arc<Mutex<Vec<(String, OwnedAttrValue)>>>,
|
||||||
groups: Vec<Arc<Mutex<WriteGroupState>>>,
|
groups: Vec<Arc<Mutex<WriteGroupState>>>,
|
||||||
@@ -80,10 +83,32 @@ impl PyFile {
|
|||||||
/// `az://`; which schemes work depends on how the wheel was built)
|
/// `az://`; which schemes work depends on how the wheel was built)
|
||||||
/// to read the file remotely with default options (see `open_url`)
|
/// to read the file remotely with default options (see `open_url`)
|
||||||
/// mode: 'r' for read (default), 'w' for write
|
/// mode: 'r' for read (default), 'w' for write
|
||||||
|
/// libver: library version bounds for mode 'w', as h5py's: one of
|
||||||
|
/// 'earliest', 'v108', 'v110', 'v112', 'v114', 'v200', 'latest' (the
|
||||||
|
/// low bound; the high bound is then 'latest') or a (low, high)
|
||||||
|
/// tuple of them. The low bound is the oldest HDF5 release whose
|
||||||
|
/// format the file uses ('v108': HDF5 1.8 can read it); the high
|
||||||
|
/// bound the newest whose features it may use. clawhdf5 cannot write
|
||||||
|
/// the pre-1.8 format, so a low bound of 'earliest' writes the 1.8
|
||||||
|
/// format (with a warning) and a high bound of 'earliest' is an
|
||||||
|
/// error. Default (None): the HDF5 1.10 format clawhdf5 has always
|
||||||
|
/// written. Ignored for reading; not supported with 'r+' / 'a'.
|
||||||
#[new]
|
#[new]
|
||||||
#[pyo3(signature = (path, mode="r"))]
|
#[pyo3(signature = (path, mode="r", libver=None))]
|
||||||
fn new(py: Python<'_>, path: &str, mode: &str) -> PyResult<Self> {
|
fn new(
|
||||||
|
py: Python<'_>,
|
||||||
|
path: &str,
|
||||||
|
mode: &str,
|
||||||
|
libver: Option<&Bound<'_, PyAny>>,
|
||||||
|
) -> PyResult<Self> {
|
||||||
let filename = path.to_string();
|
let filename = path.to_string();
|
||||||
|
let libver = libver.map(|v| parse_libver(py, v)).transpose()?;
|
||||||
|
if libver.is_some() && matches!(mode, "r+" | "a") {
|
||||||
|
return Err(PyNotImplementedError::new_err(format!(
|
||||||
|
"libver with mode '{mode}': clawhdf5's in-place editor keeps the format \
|
||||||
|
versions the file already uses"
|
||||||
|
)));
|
||||||
|
}
|
||||||
if is_url(path) {
|
if is_url(path) {
|
||||||
if mode != "r" {
|
if mode != "r" {
|
||||||
return Err(PyValueError::new_err(format!(
|
return Err(PyValueError::new_err(format!(
|
||||||
@@ -113,6 +138,7 @@ impl PyFile {
|
|||||||
// Absolute now: the file is written at close, possibly
|
// Absolute now: the file is written at close, possibly
|
||||||
// after the working directory changed.
|
// after the working directory changed.
|
||||||
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
|
path: std::path::absolute(path).unwrap_or_else(|_| PathBuf::from(path)),
|
||||||
|
libver,
|
||||||
root_datasets: Vec::new(),
|
root_datasets: Vec::new(),
|
||||||
root_attrs: Arc::new(Mutex::new(Vec::new())),
|
root_attrs: Arc::new(Mutex::new(Vec::new())),
|
||||||
groups: Vec::new(),
|
groups: Vec::new(),
|
||||||
@@ -460,10 +486,71 @@ fn parse_compression(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// One of h5py's `libver` names as a bound; `high` says which end it is.
|
||||||
|
/// `Ok(None)` is 'earliest' as the low bound: the pre-1.8 format, which
|
||||||
|
/// clawhdf5 cannot write.
|
||||||
|
fn libver_name(name: &str, high: bool) -> PyResult<Option<LibVer>> {
|
||||||
|
Ok(Some(match name {
|
||||||
|
"earliest" if high => {
|
||||||
|
return Err(PyValueError::new_err(
|
||||||
|
"libver high bound 'earliest' (the pre-1.8 format) cannot be written by \
|
||||||
|
clawhdf5; the oldest format it writes is 'v108'",
|
||||||
|
));
|
||||||
|
}
|
||||||
|
"earliest" => return Ok(None),
|
||||||
|
"v108" => LibVer::V18,
|
||||||
|
"v110" => LibVer::V110,
|
||||||
|
"v112" => LibVer::V112,
|
||||||
|
"v114" => LibVer::V114,
|
||||||
|
"v200" => LibVer::V200,
|
||||||
|
"latest" => LibVer::Latest,
|
||||||
|
other => {
|
||||||
|
return Err(PyValueError::new_err(format!(
|
||||||
|
"unknown libver '{other}'; expected 'earliest', 'v108', 'v110', 'v112', \
|
||||||
|
'v114', 'v200' or 'latest'"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
}))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// h5py's `libver=`: a name (the low bound, high bound 'latest') or a
|
||||||
|
/// `(low, high)` pair.
|
||||||
|
fn parse_libver(py: Python<'_>, v: &Bound<'_, PyAny>) -> PyResult<(LibVer, LibVer)> {
|
||||||
|
let (low, high): (String, String) = match v.extract::<String>() {
|
||||||
|
Ok(name) => (name, "latest".into()),
|
||||||
|
Err(_) => v.extract().map_err(|_| {
|
||||||
|
PyValueError::new_err("libver must be a string or a (low, high) tuple of strings")
|
||||||
|
})?,
|
||||||
|
};
|
||||||
|
let Some(high) = libver_name(&high, true)? else {
|
||||||
|
unreachable!("a high bound is never None")
|
||||||
|
};
|
||||||
|
let low = match libver_name(&low, false)? {
|
||||||
|
Some(low) => low,
|
||||||
|
None => {
|
||||||
|
PyModule::import(py, "warnings")?.getattr("warn")?.call1((
|
||||||
|
"libver 'earliest': clawhdf5 cannot write the pre-1.8 format; the file \
|
||||||
|
is written in the HDF5 1.8 format ('v108') instead",
|
||||||
|
py.get_type::<pyo3::exceptions::PyUserWarning>(),
|
||||||
|
))?;
|
||||||
|
LibVer::V18
|
||||||
|
}
|
||||||
|
};
|
||||||
|
if low > high {
|
||||||
|
return Err(PyValueError::new_err(format!(
|
||||||
|
"libver low bound {low} is newer than the high bound {high}"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
Ok((low, high))
|
||||||
|
}
|
||||||
|
|
||||||
/// Build and write the HDF5 file from accumulated write state.
|
/// Build and write the HDF5 file from accumulated write state.
|
||||||
fn finalize_write(state: WriteState) -> PyResult<()> {
|
fn finalize_write(state: WriteState) -> PyResult<()> {
|
||||||
crate::no_panic(|| {
|
crate::no_panic(|| {
|
||||||
let mut builder = clawhdf5_rs::FileBuilder::new();
|
let mut builder = clawhdf5_rs::FileBuilder::new();
|
||||||
|
if let Some((low, high)) = state.libver {
|
||||||
|
builder.libver_bounds(low, high);
|
||||||
|
}
|
||||||
|
|
||||||
// Root attributes
|
// Root attributes
|
||||||
let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner());
|
let root_attrs = state.root_attrs.lock().unwrap_or_else(|e| e.into_inner());
|
||||||
@@ -520,6 +607,7 @@ mod tests {
|
|||||||
|
|
||||||
let state = WriteState {
|
let state = WriteState {
|
||||||
path: path.clone(),
|
path: path.clone(),
|
||||||
|
libver: None,
|
||||||
root_datasets: vec![DatasetSpec {
|
root_datasets: vec![DatasetSpec {
|
||||||
name: "data".into(),
|
name: "data".into(),
|
||||||
data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]),
|
data: crate::DatasetData::F64(vec![1.0, 2.0, 3.0]),
|
||||||
|
|||||||
@@ -189,12 +189,7 @@ pub(crate) fn chunk_shape(file: &File, hdr: &ObjectHeader, rank: usize) -> Optio
|
|||||||
{
|
{
|
||||||
clawhdf5_format::data_layout::DataLayout::Chunked {
|
clawhdf5_format::data_layout::DataLayout::Chunked {
|
||||||
chunk_dimensions, ..
|
chunk_dimensions, ..
|
||||||
} if chunk_dimensions.len() >= rank => Some(
|
} if chunk_dimensions.len() >= rank => Some(chunk_dimensions[..rank].to_vec()),
|
||||||
chunk_dimensions[..rank]
|
|
||||||
.iter()
|
|
||||||
.map(|&d| u64::from(d))
|
|
||||||
.collect(),
|
|
||||||
),
|
|
||||||
_ => None,
|
_ => None,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,143 @@
|
|||||||
|
"""`clawhdf5.File(path, 'w', libver=...)`: h5py's library version bounds.
|
||||||
|
|
||||||
|
A file written with libver='v108' must open in HDF5 1.8. Its h5dump is
|
||||||
|
found through CLAWHDF5_H5DUMP18 or at ~/.cache/hdf5-1.8.23/bin/h5dump
|
||||||
|
(scripts/build-hdf5-1.8.sh builds it); without it that check is skipped.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import os
|
||||||
|
import subprocess
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
import clawhdf5
|
||||||
|
|
||||||
|
|
||||||
|
def superblock_version(path):
|
||||||
|
with open(path, "rb") as f:
|
||||||
|
head = f.read(9)
|
||||||
|
assert head[:8] == b"\x89HDF\r\n\x1a\n"
|
||||||
|
return head[8]
|
||||||
|
|
||||||
|
|
||||||
|
def h5dump18():
|
||||||
|
path = os.environ.get("CLAWHDF5_H5DUMP18") or os.path.expanduser(
|
||||||
|
"~/.cache/hdf5-1.8.23/bin/h5dump"
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
out = subprocess.run([path, "--version"], capture_output=True, text=True)
|
||||||
|
except OSError:
|
||||||
|
return None
|
||||||
|
return path if "1.8." in out.stdout else None
|
||||||
|
|
||||||
|
|
||||||
|
def write_sample(path, libver):
|
||||||
|
with clawhdf5.File(path, "w", libver=libver) as f:
|
||||||
|
f.create_dataset("x", data=np.arange(10, dtype=np.float64))
|
||||||
|
f.create_dataset(
|
||||||
|
"chunked",
|
||||||
|
data=np.arange(100, dtype=np.int32),
|
||||||
|
chunks=(30,),
|
||||||
|
compression="gzip",
|
||||||
|
)
|
||||||
|
f.create_dataset("z", data=np.array([1 + 2j, -3j], dtype=np.complex128))
|
||||||
|
g = f.create_group("g")
|
||||||
|
g.create_dataset("y", data=np.ones(3, dtype=np.float32))
|
||||||
|
f.attrs["version"] = 1
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"libver, sb",
|
||||||
|
[
|
||||||
|
(None, 3),
|
||||||
|
("v108", 2),
|
||||||
|
(("v108", "v108"), 2),
|
||||||
|
(("v108", "latest"), 2),
|
||||||
|
("v110", 3),
|
||||||
|
("v112", 3),
|
||||||
|
("v114", 3),
|
||||||
|
("v200", 3),
|
||||||
|
("latest", 3),
|
||||||
|
(("v110", "v200"), 3),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_libver_bounds(h5py, tmp_path, libver, sb):
|
||||||
|
path = str(tmp_path / "libver.h5")
|
||||||
|
write_sample(path, libver)
|
||||||
|
# v108 low bound: HDF5 1.8's version-2 superblock; 1.10 and later: 3.
|
||||||
|
assert superblock_version(path) == sb
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
|
||||||
|
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
|
||||||
|
np.testing.assert_array_equal(f["z"][:], [1 + 2j, -3j])
|
||||||
|
np.testing.assert_array_equal(f["g/y"][:], np.ones(3))
|
||||||
|
assert f.attrs["version"] == 1
|
||||||
|
with clawhdf5.File(path, "r") as f:
|
||||||
|
np.testing.assert_array_equal(f["chunked"][:], np.arange(100))
|
||||||
|
|
||||||
|
|
||||||
|
def test_libver_v108_opens_in_hdf5_1_8(tmp_path):
|
||||||
|
h5dump = h5dump18()
|
||||||
|
if h5dump is None:
|
||||||
|
pytest.skip("no HDF5 1.8 h5dump (CLAWHDF5_H5DUMP18)")
|
||||||
|
path = str(tmp_path / "v108.h5")
|
||||||
|
write_sample(path, "v108")
|
||||||
|
out = subprocess.run([h5dump, path], capture_output=True, text=True)
|
||||||
|
assert out.returncode == 0, out.stderr
|
||||||
|
assert "h5dump error" not in out.stderr
|
||||||
|
assert 'DATASET "y"' in out.stdout
|
||||||
|
# The chunked, deflated dataset (a version-1 B-tree) reads in full.
|
||||||
|
out = subprocess.run(
|
||||||
|
[h5dump, "-w", "0", "-d", "/chunked", path], capture_output=True, text=True
|
||||||
|
)
|
||||||
|
assert out.returncode == 0, out.stderr
|
||||||
|
assert "(0): " + ", ".join(str(k) for k in range(100)) + "\n" in out.stdout
|
||||||
|
# The default (1.10 format) does not open in 1.8: the bound matters.
|
||||||
|
path = str(tmp_path / "default.h5")
|
||||||
|
write_sample(path, None)
|
||||||
|
out = subprocess.run([h5dump, path], capture_output=True, text=True)
|
||||||
|
assert out.returncode != 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_libver_earliest_writes_v108_with_a_warning(h5py, tmp_path):
|
||||||
|
path = str(tmp_path / "earliest.h5")
|
||||||
|
with pytest.warns(UserWarning, match="pre-1.8"):
|
||||||
|
write_sample(path, "earliest")
|
||||||
|
assert superblock_version(path) == 2
|
||||||
|
with h5py.File(path, "r") as f:
|
||||||
|
np.testing.assert_array_equal(f["x"][:], np.arange(10.0))
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"libver, match",
|
||||||
|
[
|
||||||
|
(("v108", "earliest"), "high bound 'earliest'"),
|
||||||
|
("v109", "unknown libver 'v109'"),
|
||||||
|
(("latest", "v108"), "newer than the high bound"),
|
||||||
|
(("v108",), "tuple"),
|
||||||
|
(3, "tuple"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_libver_errors(tmp_path, libver, match):
|
||||||
|
with pytest.raises(ValueError, match=match):
|
||||||
|
clawhdf5.File(str(tmp_path / "bad.h5"), "w", libver=libver)
|
||||||
|
|
||||||
|
|
||||||
|
def test_libver_v108_high_bound_writes_everything(tmp_path):
|
||||||
|
# Native complex numbers need HDF5 2.0; the Python writer stores complex
|
||||||
|
# as h5py's {r, i} compound, which 1.8 reads, so the 1.8 high bound is
|
||||||
|
# fine for everything clawhdf5.File writes.
|
||||||
|
path = str(tmp_path / "v18only.h5")
|
||||||
|
write_sample(path, ("v108", "v108"))
|
||||||
|
assert superblock_version(path) == 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_libver_not_for_editing(tmp_path):
|
||||||
|
path = str(tmp_path / "e.h5")
|
||||||
|
write_sample(path, None)
|
||||||
|
with pytest.raises(NotImplementedError, match="libver"):
|
||||||
|
clawhdf5.File(path, "r+", libver="latest")
|
||||||
|
# Reading ignores it, as the bounds only affect what is written.
|
||||||
|
with clawhdf5.File(path, "r", libver="v108") as f:
|
||||||
|
assert f["x"].shape == (10,)
|
||||||
@@ -873,9 +873,9 @@ impl Checker<'_> {
|
|||||||
self.counts.chunk_index_checksummed += 1;
|
self.counts.chunk_index_checksummed += 1;
|
||||||
}
|
}
|
||||||
let filtered = matches!(&info.filters, Ok(Some(p)) if !p.filters.is_empty());
|
let filtered = matches!(&info.filters, Ok(Some(p)) if !p.filters.is_empty());
|
||||||
let chunk_bytes = cdims.iter().try_fold(u64::from(dt.type_size()), |a, &d| {
|
let chunk_bytes = cdims
|
||||||
a.checked_mul(u64::from(d))
|
.iter()
|
||||||
});
|
.try_fold(u64::from(dt.type_size()), |a, &d| a.checked_mul(d));
|
||||||
let max = ds.max_dimensions.clone();
|
let max = ds.max_dimensions.clone();
|
||||||
let mut seen: HashSet<Vec<u64>> = HashSet::with_capacity(chunks.len().min(1 << 20));
|
let mut seen: HashSet<Vec<u64>> = HashSet::with_capacity(chunks.len().min(1 << 20));
|
||||||
let mut reported = 0usize;
|
let mut reported = 0usize;
|
||||||
@@ -893,7 +893,7 @@ impl Checker<'_> {
|
|||||||
));
|
));
|
||||||
} else {
|
} else {
|
||||||
for (i, (&o, &cd)) in c.offsets.iter().zip(cdims).enumerate() {
|
for (i, (&o, &cd)) in c.offsets.iter().zip(cdims).enumerate() {
|
||||||
if o % u64::from(cd) != 0 {
|
if o % cd != 0 {
|
||||||
bad.push(format!(
|
bad.push(format!(
|
||||||
"offset {o} in dimension {i} is not a multiple of the chunk size {cd}"
|
"offset {o} in dimension {i} is not a multiple of the chunk size {cd}"
|
||||||
));
|
));
|
||||||
|
|||||||
@@ -603,7 +603,22 @@ impl Diff {
|
|||||||
let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) {
|
let dif = match (int_of(&va), int_of(&vb), number(&va), number(&vb)) {
|
||||||
(Some(x), Some(y), ..) => x.abs_diff(y).to_string(),
|
(Some(x), Some(y), ..) => x.abs_diff(y).to_string(),
|
||||||
(_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64),
|
(_, _, Some(x), Some(y)) => value::fmt_float((x - y).abs(), 64),
|
||||||
_ => String::new(),
|
// Each part's absolute difference, as h5diff 2.x prints
|
||||||
|
// it (`1+0i`).
|
||||||
|
_ => match (&va, &vb) {
|
||||||
|
(
|
||||||
|
Value::Complex { re: ar, im: ai, .. },
|
||||||
|
Value::Complex { re: br, im: bi, .. },
|
||||||
|
) => match (number(ar), number(ai), number(br), number(bi)) {
|
||||||
|
(Some(ar), Some(ai), Some(br), Some(bi)) => format!(
|
||||||
|
"{}+{}i",
|
||||||
|
value::fmt_float((ar - br).abs(), 64),
|
||||||
|
value::fmt_float((ai - bi).abs(), 64)
|
||||||
|
),
|
||||||
|
_ => String::new(),
|
||||||
|
},
|
||||||
|
_ => String::new(),
|
||||||
|
},
|
||||||
};
|
};
|
||||||
rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}"));
|
rows.push(format!("{pos:<24}{ta:<24}{tb:<24}{dif}"));
|
||||||
}
|
}
|
||||||
@@ -713,6 +728,9 @@ impl Diff {
|
|||||||
(Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => {
|
(Value::Array(p), Value::Array(q)) | (Value::Seq(p), Value::Seq(q)) => {
|
||||||
p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v))
|
p.len() == q.len() && p.iter().zip(q).all(|(u, v)| self.equal(a, b, u, v))
|
||||||
}
|
}
|
||||||
|
(Value::Complex { re: pr, im: pi, .. }, Value::Complex { re: qr, im: qi, .. }) => {
|
||||||
|
self.equal(a, b, pr, qr) && self.equal(a, b, pi, qi)
|
||||||
|
}
|
||||||
(Value::Ref(None), Value::Ref(None)) => true,
|
(Value::Ref(None), Value::Ref(None)) => true,
|
||||||
(Value::Ref(Some(p)), Value::Ref(Some(q))) => {
|
(Value::Ref(Some(p)), Value::Ref(Some(q))) => {
|
||||||
// Addresses mean nothing across files: compare the paths the
|
// Addresses mean nothing across files: compare the paths the
|
||||||
@@ -846,7 +864,10 @@ fn types_comparable(a: &Datatype, b: &Datatype) -> bool {
|
|||||||
(
|
(
|
||||||
Datatype::Enumeration { base_type: x, .. },
|
Datatype::Enumeration { base_type: x, .. },
|
||||||
Datatype::Enumeration { base_type: y, .. },
|
Datatype::Enumeration { base_type: y, .. },
|
||||||
) => types_comparable(x, y),
|
)
|
||||||
|
| (Datatype::Complex { base_type: x, .. }, Datatype::Complex { base_type: y, .. }) => {
|
||||||
|
types_comparable(x, y)
|
||||||
|
}
|
||||||
_ => true,
|
_ => true,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -135,9 +135,16 @@ pub fn short(dt: &Datatype) -> String {
|
|||||||
return f.short.into();
|
return f.short.into();
|
||||||
}
|
}
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::Complex { size, base_type } => {
|
// numpy's names (`complex64` is two `float32`) for IEEE parts,
|
||||||
short(&Datatype::complex_as_compound(*size, base_type))
|
// `complex<part>` otherwise.
|
||||||
}
|
Datatype::Complex { size, base_type } => match base_type.as_ref() {
|
||||||
|
Datatype::FloatingPoint { byte_order, .. } if is_ieee(base_type) => format!(
|
||||||
|
"complex{}{}",
|
||||||
|
u64::from(*size) * 8,
|
||||||
|
if be(byte_order) { "-be" } else { "" }
|
||||||
|
),
|
||||||
|
_ => format!("complex<{}>", short(base_type)),
|
||||||
|
},
|
||||||
Datatype::FixedPoint {
|
Datatype::FixedPoint {
|
||||||
size,
|
size,
|
||||||
signed,
|
signed,
|
||||||
@@ -216,8 +223,9 @@ pub fn long(dt: &Datatype) -> String {
|
|||||||
return f.long.into();
|
return f.long.into();
|
||||||
}
|
}
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::Complex { size, base_type } => {
|
// As h5ls 2.x prints a complex type it has no native name for.
|
||||||
long(&Datatype::complex_as_compound(*size, base_type))
|
Datatype::Complex { base_type, .. } => {
|
||||||
|
format!("complex number of\n {}", long(base_type))
|
||||||
}
|
}
|
||||||
Datatype::FixedPoint {
|
Datatype::FixedPoint {
|
||||||
size,
|
size,
|
||||||
@@ -361,6 +369,19 @@ fn atomic_ddl(dt: &Datatype) -> Option<String> {
|
|||||||
)),
|
)),
|
||||||
// As h5dump 2.x names them (checked against h5dump 2.2.0).
|
// As h5dump 2.x names them (checked against h5dump 2.2.0).
|
||||||
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
|
Datatype::FloatingPoint { .. } => small_float(dt).map(|f| f.ddl.to_string()),
|
||||||
|
// h5dump 2.x's predefined complex names: IEEE binary16/32/64 parts.
|
||||||
|
Datatype::Complex { base_type, .. } => match base_type.as_ref() {
|
||||||
|
Datatype::FloatingPoint {
|
||||||
|
size, byte_order, ..
|
||||||
|
} if is_ieee(base_type) && !matches!(byte_order, DatatypeByteOrder::Vax) => {
|
||||||
|
Some(format!(
|
||||||
|
"H5T_COMPLEX_IEEE_F{}{}",
|
||||||
|
u64::from(*size) * 8,
|
||||||
|
order_suffix(byte_order)
|
||||||
|
))
|
||||||
|
}
|
||||||
|
_ => None,
|
||||||
|
},
|
||||||
Datatype::BitField {
|
Datatype::BitField {
|
||||||
size, byte_order, ..
|
size, byte_order, ..
|
||||||
} => Some(format!(
|
} => Some(format!(
|
||||||
@@ -418,6 +439,10 @@ pub fn ddl(dt: &Datatype, ind: usize) -> String {
|
|||||||
Datatype::VariableLength { base_type, .. } => {
|
Datatype::VariableLength { base_type, .. } => {
|
||||||
format!("H5T_VLEN {{ {} }}", ddl(base_type, ind))
|
format!("H5T_VLEN {{ {} }}", ddl(base_type, ind))
|
||||||
}
|
}
|
||||||
|
// Parts h5dump has no predefined complex name for.
|
||||||
|
Datatype::Complex { base_type, .. } => {
|
||||||
|
format!("H5T_COMPLEX {{ {} }}", ddl(base_type, ind))
|
||||||
|
}
|
||||||
Datatype::Opaque { size, tag } => {
|
Datatype::Opaque { size, tag } => {
|
||||||
let tag = String::from_utf8_lossy(tag);
|
let tag = String::from_utf8_lossy(tag);
|
||||||
format!(
|
format!(
|
||||||
@@ -498,6 +523,9 @@ fn string_ddl(
|
|||||||
/// hdf5-json type object.
|
/// hdf5-json type object.
|
||||||
pub fn json(dt: &Datatype) -> J {
|
pub fn json(dt: &Datatype) -> J {
|
||||||
match dt {
|
match dt {
|
||||||
|
// hdf5-json (h5json 2.0.0) has no complex class: a native complex
|
||||||
|
// is written as the `{r, i}` compound h5py uses for complex numbers,
|
||||||
|
// its values as `[re, im]` pairs.
|
||||||
Datatype::Complex { size, base_type } => {
|
Datatype::Complex { size, base_type } => {
|
||||||
json(&Datatype::complex_as_compound(*size, base_type))
|
json(&Datatype::complex_as_compound(*size, base_type))
|
||||||
}
|
}
|
||||||
@@ -598,7 +626,7 @@ fn pad_json(p: &StringPadding) -> &'static str {
|
|||||||
/// Class name used to decide whether two datatypes can be compared.
|
/// Class name used to decide whether two datatypes can be compared.
|
||||||
pub fn class(dt: &Datatype) -> &'static str {
|
pub fn class(dt: &Datatype) -> &'static str {
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::Complex { .. } => "compound",
|
Datatype::Complex { .. } => "complex",
|
||||||
Datatype::FixedPoint { .. } => "integer",
|
Datatype::FixedPoint { .. } => "integer",
|
||||||
Datatype::FloatingPoint { .. } => "float",
|
Datatype::FloatingPoint { .. } => "float",
|
||||||
Datatype::Time { .. } => "time",
|
Datatype::Time { .. } => "time",
|
||||||
|
|||||||
@@ -292,7 +292,7 @@ impl Ls<'_> {
|
|||||||
.unwrap_or(0);
|
.unwrap_or(0);
|
||||||
let bytes = dims
|
let bytes = dims
|
||||||
.iter()
|
.iter()
|
||||||
.try_fold(esize, |a, &d| a.checked_mul(u64::from(d)))
|
.try_fold(esize, |a, &d| a.checked_mul(d))
|
||||||
.map(|b| b.to_string())
|
.map(|b| b.to_string())
|
||||||
.unwrap_or_else(|| "?".into());
|
.unwrap_or_else(|| "?".into());
|
||||||
writeln!(
|
writeln!(
|
||||||
|
|||||||
@@ -27,6 +27,15 @@ pub enum Value {
|
|||||||
/// An enum member (name, when the value matches one) and its value.
|
/// An enum member (name, when the value matches one) and its value.
|
||||||
Enum(Option<String>, i128),
|
Enum(Option<String>, i128),
|
||||||
Compound(Vec<(String, Value)>),
|
Compound(Vec<(String, Value)>),
|
||||||
|
/// A native complex number (HDF5 2.0 class 11): real and imaginary
|
||||||
|
/// part. `sign_free` is set when h5dump joins the parts with a bare `+`
|
||||||
|
/// (parts it has no native C type for: binary16, non-IEEE), giving
|
||||||
|
/// `1+-2i`; otherwise the imaginary part carries its sign (`1-2i`).
|
||||||
|
Complex {
|
||||||
|
re: Box<Value>,
|
||||||
|
im: Box<Value>,
|
||||||
|
sign_free: bool,
|
||||||
|
},
|
||||||
Array(Vec<Value>),
|
Array(Vec<Value>),
|
||||||
/// A variable-length sequence.
|
/// A variable-length sequence.
|
||||||
Seq(Vec<Value>),
|
Seq(Vec<Value>),
|
||||||
@@ -183,11 +192,18 @@ impl<'a> Decoder<'a> {
|
|||||||
return Value::Error("short element".into());
|
return Value::Error("short element".into());
|
||||||
};
|
};
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::Complex { size, base_type } => self.decode(
|
Datatype::Complex { base_type, .. } => {
|
||||||
&Datatype::complex_as_compound(*size, base_type),
|
let part = base_type.type_size() as usize;
|
||||||
b,
|
let (re, im) = b.split_at(part.min(b.len()));
|
||||||
depth + 1,
|
Value::Complex {
|
||||||
),
|
re: Box::new(self.decode(base_type, re, depth + 1)),
|
||||||
|
im: Box::new(self.decode(base_type, im, depth + 1)),
|
||||||
|
// h5dump prints `float`/`double` complex (whatever the
|
||||||
|
// byte order) with `%g%+gi`, anything else as
|
||||||
|
// `<re>+<im>i`.
|
||||||
|
sign_free: !(dtype::is_ieee(base_type) && matches!(part, 4 | 8)),
|
||||||
|
}
|
||||||
|
}
|
||||||
Datatype::FixedPoint { .. } => match decode_int(dt, b) {
|
Datatype::FixedPoint { .. } => match decode_int(dt, b) {
|
||||||
Some(v) => Value::Int(v),
|
Some(v) => Value::Int(v),
|
||||||
None => Value::Bytes(b.to_vec()),
|
None => Value::Bytes(b.to_vec()),
|
||||||
@@ -354,6 +370,15 @@ pub fn text(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> String {
|
|||||||
.collect::<Vec<_>>()
|
.collect::<Vec<_>>()
|
||||||
.join(", ")
|
.join(", ")
|
||||||
),
|
),
|
||||||
|
Value::Complex { re, im, sign_free } => {
|
||||||
|
let (re, im) = (text(re, h5paths), text(im, h5paths));
|
||||||
|
let plus = if *sign_free || !im.starts_with('-') {
|
||||||
|
"+"
|
||||||
|
} else {
|
||||||
|
""
|
||||||
|
};
|
||||||
|
format!("{re}{plus}{im}i")
|
||||||
|
}
|
||||||
Value::Array(vs) => format!(
|
Value::Array(vs) => format!(
|
||||||
"[ {} ]",
|
"[ {} ]",
|
||||||
vs.iter()
|
vs.iter()
|
||||||
@@ -408,6 +433,7 @@ pub fn to_json(v: &Value, h5paths: &dyn Fn(u64) -> Option<String>) -> J {
|
|||||||
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
|
Value::Bytes(b) | Value::OtherRef(b) => J::from(hex(b)),
|
||||||
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
|
Value::Enum(_, i) => to_json(&Value::Int(*i), h5paths),
|
||||||
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
|
Value::Compound(ms) => J::Array(ms.iter().map(|(_, v)| to_json(v, h5paths)).collect()),
|
||||||
|
Value::Complex { re, im, .. } => J::Array(vec![to_json(re, h5paths), to_json(im, h5paths)]),
|
||||||
Value::Array(vs) | Value::Seq(vs) => {
|
Value::Array(vs) | Value::Seq(vs) => {
|
||||||
J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect())
|
J::Array(vs.iter().map(|v| to_json(v, h5paths)).collect())
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,202 @@
|
|||||||
|
//! `h5rs dump` and `h5rs ls` of HDF5 2.0 native complex numbers (datatype
|
||||||
|
//! class 11), against the output of h5dump 2.2.0 and h5ls 2.2.0 (the Debian
|
||||||
|
//! h5dump CI installs, 1.14.x, predates the type). The h5dump output is
|
||||||
|
//! stored next to the fixture; see
|
||||||
|
//! `crates/clawhdf5/tests/fixtures/gen_native_complex.py`.
|
||||||
|
|
||||||
|
use std::path::PathBuf;
|
||||||
|
use std::process::Command;
|
||||||
|
|
||||||
|
fn fixture(name: &str) -> PathBuf {
|
||||||
|
PathBuf::from(env!("CARGO_MANIFEST_DIR"))
|
||||||
|
.join("../clawhdf5/tests/fixtures")
|
||||||
|
.join(name)
|
||||||
|
}
|
||||||
|
|
||||||
|
fn h5rs(args: &[&str]) -> String {
|
||||||
|
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
|
||||||
|
.args(args)
|
||||||
|
.output()
|
||||||
|
.unwrap();
|
||||||
|
assert!(out.status.success(), "h5rs {args:?}: {out:?}");
|
||||||
|
String::from_utf8(out.stdout).unwrap()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Split a dump into its lines, with the data lines (`(i): ...`) replaced
|
||||||
|
/// by their index, and the data values in order.
|
||||||
|
fn split(ddl: &str) -> (Vec<String>, Vec<String>) {
|
||||||
|
let (mut frame, mut data) = (Vec::new(), Vec::new());
|
||||||
|
for line in ddl.lines() {
|
||||||
|
match line.split_once("):") {
|
||||||
|
Some((idx, body)) if idx.trim_start().starts_with('(') => {
|
||||||
|
frame.push(format!("{idx}):"));
|
||||||
|
// Values are split at `, `; array elements come bracketed.
|
||||||
|
data.extend(
|
||||||
|
body.split(", ")
|
||||||
|
.map(|s| s.trim().trim_matches(|c| c == '[' || c == ']' || c == ','))
|
||||||
|
.map(str::trim)
|
||||||
|
.filter(|s| !s.is_empty() && *s != "{")
|
||||||
|
.map(String::from),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
_ => {
|
||||||
|
let t = line.trim().trim_end_matches(',');
|
||||||
|
if t.ends_with('i') || t.parse::<f64>().is_ok() {
|
||||||
|
// A compound member's value on its own line.
|
||||||
|
data.push(t.to_string());
|
||||||
|
} else {
|
||||||
|
frame.push(line.to_string());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
(frame, data)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `a+bi`, `a-bi` or `a+-bi` as its two parts (a real number as one).
|
||||||
|
fn parts(v: &str) -> Vec<f64> {
|
||||||
|
let Some(z) = v.strip_suffix('i') else {
|
||||||
|
return vec![v.parse().unwrap()];
|
||||||
|
};
|
||||||
|
let b = z.as_bytes();
|
||||||
|
let at = (1..b.len())
|
||||||
|
.find(|&k| (b[k] == b'+' || b[k] == b'-') && !matches!(b[k - 1], b'e' | b'E' | b'+'))
|
||||||
|
.unwrap_or_else(|| panic!("not a complex value: {v}"));
|
||||||
|
let (re, im) = z.split_at(at);
|
||||||
|
let im = im.strip_prefix('+').unwrap_or(im);
|
||||||
|
vec![re.parse().unwrap(), im.parse().unwrap()]
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn dump_matches_h5dump_2_2() {
|
||||||
|
let path = fixture("native_complex_hdf5_2.h5");
|
||||||
|
let ours = h5rs(&["dump", path.to_str().unwrap()]);
|
||||||
|
let reference = std::fs::read_to_string(fixture("native_complex_hdf5_2.ddl")).unwrap();
|
||||||
|
let (mut our_frame, our_data) = split(&ours);
|
||||||
|
let (mut ref_frame, ref_data) = split(&reference);
|
||||||
|
// The first line names the file as it was given.
|
||||||
|
our_frame.remove(0);
|
||||||
|
ref_frame.remove(0);
|
||||||
|
// Everything but the values is h5dump's byte for byte: the types print
|
||||||
|
// as H5T_COMPLEX_IEEE_F64LE, H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }, ...
|
||||||
|
assert_eq!(our_frame, ref_frame);
|
||||||
|
assert_eq!(our_data.len(), ref_data.len(), "{our_data:?}\n{ref_data:?}");
|
||||||
|
assert_eq!(our_data.len(), 36);
|
||||||
|
// Values: h5dump prints `%g` (6 digits), inf/nan, and the binary16
|
||||||
|
// parts it has no C type for as `<re>+<im>i` (`1+-2i`); h5rs prints
|
||||||
|
// the shortest round-trip string and Inf/NaN, joined the same way.
|
||||||
|
for (i, (o, r)) in our_data.iter().zip(&ref_data).enumerate() {
|
||||||
|
assert_eq!(o.contains("+-"), r.contains("+-"), "[{i}] {o} vs {r}");
|
||||||
|
let (op, rp) = (parts(o), parts(r));
|
||||||
|
assert_eq!(op.len(), rp.len(), "[{i}] {o} vs {r}");
|
||||||
|
for (x, y) in op.iter().zip(&rp) {
|
||||||
|
let same = if y.is_nan() || y.is_infinite() {
|
||||||
|
x.is_nan() == y.is_nan() && (y.is_nan() || x == y)
|
||||||
|
} else {
|
||||||
|
*x as f32 == *y as f32 || (x - y).abs() <= 5e-6 * y.abs()
|
||||||
|
};
|
||||||
|
assert!(same, "[{i}] h5rs {o}, h5dump {r}");
|
||||||
|
assert_eq!(
|
||||||
|
x.is_sign_negative(),
|
||||||
|
y.is_sign_negative(),
|
||||||
|
"[{i}] {o} vs {r}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn ls_describes_complex_like_h5ls_2_2() {
|
||||||
|
let path = fixture("native_complex_hdf5_2.h5");
|
||||||
|
// h5ls 2.2.0 -v, `Type:` of the types it has no native C name for (it
|
||||||
|
// prints `native double _Complex` for the others, as it prints `native
|
||||||
|
// double` where h5rs prints `IEEE 64-bit little-endian float`).
|
||||||
|
for (name, long) in [
|
||||||
|
(
|
||||||
|
"c128be",
|
||||||
|
"complex number of\n IEEE 64-bit big-endian float",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"c128",
|
||||||
|
"complex number of\n IEEE 64-bit little-endian float",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"c32",
|
||||||
|
"complex number of\n IEEE 16-bit little-endian float",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"array",
|
||||||
|
"[2] complex number of\n IEEE 32-bit little-endian float",
|
||||||
|
),
|
||||||
|
] {
|
||||||
|
let target = format!("{}/{name}", path.display());
|
||||||
|
let verbose = h5rs(&["ls", "-v", &target]);
|
||||||
|
assert!(
|
||||||
|
verbose.contains(&format!(" Type: {long}\n")),
|
||||||
|
"{name}:\n{verbose}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
let listing = h5rs(&["ls", path.to_str().unwrap()]);
|
||||||
|
for (name, short) in [
|
||||||
|
("c64", "complex64"),
|
||||||
|
("c128", "complex128"),
|
||||||
|
("c128be", "complex128-be"),
|
||||||
|
("c32", "complex32"),
|
||||||
|
("scalar", "complex128"),
|
||||||
|
("array", "array[2]<complex64>"),
|
||||||
|
("compound", "compound{z: complex128, k: int64}"),
|
||||||
|
] {
|
||||||
|
assert!(
|
||||||
|
listing
|
||||||
|
.lines()
|
||||||
|
.any(|l| l.starts_with(&format!("{name} ")) && l.ends_with(&format!(" {short}"))),
|
||||||
|
"{name}:\n{listing}"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn json_dump_uses_the_r_i_compound() {
|
||||||
|
// hdf5-json (h5json 2.0.0) has no complex class: the type is written as
|
||||||
|
// the `{r, i}` compound h5py uses, the values as `[re, im]` pairs.
|
||||||
|
let path = fixture("native_complex_hdf5_2.h5");
|
||||||
|
let j: serde_json::Value =
|
||||||
|
serde_json::from_str(&h5rs(&["dump", "--json", path.to_str().unwrap()])).unwrap();
|
||||||
|
let ds = j["datasets"]
|
||||||
|
.as_object()
|
||||||
|
.unwrap()
|
||||||
|
.values()
|
||||||
|
.find(|d| d["alias"][0] == "/scalar")
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(ds["type"]["class"], "H5T_COMPOUND");
|
||||||
|
assert_eq!(ds["type"]["fields"][0]["name"], "r");
|
||||||
|
assert_eq!(ds["type"]["fields"][1]["type"]["base"], "H5T_IEEE_F64LE");
|
||||||
|
assert_eq!(ds["value"], serde_json::json!([2.5, -0.5]));
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn diff_compares_complex_values() {
|
||||||
|
let path = fixture("native_complex_hdf5_2.h5");
|
||||||
|
let p = path.to_str().unwrap();
|
||||||
|
// Same values, different byte order: no differences.
|
||||||
|
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
|
||||||
|
.args(["diff", p, p, "/c128", "/c128be"])
|
||||||
|
.output()
|
||||||
|
.unwrap();
|
||||||
|
assert!(out.status.success(), "{out:?}");
|
||||||
|
// One element differs in its imaginary part: one difference, exit
|
||||||
|
// status 1, as h5diff 2.2.0 reports it.
|
||||||
|
let out = Command::new(env!("CARGO_BIN_EXE_h5rs"))
|
||||||
|
.args(["diff", "-r", p, p, "/c128", "/c128x"])
|
||||||
|
.output()
|
||||||
|
.unwrap();
|
||||||
|
assert_eq!(out.status.code(), Some(1), "{out:?}");
|
||||||
|
let text = String::from_utf8_lossy(&out.stdout);
|
||||||
|
assert!(text.contains("1 difference(s) found"), "{text}");
|
||||||
|
// h5diff 2.2.0: `[ 1 1 ] 0+1i 1+1i 1+0i`.
|
||||||
|
let row = text.lines().find(|l| l.starts_with("[ 1 1 ]")).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
row.split_whitespace().collect::<Vec<_>>(),
|
||||||
|
["[", "1", "1", "]", "0+1i", "1+1i", "1+0i"]
|
||||||
|
);
|
||||||
|
}
|
||||||
@@ -312,11 +312,14 @@ impl Reader {
|
|||||||
let base = array_base(dt);
|
let base = array_base(dt);
|
||||||
let is_array = !std::ptr::eq(base, dt);
|
let is_array = !std::ptr::eq(base, dt);
|
||||||
Ok(match base {
|
Ok(match base {
|
||||||
|
// Read as the part type: an array or complex element is its
|
||||||
|
// parts stored one after another (a complex number's real, then
|
||||||
|
// imaginary part).
|
||||||
Datatype::FloatingPoint { size, .. } if *size <= 4 => {
|
Datatype::FloatingPoint { size, .. } if *size <= 4 => {
|
||||||
Data::F32(data_read::read_as_f32(raw, dt).map_err(err)?)
|
Data::F32(data_read::read_as_f32(raw, base).map_err(err)?)
|
||||||
}
|
}
|
||||||
Datatype::FloatingPoint { .. } => {
|
Datatype::FloatingPoint { .. } => {
|
||||||
Data::F64(data_read::read_as_f64(raw, dt).map_err(err)?)
|
Data::F64(data_read::read_as_f64(raw, base).map_err(err)?)
|
||||||
}
|
}
|
||||||
Datatype::FixedPoint { size, signed, .. } => {
|
Datatype::FixedPoint { size, signed, .. } => {
|
||||||
let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err);
|
let signed_ints = || data_read::read_as_i64(raw, dt).map_err(err);
|
||||||
@@ -393,10 +396,13 @@ fn narrow<S: Copy + std::fmt::Display, T: TryFrom<S>>(v: Vec<S>) -> Result<Vec<T
|
|||||||
.collect()
|
.collect()
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Innermost element type of (possibly nested) array datatypes.
|
/// Innermost element type of (possibly nested) array datatypes; for a
|
||||||
|
/// native complex type, its part type (each element holds two).
|
||||||
fn array_base(dt: &Datatype) -> &Datatype {
|
fn array_base(dt: &Datatype) -> &Datatype {
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::Array { base_type, .. } => array_base(base_type),
|
Datatype::Array { base_type, .. } | Datatype::Complex { base_type, .. } => {
|
||||||
|
array_base(base_type)
|
||||||
|
}
|
||||||
_ => dt,
|
_ => dt,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -411,6 +417,12 @@ fn element_shape(dt: &Datatype) -> Vec<u64> {
|
|||||||
dims.extend(element_shape(base_type));
|
dims.extend(element_shape(base_type));
|
||||||
dims
|
dims
|
||||||
}
|
}
|
||||||
|
// `[re, im]`: a complex element reads as its two parts.
|
||||||
|
Datatype::Complex { base_type, .. } => {
|
||||||
|
let mut dims = vec![2];
|
||||||
|
dims.extend(element_shape(base_type));
|
||||||
|
dims
|
||||||
|
}
|
||||||
_ => Vec::new(),
|
_ => Vec::new(),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -487,9 +499,7 @@ pub fn describe(dt: &Datatype) -> String {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
match dt {
|
match dt {
|
||||||
Datatype::Complex { size, base_type } => {
|
Datatype::Complex { base_type, .. } => format!("complex<{}>", describe(base_type)),
|
||||||
describe(&Datatype::complex_as_compound(*size, base_type))
|
|
||||||
}
|
|
||||||
Datatype::FixedPoint {
|
Datatype::FixedPoint {
|
||||||
size,
|
size,
|
||||||
signed,
|
signed,
|
||||||
@@ -681,6 +691,74 @@ mod tests {
|
|||||||
assert!(e.contains("not supported"), "{e}");
|
assert!(e.contains("not supported"), "{e}");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn native_complex_reads_as_re_im_pairs() {
|
||||||
|
let z = [[1.0f64, -2.0], [0.5, 0.0], [-0.0, 3.25], [f64::MAX, 1e-300]];
|
||||||
|
let mut b = FileBuilder::new();
|
||||||
|
b.create_dataset("z")
|
||||||
|
.with_native_complex_f64_data(&z)
|
||||||
|
.with_shape(&[2, 2]);
|
||||||
|
b.create_dataset("z32")
|
||||||
|
.with_native_complex_f32_data(&[[1.5f32, -2.5]]);
|
||||||
|
let r = Reader::open(b.finish().unwrap()).unwrap();
|
||||||
|
let i = r.info("z").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
(i.shape, i.dtype, i.element_shape),
|
||||||
|
(vec![2, 2], "complex<f64>".to_string(), vec![2])
|
||||||
|
);
|
||||||
|
let all = r.read("z", None).unwrap();
|
||||||
|
assert_eq!(all.shape, vec![2, 2, 2]);
|
||||||
|
assert_eq!(all.data, Data::F64(z.iter().flatten().copied().collect()));
|
||||||
|
let slab = Hyperslab {
|
||||||
|
start: vec![1, 0],
|
||||||
|
count: vec![1, 2],
|
||||||
|
stride: None,
|
||||||
|
block: None,
|
||||||
|
};
|
||||||
|
let part = r.read("z", Some(&slab)).unwrap();
|
||||||
|
assert_eq!(part.shape, vec![1, 2, 2]);
|
||||||
|
assert_eq!(part.data, Data::F64(vec![-0.0, 3.25, f64::MAX, 1e-300]));
|
||||||
|
let z32 = r.read("z32", None).unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
(z32.shape, z32.data),
|
||||||
|
(vec![1, 2], Data::F32(vec![1.5, -2.5]))
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
#[test]
|
||||||
|
fn native_complex_written_by_libhdf5_2() {
|
||||||
|
// Written by h5py 3.16 / libhdf5 2.0.0 (gen_native_complex.py).
|
||||||
|
let path = concat!(
|
||||||
|
env!("CARGO_MANIFEST_DIR"),
|
||||||
|
"/../clawhdf5/tests/fixtures/native_complex_hdf5_2.h5"
|
||||||
|
);
|
||||||
|
let r = Reader::open(std::fs::read(path).unwrap()).unwrap();
|
||||||
|
let be = r.read("c128be", None).unwrap();
|
||||||
|
assert_eq!(be.shape, vec![2, 3, 2]);
|
||||||
|
assert_eq!(be.data, r.read("c128", None).unwrap().data);
|
||||||
|
assert_eq!(r.info("c128be").unwrap().dtype, "complex<f64 (big-endian)>");
|
||||||
|
// binary16 parts widen to f32.
|
||||||
|
assert_eq!(
|
||||||
|
r.read("c32", None).unwrap().data,
|
||||||
|
Data::F32(vec![1.0, -2.0, 0.5, 65504.0, 0.0, -0.0])
|
||||||
|
);
|
||||||
|
// An array of complex: array dimensions, then the two parts.
|
||||||
|
let i = r.info("array").unwrap();
|
||||||
|
assert_eq!(
|
||||||
|
(i.dtype.as_str(), i.element_shape),
|
||||||
|
("array[2]<complex<f32>>", vec![2, 2])
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
r.read("array", None).unwrap().data,
|
||||||
|
Data::F32(vec![1.0, 2.0, 3.0, 4.0, -1.0, -1.0, 0.0, 0.0])
|
||||||
|
);
|
||||||
|
let s = r.read("scalar", None).unwrap();
|
||||||
|
assert_eq!((s.shape, s.data), (vec![2], Data::F64(vec![2.5, -0.5])));
|
||||||
|
// A compound holding a complex member is still refused.
|
||||||
|
let e = r.read("compound", None).unwrap_err();
|
||||||
|
assert!(e.contains("compound{z: complex<f64>, k: i64}"), "{e}");
|
||||||
|
}
|
||||||
|
|
||||||
#[test]
|
#[test]
|
||||||
fn garbage_is_an_error() {
|
fn garbage_is_an_error() {
|
||||||
assert!(Reader::open(vec![0u8; 64]).is_err());
|
assert!(Reader::open(vec![0u8; 64]).is_err());
|
||||||
|
|||||||
@@ -0,0 +1,46 @@
|
|||||||
|
//! Write one dataset of `u8` stored as a single chunk, for measuring the
|
||||||
|
//! writer's peak memory (`/usr/bin/time -v`):
|
||||||
|
//!
|
||||||
|
//! ```text
|
||||||
|
//! cargo run --release -p clawhdf5 --example write_one_chunk -- BYTES OUT.h5 [slice|owned] [deflate|lz4|zstd]
|
||||||
|
//! ```
|
||||||
|
//!
|
||||||
|
//! `slice` passes the data by reference (`with_u8_data`, which copies it),
|
||||||
|
//! `owned` hands the builder the vector (`with_u8_data_owned`). The data is
|
||||||
|
//! `i % 251`, one byte per element; with no filter the chunk is stored raw.
|
||||||
|
fn main() {
|
||||||
|
let args: Vec<String> = std::env::args().collect();
|
||||||
|
let n: u64 = args[1].parse().expect("BYTES");
|
||||||
|
let out = &args[2];
|
||||||
|
let mode = args.get(3).map_or("slice", String::as_str);
|
||||||
|
let filter = args.get(4).map(String::as_str);
|
||||||
|
let data: Vec<u8> = (0..n).map(|i| (i % 251) as u8).collect();
|
||||||
|
let mut b = clawhdf5::FileBuilder::new();
|
||||||
|
let ds = b.create_dataset("x");
|
||||||
|
match mode {
|
||||||
|
"slice" => {
|
||||||
|
ds.with_u8_data(&data);
|
||||||
|
}
|
||||||
|
"owned" => {
|
||||||
|
ds.with_u8_data_owned(data);
|
||||||
|
}
|
||||||
|
m => panic!("mode {m}"),
|
||||||
|
}
|
||||||
|
ds.with_chunks(&[n]);
|
||||||
|
match filter {
|
||||||
|
None => {}
|
||||||
|
Some("deflate") => {
|
||||||
|
ds.with_deflate(1);
|
||||||
|
}
|
||||||
|
#[cfg(feature = "lz4")]
|
||||||
|
Some("lz4") => {
|
||||||
|
ds.with_lz4();
|
||||||
|
}
|
||||||
|
#[cfg(feature = "zstd")]
|
||||||
|
Some("zstd") => {
|
||||||
|
ds.with_zstd(1);
|
||||||
|
}
|
||||||
|
Some(f) => panic!("filter {f} (not built in?)"),
|
||||||
|
}
|
||||||
|
b.write(out).unwrap();
|
||||||
|
}
|
||||||
@@ -408,10 +408,7 @@ impl<'t> ChunkedEdit<'t> {
|
|||||||
"chunk rank differs from the dataspace".into(),
|
"chunk rank differs from the dataspace".into(),
|
||||||
));
|
));
|
||||||
}
|
}
|
||||||
let cd: Vec<u64> = chunk_dimensions[..rank]
|
let cd: Vec<u64> = chunk_dimensions[..rank].to_vec();
|
||||||
.iter()
|
|
||||||
.map(|&c| u64::from(c))
|
|
||||||
.collect();
|
|
||||||
// Chunks of 4 GiB or more (HDF5 2.0 writes them with layout message
|
// Chunks of 4 GiB or more (HDF5 2.0 writes them with layout message
|
||||||
// version 5) are read, but not rewritten: the editor holds a chunk
|
// version 5) are read, but not rewritten: the editor holds a chunk
|
||||||
// it rewrites in memory, compressed and not. Writing values, and a
|
// it rewrites in memory, compressed and not. Writing values, and a
|
||||||
|
|||||||
@@ -1342,7 +1342,10 @@ impl<'f> Dataset<'f> {
|
|||||||
) -> Result<(Vec<T>, Vec<T>), Error> {
|
) -> Result<(Vec<T>, Vec<T>), Error> {
|
||||||
let dt = self.datatype()?;
|
let dt = self.datatype()?;
|
||||||
let is_complex = match &dt {
|
let is_complex = match &dt {
|
||||||
// Class 11 parses to this same `{r, i}` compound.
|
Datatype::Complex { size, base_type } => {
|
||||||
|
matches!(base_type.as_ref(), Datatype::FloatingPoint { .. })
|
||||||
|
&& *size == 2 * base_type.type_size()
|
||||||
|
}
|
||||||
Datatype::Compound { size, members } => {
|
Datatype::Compound { size, members } => {
|
||||||
matches!(members.as_slice(), [r, i]
|
matches!(members.as_slice(), [r, i]
|
||||||
if r.name == "r" && i.name == "i"
|
if r.name == "r" && i.name == "i"
|
||||||
|
|||||||
@@ -41,6 +41,13 @@ pub enum DType {
|
|||||||
Array(Box<DType>, Vec<u32>),
|
Array(Box<DType>, Vec<u32>),
|
||||||
/// Variable-length UTF-8 string stored via a global heap reference.
|
/// Variable-length UTF-8 string stored via a global heap reference.
|
||||||
VariableLengthString,
|
VariableLengthString,
|
||||||
|
/// HDF5 2.0 native complex number (datatype class 11) over the given
|
||||||
|
/// floating-point part type: `Complex(F32)` is numpy `complex64`,
|
||||||
|
/// `Complex(F64)` `complex128`, and a half-precision part is
|
||||||
|
/// `Complex(Other("float16"))`. h5py's default complex encoding, the
|
||||||
|
/// compound `{r, i}`, stays a [`DType::Compound`]; `read_complex_f64`
|
||||||
|
/// and `read_complex_f32` read both.
|
||||||
|
Complex(Box<DType>),
|
||||||
/// Catch-all for HDF5 datatypes that do not map to a specific variant.
|
/// Catch-all for HDF5 datatypes that do not map to a specific variant.
|
||||||
Other(String),
|
Other(String),
|
||||||
}
|
}
|
||||||
@@ -60,6 +67,7 @@ impl fmt::Display for DType {
|
|||||||
DType::U64 => write!(f, "u64"),
|
DType::U64 => write!(f, "u64"),
|
||||||
DType::String => write!(f, "string"),
|
DType::String => write!(f, "string"),
|
||||||
DType::VariableLengthString => write!(f, "vlen_string"),
|
DType::VariableLengthString => write!(f, "vlen_string"),
|
||||||
|
DType::Complex(base) => write!(f, "complex<{base}>"),
|
||||||
DType::Compound(fields) => {
|
DType::Compound(fields) => {
|
||||||
write!(f, "compound{{")?;
|
write!(f, "compound{{")?;
|
||||||
for (i, (name, dt)) in fields.iter().enumerate() {
|
for (i, (name, dt)) in fields.iter().enumerate() {
|
||||||
@@ -148,6 +156,9 @@ pub(crate) fn classify_datatype(dt: &clawhdf5_format::datatype::Datatype) -> DTy
|
|||||||
base_type,
|
base_type,
|
||||||
dimensions,
|
dimensions,
|
||||||
} => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()),
|
} => DType::Array(Box::new(classify_datatype(base_type)), dimensions.clone()),
|
||||||
|
Datatype::Complex { base_type, .. } => {
|
||||||
|
DType::Complex(Box::new(classify_datatype(base_type)))
|
||||||
|
}
|
||||||
_ => DType::Other(format!("{dt:?}")),
|
_ => DType::Other(format!("{dt:?}")),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -137,10 +137,15 @@ impl FileBuilder {
|
|||||||
Ok(self.writer.finish()?)
|
Ok(self.writer.finish()?)
|
||||||
}
|
}
|
||||||
|
|
||||||
/// Serialize and write the file to the given path.
|
/// Serialize and write the file to the given path. The file is written
|
||||||
|
/// as it is serialized, not assembled in memory first: unfiltered
|
||||||
|
/// chunks and contiguous data go straight from the datasets' data.
|
||||||
pub fn write<P: AsRef<std::path::Path>>(self, path: P) -> Result<(), Error> {
|
pub fn write<P: AsRef<std::path::Path>>(self, path: P) -> Result<(), Error> {
|
||||||
let bytes = self.finish()?;
|
use std::io::Write;
|
||||||
write_file_atomically(path.as_ref(), &bytes).map_err(Error::Io)
|
write_file_atomically_with(path.as_ref(), |f| {
|
||||||
|
self.writer
|
||||||
|
.finish_with(|bytes| f.write_all(bytes).map_err(Error::Io))
|
||||||
|
})
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -259,8 +264,19 @@ pub fn create_datasets_parallel(specs: Vec<DatasetSpec>) -> Result<Vec<u8>, Erro
|
|||||||
/// file or the complete new one — never a truncated mix. `std::fs::write`
|
/// file or the complete new one — never a truncated mix. `std::fs::write`
|
||||||
/// truncates the destination first, so dying mid-write used to destroy the
|
/// truncates the destination first, so dying mid-write used to destroy the
|
||||||
/// existing file.
|
/// existing file.
|
||||||
|
#[cfg(test)]
|
||||||
fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Result<()> {
|
fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Result<()> {
|
||||||
use std::io::Write;
|
use std::io::Write;
|
||||||
|
write_file_atomically_with(path, |f| f.write_all(bytes))
|
||||||
|
}
|
||||||
|
|
||||||
|
/// [`write_file_atomically`] with the contents written by `write` (into a
|
||||||
|
/// buffered writer over the temporary file).
|
||||||
|
fn write_file_atomically_with<E: From<std::io::Error>>(
|
||||||
|
path: &std::path::Path,
|
||||||
|
write: impl FnOnce(&mut std::io::BufWriter<std::fs::File>) -> Result<(), E>,
|
||||||
|
) -> Result<(), E> {
|
||||||
|
use std::io::Write;
|
||||||
|
|
||||||
// Same directory as the target, so the rename stays on one filesystem.
|
// Same directory as the target, so the rename stays on one filesystem.
|
||||||
let mut tmp_name = path
|
let mut tmp_name = path
|
||||||
@@ -272,11 +288,18 @@ fn write_file_atomically(path: &std::path::Path, bytes: &[u8]) -> std::io::Resul
|
|||||||
tmp_name.push(format!(".tmp-{}", std::process::id()));
|
tmp_name.push(format!(".tmp-{}", std::process::id()));
|
||||||
let tmp_path = path.with_file_name(tmp_name);
|
let tmp_path = path.with_file_name(tmp_name);
|
||||||
|
|
||||||
let result = (|| {
|
let result = (|| -> Result<(), E> {
|
||||||
let mut f = std::fs::File::create(&tmp_path)?;
|
// 64 KiB, below glibc's 128 KiB mmap threshold: a 1 MiB buffer was
|
||||||
f.write_all(bytes)?;
|
// mapped afresh by each write once the threshold had not risen, and
|
||||||
f.sync_all()?;
|
// its page faults made a 1 MiB chunked write 1.6x slower (criterion
|
||||||
std::fs::rename(&tmp_path, path)
|
// `write_2d_chunked/512x512` after the smaller cases, 2026-09-29).
|
||||||
|
// Writes larger than the buffer go straight to the file.
|
||||||
|
let mut f = std::io::BufWriter::with_capacity(64 << 10, std::fs::File::create(&tmp_path)?);
|
||||||
|
write(&mut f)?;
|
||||||
|
f.flush()?;
|
||||||
|
f.get_ref().sync_all()?;
|
||||||
|
std::fs::rename(&tmp_path, path)?;
|
||||||
|
Ok(())
|
||||||
})();
|
})();
|
||||||
if result.is_err() {
|
if result.is_err() {
|
||||||
let _ = std::fs::remove_file(&tmp_path);
|
let _ = std::fs::remove_file(&tmp_path);
|
||||||
|
|||||||
+57
-2
@@ -2,6 +2,12 @@
|
|||||||
|
|
||||||
python gen_huge_chunks.py filtered OUT.h5 # huge_chunks_filtered.h5
|
python gen_huge_chunks.py filtered OUT.h5 # huge_chunks_filtered.h5
|
||||||
python gen_huge_chunks.py unfiltered OUT.h5 # generated at test time
|
python gen_huge_chunks.py unfiltered OUT.h5 # generated at test time
|
||||||
|
python gen_huge_chunks.py dims OUT.h5 # huge_chunk_dims.h5
|
||||||
|
python gen_huge_chunks.py dims-unfiltered OUT.h5 # generated at test time
|
||||||
|
|
||||||
|
`dims` and `dims-unfiltered` (see `make_dims`) hold `u1` datasets whose
|
||||||
|
chunk dimension is N8 = 2**32 + 7: a chunk dimension of 2**32 or more,
|
||||||
|
which a layout message of version 5 stores in 5 bytes.
|
||||||
|
|
||||||
libhdf5 2.x writes a chunk of more than 0xFFFFFFFF bytes with layout message
|
libhdf5 2.x writes a chunk of more than 0xFFFFFFFF bytes with layout message
|
||||||
version 5 (`H5D__chunk_construct`: "chunk size > 4GB requires
|
version 5 (`H5D__chunk_construct`: "chunk size > 4GB requires
|
||||||
@@ -38,7 +44,8 @@ allocated early (see `make`).
|
|||||||
libhdf5 holds a whole filtered chunk in memory while it writes it, so the
|
libhdf5 holds a whole filtered chunk in memory while it writes it, so the
|
||||||
`filtered` run needs about 4 GiB of memory and 100 s (tank).
|
`filtered` run needs about 4 GiB of memory and 100 s (tank).
|
||||||
|
|
||||||
Generated 2026-09-28 with h5py 3.16.0 (libhdf5 2.0.0).
|
huge_chunks_filtered.h5 was generated 2026-09-28 and huge_chunk_dims.h5
|
||||||
|
2026-09-29, both with h5py 3.16.0 (libhdf5 2.0.0).
|
||||||
"""
|
"""
|
||||||
import sys
|
import sys
|
||||||
|
|
||||||
@@ -92,12 +99,60 @@ def make(fid, name, index, filtered):
|
|||||||
dsid.close()
|
dsid.close()
|
||||||
|
|
||||||
|
|
||||||
|
N8 = 2**32 + 7
|
||||||
|
|
||||||
|
|
||||||
|
def make_dims(fid, name, index, filtered):
|
||||||
|
"""A `u1` dataset whose chunks are N8 elements (a chunk dimension of
|
||||||
|
2**32 or more). `d[0:10] = 0..9` and `d[N8-10:N8] = 100..109`; where
|
||||||
|
there is a second chunk (`farray`, `earray`: shape N8 + 10),
|
||||||
|
`d[N8:N8+10] = 200..209`.
|
||||||
|
|
||||||
|
filtered (`dims`, committed): deflate level 9 twice, fill value 7;
|
||||||
|
`single` (Single Chunk) and `earray` (Extensible Array, maxshape None).
|
||||||
|
Writing it holds a 4 GiB chunk: 4.1 GiB peak resident memory and about
|
||||||
|
a minute (tank, 2026-09-29).
|
||||||
|
unfiltered (`dims-unfiltered`): fill time "never", sparse; `single`
|
||||||
|
(allocated early, as in `make`) and `farray` (Fixed Array).
|
||||||
|
"""
|
||||||
|
dcpl = h5py.h5p.create(h5py.h5p.DATASET_CREATE)
|
||||||
|
if index == "single":
|
||||||
|
shape, maxshape = (N8,), (N8,)
|
||||||
|
elif index == "farray":
|
||||||
|
shape, maxshape = (N8 + 10,), (N8 + 10,)
|
||||||
|
elif index == "earray":
|
||||||
|
shape, maxshape = (N8 + 10,), (UNLIM,)
|
||||||
|
dcpl.set_chunk((N8,))
|
||||||
|
if filtered:
|
||||||
|
dcpl.set_deflate(9)
|
||||||
|
dcpl.set_deflate(9)
|
||||||
|
dcpl.set_fill_value(np.array(7, dtype="u1"))
|
||||||
|
else:
|
||||||
|
dcpl.set_fill_time(h5py.h5d.FILL_TIME_NEVER)
|
||||||
|
if index == "single":
|
||||||
|
dcpl.set_alloc_time(h5py.h5d.ALLOC_TIME_EARLY)
|
||||||
|
space = h5py.h5s.create_simple(shape, maxshape)
|
||||||
|
dsid = h5py.h5d.create(fid, name.encode(), h5py.h5t.STD_U8LE, space, dcpl=dcpl)
|
||||||
|
ds = h5py.Dataset(dsid)
|
||||||
|
ds[0:10] = np.arange(10, dtype="u1")
|
||||||
|
ds[N8 - 10 : N8] = np.arange(100, 110, dtype="u1")
|
||||||
|
if shape[0] > N8:
|
||||||
|
ds[N8 : N8 + 10] = np.arange(200, 210, dtype="u1")
|
||||||
|
dsid.close()
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
mode, out = sys.argv[1], sys.argv[2]
|
mode, out = sys.argv[1], sys.argv[2]
|
||||||
filtered = {"filtered": True, "unfiltered": False}[mode]
|
|
||||||
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
|
fapl = h5py.h5p.create(h5py.h5p.FILE_ACCESS)
|
||||||
fapl.set_libver_bounds(h5py.h5f.LIBVER_EARLIEST, h5py.h5f.LIBVER_V200)
|
fapl.set_libver_bounds(h5py.h5f.LIBVER_EARLIEST, h5py.h5f.LIBVER_V200)
|
||||||
fid = h5py.h5f.create(out.encode(), h5py.h5f.ACC_TRUNC, fapl=fapl)
|
fid = h5py.h5f.create(out.encode(), h5py.h5f.ACC_TRUNC, fapl=fapl)
|
||||||
|
if mode in ("dims", "dims-unfiltered"):
|
||||||
|
filtered = mode == "dims"
|
||||||
|
for index in ["single", "earray"] if filtered else ["single", "farray"]:
|
||||||
|
make_dims(fid, index, index, filtered)
|
||||||
|
fid.close()
|
||||||
|
return
|
||||||
|
filtered = {"filtered": True, "unfiltered": False}[mode]
|
||||||
indexes = ["single", "farray", "earray", "btree2"]
|
indexes = ["single", "farray", "earray", "btree2"]
|
||||||
if not filtered:
|
if not filtered:
|
||||||
indexes.insert(1, "implicit")
|
indexes.insert(1, "implicit")
|
||||||
|
|||||||
@@ -0,0 +1,75 @@
|
|||||||
|
"""Generate native_complex_hdf5_2.h5: HDF5 2.0 native complex (class 11).
|
||||||
|
|
||||||
|
Datasets (every one class 11, or a type holding it):
|
||||||
|
|
||||||
|
c64 H5T_COMPLEX_IEEE_F32LE, 4 values incl. -0, inf, nan, subnormal
|
||||||
|
c128 H5T_COMPLEX_IEEE_F64LE, 2 x 3
|
||||||
|
c128be H5T_COMPLEX_IEEE_F64BE, the same values
|
||||||
|
c128x H5T_COMPLEX_IEEE_F64LE, c128 with element (1,1) changed
|
||||||
|
c32 H5T_COMPLEX_IEEE_F16LE, 3 values
|
||||||
|
scalar H5T_COMPLEX_IEEE_F64LE, scalar dataspace
|
||||||
|
compound {z: H5T_COMPLEX_IEEE_F64LE @0, k: H5T_STD_I64LE @16}
|
||||||
|
array H5T_ARRAY [2] of H5T_COMPLEX_IEEE_F32LE, 2 elements
|
||||||
|
|
||||||
|
plus the root attribute `zattr` (H5T_COMPLEX_IEEE_F64LE, 2 values).
|
||||||
|
|
||||||
|
Written with h5py 3.16 (libhdf5 2.0.0) through its low-level API, the file
|
||||||
|
type as the memory type (no conversion). Re-run only if the fixture ever
|
||||||
|
needs regenerating (generated 2026-09-29):
|
||||||
|
|
||||||
|
python gen_native_complex.py native_complex_hdf5_2.h5
|
||||||
|
|
||||||
|
The h5dump / h5ls 2.2.0 reference output the tests compare with was made
|
||||||
|
with the tools of the libhdf5 2.2.0 build described in gen_mx_floats.py:
|
||||||
|
|
||||||
|
h5dump native_complex_hdf5_2.h5 > native_complex_hdf5_2.ddl
|
||||||
|
h5ls -r -v native_complex_hdf5_2.h5 > native_complex_hdf5_2.ls
|
||||||
|
"""
|
||||||
|
import sys
|
||||||
|
|
||||||
|
import h5py
|
||||||
|
import numpy as np
|
||||||
|
from h5py import h5a, h5d, h5s, h5t
|
||||||
|
|
||||||
|
out = sys.argv[1]
|
||||||
|
f = h5py.File(out, "w", libver="latest")
|
||||||
|
|
||||||
|
|
||||||
|
def ds(name, tid, arr, shape=None):
|
||||||
|
shape = arr.shape if shape is None else shape
|
||||||
|
sp = h5s.create_simple(shape) if shape != () else h5s.create(h5s.SCALAR)
|
||||||
|
d = h5d.create(f.id, name.encode(), tid, sp)
|
||||||
|
d.write(h5s.ALL, h5s.ALL, np.ascontiguousarray(arr), mtype=tid)
|
||||||
|
|
||||||
|
|
||||||
|
nan, inf = float("nan"), float("inf")
|
||||||
|
c64 = np.array(
|
||||||
|
[1.5 - 2j, complex(0.0, -0.0), complex(inf, nan), complex(-inf, 1e-40)],
|
||||||
|
dtype=np.complex64,
|
||||||
|
)
|
||||||
|
ds("c64", h5t.COMPLEX_IEEE_F32LE, c64)
|
||||||
|
c128 = np.array(
|
||||||
|
[[1 + 2j, -3.5 + 4e300j, 0.1 + 0.2j], [5e-324 - 0j, 1j, -1]], dtype=np.complex128
|
||||||
|
)
|
||||||
|
ds("c128", h5t.COMPLEX_IEEE_F64LE, c128)
|
||||||
|
ds("c128be", h5t.COMPLEX_IEEE_F64BE, c128.astype(">c16"))
|
||||||
|
c128x = c128.copy()
|
||||||
|
c128x[1, 1] = 1 + 1j
|
||||||
|
ds("c128x", h5t.COMPLEX_IEEE_F64LE, c128x)
|
||||||
|
half = np.array([1.0, -2.0, 0.5, 65504.0, 0.0, -0.0], dtype=np.float16)
|
||||||
|
ds("c32", h5t.COMPLEX_IEEE_F16LE, half, shape=(3,))
|
||||||
|
ds("scalar", h5t.COMPLEX_IEEE_F64LE, np.array(2.5 - 0.5j), shape=())
|
||||||
|
|
||||||
|
ct = h5t.create(h5t.COMPOUND, 24)
|
||||||
|
ct.insert(b"z", 0, h5t.COMPLEX_IEEE_F64LE)
|
||||||
|
ct.insert(b"k", 16, h5t.STD_I64LE)
|
||||||
|
cd = np.array([(1 + 1j, 7), (-2.5 + 0j, -8)], dtype=[("z", "<c16"), ("k", "<i8")])
|
||||||
|
ds("compound", ct, cd)
|
||||||
|
|
||||||
|
at = h5t.array_create(h5t.COMPLEX_IEEE_F32LE, (2,))
|
||||||
|
ad = np.array([[1 + 2j, 3 + 4j], [-1 - 1j, 0j]], dtype=np.complex64)
|
||||||
|
ds("array", at, ad, shape=(2,))
|
||||||
|
|
||||||
|
a = h5a.create(f.id, b"zattr", h5t.COMPLEX_IEEE_F64LE, h5s.create_simple((2,)))
|
||||||
|
a.write(np.array([0.5 - 1.5j, 2j]), mtype=h5t.COMPLEX_IEEE_F64LE)
|
||||||
|
f.close()
|
||||||
Binary file not shown.
@@ -0,0 +1,80 @@
|
|||||||
|
HDF5 "native_complex_hdf5_2.h5" {
|
||||||
|
GROUP "/" {
|
||||||
|
ATTRIBUTE "zattr" {
|
||||||
|
DATATYPE H5T_COMPLEX_IEEE_F64LE
|
||||||
|
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
|
||||||
|
DATA {
|
||||||
|
(0): 0.5-1.5i, 0+2i
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "array" {
|
||||||
|
DATATYPE H5T_ARRAY { [2] H5T_COMPLEX_IEEE_F32LE }
|
||||||
|
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
|
||||||
|
DATA {
|
||||||
|
(0): [ 1+2i, 3+4i ], [ -1-1i, 0+0i ]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "c128" {
|
||||||
|
DATATYPE H5T_COMPLEX_IEEE_F64LE
|
||||||
|
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
|
||||||
|
DATA {
|
||||||
|
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
|
||||||
|
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "c128be" {
|
||||||
|
DATATYPE H5T_COMPLEX_IEEE_F64BE
|
||||||
|
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
|
||||||
|
DATA {
|
||||||
|
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
|
||||||
|
(1,0): 4.94066e-324-0i, 0+1i, -1+0i
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "c128x" {
|
||||||
|
DATATYPE H5T_COMPLEX_IEEE_F64LE
|
||||||
|
DATASPACE SIMPLE { ( 2, 3 ) / ( 2, 3 ) }
|
||||||
|
DATA {
|
||||||
|
(0,0): 1+2i, -3.5+4e+300i, 0.1+0.2i,
|
||||||
|
(1,0): 4.94066e-324-0i, 1+1i, -1+0i
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "c32" {
|
||||||
|
DATATYPE H5T_COMPLEX_IEEE_F16LE
|
||||||
|
DATASPACE SIMPLE { ( 3 ) / ( 3 ) }
|
||||||
|
DATA {
|
||||||
|
(0): 1+-2i, 0.5+65504i, 0+-0i
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "c64" {
|
||||||
|
DATATYPE H5T_COMPLEX_IEEE_F32LE
|
||||||
|
DATASPACE SIMPLE { ( 4 ) / ( 4 ) }
|
||||||
|
DATA {
|
||||||
|
(0): 1.5-2i, 0-0i, inf+nani, -inf+9.99995e-41i
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "compound" {
|
||||||
|
DATATYPE H5T_COMPOUND {
|
||||||
|
H5T_COMPLEX_IEEE_F64LE "z";
|
||||||
|
H5T_STD_I64LE "k";
|
||||||
|
}
|
||||||
|
DATASPACE SIMPLE { ( 2 ) / ( 2 ) }
|
||||||
|
DATA {
|
||||||
|
(0): {
|
||||||
|
1+1i,
|
||||||
|
7
|
||||||
|
},
|
||||||
|
(1): {
|
||||||
|
-2.5+0i,
|
||||||
|
-8
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
DATASET "scalar" {
|
||||||
|
DATATYPE H5T_COMPLEX_IEEE_F64LE
|
||||||
|
DATASPACE SCALAR
|
||||||
|
DATA {
|
||||||
|
(0): 2.5-0.5i
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
Binary file not shown.
@@ -81,8 +81,15 @@ fn layout_of(
|
|||||||
let lm = msg(MessageType::DataLayout);
|
let lm = msg(MessageType::DataLayout);
|
||||||
let layout = DataLayout::parse(&lm.data, sb.offset_size, sb.length_size).unwrap();
|
let layout = DataLayout::parse(&lm.data, sb.offset_size, sb.length_size).unwrap();
|
||||||
let space = Dataspace::parse(&msg(MessageType::Dataspace).data, sb.length_size).unwrap();
|
let space = Dataspace::parse(&msg(MessageType::Dataspace).data, sb.length_size).unwrap();
|
||||||
|
// The element size is the layout's last dimension.
|
||||||
|
let elem = match &layout {
|
||||||
|
DataLayout::Chunked {
|
||||||
|
chunk_dimensions, ..
|
||||||
|
} => *chunk_dimensions.last().unwrap() as usize,
|
||||||
|
_ => 8,
|
||||||
|
};
|
||||||
let (chunks, _) =
|
let (chunks, _) =
|
||||||
list_chunks(data, &layout, &space, 8, sb.offset_size, sb.length_size).unwrap();
|
list_chunks(data, &layout, &space, elem, sb.offset_size, sb.length_size).unwrap();
|
||||||
(lm.data[0], layout, space, chunks)
|
(lm.data[0], layout, space, chunks)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -516,3 +523,343 @@ print("ok")
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ---- Chunk dimensions of 2^32 or more ----
|
||||||
|
|
||||||
|
/// Elements per chunk of the `u8` datasets below: a chunk dimension of
|
||||||
|
/// 2^32 or more (4 GiB + 7 bytes), which a version-5 layout message stores
|
||||||
|
/// in 5 bytes.
|
||||||
|
const N8: u64 = (1 << 32) + 7;
|
||||||
|
|
||||||
|
fn dims_fixture() -> PathBuf {
|
||||||
|
Path::new(env!("CARGO_MANIFEST_DIR")).join("tests/fixtures/huge_chunk_dims.h5")
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `u8` values of `sel` in dataset `name` of `file`.
|
||||||
|
fn sel_u8(file: &File, name: &str, start: u64, count: u64) -> Vec<u8> {
|
||||||
|
let ds = file.dataset(name).unwrap();
|
||||||
|
let s = Selection::Hyperslab {
|
||||||
|
start: vec![start],
|
||||||
|
stride: vec![1],
|
||||||
|
count: vec![count],
|
||||||
|
block: vec![1],
|
||||||
|
};
|
||||||
|
ds.read_selection(&s).unwrap()
|
||||||
|
}
|
||||||
|
|
||||||
|
fn u8s(range: std::ops::Range<u8>) -> Vec<u8> {
|
||||||
|
range.collect()
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `fixtures/huge_chunk_dims.h5` (libhdf5 2.0.0 through h5py 3.16,
|
||||||
|
/// `gen_huge_chunks.py dims`): `u8` datasets chunked by N8 = 2^32 + 7, in a
|
||||||
|
/// Single Chunk and an Extensible Array index. Their layouts parse to the
|
||||||
|
/// whole dimension (it was refused: chunk dimensions were `u32`) and their
|
||||||
|
/// chunks list at the right origins.
|
||||||
|
#[test]
|
||||||
|
fn huge_chunk_dims_list() {
|
||||||
|
let data = std::fs::read(dims_fixture()).unwrap();
|
||||||
|
for (name, index, shape, origins) in [
|
||||||
|
("single", 1u8, N8, &[0u64][..]),
|
||||||
|
("earray", 4, N8 + 10, &[0, N8]),
|
||||||
|
] {
|
||||||
|
let (version, layout, space, mut chunks) = layout_of(&data, name);
|
||||||
|
assert_eq!(version, 5, "{name}");
|
||||||
|
let DataLayout::Chunked {
|
||||||
|
chunk_dimensions,
|
||||||
|
chunk_index_type,
|
||||||
|
..
|
||||||
|
} = &layout
|
||||||
|
else {
|
||||||
|
panic!("{name}: {layout:?}");
|
||||||
|
};
|
||||||
|
assert_eq!(chunk_dimensions, &[N8, 1], "{name}");
|
||||||
|
assert_eq!(*chunk_index_type, Some(index), "{name}");
|
||||||
|
assert_eq!(space.dimensions, [shape], "{name}");
|
||||||
|
chunks.sort_by_key(|c| c.offsets[0]);
|
||||||
|
let got: Vec<u64> = chunks.iter().map(|c| c.offsets[0]).collect();
|
||||||
|
assert_eq!(got, origins, "{name}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Decoding the fixture's chunks (4 GiB each, one at a time): written
|
||||||
|
/// values, the fill value (7) around them, and the second chunk.
|
||||||
|
#[test]
|
||||||
|
fn huge_chunk_dims_read() {
|
||||||
|
if !heavy() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let file = File::open(dims_fixture()).unwrap();
|
||||||
|
let mut first = u8s(0..10);
|
||||||
|
first.extend([7, 7]);
|
||||||
|
let mut last = vec![7, 7];
|
||||||
|
last.extend(u8s(100..110));
|
||||||
|
for name in ["single", "earray"] {
|
||||||
|
assert_eq!(sel_u8(&file, name, 0, 12), first, "{name}");
|
||||||
|
assert_eq!(sel_u8(&file, name, N8 - 12, 12), last, "{name}");
|
||||||
|
}
|
||||||
|
let mut across = u8s(108..110);
|
||||||
|
across.extend(u8s(200..210));
|
||||||
|
assert_eq!(sel_u8(&file, "earray", N8 - 2, 12), across);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Unfiltered chunks of N8 bytes that libhdf5 2.0.0 wrote (a sparse file,
|
||||||
|
/// `gen_huge_chunks.py dims-unfiltered`), in a Single Chunk and a Fixed
|
||||||
|
/// Array index: read through a memory map, and through positioned reads
|
||||||
|
/// that fetch only the rows selected.
|
||||||
|
#[test]
|
||||||
|
fn huge_chunk_dims_unfiltered_read() {
|
||||||
|
if !heavy() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = scratch();
|
||||||
|
let Some(path) = generate(dir.path(), "dims-unfiltered") else {
|
||||||
|
return;
|
||||||
|
};
|
||||||
|
let check = |file: &File| {
|
||||||
|
let mut first = u8s(0..10);
|
||||||
|
first.extend([0, 0]);
|
||||||
|
let mut last = vec![0, 0];
|
||||||
|
last.extend(u8s(100..110));
|
||||||
|
for name in ["single", "farray"] {
|
||||||
|
assert_eq!(sel_u8(file, name, 0, 12), first, "{name}");
|
||||||
|
assert_eq!(sel_u8(file, name, N8 - 12, 12), last, "{name}");
|
||||||
|
}
|
||||||
|
let mut across = u8s(108..110);
|
||||||
|
across.extend(u8s(200..210));
|
||||||
|
assert_eq!(sel_u8(file, "farray", N8 - 2, 12), across);
|
||||||
|
};
|
||||||
|
check(&File::open(&path).unwrap());
|
||||||
|
let file = std::fs::File::open(&path).unwrap();
|
||||||
|
let len = file.metadata().unwrap().len();
|
||||||
|
let storage = std::sync::Arc::new(Counting {
|
||||||
|
file,
|
||||||
|
len,
|
||||||
|
read: 0.into(),
|
||||||
|
});
|
||||||
|
check(&File::open_storage(storage.clone()).unwrap());
|
||||||
|
let read = storage.read.load(std::sync::atomic::Ordering::Relaxed);
|
||||||
|
assert!(read < 1 << 20, "{read} bytes read");
|
||||||
|
}
|
||||||
|
|
||||||
|
/// Run `script` under h5py with `path` as `sys.argv[1]`; skipped without
|
||||||
|
/// h5py (a failure under `CLAWHDF5_REQUIRE_INTEROP=1`).
|
||||||
|
fn h5py_check(script: &str, path: &Path) {
|
||||||
|
if !python_available() {
|
||||||
|
assert!(!interop_required(), "h5py not available");
|
||||||
|
eprintln!("SKIP: python3 with h5py not available");
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let out = Command::new(python())
|
||||||
|
.arg("-c")
|
||||||
|
.arg(script)
|
||||||
|
.arg(path)
|
||||||
|
.output()
|
||||||
|
.unwrap();
|
||||||
|
assert!(
|
||||||
|
out.status.success(),
|
||||||
|
"h5py:\n{}",
|
||||||
|
String::from_utf8_lossy(&out.stderr)
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
/// `h5dump` 2.x's output for `args` on `path`, checked to succeed without
|
||||||
|
/// errors; `None` without `CLAWHDF5_H5DUMP2`.
|
||||||
|
fn h5dump2_text(args: &[&str], path: &Path) -> Option<String> {
|
||||||
|
let h5dump = h5dump2()?;
|
||||||
|
let out = Command::new(&h5dump).args(args).arg(path).output().unwrap();
|
||||||
|
let text = format!(
|
||||||
|
"{}{}",
|
||||||
|
String::from_utf8_lossy(&out.stdout),
|
||||||
|
String::from_utf8_lossy(&out.stderr)
|
||||||
|
);
|
||||||
|
assert!(out.status.success(), "h5dump {args:?}:\n{text}");
|
||||||
|
assert!(!text.contains("rror"), "h5dump {args:?}:\n{text}");
|
||||||
|
Some(text)
|
||||||
|
}
|
||||||
|
|
||||||
|
/// clawhdf5 writes deflated `u8` chunks of N8 elements (a chunk dimension
|
||||||
|
/// of 2^32 or more) in a Single Chunk and an Extensible Array index; it,
|
||||||
|
/// h5py (libhdf5 2.0.0) and h5dump 2.2.0 (layouts only: it cannot inflate
|
||||||
|
/// chunks over 4 GiB, see known-issues) read them.
|
||||||
|
#[test]
|
||||||
|
fn writer_huge_chunk_dims_round_trip() {
|
||||||
|
if !heavy() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = scratch();
|
||||||
|
let path = dir.path().join("dims.h5");
|
||||||
|
let values = u8s(0..10);
|
||||||
|
let mut b = clawhdf5::FileBuilder::new();
|
||||||
|
for (name, max) in [("single", N8), ("earray", u64::MAX)] {
|
||||||
|
b.create_dataset(name)
|
||||||
|
.with_u8_data(&values)
|
||||||
|
.with_maxshape(&[max])
|
||||||
|
.with_chunks(&[N8])
|
||||||
|
.with_deflate(1);
|
||||||
|
}
|
||||||
|
b.write(&path).unwrap();
|
||||||
|
|
||||||
|
let data = std::fs::read(&path).unwrap();
|
||||||
|
// Deflate level 1 leaves about 40 MiB of each 4 GiB chunk of zeros.
|
||||||
|
assert!(data.len() < 256 << 20, "{} bytes", data.len());
|
||||||
|
for (name, index) in [("single", 1), ("earray", 4)] {
|
||||||
|
let (version, layout, _, chunks) = layout_of(&data, name);
|
||||||
|
assert_eq!(version, 5, "{name}");
|
||||||
|
assert!(
|
||||||
|
matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(t), .. }
|
||||||
|
if *t == index && chunk_dimensions == &[N8, 1]),
|
||||||
|
"{name}: {layout:?}"
|
||||||
|
);
|
||||||
|
assert_eq!(chunks.len(), 1, "{name}");
|
||||||
|
}
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
for name in ["single", "earray"] {
|
||||||
|
assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}");
|
||||||
|
}
|
||||||
|
drop(file);
|
||||||
|
|
||||||
|
h5py_check(
|
||||||
|
r#"
|
||||||
|
import sys, h5py
|
||||||
|
with h5py.File(sys.argv[1], "r") as f:
|
||||||
|
for name in ("single", "earray"):
|
||||||
|
d = f[name]
|
||||||
|
assert d.chunks == (2**32 + 7,), (name, d.chunks)
|
||||||
|
assert list(d[:]) == list(range(10)), (name, d[:])
|
||||||
|
"#,
|
||||||
|
&path,
|
||||||
|
);
|
||||||
|
for name in ["single", "earray"] {
|
||||||
|
if let Some(text) = h5dump2_text(&["-H", "-p", "-d", name], &path) {
|
||||||
|
assert!(text.contains("CHUNKED ( 4294967303 )"), "{name}:\n{text}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// An unfiltered chunk of N8 bytes (a chunk dimension of 2^32 or more)
|
||||||
|
/// written by clawhdf5 from data it was handed (`with_u8_data_owned`):
|
||||||
|
/// `FileBuilder::write` streams it to the file without copying it, so the
|
||||||
|
/// write needs about the data's own 4 GiB. Read back by clawhdf5, h5py
|
||||||
|
/// (a few elements) and h5dump 2.2.0.
|
||||||
|
#[test]
|
||||||
|
fn writer_unfiltered_huge_chunk_round_trip() {
|
||||||
|
if !heavy() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = scratch();
|
||||||
|
let path = dir.path().join("raw.h5");
|
||||||
|
let value = |i: u64| (i % 251) as u8;
|
||||||
|
let mut b = clawhdf5::FileBuilder::new();
|
||||||
|
b.create_dataset("single")
|
||||||
|
.with_u8_data_owned((0..N8).map(value).collect())
|
||||||
|
.with_chunks(&[N8]);
|
||||||
|
b.write(&path).unwrap();
|
||||||
|
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
let data = file.as_bytes();
|
||||||
|
let (version, layout, space, chunks) = layout_of(data, "single");
|
||||||
|
assert_eq!(version, 5);
|
||||||
|
assert!(
|
||||||
|
matches!(&layout, DataLayout::Chunked { chunk_dimensions, chunk_index_type: Some(1), .. }
|
||||||
|
if chunk_dimensions == &[N8, 1]),
|
||||||
|
"{layout:?}"
|
||||||
|
);
|
||||||
|
assert_eq!(space.dimensions, [N8]);
|
||||||
|
assert_eq!(chunks.len(), 1);
|
||||||
|
assert_eq!(chunks[0].chunk_size, N8);
|
||||||
|
for start in [0, (1 << 32) - 3, N8 - 6] {
|
||||||
|
let want: Vec<u8> = (start..start + 6).map(value).collect();
|
||||||
|
assert_eq!(sel_u8(&file, "single", start, 6), want, "at {start}");
|
||||||
|
}
|
||||||
|
drop(file);
|
||||||
|
|
||||||
|
h5py_check(
|
||||||
|
r#"
|
||||||
|
import sys, h5py, numpy as np
|
||||||
|
n = 2**32 + 7
|
||||||
|
with h5py.File(sys.argv[1], "r") as f:
|
||||||
|
d = f["single"]
|
||||||
|
assert d.chunks == (n,) and d.shape == (n,) and d.compression is None
|
||||||
|
for start in (0, 2**32 - 3, n - 6):
|
||||||
|
want = [i % 251 for i in range(start, start + 6)]
|
||||||
|
assert list(d[start:start + 6]) == want, (start, d[start:start + 6])
|
||||||
|
"#,
|
||||||
|
&path,
|
||||||
|
);
|
||||||
|
let start = (1u64 << 32) - 3;
|
||||||
|
if let Some(text) = h5dump2_text(
|
||||||
|
&["-d", "single", "-s", &start.to_string(), "-c", "6"],
|
||||||
|
&path,
|
||||||
|
) {
|
||||||
|
let want: Vec<String> = (start..start + 6).map(|i| value(i).to_string()).collect();
|
||||||
|
assert!(text.contains(&want.join(", ")), "{text}");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/// LZ4 and Zstandard chunks of 4 GiB or more: written by clawhdf5, read
|
||||||
|
/// back by it and by h5py (through `hdf5plugin`'s LZ4 and Zstandard
|
||||||
|
/// filters when installed).
|
||||||
|
#[cfg(any(feature = "lz4", feature = "zstd"))]
|
||||||
|
#[test]
|
||||||
|
fn writer_huge_chunks_lz4_zstd_round_trip() {
|
||||||
|
if !heavy() {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
let dir = scratch();
|
||||||
|
let path = dir.path().join("codecs.h5");
|
||||||
|
let values = u8s(0..10);
|
||||||
|
let mut b = clawhdf5::FileBuilder::new();
|
||||||
|
let mut names = Vec::new();
|
||||||
|
#[cfg(feature = "lz4")]
|
||||||
|
{
|
||||||
|
b.create_dataset("lz4")
|
||||||
|
.with_u8_data(&values)
|
||||||
|
.with_maxshape(&[N8])
|
||||||
|
.with_chunks(&[N8])
|
||||||
|
.with_lz4();
|
||||||
|
names.push("lz4");
|
||||||
|
}
|
||||||
|
#[cfg(feature = "zstd")]
|
||||||
|
{
|
||||||
|
b.create_dataset("zstd")
|
||||||
|
.with_u8_data(&values)
|
||||||
|
.with_maxshape(&[N8])
|
||||||
|
.with_chunks(&[N8])
|
||||||
|
.with_zstd(3);
|
||||||
|
names.push("zstd");
|
||||||
|
}
|
||||||
|
b.write(&path).unwrap();
|
||||||
|
let data = std::fs::read(&path).unwrap();
|
||||||
|
assert!(data.len() < 256 << 20, "{} bytes", data.len());
|
||||||
|
for name in &names {
|
||||||
|
let (version, _, _, chunks) = layout_of(&data, name);
|
||||||
|
assert_eq!(version, 5, "{name}");
|
||||||
|
assert_eq!(chunks.len(), 1, "{name}");
|
||||||
|
}
|
||||||
|
drop(data);
|
||||||
|
let file = File::open(&path).unwrap();
|
||||||
|
for name in &names {
|
||||||
|
assert_eq!(sel_u8(&file, name, 0, 10), values, "{name}");
|
||||||
|
}
|
||||||
|
drop(file);
|
||||||
|
h5py_check(
|
||||||
|
&format!(
|
||||||
|
r#"
|
||||||
|
import sys, h5py
|
||||||
|
try:
|
||||||
|
import hdf5plugin
|
||||||
|
except ImportError:
|
||||||
|
print("SKIP: no hdf5plugin")
|
||||||
|
sys.exit(0)
|
||||||
|
with h5py.File(sys.argv[1], "r") as f:
|
||||||
|
for name in {names:?}:
|
||||||
|
d = f[name]
|
||||||
|
assert d.chunks == (2**32 + 7,), (name, d.chunks)
|
||||||
|
assert list(d[0:10]) == list(range(10)), (name, d[0:10])
|
||||||
|
"#,
|
||||||
|
names = names
|
||||||
|
),
|
||||||
|
&path,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|||||||
@@ -1064,12 +1064,26 @@ fn complex_datasets_and_attributes_round_trip() {
|
|||||||
let ds = file.dataset(name).unwrap();
|
let ds = file.dataset(name).unwrap();
|
||||||
assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}");
|
assert_eq!(ds.read_complex_f32().unwrap(), z64, "{name}");
|
||||||
assert_eq!(ds.shape().unwrap(), vec![3]);
|
assert_eq!(ds.shape().unwrap(), vec![3]);
|
||||||
// Both encodings read as the same {r, i} compound.
|
// h5py's encoding is a compound, the native one a complex type.
|
||||||
assert_eq!(
|
let want = if name == "native64" {
|
||||||
ds.dtype().unwrap(),
|
DType::Complex(Box::new(DType::F32))
|
||||||
|
} else {
|
||||||
DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)])
|
DType::Compound(vec![("r".into(), DType::F32), ("i".into(), DType::F32)])
|
||||||
);
|
};
|
||||||
|
assert_eq!(ds.dtype().unwrap(), want, "{name}");
|
||||||
}
|
}
|
||||||
|
assert_eq!(
|
||||||
|
file.dataset("native128").unwrap().dtype().unwrap(),
|
||||||
|
DType::Complex(Box::new(DType::F64))
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
file.dataset("native128").unwrap().raw_datatype().unwrap(),
|
||||||
|
clawhdf5::make_native_complex_f64_type()
|
||||||
|
);
|
||||||
|
assert_eq!(
|
||||||
|
DType::Complex(Box::new(DType::F64)).to_string(),
|
||||||
|
"complex<f64>"
|
||||||
|
);
|
||||||
for name in ["compound128", "native128"] {
|
for name in ["compound128", "native128"] {
|
||||||
same(
|
same(
|
||||||
file.dataset(name).unwrap().read_complex_f64().unwrap(),
|
file.dataset(name).unwrap().read_complex_f64().unwrap(),
|
||||||
@@ -1094,7 +1108,7 @@ fn complex_datasets_and_attributes_round_trip() {
|
|||||||
data,
|
data,
|
||||||
}) => {
|
}) => {
|
||||||
assert!(shape.is_empty());
|
assert!(shape.is_empty());
|
||||||
assert_eq!(datatype.type_size(), 16);
|
assert_eq!(datatype, clawhdf5::make_native_complex_f64_type());
|
||||||
let fields =
|
let fields =
|
||||||
clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap();
|
clawhdf5_format::data_read::read_compound_fields(&data, &datatype).unwrap();
|
||||||
let part = |k: usize| {
|
let part = |k: usize| {
|
||||||
|
|||||||
+119
-31
@@ -17,6 +17,11 @@ Checked against `main` at `9b5803f` on 2026-09-28.
|
|||||||
|---|---|---|
|
|---|---|---|
|
||||||
| [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 |
|
| [NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c) | deliberate (floats of other than 4 or 8 bytes, axes netCDF-C leaves without a dimension), and 9 corpus files (external links, values the HDF5 reader refuses) | 2026-09-29 |
|
||||||
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
|
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (chunk dimensions of 2^32 or more, five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
|
||||||
|
|
||||||
|
|
||||||
|
| [NetCDF-4: variables' dimensions are guessed from sizes](#netcdf-4-variables-dimensions-are-guessed-from-sizes) | **wrong metadata** (`Variable::dimensions`, pure dimension scales listed as variables; values and dimension sizes are right) | 2026-09-28 |
|
||||||
|
| [Chunks of 4 GiB or more: limits](#chunks-of-4-gib-or-more-limits) | refused (five filters, editor rewrites), memory (a decoded chunk is held whole) | 2026-09-28 |
|
||||||
|
| [NetCDF-4: an unlimited dimension reports size 0](#netcdf-4-an-unlimited-dimension-reports-size-0) | **wrong metadata** (dimension size; variable shapes and values are right) | 2026-09-28 |
|
||||||
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
|
| [Small floats decode as libhdf5 does, not as the OCP MX specification](#small-floats-decode-as-libhdf5-does-not-as-the-ocp-mx-specification) | deliberate: libhdf5's values (FP4/FP6/FP8 E4M3 all-ones exponent is inf/NaN) | 2026-09-28 |
|
||||||
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
|
| [In-place modification (`FileEditor`) limits](#in-place-modification-fileeditor-limits) | refused edits (`Error::Unsupported`), space reuse per editor, no journal | 2026-09-26 |
|
||||||
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
|
| [Python in-place editing limits](#python-in-place-editing-clawhdf5filepath-r-limits) | refused writes (`NotImplementedError`), deliberate conversion differences | 2026-09-27 |
|
||||||
@@ -241,7 +246,10 @@ Values are correct in every case; this is cost only. Selections other than
|
|||||||
|
|
||||||
## Chunks of 4 GiB or more: limits
|
## Chunks of 4 GiB or more: limits
|
||||||
|
|
||||||
**Status:** open (documented 2026-09-28). HDF5 2.0 writes a chunk of more
|
**Status:** open (documented 2026-09-28; updated 2026-09-29, when
|
||||||
|
[chunk dimensions of 2^32 or more](#chunk-dimensions-of-232-or-more-were-refused)
|
||||||
|
and [unfiltered writes](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)
|
||||||
|
were fixed). HDF5 2.0 writes a chunk of more
|
||||||
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
|
than 0xFFFFFFFF bytes with layout message version 5 (`H5D__chunk_construct`
|
||||||
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
|
in libhdf5 2.2.0: "chunk size > 4GB requires H5F_LIBVER_V200"), which
|
||||||
also makes a filtered chunk index element store the chunk's size in "size
|
also makes a filtered chunk index element store the chunk's size in "size
|
||||||
@@ -253,34 +261,40 @@ The [fix](#chunks-of-4-gib-or-more-could-not-be-read) is tested by
|
|||||||
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
|
`crates/clawhdf5/tests/huge_chunks_interop.rs`; its decoding tests are
|
||||||
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
|
opt-in (`CLAWHDF5_HUGE_CHUNKS=1`, one at a time: `--test-threads=1`).
|
||||||
What remains:
|
What remains:
|
||||||
- **Chunk dimensions of 2^32 or more** (which libhdf5 2.x also allows under
|
|
||||||
`H5F_LIBVER_V200`) are refused, when read (`InvalidChunkDimensions`)
|
|
||||||
and when written: `DataLayout::Chunked::chunk_dimensions` is `Vec<u32>`.
|
|
||||||
A 4 GiB chunk of 1-byte elements needs one.
|
|
||||||
- **Filters the writer refuses for such a chunk** (`FilterError`, before
|
- **Filters the writer refuses for such a chunk** (`FilterError`, before
|
||||||
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
|
anything is written): LZF, bitshuffle, bzip2 and Blosc (their HDF5
|
||||||
filters record sizes or block lengths in 32 bits, or cannot take such a
|
filters record sizes or block lengths in 32 bits, or cannot take such a
|
||||||
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
|
buffer) and pcodec. Deflate, shuffle, Fletcher-32, LZ4 and Zstd are
|
||||||
written; only shuffle + deflate is tested end to end (LZ4's framing and
|
written and tested end to end (`writer_huge_chunks_round_trip`,
|
||||||
the 256 MiB limit it had are unit-tested; Zstd is not tested at this
|
`writer_huge_chunk_dims_round_trip`,
|
||||||
size).
|
`writer_huge_chunks_lz4_zstd_round_trip`: read back by clawhdf5 and
|
||||||
- **Unfiltered chunks this size are written but not tested end to end:**
|
h5py 3.16, LZ4 and Zstd through hdf5plugin 7.1.0).
|
||||||
`FileBuilder` assembles the whole file in memory, several copies of the
|
|
||||||
chunk (well over the 12 GiB the tests may use). Their layout messages
|
|
||||||
are unit-tested (`chunked_write::tests::huge_chunk_layout_messages_and_index_elements`).
|
|
||||||
- **`FileEditor`** refuses to rewrite such chunks (see
|
- **`FileEditor`** refuses to rewrite such chunks (see
|
||||||
[its limits](#in-place-modification-fileeditor-limits)).
|
[its limits](#in-place-modification-fileeditor-limits)).
|
||||||
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
|
- **Memory:** decoding a filtered chunk holds all of it (4 GiB and more),
|
||||||
twice when it is shuffled (the inflated and the unshuffled copy, as
|
twice when it is shuffled (the inflated and the unshuffled copy, as
|
||||||
libhdf5's filters do too); writing one holds the chunk and its shuffled
|
libhdf5's filters do too). Writing a filtered chunk holds its compressed
|
||||||
copy. A selection decodes only the chunks it touches; one of an
|
copy and, where it is not a contiguous run of the dataset's data (a
|
||||||
unfiltered chunk reads only the rows it selects, through a memory map or
|
chunk past the dataset's edge is zero-padded), the raw chunk; shuffled,
|
||||||
|
the shuffled copy too. An unfiltered chunk is not copied when it is such
|
||||||
|
a run (a dataset stored as one chunk of its own shape): writing one
|
||||||
|
unfiltered chunk of 2^32 + 7 bytes with `with_u8_data_owned` and
|
||||||
|
`FileBuilder::write` peaked at 4.00 GiB resident, the data itself
|
||||||
|
(`write_one_chunk 4294967303 OUT owned`, tank, 2026-09-29, commit
|
||||||
|
`0065c6b`; see [the fix](#writing-a-4-gib-unfiltered-chunk-held-several-copies-of-it)).
|
||||||
|
`FileWriter::finish` (a `Vec`) holds the data and the file.
|
||||||
|
A selection decodes only the chunks it touches; one of an unfiltered
|
||||||
|
chunk reads only the rows it selects, through a memory map or
|
||||||
positioned reads (`File::open_storage`): every selection of
|
positioned reads (`File::open_storage`): every selection of
|
||||||
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
|
`unfiltered_huge_chunks_read` together read under 1 MiB. Peak resident
|
||||||
memory of `cargo test --release -p clawhdf5 --test huge_chunks_interop
|
memory of `cargo test --release -p clawhdf5 --features lz4,zstd --test
|
||||||
-- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1` (all six tests, h5py
|
huge_chunks_interop -- --test-threads=1` with `CLAWHDF5_HUGE_CHUNKS=1`
|
||||||
included) was 8.06 GiB (tank, 2026-09-28, commit `9143737`); selection
|
(all twelve tests, h5py and h5dump included) was 10.04 GiB (tank,
|
||||||
reads of the double-deflated fixture alone, 4.0 GiB.
|
2026-09-29, commit `0065c6b`), in `writer_huge_chunks_lz4_zstd_round_trip`
|
||||||
|
(the other tests, 8.2 GiB or less; the first run of the suite, six
|
||||||
|
tests, 8.06 GiB on 2026-09-28). Which step of that test holds 10 GiB
|
||||||
|
(clawhdf5's LZ4 or Zstd write or read, or hdf5plugin's) was not
|
||||||
|
measured separately.
|
||||||
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
|
- **32-bit targets (wasm32):** such a chunk cannot be held in memory:
|
||||||
reading or writing one is `FormatError::Overflow` ("exceeds the
|
reading or writing one is `FormatError::Overflow` ("exceeds the
|
||||||
addressable size" / "address space"); the file still opens and lists
|
addressable size" / "address space"); the file still opens and lists
|
||||||
@@ -315,15 +329,12 @@ wrong data.
|
|||||||
decoded, and external references are an error (object references
|
decoded, and external references are an error (object references
|
||||||
decode). A multi-dimensional numeric attribute is returned as a flat
|
decode). A multi-dimensional numeric attribute is returned as a flat
|
||||||
array (its shape is not reported; `AttrValue::Raw` carries the shape).
|
array (its shape is not reported; `AttrValue::Raw` carries the shape).
|
||||||
HDF5 2.0's native complex type (class 11) is read as the equivalent
|
HDF5 2.0's native complex type (class 11) is read as its own type since
|
||||||
`{r, i}` compound: values are right (`Dataset::read_complex_f64`, and
|
2026-09-29 ([fixed](#hdf5-20-native-complex-numbers-read-as-a-r-i-compound));
|
||||||
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs
|
writing it is opt-in (`with_native_complex_f64_data`,
|
||||||
ls`/`dump` and the browser reader show a compound where h5dump 2.x
|
`make_native_complex_f64_type`), and the Python bindings write complex
|
||||||
prints `H5T_COMPLEX_IEEE_F64LE`, so `h5rs dump` of such a file does not
|
arrays as h5py's compound only. `h5rs dump --json` writes it as the
|
||||||
match h5dump 2.x (h5dump 1.14 cannot read it at all). Writing class 11
|
`{r, i}` compound: hdf5-json (h5json 2.0.0) has no complex class.
|
||||||
(added 2026-09-28) is opt-in (`with_native_complex_f64_data`,
|
|
||||||
`make_native_complex_f64_type`); the Python bindings write complex
|
|
||||||
arrays as h5py's compound only.
|
|
||||||
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
|
- **Metadata cache images** (read since 2026-09-26) differ from libhdf5 in
|
||||||
that: libhdf5 fails only the first metadata read of an image it cannot
|
that: libhdf5 fails only the first metadata read of an image it cannot
|
||||||
load and then reads the file's own (possibly stale) metadata, where we
|
load and then reads the file's own (possibly stale) metadata, where we
|
||||||
@@ -533,8 +544,10 @@ browser (`clawhdf5-wasm`'s `openUrl`) open URLs since 2026-09-27.
|
|||||||
again retries it; what was fetched stays cached).
|
again retries it; what was fetched stays cached).
|
||||||
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
- Tested under Node 22 and headless Chromium (Playwright's build) against
|
||||||
a local server, cross-origin included; not in Firefox or Safari.
|
a local server, cross-origin included; not in Firefox or Safari.
|
||||||
- Compound, reference, opaque, bitfield, time and VL-sequence datasets are
|
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
|
||||||
refused with an error naming the type; attributes of those types come back
|
reference, opaque, bitfield, time and VL-sequence datasets are
|
||||||
|
refused with an error naming the type (HDF5 2.0 native complex datasets
|
||||||
|
read, as `[re, im]` pairs); attributes of those types come back
|
||||||
as `value: null` with their `dtype`. (VL strings read, with h5py's
|
as `value: null` with their `dtype`. (VL strings read, with h5py's
|
||||||
values, through the same `VlResolver` as `File` and `h5rs`.)
|
values, through the same `VlResolver` as `File` and `h5rs`.)
|
||||||
- No Zstd or SZIP (both link C): such datasets fail with
|
- No Zstd or SZIP (both link C): such datasets fail with
|
||||||
@@ -616,7 +629,7 @@ given.
|
|||||||
|
|
||||||
## NetCDF-4: files without dimension scales, types and order differed from netCDF-C
|
## NetCDF-4: files without dimension scales, types and order differed from netCDF-C
|
||||||
|
|
||||||
**Status:** fixed 2026-09-29 (branch `fix/netcdf-phony-dims`). Affected
|
**Status:** fixed 2026-09-29 (#30). Affected
|
||||||
every release (v2.1.0 to v2.7.0) and `main` after #27. Wrong metadata only
|
every release (v2.1.0 to v2.7.0) and `main` after #27. Wrong metadata only
|
||||||
(names, order and types of dimensions and variables); values were read
|
(names, order and types of dimensions and variables); values were read
|
||||||
right. Users: the dimensions of a file without dimension scales are now
|
right. Users: the dimensions of a file without dimension scales are now
|
||||||
@@ -669,6 +682,81 @@ netCDF4-python-written files with netCDF-C itself, and
|
|||||||
`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in
|
`corpus_vs_netcdf_c` the corpus (420 of 429 match; the rest are in
|
||||||
[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)).
|
[NetCDF-4: differences from netCDF-C](#netcdf-4-differences-from-netcdf-c)).
|
||||||
|
|
||||||
|
## HDF5 2.0 native complex numbers read as a `{r, i}` compound
|
||||||
|
|
||||||
|
**Status:** fixed 2026-09-29 (#31). Affected v2.2.0 to v2.7.0 (class 11 has parsed since v2.2.0)
|
||||||
|
and `main` until then. Listed until now under
|
||||||
|
[HDF5 features still unsupported](#hdf5-features-still-unsupported).
|
||||||
|
|
||||||
|
A class-11 datatype (`H5T_COMPLEX_IEEE_F64LE`, ...) parsed into the
|
||||||
|
equivalent compound `{r, i}`. Values were right (`read_complex_f64`,
|
||||||
|
numpy complex in Python), but `raw_datatype()`, `dtype()`, `h5rs ls`/`dump`
|
||||||
|
and the browser reader showed a compound, so `h5rs dump` did not match
|
||||||
|
h5dump 2.x, the browser refused the dataset, and a parsed type written
|
||||||
|
back out became a compound. `Datatype::parse` now returns
|
||||||
|
`Datatype::Complex` (in compounds, arrays and VL types too),
|
||||||
|
`Dataset::dtype()` is `DType::Complex(..)`, `h5rs` prints what h5dump,
|
||||||
|
h5ls and h5diff 2.2.0 print (checked on tank, 2026-09-29, against a file
|
||||||
|
written by h5py 3.16 / libhdf5 2.0.0:
|
||||||
|
`cargo test -p clawhdf5-tools --test native_complex_dump`), and the
|
||||||
|
browser reads `[re, im]` pairs.
|
||||||
|
|
||||||
|
**What users must do:** code that matched a native complex type as a
|
||||||
|
`Datatype::Compound` (or a `DType::Compound` of `r`/`i`) must match
|
||||||
|
`Datatype::Complex` / `DType::Complex`, or take the compound view with
|
||||||
|
`Datatype::complex_as_compound`; exhaustive `match`es over `DType` need a
|
||||||
|
`Complex` arm. Files are unchanged; nothing needs rewriting.
|
||||||
|
|
||||||
|
## Chunk dimensions of 2^32 or more were refused
|
||||||
|
|
||||||
|
**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). An error, never wrong
|
||||||
|
data. Nothing for users to do but upgrade; code that matches on
|
||||||
|
`DataLayout::Chunked::chunk_dimensions` gets `u64`s now.
|
||||||
|
|
||||||
|
libhdf5 2.x writes a chunk larger than 4 GiB with layout message version
|
||||||
|
5, whose dimensions take up to 8 bytes each, so a chunk dimension can be
|
||||||
|
2^32 or more (a 4 GiB chunk of 1-byte elements needs one).
|
||||||
|
`DataLayout::Chunked::chunk_dimensions` was `Vec<u32>`: such a file was
|
||||||
|
refused when opened for reading (`InvalidChunkDimensions`, "larger than
|
||||||
|
2^32 - 1"), and the writer refused such a chunk. Chunk dimensions are
|
||||||
|
`u64` throughout (format crate, facade, editor, `h5rs`, Python bindings);
|
||||||
|
a version-3 layout, whose dimensions are 4 bytes, is never written for
|
||||||
|
them (a chunk that large always takes version 5, as in libhdf5). Tests:
|
||||||
|
`huge_chunks_interop` (`huge_chunk_dims_list` always, over
|
||||||
|
`fixtures/huge_chunk_dims.h5`, which libhdf5 2.0.0 wrote: `u8` chunks of
|
||||||
|
2^32 + 7 in a Single Chunk and an Extensible Array index; with
|
||||||
|
`CLAWHDF5_HUGE_CHUNKS=1`, reads of it and of an unfiltered sparse file
|
||||||
|
h5py writes, and clawhdf5's own such chunks read back by h5py 3.16 and
|
||||||
|
h5dump 2.2.0), unit tests of the 5-byte encoding, and the wasm package
|
||||||
|
test (the file lists; reading such a chunk on wasm32 is
|
||||||
|
`FormatError::Overflow`).
|
||||||
|
|
||||||
|
## Writing a 4 GiB unfiltered chunk held several copies of it
|
||||||
|
|
||||||
|
**Status:** fixed 2026-09-29 (#32); affected every release (v2.1.0 to v2.7.0). Memory only. To write
|
||||||
|
large data with the least memory, hand the builder the data
|
||||||
|
(`DatasetBuilder::with_u8_data_owned`) and write with
|
||||||
|
`FileBuilder::write` (or `FileWriter::finish_with`).
|
||||||
|
|
||||||
|
`FileWriter::finish` copied every chunk into a per-chunk buffer, laid the
|
||||||
|
chunks out into one buffer per pass (two passes, the first kept), then
|
||||||
|
copied everything into the file's buffer, and `FileBuilder::write` wrote
|
||||||
|
that buffer out: about five times a dataset's size for one unfiltered
|
||||||
|
chunk, six with `with_u8_data` (which copies its argument). Measured
|
||||||
|
before and after on tank, 2026-09-29, peak resident memory of
|
||||||
|
`write_one_chunk 1073741824 OUT <mode>` (`crates/clawhdf5/examples`, one
|
||||||
|
unfiltered 1 GiB chunk of `u8`): `owned` 5.0 GiB before (the writer at
|
||||||
|
`4260af4`), 1.0 GiB after (commit `0065c6b`); `slice` 6.0 GiB before,
|
||||||
|
2.0 GiB after. With 4 GiB + 7 bytes and `owned`: 4.00 GiB after. The
|
||||||
|
files written are byte for byte the same. Now a chunk that is a contiguous
|
||||||
|
run of the dataset's data is borrowed (and a filtered one compressed
|
||||||
|
straight from it), the layout refers to chunks instead of copying them,
|
||||||
|
and `FileBuilder::write` streams the file to disk
|
||||||
|
(`FileWriter::finish_with`); contiguous datasets are not copied either.
|
||||||
|
Test: `writer_unfiltered_huge_chunk_round_trip` (with
|
||||||
|
`CLAWHDF5_HUGE_CHUNKS=1`: 4 GiB + 7 bytes written and read back by
|
||||||
|
clawhdf5, h5py 3.16 and h5dump 2.2.0; the test process peaked at 4.00 GiB).
|
||||||
|
|
||||||
## NetCDF-4: variables' dimensions are guessed from sizes
|
## NetCDF-4: variables' dimensions are guessed from sizes
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -123,8 +123,12 @@ else throws an `Error` naming the type.
|
|||||||
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
|
`readHyperslab`). Files of 4 GiB or more are refused at open (wasm32).
|
||||||
Every response body is cut off past the length asked for. More in
|
Every response body is cut off past the length asked for. More in
|
||||||
`docs/known-issues.md`.
|
`docs/known-issues.md`.
|
||||||
- Compound, reference, opaque and variable-length-sequence datasets are
|
- HDF5 2.0 native complex datasets read as `[re, im]` pairs: `dtype`
|
||||||
refused with an error. Attributes of those types are listed with
|
`complex<f64>`, `elementShape` ending in `2`, the parts interleaved in
|
||||||
|
the typed array.
|
||||||
|
- Compound (h5py's complex numbers, a compound `{r, i}`, included),
|
||||||
|
reference, opaque and variable-length-sequence datasets are refused with
|
||||||
|
an error. Attributes of those types are listed with
|
||||||
`value: null` and their `dtype`.
|
`value: null` and their `dtype`.
|
||||||
- No Zstd or SZIP filters (they link C): such a dataset fails with
|
- No Zstd or SZIP filters (they link C): such a dataset fails with
|
||||||
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
|
`unsupported filter`. Deflate, shuffle, Fletcher-32, LZ4, N-Bit and
|
||||||
|
|||||||
@@ -93,6 +93,12 @@ expect $H5 "/vlen_str" '<td>"двa"</td>' '<dd>vlen string</dd>'
|
|||||||
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]<i32></dd>'
|
expect $H5 "/pairs" '<td>[2, 3]</td>' '<dd>array[2]<i32></dd>'
|
||||||
# 3-D: leading dimension held at 0, window over the last two.
|
# 3-D: leading dimension held at 0, window over the last two.
|
||||||
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
|
expect $H5 "/cube" '<dd>(2, 5, 6)</dd>' '<td>29</td>' 'dim 0'
|
||||||
|
# HDF5 2.0 native complex (written when h5py's libhdf5 is 2.0 or later):
|
||||||
|
# each cell is the [re, im] pair.
|
||||||
|
if "$PY" -c 'import sys, h5py; sys.exit(not getattr(h5py.get_config(), "has_native_complex", False))'; then
|
||||||
|
expect $H5 "/native_c128" '<dd>complex<f64></dd>' '<dd>(3, 2)</dd>' '<td>[1, -1.5]</td>' \
|
||||||
|
'<td>[5, -7.5]</td>'
|
||||||
|
fi
|
||||||
# Unsupported type: an error, not values.
|
# Unsupported type: an error, not values.
|
||||||
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
|
expect $H5 "/table" 'class="error"' 'reading compound{x: f64, n: i32} datasets is not supported'
|
||||||
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
|
# Cross-origin: the page is on 127.0.0.1, the file on localhost. CORS that
|
||||||
|
|||||||
@@ -80,6 +80,15 @@ with h5py.File(h5, "w") as f:
|
|||||||
**hdf5plugin.Zstd())
|
**hdf5plugin.Zstd())
|
||||||
comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")])
|
comp = np.zeros(2, dtype=[("x", "<f8"), ("n", "<i4")])
|
||||||
f.create_dataset("table", data=comp)
|
f.create_dataset("table", data=comp)
|
||||||
|
if getattr(h5py.get_config(), "has_native_complex", False):
|
||||||
|
# HDF5 2.0 native complex (class 11): read as [re, im] pairs.
|
||||||
|
from h5py import h5d, h5s, h5t
|
||||||
|
|
||||||
|
for name, t, dt in [(b"native_c128", h5t.COMPLEX_IEEE_F64LE, "<c16"),
|
||||||
|
(b"native_c64_be", h5t.COMPLEX_IEEE_F32BE, ">c8")]:
|
||||||
|
z = (np.arange(6) - 1.5j * np.arange(6)).reshape(3, 2).astype(dt)
|
||||||
|
d = h5d.create(f.id, name, t, h5s.create_simple(z.shape))
|
||||||
|
d.write(h5s.ALL, h5s.ALL, z, mtype=t)
|
||||||
g = f.create_group("sensors")
|
g = f.create_group("sensors")
|
||||||
g.attrs["location"] = "lab"
|
g.attrs["location"] = "lab"
|
||||||
g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4"))
|
g.create_dataset("temp", data=np.array([21.5, 22.0, 22.25], dtype="<f4"))
|
||||||
@@ -105,6 +114,8 @@ def kind(dt):
|
|||||||
return "strings"
|
return "strings"
|
||||||
if dt.subdtype is not None:
|
if dt.subdtype is not None:
|
||||||
return kind(dt.subdtype[0])
|
return kind(dt.subdtype[0])
|
||||||
|
if dt.kind == "c":
|
||||||
|
return "f64" if dt.itemsize == 16 else "f32"
|
||||||
if dt.kind == "f":
|
if dt.kind == "f":
|
||||||
return "f64" if dt.itemsize == 8 else "f32"
|
return "f64" if dt.itemsize == 8 else "f32"
|
||||||
if dt.kind in "iu":
|
if dt.kind in "iu":
|
||||||
@@ -120,20 +131,25 @@ def flat(a, k):
|
|||||||
return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()]
|
return [x.decode() if isinstance(x, bytes) else str(x) for x in a.ravel()]
|
||||||
if k.startswith(("i", "u")):
|
if k.startswith(("i", "u")):
|
||||||
return [str(int(x)) for x in a.ravel()]
|
return [str(int(x)) for x in a.ravel()]
|
||||||
|
if a.dtype.kind == "c":
|
||||||
|
# [re, im] pairs, in order.
|
||||||
|
a = np.stack([a.real, a.imag], axis=-1)
|
||||||
return [float(x) for x in a.ravel()]
|
return [float(x) for x in a.ravel()]
|
||||||
|
|
||||||
|
|
||||||
def entry(ds, slab=None):
|
def entry(ds, slab=None):
|
||||||
k = kind(ds.dtype)
|
k = kind(ds.dtype)
|
||||||
data = ds[()]
|
data = ds[()]
|
||||||
e = {"kind": k, "shape": list(np.shape(data)), "values": flat(data, k)}
|
# A complex element reads as its two parts: one more dimension.
|
||||||
|
pair = [2] if ds.dtype.kind == "c" else []
|
||||||
|
e = {"kind": k, "shape": list(np.shape(data)) + pair, "values": flat(data, k)}
|
||||||
if slab:
|
if slab:
|
||||||
start, count, stride = slab
|
start, count, stride = slab
|
||||||
idx = tuple(slice(s, s + (c - 1) * st + 1, st)
|
idx = tuple(slice(s, s + (c - 1) * st + 1, st)
|
||||||
for s, c, st in zip(start, count, stride))
|
for s, c, st in zip(start, count, stride))
|
||||||
part = ds[idx]
|
part = ds[idx]
|
||||||
e["slab"] = {"start": start, "count": count, "stride": stride,
|
e["slab"] = {"start": start, "count": count, "stride": stride,
|
||||||
"shape": list(part.shape), "values": flat(part, k)}
|
"shape": list(part.shape) + pair, "values": flat(part, k)}
|
||||||
return e
|
return e
|
||||||
|
|
||||||
|
|
||||||
@@ -176,7 +192,10 @@ def describe(path):
|
|||||||
"datasets": sorted(sets)}
|
"datasets": sorted(sets)}
|
||||||
for n, o in members.items():
|
for n, o in members.items():
|
||||||
walk(key.rstrip("/") + "/" + n, o)
|
walk(key.rstrip("/") + "/" + n, o)
|
||||||
elif obj.dtype.names:
|
elif obj.dtype.names or (
|
||||||
|
obj.dtype.kind == "c"
|
||||||
|
and obj.id.get_type().get_class() != getattr(h5py.h5t, "COMPLEX", None)):
|
||||||
|
# h5py's own complex numbers are a compound {r, i}: refused.
|
||||||
expected["errors"][key] = "compound"
|
expected["errors"][key] = "compound"
|
||||||
else:
|
else:
|
||||||
expected["datasets"][key] = entry(obj, slab_for(obj))
|
expected["datasets"][key] = entry(obj, slab_for(obj))
|
||||||
|
|||||||
@@ -399,7 +399,7 @@ async function remoteTests() {
|
|||||||
await fails(async () => {
|
await fails(async () => {
|
||||||
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
const f = await pkg.openUrl(`${base}/fix/fixture.h5`, {
|
||||||
blockSize: 512,
|
blockSize: 512,
|
||||||
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": "bytes 0-511/25752" })),
|
fetch: tamper(async (r) => withHeaders(r, { "Content-Range": `bytes 0-511/${statSync(join(fixDir, "fixture.h5")).size}` })),
|
||||||
});
|
});
|
||||||
await f.read("/grid");
|
await f.read("/grid");
|
||||||
}, /the server sent 0-511/, "wrong range");
|
}, /the server sent 0-511/, "wrong range");
|
||||||
@@ -514,6 +514,15 @@ async function limitTests() {
|
|||||||
`huge chunk: ${name}`);
|
`huge chunk: ${name}`);
|
||||||
}
|
}
|
||||||
hc.free();
|
hc.free();
|
||||||
|
// Likewise chunks whose dimension is 2^32 or more (u8 chunks of 2^32 + 7).
|
||||||
|
const dims = pkg.open(new Uint8Array(readFileSync(join(import.meta.dirname, "..", "..", "..",
|
||||||
|
"crates/clawhdf5/tests/fixtures/huge_chunk_dims.h5"))));
|
||||||
|
eq(dims.list("/").map((e) => e.name), ["earray", "single"], "huge chunk dims: list");
|
||||||
|
for (const name of ["single", "earray"]) {
|
||||||
|
await fails(() => dims.readHyperslab(`/${name}`, [0], [4]), /exceeds this platform.s address space/,
|
||||||
|
`huge chunk dims: ${name}`);
|
||||||
|
}
|
||||||
|
dims.free();
|
||||||
}
|
}
|
||||||
|
|
||||||
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
|
// A body of `total` bytes in 64 KiB pieces, made as they are read; `pulled()`
|
||||||
@@ -550,7 +559,7 @@ async function floodTests() {
|
|||||||
const range = new Headers(init.headers).get("Range");
|
const range = new Headers(init.headers).get("Range");
|
||||||
if (range === "bytes=0-511") return fetch(url, init);
|
if (range === "bytes=0-511") return fetch(url, init);
|
||||||
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
|
const m = /^bytes=(\d+)-(\d+)$/.exec(range);
|
||||||
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/25752` } });
|
return new Response(f.body, { status: 206, headers: { "Content-Range": `bytes ${m[1]}-${m[2]}/${statSync(join(fixDir, "fixture.h5")).size}` } });
|
||||||
},
|
},
|
||||||
}).catch((e) => e);
|
}).catch((e) => e);
|
||||||
// The open itself may need a second range: flooded either way.
|
// The open itself may need a second range: flooded either way.
|
||||||
|
|||||||
Reference in New Issue
Block a user